CGO 2026
2026 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
Powered by
Conference Publishing Consulting

2026 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), January 31 – February 4, 2026, Sydney, Australia

CGO 2026 – Proceedings

Contents - Abstracts - Authors

Frontmatter

Title Page
Article: cgo26foreword-fm000-p (type: Frontmatter) doi:
Welcome from the General Chair
Article: cgo26foreword-fm001-p (type: Frontmatter) doi:
Welcome from the Program Chairs
Article: cgo26foreword-fm004-p (type: Frontmatter) doi:
CGO 2026 Organization
Article: cgo26foreword-fm002-p (type: Frontmatter) doi:
CGO 2026 Sponsors and Supporters
Article: cgo26foreword-fm003-p (type: Frontmatter) doi:

Compiling for ML 1

Enabling Spill-Free Compilation via Affine-Based Live Range Reduction Optimization
Prasanth Chatarasi, Alex Gatea, Wei Wang, Chris Bowler, Shubham Jain, Masoud Ataei Jaliseh, Nicole Khoun, Alberto Mannari, Bardia Mahjour, Viji Srinivasan, and Swagath Venkataramani
(IBM Research, USA; IBM, Canada; IBM, Switzerland)
Publisher's Version Archive submitted (72 kB) Article: cgo26main-p24-p (type: Full Paper) doi:
GRANII: Selection and Ordering of Primitives in GRAph Neural Networks using Input Inspection
Damitha Lenadora, Vimarsh Sathia, Gerasimos Gerogiannis, Serif Yesil, Josep Torrellas, and Charith Mendis
(University of Illinois at Urbana-Champaign, USA; NVIDIA, USA)
Publisher's Version Archive submitted (270 kB) Artifacts Functional Article: cgo26main-p52-p (type: Full Paper) doi:
Fast Autoscheduling for Sparse ML Frameworks
Bobby Yan, Alexander J Root, Trevor Gale, David Broman, and Fredrik Kjolstad
(Stanford University, USA; KTH Royal Institute of Technology, Sweden)
Publisher's Version Archive submitted (240 kB) Article: cgo26main-p275-p (type: Full Paper) doi:
Eliminating Redundancy: Ultra-compact Code Generation for Programmable Dataflow Accelerators
Prasanth Chatarasi, Alex Gatea, Bardia Mahjour, Jintao Zhang, Alberto Mannari, Chris Bowler, Shubham Jain, Masoud Ataei Jaliseh, Nicole Khoun, Kamlesh Kumar, Viji Srinivasan, and Swagath Venkataramani
(IBM Research, USA; IBM, Canada; IBM, Switzerland)
Publisher's Version Article: cgo26main-p29-p (type: Full Paper) doi:

Security

PriTran: Privacy-Preserving Inference for Transformer-Based Language Models under Fully Homomorphic Encryption
Yuechen Mu, Guangli Li, Shiping Chen, and Jingling Xue
(UNSW, Australia; Institute of Computing Technology at Chinese Academy of Sciences, China; CSIRO’s Data61, Australia)
Publisher's Version Article: cgo26main-p35-p (type: Full Paper) doi:
FHEFusion: Enabling Operator Fusion in FHE Compilers for Depth-Efficient DNN Inference
Tianxiang Sui, Jianxin Lai, Long Li, Peng Yuan, Yan Liu, Qing Zhu, Xiaojing Zhang, Linjie Xiao, Mingzhe Zhang, and Jingling Xue
(Ant Group, China; UNSW, Australia)
Publisher's Version Published Artifact Archive submitted (140 kB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p45-p (type: Full Paper) doi:
Reproduction Package for Article 'FHEFusion: Enabling Operator Fusion in FHE Compilers for Depth-Efficient DNN Inference' (doi:10.5281/zenodo.17630291): It contains artifact for the FHEFusion, include compiler source code, test models and scripts.
Towards Path-Aware Coverage-Guided Fuzzing
Giacomo Priamo, Daniele Cono D'Elia, Mathias Payer, and Leonardo Querzoni
(Sapienza University of Rome, Italy; EPFL, Switzerland)
Publisher's Version Published Artifact Archive submitted (140 kB) Artifacts Available Artifacts Reusable Article: cgo26main-p50-p (type: Full Paper) doi:
Artifact for "Towards Path-aware Coverage-guided Fuzzing" (doi:10.6084/m9.figshare.30646583.v2): This artifact provides the source code and evaluation scripts used in the "Towards Path-aware Coverage-guided Fuzzing" paper. The purpose of this work is to improve coverage-guided fuzzing by introducing a path-aware instrumentation: an intra-procedural execution-path-based feedback mechanism which increases ...
SecSwift, a Compiler-Based Framework for Software Countermeasures in Cybersecurity
François de Ferrière, Yves Janin, and Sirine Mechmech
(STMICROELECTRONICS, France; Grenoble INP, France)
Publisher's Version Article: cgo26main-p121-p (type: Full Paper) doi:

Abstractions

Partial-Evaluation Templates: Accelerating Partial Evaluation with Pre-compiled Templates
Florian Huemer, Aleksandar Prokopec, David Leopoldseder, Raphael Mosaner, and Hanspeter Mössenböck
(JKU Linz, Austria; Oracle Labs, Zurich, Switzerland; Oracle Labs, Vienna, Austria; Oracle Labs, Linz, Austria)
Publisher's Version Article: cgo26main-p38-p (type: Full Paper) doi:
Pyls: Enabling Python Hardware Synthesis with Dynamic Polymorphism via LCRS Encoding
Bolei Tong, Yongyan Fang, Chaorui Wang, Qingan Li, Jingling Xue, and Yuan Mengting
(Wuhan University, China; UNSW, Australia)
Publisher's Version Article: cgo26main-p143-p (type: Full Paper) doi:
SkeleShare: Algorithmic Skeletons and Equality Saturation for Hardware Resource Sharing
Jonathan Van der Cruysse, Tzung-Han Juang, Shakiba Bolbolian Khah, and Christophe Dubach
(McGill University, Canada; Mila, Canada)
Publisher's Version Published Artifact Archive submitted (190 kB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p189-p (type: Full Paper) doi:
SkeleShare: Algorithmic Skeletons and Equality Saturation for Hardware Resource Sharing (doi:10.5281/zenodo.17925912): This artifact contains logic to evaluate the effectiveness of SkeleShare, as presented in SkeleShare: Algorithmic Skeletons and Equality Saturation for Hardware Resource Sharing, an upcoming CGO'26 paper. SkeleShare is a fully automated system for resource allocation and hardware sharing in functional FPGA ...
Ember: A Compiler for Embedding Operations on Decoupled Access-Execute Architectures
Marco Siracusa, Olivia Hsu, Víctor Soria-Pardos, Joshua Randall, Arnaud Grasset, Eric Biscondi, Doug Joseph, Randy Allen, Fredrik Kjolstad, Miquel Moretó Planas, and Adrià Armejach
(Barcelona Supercomputing Center, Spain; Stanford University, USA; Carnegie Mellon University, USA; Arm, USA; Universitat Politècnica de Catalunya, Spain)
Publisher's Version Published Artifact Archive submitted (100 kB) Artifacts Available Article: cgo26main-p78-p (type: Full Paper) doi:
Ember: A Compiler for Embedding Operations on Decoupled Access-Execute Architectures (doi:10.5281/zenodo.17636956): This repo contains an MLIR implementation of Ember, a compiler to lower PyTorch and TensorFlow embedding operations to Decoupled Access-Execute (DAE) architectures.

Memory

Flow-Graph-Aware Tiling and Rescheduling for Memory-Efficient On-Device Inference
Yeonoh Jeong, Taehyeong Park, and Yongjun Park
(Yonsei University, Republic of Korea)
Publisher's Version Article: cgo26main-p136-p (type: Full Paper) doi:
VFlatten: Selective Value-Object Flattening using Hybrid Static and Dynamic Analysis
Arjun H. Kumar, Bhavya Hirani, Hang Shao, Tobi Ajila, Vijay Sundaresan, Daryl Maier, and Manas Thakur
(IIT Mandi, India; Sardar Vallabhbhai National Institute of Technology, Surat, India; IBM, Canada; IIT Bombay, India)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p164-p (type: Full Paper) doi:
VFlatten: Selective Value-Object Flattening using Hybrid Static and Dynamic Analysis (doi:10.5281/zenodo.17212535): Artifact of VFlatten: Selective Value-Object Flattening using Hybrid Static and Dynamic Analysis, published at CGO 2026.
FRUGAL: Pushing GPU Applications beyond Memory Limits
Lingqi Zhang, Tengfei Wang, Jiajun Huang, Chen Zhuang, Ivan R. Ivanov, Peng Chen, Toshio Endo, and Mohamed Wahib
(RIKEN RCCS, Japan; Google Cloud, Japan; University of South Florida, USA; Institute of Science Tokyo, Japan)
Publisher's Version Archive submitted (510 kB) Article: cgo26main-p193-p (type: Full Paper) doi:
Automatic Data Enumeration for Fast Collections
Tommy McMichen and Simone Campanoni
(Northwestern University, USA; Google, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p20-p (type: Full Paper) doi:
Automatic Data Enumeration for Fast Collections (doi:10.5281/zenodo.17872601): Artifact for "Automatic Data Enumeration for Fast Collections", accepted at CGO'26 Abstract: Data collections provide a powerful abstraction to organize data, simplifying development and maintenance. Choosing an implementation for each collection is a critical decision, with performance, memory and energy tradeoffs ...

DSLs

FORTE: Online DataFrame Query Optimizer
Yoonho Choi, Kyoungtae Lee, Minji Kim, Hyungsoo Jung, and Hyojin Sung
(POSTECH, Republic of Korea; Seoul National University, Republic of Korea; Ewha Womans University, Republic of Korea)
Publisher's Version Article: cgo26main-p98-p (type: Full Paper) doi:
LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
Amir Mohammad Tavakkoli, Cosmin E. Oancea, and Mary Hall
(University of Utah, USA; University of Copenhagen, Denmark)
Publisher's Version Published Artifact Info Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p293-p (type: Full Paper) doi:
Reproduction Package for "LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping" (doi:10.5281/zenodo.17633994): This artifact contains the source code of the LEGO framework and the scripts used to execute and evaluate all benchmarks in the paper. LEGO provides an algebraic, compileragnostic framework for specifying and transforming memory layouts. Through integrations with Triton, CUDA, and MLIR, we compare LEGO-generated ...
Pushing Tensor Accelerators beyond MatMul in a User-Schedulable Language
Yihong Zhang, Derek Gerstmann, Andrew Adams, and Maaz Bin Safeer Ahmad
(University of Washington, USA; Adobe, USA)
Publisher's Version Published Artifact Artifacts Available Article: cgo26main-p331-p (type: Full Paper) doi:
Artifact of CGO 2026 "Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language" (doi:10.5281/zenodo.17810573): The HardBoiled implementation as described in the paper
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
Hongzheng Chen, Bin Fan, Alexander Collins, Bastian Hagedorn, Evghenii Gaburov, Masahiro Masuda, Matthew Brookhart, Chris Sullivan, Jason Knight, Zhiru Zhang, and Vinod Grover
(Cornell University, USA; NVIDIA, USA; NVIDIA, UK; NVIDIA, Germany)
Publisher's Version Article: cgo26main-p90-p (type: Full Paper) doi:

Quantum / HLS

Dependence-Driven, Scalable Quantum Circuit Mapping with Affine Abstractions
Marouane Benbetka, Merwan Bekkar, Riyadh Baghdadi, and Martin Kong
(NYU Abu Dhabi, United Arab Emirates; Ohio State University, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p61-p (type: Full Paper) doi:
Reproduction Package for Article 'Dependence-Driven, Scalable Quantum Circuit Mapping with Affine Abstractions' (doi:10.5281/zenodo.16899145): This artifact provides an implementation of Qlosure, a dependence-driven quantum circuit mapping algorithm that uses affine abstractions to systematically uncover and exploit transitive dependences during qubit routing. The artifact reproduces all key experimental results reported in the paper “Dependence-Driven, ...
Space-Time Optimisations for Early Fault-Tolerant Quantum Computation
Sanaa Sharma and Prakash Murali
(University of Cambridge, UK)
Publisher's Version Published Artifact Artifacts Available Article: cgo26main-p151-p (type: Full Paper) doi:
Space-Time Optimisations for Early Fault-Tolerant Quantum Computation (doi:10.5281/zenodo.17632718): Space-Time Optimisations for Early Fault-Tolerant Quantum Computation Based on work described in https://doi.org/10.48550/arXiv.2511.08848 To run the experiments, use the script.py file. The default hamiltonian is set to ising_model. The experiments are done for data qubits ranging from n=4 to n=100. The initial ...
OpenQudit: Extensible and Accelerated Numerical Quantum Compilation via a JIT-Compiled DSL
Ed Younis
(Lawrence Berkeley National Laboratory, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p308-p (type: Full Paper) doi:
OpenQudit: Extensible and Accelerated Numerical Quantum Compilation via a JIT-Compiled DSL (Artifact) (doi:10.5281/zenodo.17574908): This artifact contains the frozen-in-time source code of OpenQudit used in the evaluation, along with the scripts used to install dependencies, gather data, and plot the results shown in Figures 4, 6, and 7. An automated pipeline is provided to build, run, and plot the experiments using Docker, which is designed to be ...
Selene: Cross-Level Barrier-Free Pipelining for Irregular Nested Loops in High-Level Synthesis
Sungwoo Yun, Seonyoung Cheon, Dongkwan Kim, Heelim Choi, Kunmo Jeong, Chan Lee, Yongwoo Lee, and Hanjun Kim
(Yonsei University, Republic of Korea; DGIST, Republic of Korea)
Publisher's Version Article: cgo26main-p76-p (type: Full Paper) doi:

Parallelization / Vectorization

Enabling Automatic Compiler-Driven Vectorization of Transformers
Shreya Alladi, Alberto Ros, and Alexandra Jimborean
(University of Murcia, Spain)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p71-p (type: Full Paper) doi:
OML-vect: Enabling Automatic Compiler-Driven Vectorization of Transformers (doi:10.5281/zenodo.18006548): This is the supporting artifact for the paper titled "Enabling Automatic Compiler-Driven Vectorization of Transformers" as published in CGO 2026. It includes the source code for the proposed tool (oml-vect), along with scripts to compile and reproduce all experimental results presented in the paper. We provide a ...
Unlocking Python Multithreading Capabilities using OpenMP-Based Programming with OMP4Py
César Piñeiro and Juan C. Pichel
(University of Santiago de Compostela, Spain)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p47-p (type: Full Paper) doi:
Material for article "Unlocking Python Multithreading Capabilities using OpenMP-Based Programming with OMP4Py" (doi:10.5281/zenodo.17987713): Artifact for the OMP4Py project submitted to the CGO 2026 conference for artifact evaluation. This repository contains the materials, code, and supplementary resources required to reproduce and validate the experimental results presented in the submission.
The Parallel-Semantics Program Dependence Graph for Parallel Optimization
Yian Su, Brian Homerding, Haocheng Gao, Federico Sossai, Yebin Chon, David I. August, and Simone Campanoni
(Northwestern University, USA; Princeton University, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Results Reproduced Article: cgo26main-p58-p (type: Full Paper) doi:
The Parallel-Semantics Program Dependence Graph for Parallel Optimization (doi:10.5281/zenodo.17633089): The artifact includes a fully automatic workflow to generate all experimental results included in the paper. The workflow runs GINO to compile all benchmarks evaluated in the paper. Then, the generated binaries are invoked to generate the raw results (e.g., the execution times of multiple runs of a benchmark). Then, ...
From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization
Shuaijiang Li, Jiacheng Zhao, Ying Liu, Shuoming Zhang, Lei Chen, Yijin Li, Yangyu Zhang, Zhicheng Li, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng, and Huimin Cui
(Institute of Computing Technology at Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, China; University of Leeds, UK; University of Aberdeen, UK; XCORESIGMA, China)
Publisher's Version Article: cgo26main-p77-p (type: Full Paper) doi:

Binary / JIT

Binary Diffing via Library Signatures
Andrei Rimsa, Anderson Faustino da Silva, Camilo Santana, and Fernando Magno Quintão Pereira
(CEFET-MG, Brazil; State University of Maringá, Brazil; Federal University of Minas Gerais, Brazil)
Publisher's Version Published Artifact Info Artifacts Available Artifacts Functional Results Reproduced Article: cgo26main-p34-p (type: Full Paper) doi:
LIbSIG's artifacts (doi:10.5281/zenodo.17082032): This artifact reproduces the experiments conducted in IV. A docker container with scripts to rebuild the paper results, perform the game and produce all the figures and tables automatically can be found in Zenodo.
BIT: Empowering Binary Analysis through the LLVM Toolchain
Puzhuo Liu, Peng Di, Jingling Xue, and Yu Jiang
(Ant Group, China; Tsinghua University, China; UNSW, Australia)
Publisher's Version Article: cgo26main-p14-p (type: Full Paper) doi:
Dr.avx: A Dynamic Compilation System for Seamlessly Executing Hardware-Unsupported Vectorization Instructions
Yue Tang, Mianzhi Wu, Yufeng Li, Haoyu Liao, Jianmei Guo, and Bo Huang
(East China Normal University, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Article: cgo26main-p33-p (type: Full Paper) doi:
Dr.avx AE (doi:10.5281/zenodo.17902199): This is an artifact containing a pre-configured Docker environment and automated pipeline scripts to facilitate the reproduction of the key results in the paper, from executing experiments to generating performance plots.
Practical: Are Abstract-Interpreter Baseline JITs Worth It? An Empirical Evaluation through Metacompilation
Nahuel Palumbo, Guillermo Polito, Stéphane Ducasse, and Pablo Tesone
(Univ. Lille - Inria - CNRS - Centrale Lille - UMR 9189 CRIStAL, France)
Publisher's Version Archive submitted (77 kB) Article: cgo26main-p313-p (type: Full Paper) doi:

Code Generation

TPDE: A Fast Adaptable Compiler Back-End Framework
Tobias Schwarz, Tobias Kamm, and Alexis Engelke
(TU Munich, Germany)
Publisher's Version Published Artifact Archive submitted (70 kB) Info Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p144-p (type: Full Paper) doi:
Artifact for "TPDE: A Fast Adaptable Compiler Back-End Framework" (doi:10.5281/zenodo.17867601): Artifact containing all tools and scripts to obtain the performance results and plots shown in the paper. SPEC CPU 2017 is not included; we therefore include the LLVM test-suite benchmarks as freely available substitute.
Synthesizing Instruction Selection Back-Ends from ISA Specifications Made Practical
Florian Drescher and Alexis Engelke
(TU Munich, Germany)
Publisher's Version Article: cgo26main-p173-p (type: Full Paper) doi:
SparseX: Synergizing GPU Libraries for Sparse Matrix Multiplication on Heterogeneous Processors
Ruifeng Zhang, Xiangwei Wang, Ang Li, and Xipeng Shen
(North Carolina State University, USA; Pacific Northwest National Laboratory, USA; University of Washington, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p261-p (type: Full Paper) doi:
Reproduction Package for Article "SparseX: Synergizing GPU Libraries for Sparse Matrix Multiplication on Heterogeneous Processors" (doi:10.6084/m9.figshare.30639536.v1): Artifact for reproducing the key results of the paper "SparseX: Synergizing GPU Libraries for Sparse Matrix Multiplication on Heterogeneous Processors." SparseX is an adaptive SpMM framework that selects and combines GPU kernels from multiple libraries (cuSPARSE, cuBLAS, Sputnik, CLASP, Jigsaw) to exploit ...
Compilation of Generalized Matrix Chains with Symbolic Sizes
Francisco López, Lars Karlsson, and Paolo Bientinesi
(Umeå University, Sweden)
Publisher's Version Published Artifact Archive submitted (350 kB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p180-p (type: Full Paper) doi:
Compilation of Generalized Matrix Chains with Symbolic Sizes (doi:10.5281/zenodo.17808079): The artifact presents a code generator for Generalized Matrix Chains with symbolic sizes that relies on novel theoretical results presented in the paper.

Profiling / Instrumentation

TRACE4J: A Lightweight, Flexible, and Insightful Performance Tracing Tool for Java
Haide He and Pengfei Su
(University of California at Merced, USA)
Publisher's Version Published Artifact Info Artifacts Available Artifacts Functional Results Reproduced Article: cgo26main-p23-p (type: Full Paper) doi:
Reproduction Package for Article `TRACE4J: A Lightweight, Flexible, and Insightful Performance Tracing Tool for Java' (doi:10.5281/zenodo.17181411): The artifact package contains TRACE4J and the benchmark suites, together with step-by-step instructions for reproducing the results shown in the article. For convenience, a prebuilt Docker image with all dependencies is provided.
Proton: Towards Multi-level, Adaptive Profiling for Triton
Keren Zhou, Tianle Zhong, Hao Wu, Jihyeong Lee, Yue Guan, Yufei Ding, Corbin Robeck, Yuanwei Fang, Jeff Niu, and Philippe Tillet
(George Mason University, USA; OpenAI, USA; University of Virginia, USA; University of California at San Diego, USA; Meta, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p46-p (type: Full Paper) doi:
Proton: Towards Multi-level, Adaptive Profiling for Triton (doi:10.5281/zenodo.17972732): The submitted artifact consists of 2 parts. First, the source code of Proton implementation. Second, the scripts demonstrate the usage of Proton and reproduce the key results presented in the paper. The artifact aims to reproduce the results in sections 5 and 6.
On the Precision of Dynamic Program Fingerprints Based on Performance Counters
Anderson Faustino da Silva, Marcelo Borges Nogueira, Sérgio Queiroz de Medeiros, Jeronimo Castrillon, and Fernando Magno Quintão Pereira
(State University of Maringá, Brazil; Federal University of Rio Grande do Norte, Brazil; TU Dresden, Germany; Federal University of Minas Gerais, Brazil)
Publisher's Version Published Artifact Artifacts Available Article: cgo26main-p18-p (type: Full Paper) doi:
On the Precision of Dynamic Program Fingerprints based on Performance Counters (doi:10.5281/zenodo.17574396): Scripts and benchmarks necessary to reproduce the experiments in the paper "On the Precision of Dynamic Program Fingerprints based on Performance Counters".
PASTA: A Modular Program Analysis Tool Framework for Accelerators
Mao Lin, Hyeran Jeon, and Keren Zhou
(University of California at Merced, USA; George Mason University, USA; OpenAI, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Results Reproduced Article: cgo26main-p69-p (type: Full Paper) doi:
PASTA: A Modular Program Analysis Tool Framework for Accelerators (doi:10.5281/zenodo.17547322): This repository contains the artifact for the CGO 2026 paper entitled "PASTA: A Modular Program Analysis Tool Framework for Accelerators".

Analysis

PIP: Making Andersen’s Points-to Analysis Sound and Practical for Incomplete C Programs
Håvard Rognebakke Krogstie, Helge Bahmann, Magnus Själander, and Nico Reissmann
(NTNU, Norway; Independent Researcher, Switzerland; Independent Researcher, Norway)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p36-p (type: Full Paper) doi:
PIP: Making Andersen's Points-to Analysis Sound and Practical for Incomplete C Programs (Artifact) (doi:10.5281/zenodo.16900791): This artifact provides the source code for the `jlm` compiler, including an implementation of the Andersen-style analysis presented in the paper. It also includes the source code for most of the programs benchmarked in the paper, and scripts for performing these benchmarks and producing the figures and tables from the ...
Thinking Fast and Correct: Automated Rewriting of Numerical Code through Compiler Augmentation
Siyuan Brant Qian, Vimarsh Sathia, Ivan R. Ivanov, Jan Hückelheim, Paul Hovland, and William S. Moses
(University of Illinois at Urbana-Champaign, USA; Institute of Science Tokyo, Japan; RIKEN RCCS, Japan; Argonne National Laboratory, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p116-p (type: Full Paper) doi:
Reproduction Artifact for "Thinking Fast and Correct: Automated Rewriting of Numerical Code through Compiler Augmentation" (doi:10.5281/zenodo.18004991): This is the artifact for the paper "Thinking Fast and Correct: Automated Rewriting of Numerical Code through Compiler Augmentation" (CGO 2026) by Siyuan Brant Qian, Vimarsh Sathia, Ivan R. Ivanov, Jan Hückelheim, Paul Hovland, and William S. Moses. The latest version of this artifact is available at the following ...
PolyUFC: Polyhedral Compilation Meets Roofline Analysis for Uncore Frequency Capping
Nilesh Rajendra Shah, M V V S Manoj Kumar, Dhairya Baxi, and Ramakrishna Upadrasta
(IIT Hyderabad, India)
Publisher's Version Info Article: cgo26main-p139-p (type: Full Paper) doi:
Accelerating App Recompilation across Android System Updates by Code Reusing
Hongtao Wu, Yu Chen, Mengfei Xie, Futeng Yang, Jun Yan, Jiang Ma, Jianming Fu, Chun Jason Xue, and Qingan Li
(Wuhan University, China; Guangdong OPPO Mobile Telecommunications, China; Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates)
Publisher's Version Article: cgo26main-p177-p (type: Full Paper) doi:

Compiling for ML 2

QIGen: A Kernel Generator for Inference on Nonuniformly Quantized Large Language Models
Tommaso Pegolotti, Dan Alistarh, and Markus Püschel
(ETH Zurich, Switzerland; IST Austria, Austria)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Article: cgo26main-p70-p (type: Full Paper) doi:
QIGen: A Kernel Generator for Inference on Nonuniformly Quantized Large Language Models (doi:10.5281/zenodo.16894301): Our artifact provides the source code for the QIGen code generator, quantized models to evaluate the performance of the generated kernels, and scripts for reproducing main experiments. The artifact is in the form of a Docker container which provides all required dependencies. More specifically, our artifact consists ...
DyPARS: Dynamic-Shape DNN Optimization via Pareto-Aware MCTS for Graph Variants
Hao Qian, Guangli Li, Qiuchu Yu, Xueying Wang, and Jingling Xue
(UNSW, Australia; Institute of Computing Technology at Chinese Academy of Sciences, China; Beijing University of Posts and Telecommunications, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p118-p (type: Full Paper) doi:
DyPARS: Dynamic-Shape DNN Optimization via Pareto-Aware MCTS for Graph Variants (doi:10.5281/zenodo.16758227): The research artifact enables reproduction of the main experimental results presented in this paper. In this evaluation, it demonstrates that (1) DyPARS improves the inference performance of dynamic-shape DNNs compared with existing deep learning compilers, including PyTorch and BladeDisc, and (2)DyPARS efficiently ...
Compiler-Runtime Co-operative Chain of Verification for LLM-Based Code Optimization
Hyunho Kwon, Sanggyu Shin, Ju Min Lee, Hoyun Youm, Seungbin Song, Seongho Kim, Hanwoong Jung, Seungwon Lee, and Hanjun Kim
(Yonsei University, Republic of Korea; SAIT, Republic of Korea)
Publisher's Version Article: cgo26main-p25-p (type: Full Paper) doi:
Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs
Xiao Zhang, Yaoyao Ding, Bolin Sun, Yang Hu, Tatiana Shpeisman, and Gennady Pekhimenko
(University of Toronto, Canada; NVIDIA, Canada; Vector Institute, Canada)
Publisher's Version Published Artifact Archive submitted (3.6 MB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p5-p (type: Full Paper) doi:
Reproduction Package for "Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs" (doi:10.5281/zenodo.18005106): This artifact contains the benchmark scripts for our paper: **Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs** The structure of this artifact is organized as follows: - `docker`: Contains the Dockerfiles for building the Docker containers needed for the runtime environment. - `hidet`: ...

Tensor Optimization

Multidirectional Propagation of Sparsity Information across Tensor Slices
Kaio Henrique Andrade Ananias, Danila Seliayeu, J. Nelson Amaral, and Fernando Magno Quintão Pereira
(Federal University of Minas Gerais, Brazil; University of Alberta, Canada)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p15-p (type: Full Paper) doi:
Reproduction Packaged for Artical 'Multidirectional Propagation of Sparsity Information across Tensor Slices' (doi:10.5281/zenodo.17593823): A Dockerfile builds the infrastructure necessary to reproduce all the experiments from the evaluation section of the paper. A set of scripts will be automatically called to produce the data and figures. At the end, a directory named `results` should contain a folder for each figure (7-12) with its CSV and PNG files.
Synthesizing Specialized Sparse Tensor Accelerators for FPGAs via High-Level Functional Abstractions
Hamza Javed and Christophe Dubach
(McGill University, Canada; Mila, Canada)
Publisher's Version Article: cgo26main-p128-p (type: Full Paper) doi:
Progressive Low-Precision Approximation of Tensor Operators on GPUs: Enabling Greater Trade-Offs between Performance and Accuracy
Fan Luo, Guangli Li, Zhaoyang Hao, Xueying Wang, Xiaobing Feng, Huimin Cui, and Jingling Xue
(Institute of Computing Technology at Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, China; UNSW, Australia; Beijing University of Posts and Telecommunications, China)
Publisher's Version Article: cgo26main-p140-p (type: Full Paper) doi:
Tensor Program Superoptimization through Cost-Guided Symbolic Program Synthesis
Alexander Brauckmann, Aarsh Chaube, José Wesley de Souza Magalhães, Elizabeth Polgreen, and Michael F. P. O’Boyle
(University of Edinburgh, UK)
Publisher's Version Published Artifact Archive submitted (44 kB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p1649-p (type: Full Paper) doi:
STENSO: Tensor Program Superoptimization through Cost-Guided Symbolic Program Synthesis (Artifact) (doi:10.5281/zenodo.17638077): This artifact provides the complete source code, evaluation benchmarks, and analysis scripts for STENSO, a sketch-based program synthesizer designed for superoptimizing Tensor DSL programs. The artifact includes an automated Docker-based pipeline to: - Build and run the STENSO synthesizer and a bottom-up synthesizer ...

Optimization

A Reinforcement Learning Environment for Automatic Code Optimization in the MLIR Compiler
Mohammed Tirichine, Nassim Ameur, Nazim Bendib, Iheb Nassim Aouadj, Djad Bouchama, Rafik Bouloudene, and Riyadh Baghdadi
(NYU Abu Dhabi, United Arab Emirates; École Nationale Supérieure d’Informatique, Algeria; University of Science and Technology Houari Boumediene, Algeria)
Publisher's Version Published Artifact Archive submitted (130 kB) Artifacts Available Artifacts Reusable Results Reproduced Article: cgo26main-p529-p (type: Full Paper) doi:
MLIR RL Artifact (doi:10.5281/zenodo.17987133): This artifact contains MLIR RL, a deep reinforcement learning (RL) system that optimizes loop nests in MLIR. It includes the RL environment, pre-trained models, MLIR benchmarks, and scripts required to reproduce all results for MLIR RL (Figure 5 and Tables III–IV) and the PyTorch / PyTorch compiler results in Figure 5 ...
Towards Threading the Needle of Debuggable Optimized Binaries
Cristian Assaiante, Simone Di Biasio, Snehasish Kumar, Giuseppe Antonio Di Luna, Daniele Cono D'Elia, and Leonardo Querzoni
(Sapienza University of Rome, Italy; Google, USA)
Publisher's Version Published Artifact Archive submitted (220 kB) Artifacts Available Artifacts Reusable Article: cgo26main-p72-p (type: Full Paper) doi:
Artifact for 'Towards Threading the Needle of Debuggable Optimized Binaries' (doi:10.5281/zenodo.17865056): The artifact contains DebugTuner, a framework for tuning compilers towards the generation of more debuggable programs for a low performance overhead. DebugTuner is made of two components: the first analyzes the effect of compiler optimization passes on debug information availability, and the second ranks and selects ...
Compiler-Assisted Instruction Fusion
Ravikiran Ravindranath Reddy, Sawan Singh, Arthur Perais, Alberto Ros, and Alexandra Jimborean
(University of Murcia, Spain; Univ. Grenoble Alpes - CNRS - Grenoble INP - TIMACNRS, France)
Publisher's Version Archive submitted (37 kB) Article: cgo26main-p85-p (type: Full Paper) doi:
LLM-VeriOpt: Verification-Guided Reinforcement Learning for LLM-Based Compiler Optimization
Xiangxin Fang, Jiaqin Kang, Rodrigo Rocha, Sam Ainsworth, and Lev Mukhanov
(Queen Mary University of London, UK; University of Edinburgh, UK)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Results Reproduced Article: cgo26main-p110-p (type: Full Paper) doi:
LLMVeriOpt for CGO 2026 Artifact Evaluation (doi:10.5281/zenodo.17672452): This artifact contains all materials required to reproduce the experimental results presented in the paper. It includes trained SFT and GRPO LoRA models, evaluation datasets, inference pipelines, configuration files, figure-generation scripts, and a pre-computed summary table aggregating IR statistics, latency, ...

proc time: 0.12