PPoPP 2026
31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP 2026)
Powered by
Conference Publishing Consulting

31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP 2026), January 31 – February 4, 2026, Sydney, NSW, Australia

PPoPP 2026 – Proceedings

Contents - Abstracts - Authors

Frontmatter

Title Page
Article: ppopp26foreword-fm000-p (type: Frontmatter) doi:
Welcome from the Chairs
Article: ppopp26foreword-fm001-p (type: Frontmatter) doi:
PPoPP 2026 Organization
Article: ppopp26foreword-fm002-p (type: Frontmatter) doi:
PPoPP 2026 Sponsors and Supporters
Article: ppopp26foreword-fm003-p (type: Frontmatter) doi:

Concurrency Control

Binary Compatible Critical Section Delegation
Junyao Zhang, Zhuo Wang, and Zhe Zhou
(Fudan University, China; Alibaba Cloud Computing, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Results Reproduced Article: ppopp26main-p191-p (type: Full Paper) doi:10.1145/3774934.3786439
Reproduction Package for Binary Compatible Critical Section Delegation (doi:10.5281/zenodo.18046921): All the codes and scripts to reproduce BCD
Hapax Locks: Scalable Value-Based Mutual Exclusion
Dave Dice and Alex Kogan
(Independent, USA; Oracle Labs, USA)
Publisher's Version Article: ppopp26main-p237-p (type: Full Paper) doi:10.1145/3774934.3786443
Fixing Non-blocking Data Structures for Better Compatibility with Memory Reclamation Schemes
Md Amit Hasan Arovi and Ruslan Nikolaev
(Pennsylvania State University, USA)
Publisher's Version Published Artifact Info Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p332-p (type: Full Paper) doi:10.1145/3774934.3786455
Fixing Non-blocking Data Structures for Better Compatibility with Memory Reclamation Schemes - Artifact for PPoPP'26 (doi:10.5281/zenodo.17707898): The artifact contains a VM image (VirtualBox) with preinstalled Ubuntu 22.04 and the benchmark source code. The artifact also contains instructions for manual (bare-metal) installations. The artifact also includes our data measurements and scripts for generating plots. Please see README.txt for more details.
Multiverse: Transactional Memory with Dynamic Multiversioning
Gaetano Coccimiglio, Trevor Brown, and Srivatsan Ravi
(University of Waterloo, Canada; University of Southern California, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p170-p (type: Full Paper) doi:10.1145/3774934.3786436
Multiverse: Transactional Memory with Dynamic Multiversioning (doi:10.5281/zenodo.18099743): This is an archive of the repository corresponding to the artifact submitted to PPoPP 2026 for the the paper Multiverse: Transactional Memory with Dynamic Multiversioning. This archive includes the full source code for the implementation of the Multiverse algorithm along with implementations for the benchmark and ...

Scheduling and Load Balancing

Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
Aleix Roca and Vicenç Beltran
(Barcelona Supercomputing Center, Spain)
Publisher's Version Article: ppopp26main-p298-p (type: Full Paper) doi:10.1145/3774934.3786451
Waste-Efficient Work Stealing
Kyle Singer, Kunal Agrawal, and Tao B. Schardl
(Massachusetts Institute of Technology, USA; Washington University in St. Louis, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p307-p (type: Full Paper) doi:10.1145/3774934.3786452
Waste-Efficient Work Stealing Artifact Appendix: The artifact appendix for "Waste-Efficient Work Stealing."
Waste-Efficient Work Stealing - Software Artifact (doi:10.5281/zenodo.17706312): Software artifact for the PPoPP'26 paper "Waste-Efficient Work Stealing," by Kyle Singer (MIT CSAIL), Kunal Agrawal (Washington University in St. Louis), and Tao B. Schardl (MIT CSAIL). Contains the code for several versions of the OpenCilk runtime (each implementing a different protocol for waking/sleeping worker ...
DiggerBees: Depth First Search Leveraging Hierarchical Block-Level Stealing on GPUs
Yuyao Niu, Yuechen Lu, Weifeng Liu, and Marc Casas
(Barcelona Supercomputing Center, Spain; Universitat Politècnica de Catalunya, Spain; China University of Petroleum-Beijing, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p340-p (type: Full Paper) doi:10.1145/3774934.3786457
Appendix: Artifact Evaluation: This artifact evaluation accompanies the paper entitled ``DiggerBees: Depth First Search Leveraging Hierarchical Block-Level Stealing on GPUs'.
Reproduction Package for the paper ``DiggerBees: Depth First Search Leveraging Hierarchical Block-Level Stealing on GPUs'' (doi:10.5281/zenodo.18072817): This artifact mainly consists of (1) a quick start based on a pre-built Docker image that bundles all required software environments and dependencies. This guide offers clear setup instructions for running experiments on a small subset of graphs to verify basic correctness and performance, and (2) a detailed ...
PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUs
Haodong Bian, Youhui Zhang, Xiang Fei, Jianqiang Huang, and Xiaoying Wang
(Tsinghua University, China; Qinghai University, China; Zhongguancun Laboratory, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p355-p (type: Full Paper) doi:10.1145/3774934.3786459
PANA Artifact Appendix: We release the source code for our approach, PANA, and for the competitive baselines (CAMLB, Merge, CSR5, CVR, SpV8). We note that the MKL and AOCL implementations utilize commercial libraries from Intel and AMD. Additionally, we make available the 2,898 sparse matrices used in our tests, and the scripts necessary to ...
PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUs (doi:10.5281/zenodo.17709124): We release the source code for our approach, PANA, and for the competitive baselines (CAMLB, Merge, CSR5, CVR, SpV8). We note that the MKL and AOCL implementations utilize commercial libraries from Intel and AMD. Additionally, we make available the 2,898 sparse matrices used in our tests, and the scripts necessary to ...

Concurrent Data Structures

UFO Trees: Practical and Provably-Efficient Parallel Batch-Dynamic Trees
Quinten De Man, Atharva Sharma, Kishen N Gowda, and Laxman Dhulipala
(University of Maryland, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p141-p (type: Full Paper) doi:10.1145/3774934.3786431
ParAlg/UFOTree: Artifact Evaluation Release for PPoPP 2026 (doi:10.5281/zenodo.18141137): Code for paper "UFO Trees: Practical and Provably-Efficient Parallel Batch-Dynamic Trees" (PPoPP 2026).
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
Ajay Singh, Nikos Metaxakis, and Panagiota Fatourou
(ICS-FORTH, Greece; University of Crete, Greece)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p346-p (type: Full Paper) doi:10.1145/3774934.3786458
Reproduction package for "SEC Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks" (doi:10.5281/zenodo.18109078): code: at https://doi.org/10.5281/zenodo.18109078 This microbenchmark builds upon the setbench framework to test and evaluate concurrent stack algorithms. Specifically, this artifact supports the evaluation of SEC (Sharded Elimination and Combining), alongside the state-of-the-art stacks considered in the paper ...
Concurrent Balanced Augmented Trees
Evan Wrench, Ajay Singh, Younghun Roh, Panagiota Fatourou, Siddhartha Jayanti, Eric Ruppert, and Yuanhao Wei
(University of British Columbia, Canada; ICS-FORTH, Greece; Massachusetts Institute of Technology, USA; University of Crete, Greece; Dartmouth College, USA; York University, Canada)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p181-p (type: Full Paper) doi:10.1145/3774934.3786437
Artifact for "Concurrent Balanced Augmented Trees" (doi:10.5281/zenodo.18354864): Contains the implementation and experimental scripts for the paper "Concurrent Balanced Augmented Trees" from PPoPP 2026. Includes instructions for running the code and reproducing results found in the paper.
Parallel Dynamic Spatial Indexes
Ziyang Men, Bo Huang, Yan Gu, and Yihan Sun
(University of California at Riverside, USA)
Publisher's Version Published Artifact Info Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p3-p (type: Full Paper) doi:10.1145/3774934.3786412
Parallel Dynamic Spatial Indexes (doi:10.5281/zenodo.18011501): PSI (Parallel Spatial Indexes) is a highly optimized C++ library for parallel spatial indexing and querying on multi-core architectures. It provides efficient implementations of three spatial index structures: the kd-tree, the Orth-tree, and the R-tree, optimized for parallel processing. PSI is designed to handle ...

GPU and Heterogeneous Computing

PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation
Bing Lu, Zedong Liu, Hairui Zhao, Dejun Luo, Wenjing Huang, Yida Gu, Jinyang Liu, Guangming Tan, and Dingwen Tao
(Institute of Computing Technology at Chinese Academy of Sciences, China; Jilin University, China; University of Chinese Academy of Sciences, China; University of California at Riverside, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p186-p (type: Full Paper) doi:10.1145/3774934.3786438
AE of PRISM: This is aritifact evaluation of "PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation".
Reproduction Package for Article "PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation" (doi:10.5281/zenodo.17710112): The artifact contains the source code for our PRISM and the benchmarks used in the evaluation. PRISM is also available at \url{https://github.com/hpdps-group/PRISM.git}. It supports the results in Section 6. To validate or reproduce the results, build this artifact and check the results returned by running benchmarks.
Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications
Luke Marzen, Junhyung Shim, and Ali Jannesari
(Iowa State University, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Article: ppopp26main-p331-p (type: Full Paper) doi:10.1145/3774934.3786454
Supplemental Material: Supplemental Material for DynamDynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications. By Luke Marzen, Junhyung Shim, and Ali Jannesari.
OMPDataPerf (doi:10.5281/zenodo.18102356): This artifact provides the source code and benchmark applications for OMPDataPerf, a portable, low-overhead performance analysis tool for detecting, profiling, and attributing inefficient data-mapping patterns in OpenMP programs. The artifact includes the implementation of Algorithms 1 through 5, allowing users to ...
Root-Down Exposure for Maximal Clique Enumeration on GPUs
Zhe Pan, Peng Qu, and Youhui Zhang
(Tsinghua University, China; Zhongguancun Laboratory, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p291-p (type: Full Paper) doi:10.1145/3774934.3786449
RDMCE Artifacts (doi:10.5281/zenodo.17677584): This artifact implements \textbf{RDMCE}, a high-performance GPU-based solution for maximal clique enumeration (MCE) in large-scale graphs. The system introduces three key innovations: a root-down exposure mechanism for dynamic load balancing, a bitmap-centric MCE scheme to simplify memory layout, and an aggressive ...
ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities
Weile Luo, Yuhan Chen, Xiangrui Yu, Qiang Wang, Ruibo Fan, Hongyuan Liu, and Xiaowen Chu
(Hong Kong University of Science and Technology (Guangzhou), China; Harbin Institute of Technology, Shenzhen, China; Stevens Institute of Technology, USA; Hong Kong University of Science and Technology, Hong Kong)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p366-p (type: Full Paper) doi:10.1145/3774934.3786461
Artifact for the paper: ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities (doi:10.5281/zenodo.17709850): This is the repository for the PPoPP'26 paper: ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities. ROME is a GPU-optimized All-Pairs Shortest Path (APSP) system that regularizes irregular graph workloads via spatial computation restructuring and an asynchronous execution ...

Stencil and Sparse Matrix Computation

SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
Qiqi Gu, Chenpeng Wu, Heng Shi, and Jianguo Yao
(Shanghai Jiao Tong University, China; Shanghai Enflame Technology, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p8-p (type: Full Paper) doi:10.1145/3774934.3786414
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping (doi:10.5281/zenodo.17678931): This is the artifact for the PPoPP 2026 paper: SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping.
ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication Acceleration
Jiazhi Jiang, Xijia Yao, Jiayu Chen, Jinhui Wei, Dan Huang, and Yutong Lu
(Sun Yat-sen University, China)
Publisher's Version Article: ppopp26main-p58-p (type: Full Paper) doi:10.1145/3774934.3786422
Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores
Kaige Zhang, Hailong Yang, Xin You, Tianyu Feng, Yufan Xu, Zhongzhi Luan, Yi Liu, and Depei Qian
(Beihang University, China; Independent Researcher, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p208-p (type: Full Paper) doi:10.1145/3774934.3786441
AE description tex file for Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores: The submitted files contain the Artifact Evaluation (AE) description for "Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores", including the LaTeX source file of the AE description and the corresponding compiled PDF.
Reproduction Package for Paper "Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores" (doi:10.5281/zenodo.17709956): The artifact contains the implementation of Drawloom, a Tensor Core–accelerated SpMV framework. It includes the ArbitWeave mapping strategy, the Zig-zag Chained Format (ZCF), and a multi-stage register pipeline, along with scripts for compiling and running performance evaluations. The artifact is intended to reproduce ...
VDHA: Vector-Driven Hash Aggregation for Sparse Matrix-Sparse Vector Multiplication on GPUs
Yuchen Li, Zhe Pan, Peng Qu, and Youhui Zhang
(Tsinghua University, China; Zhongguancun Laboratory, China)
Publisher's Version Artifacts Reusable Results Reproduced Article: ppopp26main-p283-p (type: Full Paper) doi:10.1145/3774934.3786447
Supplementary Appendix: Supplementary appendix including artifact details (requirements, installation, and project structure).

Mixed Precision and Quantization

RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization
Qihao Zhang, MingLiang Tang, Mingshu Zhai, Kinman Lei, and Jidong Zhai
(Tsinghua University, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p30-p (type: Full Paper) doi:10.1145/3774934.3786419
Artifact for PPoPP'26 "RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization" (doi:10.5281/zenodo.17678080): This repository contains the code for the reproduction of the paper "RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization" at PPoPP'26. The reproduction includes Tables 1 and 2 and Figures 7, 8, 9, and 10 from the submitted version of the paper.
High-Throughput Non-uniformly Quantized 3-bit LLM Inference
YuAng Chen, Wenqi Zeng, and Jeffrey Xu Yu
(Chinese University of Hong Kong, China; Hong Kong University of Science and Technology, China; Hong Kong University of Science and Technology (Guangzhou), China)
Publisher's Version Article: ppopp26main-p59-p (type: Full Paper) doi:10.1145/3774934.3786423
JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context Inference
Chengyu Sun, Yaqi Xia, Hulin Wang, Donglin Yang, Xiaobo Zhou, and Dazhao Cheng
(Wuhan University, China; Nvidia Corporation, USA; University of Macau, China)
Publisher's Version Article: ppopp26main-p96-p (type: Full Paper) doi:10.1145/3774934.3786428
HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff
Zeyu Song, Lin Gan, Xiaohui Duan, Zhengrui Li, Jiayu Fu, Yinuo Wang, Guangzhao Li, and Guangwen Yang
(Tsinghua University, China; Shandong University, China; Institute of Software at Chinese Academy of Sciences, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p160-p (type: Full Paper) doi:10.1145/3774934.3786433
appendix: Reproduction instructions for the paper "HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff".
Reproduction Package for Article "HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff" (doi:10.5281/zenodo.17958414): This artifact contains source code and scripts to reproduce crucial figures in the paper.

Cluster and Cloud Computing

Cacheman: A Comprehensive Last-Level Cache Management System for Multi-tenant Clouds
Xiaokang Hu, Yuchao Cao, Naixuan Guan, Yifan Wu, Xishi Qiu, Shengdong Dai, Ben Luo, Sanchuan Cheng, Fudong Qiu, Yibin Shen, and Jiesheng Wu
(Alibaba Cloud Computing, China)
Publisher's Version Article: ppopp26main-p13-p (type: Full Paper) doi:10.1145/3774934.3786415
zBuffer: Zero-Copy and Metadata-Free Serialization for Fast RPC with Scatter-Gather Reflection
Xiangyu Liu, Huiba Li, Shun Gai, Youmin Chen, and Yiming Zhang
(Xiamen University, China; Alibaba Cloud, China; Shanghai Jiao Tong University, China)
Publisher's Version Article: ppopp26main-p81-p (type: Full Paper) doi:10.1145/3774934.3786426
Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU Clusters
Ruobing Han and Hyesoon Kim
(Georgia Tech, USA)
Publisher's Version Article: ppopp26main-p169-p (type: Full Paper) doi:10.1145/3774934.3786435
Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters
Yida Li, Siwei Zhang, Yiduo Niu, Yang Du, Qingxiao Sun, Zhou Jin, and Weifeng Liu
(China University of Petroleum-Beijing, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p231-p (type: Full Paper) doi:10.1145/3774934.3786442
Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters: This artifact for our PPoPP '26 submission #231, entitled "Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters," includes the following components: (1) Prerequisites, introducing the hardware and software requirements; (2) Introduction to Each Figure, analyzing data and trends from ...
Reproduction Package for Article `Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters' (doi:10.5281/zenodo.17706095): This artifact for our PPoPP '26 submission #231, entitled "Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters," includes the following components: (1) Prerequisites, introducing the hardware and software requirements; (2) Introduction to Each Figure, analyzing data and trends from ...

Distributed Training

COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM Training
Xingchen Liu, Haoran Kong, Hairui Zhao, Shengkai Lyu, Zheng Wei, Man Liu, Xingjian Tian, Liyang Zhao, Zhuohan Chen, Fakang Wang, Zizhong Chen, Zhan Wang, Guangming Tan, and Dingwen Tao
(University of Chinese Academy of Sciences, China; Shenzhen Loop Area Institute, China; Chinese University of Hong Kong, Shenzhen, China; Jilin University, China; Ant Group, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p159-p (type: Full Paper) doi:10.1145/3774934.3786432
Appendix: Artifact Description/Artifact Evaluation: The artifact description and artifact evaluation of COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM Training
Reproduction Package for Article "COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM Training" (doi:10.5281/zenodo.17955665): This artifact contains the full source code of COCCL and the complete set of benchmarks used in the experimental evaluation. It corresponds to the results presented in Sec. 5 of the paper and provides all necessary components to validate COCCL's design, performance, and accuracy characteristics.
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training
Xuanyu Wang, Fangcheng Fu, Haoyang Li, Hao Ge, Sheng Lin, Jiawen Niu, and Bin Cui
(Peking University, China; Shanghai Jiao Tong University, China)
Publisher's Version Article: ppopp26main-p268-p (type: Full Paper) doi:10.1145/3774934.3786445
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
Geng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu, and Yang You
(National University of Singapore, Singapore)
Publisher's Version Article: ppopp26main-p22-p (type: Full Paper) doi:10.1145/3774934.3786417
CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang, Hairui Zhao, Xingchen Liu, Yang Tian, Wenjing Huang, Zedong Liu, Yifan Chen, Jinwu Yang, Yueyuan Zhou, Qian Zhao, Haoxu Li, Tao Wang, Feng Yu, Zhan Wang, Guangming Tan, and Dingwen Tao
(University of Chinese Academy of Sciences, China; Ant Group, China; Jilin University, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Article: ppopp26main-p110-p (type: Full Paper) doi:10.1145/3774934.3786429
Artifact Evaluation: This appendix contains the artifact evaluation for CCL-D. It covers the use of artifacts and how experiments are conducted.
Reproduction Package for Article 'CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training' (doi:10.5281/zenodo.17696975): This artifact includes a rank communication metric measurement module and an metrics-based Slow/Hang anomaly decision analysis module. These two modules work together to quickly detect and locate six types of Slow/Hang anomalies.

Parallel Algorithms

Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based Protocols
Zhiyuan Zhang, Yanxin Cai, Wenhao Yin, Xueyu Wu, Yi Wang, Lei Ju, and Zhuoran Ji
(Shandong University, China; Quan Cheng Laboratory, China; University of Hong Kong, China; Shenzhen University, China; State Key Laboratory of Cryptography and Digital Economy Security, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p288-p (type: Full Paper) doi:10.1145/3774934.3786448
ppopp26-artifact: pipelonk artifact (doi:10.5281/zenodo.18034218): see readme
ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing
Shuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao, Ruibai Tang, Yidong Chen, Jiping Yu, Yang Li, Chao Jiang, Limin Xiao, and Jidong Zhai
(Tsinghua University, China; Qingcheng.AI, China; University of Aberdeen, UK; Lenovo Research, China)
Publisher's Version Article: ppopp26main-p27-p (type: Full Paper) doi:10.1145/3774934.3786418
Faster and Cheaper: Pushing the Sequence Alignment Throughput with Commercial CPUs
Zhonghai Zhang, Yewen Li, Ke Meng, Chunming Zhang, and Guangming Tan
(Institute of Computing Technology at Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, China; Hong Kong University of Science and Technology, China; Phil Rivers Technology, China)
Publisher's Version Article: ppopp26main-p37-p (type: Full Paper) doi:10.1145/3774934.3786421
PIM-zd-tree: A Fast Space-Partitioning Index Leveraging Processing-in-Memory
Yiwei Zhao, Hongbo Kang, Ziyang Men, Yan Gu, Guy E. Blelloch, Laxman Dhulipala, Charles McGuffey, and Phillip B. Gibbons
(Carnegie Mellon University, USA; Tsinghua University, China; University of California at Riverside, USA; University of Maryland, USA; Reed College, USA)
Publisher's Version Info Article: ppopp26main-p1-p (type: Full Paper) doi:10.1145/3774934.3786411

ML Inference

BEEMS: Boosting Machine Vision Efficiency via Computation Graph-Based Memory Smoothing
Hanjing Shen, Fangxin Liu, Jian Liu, Li Jiang, and Haibing Guan
(Shanghai Jiao Tong University, China; Beihang University, China)
Publisher's Version Article: ppopp26main-p128-p (type: Full Paper) doi:10.1145/3774934.3786430
Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving
Jianxiong Liao, Quanxing Dong, Yunkai Liang, Zhi Zhou, and Xu Chen
(Sun Yat-sen University, China)
Publisher's Version Article: ppopp26main-p7-p (type: Full Paper) doi:10.1145/3774934.3786413
MixFusion: A Patch-Level Parallel Serving System for Mixed-Resolution Diffusion Models
Desen Sun, Zepeng Zhao, and Yuke Wang
(University of Waterloo, Canada; Carnegie Mellon University, USA; Rice University, USA)
Publisher's Version Artifacts Reusable Results Reproduced Article: ppopp26main-p35-p (type: Full Paper) doi:10.1145/3774934.3786420
ChituDiffusion: A Data-Characteristic-Aware Serving System for Diffusion Models
Chengzhang Wu, Liyan Zheng, Haojie Wang, Kezhao Huang, Zixuan Ma, Dong Dong, and Jidong Zhai
(Tsinghua University, China)
Publisher's Version Article: ppopp26main-p64-p (type: Full Paper) doi:10.1145/3774934.3786424

Graphs and Graph Neural Networks

ElasGNN: An Elastic Training Framework for Distributed GNN Training
Siqi Wang, Hailong Yang, Pengbo Wang, Hongliang Cao, Yufan Xu, Xuezhu Wang, Zhongzhi Luan, Yi Liu, and Depei Qian
(Beihang University, China; Independent Researcher, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p207-p (type: Full Paper) doi:10.1145/3774934.3786440
Artifact Appendix: The artifact appendix of ElasGNN.
Reproduction Package for Article "ElasGNN : An Elastic Training Framework for Distributed GNN Training" (doi:10.5281/zenodo.18143217): Reproduction Package for Article "ElasGNN : An Elastic Training Framework for Distributed GNN Training", including the environment and scripts.
APERTURE: Algorithm-System Co-optimization for Temporal Graph Network Inference
Yiqing Wang, Hailong Yang, Enze Yu, Qingxiao Sun, Kejie Ma, Kaige Zhang, Chenhao Xie, and Depei Qian
(Beihang University, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p296-p (type: Full Paper) doi:10.1145/3774934.3786450
PPoPP26_AE_APERTURE_CODE (doi:10.5281/zenodo.17710612): This artifact provides the implementation of \textit{APERTURE}, a decoupled and memory-aware temporal graph network (TGN) inference framework. \textit{APERTURE} includes (i) a computation graph transformation engine with a global feature map, (ii) a state-based memory update manager, and (iii) an analytical ...
TAC: Cache-Based System for Accelerating Billion-Scale GNN Training on Multi-GPU Platform
Zhiqiang Liang, Hongyu Gao, Jue Wang, Fang Liu, Xingguo Shi, Junyu Gu, Peng Di, Sian Li, Lei Tang, Chunbao Zhou, Lian Zhao, Yangang Wang, and Xuebin Chi
(Computer Network Information Center at Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, China; Ant Group, China; UNSW, Australia)
Publisher's Version Article: ppopp26main-p357-p (type: Full Paper) doi:10.1145/3774934.3786460
DTMiner: A Data-Centric System for Efficient Temporal Motif Mining
Yinbo Hou, Hao Qi, Ligang He, Jin Zhao, Yu Zhang, Hui Yu, Longlong Lin, Lin Gu, Wenbin Jiang, Xiaofei Liao, and Hai Jin
(Huazhong University of Science and Technology, China; University of Warwick, UK; Hong Kong University of Science and Technology, China; Southwest University, China)
Publisher's Version Article: ppopp26main-p18-p (type: Full Paper) doi:10.1145/3774934.3786416

Optimizing Transformers

FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism
Jianxing Xu, Yuanbo Wen, Jun Bi, Ruibai Xu, Guanglin Xu, Rui Zhang, Wei Li, Ling Li, Tianshi Chen, Qi Guo, and Yunji Chen
(University of Science and Technology of China, China; Institute of Computing Technology at Chinese Academy of Sciences, China; Institute of Software at Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, China; Cambricon Technologies, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p72-p (type: Full Paper) doi:10.1145/3774934.3786425
Artifact for the PPoPP '26 paper: "FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor–Vector Parallelism." (doi:10.5281/zenodo.17673796): This is the artifact package for the PPoPP '26 paper: “FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor–Vector Parallelism.” FlashAttention-T is a prototype fused-attention implementation built on FlashAttention-2 and FlashAttention-3 that advances toward fully tensorized fused attention. The ...
Accelerating Sparse Transformer Inference on GPU
Wenhao Dai, Haodong Deng, Mengfei Rong, Xinyu Yang, Hongyu Liu, Fangxin Liu, Hailong Yang, Qianwen Cao, and Qingxiao Sun
(China University of Petroleum-Beijing, China; Beihang University, China; Baidu, China; Shanghai Jiao Tong University, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Functional Results Reproduced Article: ppopp26main-p161-p (type: Full Paper) doi:10.1145/3774934.3786434
PPoPP26_AE_STOF_CODE (doi:10.5281/zenodo.17705801): This artifact contains the system prototype of STOF (pap161) at PPoPP '26, titled "Accelerating Sparse Transformer Inference on GPU", including Figure 9_10, Figure 11, Figure 12, Figure 13 and Table 4. The folder is organized as below: data/: original data log in data/MHA_performance for Figure 9_10, ...
MetaAttention: A Unified and Performant Attention Framework across Hardware Backends
Feiyang Chen, Yu Cheng, Lei Wang, Yuqing Xia, Ziming Miao, Lingxiao Ma, Fan Yang, Jilong Xue, Zhi Yang, Mao Yang, Xingda Wei, and Haibo Chen
(Shanghai Jiao Tong University, China; Peking University, China; Microsoft Research, China)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p238-p (type: Full Paper) doi:10.1145/3774934.3786444
Reproduction Package for Article “MetaAttention: A Unified and Performant Attention Framework across Hardware Backends” (doi:10.5281/zenodo.17701680): This repository contains the artifacts for the PPoPP'26 Artifact Evaluation of paper #238: "MetaAttention: A Unified and Performant Attention Framework Across Hardware Backends". We provide detailed documentation and Docker support to build, install, and test MetaAttention. The examples/ folder contains all Attention ...

Matrix and Linear Algebra Algorithms

Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU Architectures
Lu Shi, WeiWei Xu, and Shaoshuai Zhang
(University of Electronic Science and Technology of China, China; Nanjing University of Information Science and Technology, China)
Publisher's Version Article: ppopp26main-p92-p (type: Full Paper) doi:10.1145/3774934.3786427
A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue Solvers
Xiaojian Yang, Yuhui Ni, Fan Yuan, Shengguo Li, Dezun Dong, Chuanfu Xu, Haipeng Jia, and Jie Liu
(National University of Defense Technology, China; Xiangtan University, China; University of Chinese Academy of Sciences, China)
Publisher's Version Article: ppopp26main-p271-p (type: Full Paper) doi:10.1145/3774934.3786446
A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability
Yuchen Ma, Bin Ren, and Andreas Stathopoulos
(College of William & Mary, USA)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Article: ppopp26main-p310-p (type: Full Paper) doi:10.1145/3774934.3786453
SMatVec (doi:10.5281/zenodo.18136205): This artifact contains the following components: 1. 1-comm-simulator: Communication Simulator Produces Figures 5, 7, 8, and 15 of the paper. 2. 2a-SMatVec: The main SMatVec software Produces Figures 9, 10, 12, 13, 14, 15, 16, 18 of the paper. 3. 2c-SMatVec-old-rowmajor: The old SMatVec version Produces Figure 18 of ...
Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific Computing
Yuechen Lu, Hongwei Zeng, Marc Casas, and Weifeng Liu
(China University of Petroleum-Beijing, China; Barcelona Supercomputing Center, Spain; Universitat Politècnica de Catalunya, Spain)
Publisher's Version Published Artifact Artifacts Available Artifacts Reusable Results Reproduced Article: ppopp26main-p336-p (type: Full Paper) doi:10.1145/3774934.3786456
Appendix: Artifact Evaluation: This artifact evaluation accompanies the paper entitled 'Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific Computing'.
Cubie (doi:10.5281/zenodo.17725527): The artifact for PPoPP26 submission #336: "Characterizing Matrix Multiplication Units Across General Parallel Patterns in Scientific Computing".

proc time: 0.16