Abner Law
I'm preparing for the Archtectonics Lab.
I'm Zhongdi Luo, a PhD student major in Computer Science and Engineering.
I am currently focus on performance analysis of accelerators, encompassing performance modeling, software‑hardware co‑design and evaluation, and agile accelerator design‑space exploration. In particular, the automatic generation of accelerator hardware through high‑level language compilation.
news
- Jul 24, 2026
I'm serving as Artifact Evaluation Committee member in MICRO 2026
- Jun 26, 2026
Got First Prize (10th Place) at ASC26 and 3rd Place in ISC26 SCC as Assistant Advisor!
- Apr 07, 2026
I won the Silver🥈in Architecture 2.0 Workshop in ASPLOS 2026
- Jan 04, 2026
I'm serving as Artifact Evaluation Committee member in OOPSLA 2026
- Jun 30, 2025
Received Year 2025 Provincial Outstanding Student Award
- Apr 12, 2024
Got First Prize (rank 5) as 64bit-brainstorm Captain at ASC22-23 (10th ASC)!
latest posts
- May 21, 2026Encore: My Final ASC
- Apr 22, 2025Stroll with me through the streets of Chengdu
- Nov 14, 2024Swiss Journey: Observations, Reflections, and Impressions
selected publications
@article{jing2026bfp,
title = {Precision boundary modeling for area-efficient Block Floating Point accumulation},
journal = {Journal of Systems Architecture},
volume = {173},
pages = {103704},
year = {2026},
issn = {1383-7621},
doi = {https://doi.org/10.1016/j.sysarc.2026.103704},
url = {https://www.sciencedirect.com/science/article/pii/S1383762126000226},
author = {Jun He and Jing Feng and Xin Ju and Yasong Cao and Zhongdi Luo and Jianchao Yang and Jingkui Yang and Gang Li and Jian Cheng and Dong Chen and Mei Wen},
keywords = {Block Floating Point computation, Accumulation precision, Frobenius norm Retention Ratio (FnRR), Hardware efficiency},
abstract = {Block Floating Point (BFP) is extensively employed for the low-precision quantization of deep network weights and activations to attain advantages in both hardware efficiency and performance. Nevertheless, when the precision of weights and activations is diminished to below 8 bits, the required high-precision floating-point accumulation becomes a dominant hardware bottleneck in the BFP processing element (PE). To address this challenge, we introduce a framework based on the Frobenius norm Retention Ratio (FnRR) to explore the precision boundaries for BFP accumulation, and extend it to a hierarchical chunk-based accumulation scheme. Comprehensive experiments across representative CNN and LLM models demonstrate that our predicted precision boundaries maintain performance closely matching FP32 baselines, while further precision reduction leads to substantial accuracy degradation, validating the effectiveness of our boundary determination. Guided by this analysis, we present a corresponding hardware for BFP computation. This design achieves 13.7%–25.2% improvements in area and power efficiency compared with FP32 accumulation under identical quantization settings, and delivers up to 10.3× area and 11.0× power reductions relative to conventional BFP implementations.}
}@article{zeyu2025sspmm,
author={Xue, Zeyu and Wen, Mei and Yang, Jianchao and Tang, Minjin and Luo, Zhongdi and Feng, Jing and Shi, Yang and Chen, Zhaoyun and Shen, Junzhong and Langguth, Johannes},
journal={IEEE Transactions on Parallel and Distributed Systems},
title={SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores},
year={2025},
volume={36},
number={12},
pages={2652-2667},
keywords={Tensors;Graphics processing units;Sparse matrices;Kernel;Parallel processing;Scalability;Computer architecture;Artificial intelligence;Programming;Layout;GPGPU;tensor core;CUDA;SpMM},
doi={10.1109/TPDS.2025.3616981}}
@article{yasong2024abs,
author={Cao, Yasong and Wen, Mei and Luo, Zhongdi and Ju, Xin and Huang, Haolan and Shen, Junzhong and Chen, Haiyan},
journal={IEEE Transactions on Very Large Scale Integration (VLSI) Systems},
title={ABS: Accumulation Bit-Width Scaling Method for Designing Low-Precision Tensor Core},
year={2024},
volume={32},
number={9},
pages={1590-1601},
keywords={Artificial neural networks;Training;Task analysis;Quantization (signal);Tensors;Hardware;Systolic arrays;Accumulation length;floating-point (FP);low-precision computation;systolic array (SA);variance retention ratio (VRR)},
doi={10.1109/TVLSI.2024.3414260}}