TCAD 2026: APTQ+: Attention-FFN-aware Post-Training Quantization for a Layer-wise LLM Accelerator on FPGA

Published in IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), 2026 (CCF-A), 2026

APTQ+ studies attention- and FFN-aware post-training quantization for layer-wise deployment of large language model accelerators. The work explores how sensitivity across attention and feed-forward components can guide practical mixed-precision decisions while preserving model quality and improving deployment efficiency.

Google Scholar record

Recommended citation: Ziyi Guan, Zhengfei Chen, Dongkuan Wu, Ao Shen, Yifan Guo, Yupeng Su, Graziano Chesi, Mingqiang Huang, Ngai Wong, and Hao Yu. "APTQ+: Attention-FFN-aware Post-Training Quantization for a Layer-wise LLM Accelerator on FPGA." IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2026.
Download Paper