Biography
I’m Ziyi Guan (管子义), an AI Infra Researcher with ByteDance — Seed Infra, Heterogeneous Computing Group. I completed my Ph.D. in Electrical and Electronic Engineering at The University of Hong Kong (HKU) in November 2025, supervised by Dr. Ngai Wong and Prof. Graziano Chesi. Before that, I received my Bachelor’s degree from the School of Microelectronics at the Southern University of Science and Technology in 2021, supervised by Prof. Hao Yu.
My current work focuses on large language model (LLM) inference and deployment. During my work at Seed Infra, I contributed to inference acceleration and deployment for production LLM workloads, including model/runtime integration, internal inference framework development, graph execution adaptation, and CI-based regression and precision/performance validation. More broadly, I work across model, operator, and serving-system layers on prefill/decode acceleration, sparse attention, KV-cache compression, quantization, continuous batching, and graph-based execution. The goal is to reduce latency and memory use while improving throughput, efficiency, and deployment reliability.
My earlier research covered LLM compression, multimodal and GUI agents, RAG frameworks, and hardware-efficient neural network architectures co-designed with emerging accelerators.
My publications include work appearing at DAC, EMNLP, ICCAD, DATE, and IEEE TCAD, as well as an industry technical report on Seed2.0; see my Google Scholar for the latest list.
Previously, I worked at Huawei Hong Kong Research Center (Nov 2024 – Sep 2025) on KG-RAG GUI Test Agents, enhancing multi-platform mobile app testing via retrieval-augmented reasoning and this line of work includes a paper accepted to EMNLP 2025 (Main).
Research Interests:
LLM Inference & Deployment: Prefill/decode optimization, KV-cache management and compression, continuous batching, long-context serving, and inference performance analysis.
Inference Runtime & Engineering: Graph-based execution, runtime and operator adaptation, internal inference framework development, CI/regression systems, and precision/performance validation.
Inference Optimization: Weight and KV quantization, sparsity/pruning, sparse attention, distillation, and operator/kernel co-design for efficient deployment.
Heterogeneous AI Systems: Algorithm-hardware co-design and performance optimization for domestic and emerging AI accelerators.
LLM Agents: GUI/App agents and Retrieval-Augmented Generation (RAG) for task automation.
Selected Engineering Work:
Inference acceleration and serving: Prefill/decode optimization, sparse attention, KV-cache compression, continuous batching, and long-context serving for production LLM workloads.
Quantization and memory efficiency: Low-bit weight and KV-cache quantization, mixed-precision mapping, scale/weight validation, and memory-pressure reduction for efficient deployment.
Runtime and graph execution: Model/runtime integration, internal inference framework development, graph execution adaptation, operator integration, and backend portability across heterogeneous accelerator environments.
CI and quality engineering: Precision regression suites, reference comparisons, performance benchmark automation, deployment smoke tests, and runtime reachability checks for inference features.
You can find my latest Chinese CV here Chinese CV
You can contact me by my Email or gzygwp@gmail.com
Selected Publications (*represents equal contribution)
Industry technical report:
- ByteDance Seed, “Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity”, technical report, arXiv:2607.00248 (2026). During my work at Seed Infra, I contributed to inference acceleration and deployment for the Seed2.0 model family. arXiv · Google Scholar
Academic publications, first author and co-first author:
Ziyi Guan, et al, “APTQ+: Attention-FFN-aware Post-Training Quantization for a Layer-wise LLM Accelerator on FPGA”, published in IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD 2026, CCF-A) IEEE Xplore · Google Scholar
Ziyi Guan, et al, “KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation” In Proceedings of EMNLP 2025 Main Conference (CCF-B NLP Top Conference). (EMNLP 2025 (CCF-B)) PDF
Yupeng Su, Ziyi Guan*, et al, “LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models”, In Proceedings of DAC 2025 poster: 62nd IEEE/ACM Design Automation Conference. (DAC 2025 (CCF-A)) PDF
Dingbang Liu, Ziyi Guan*, et al, “A Highly Energy-Efficient Binary BERT Model on Group Vector Systolic CIM Accelerator”, In Proceedings of DAC 2025 poster: 62nd IEEE/ACM Design Automation Conference. (DAC 2025 (CCF-A))
Ziyi Guan, et al, “APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models”, In Proceedings of DAC 2024: 61st IEEE/ACM Design Automation Conference. (DAC 2024 Oral(CCF-A)), San Francisco, CA, June 23-27, 2024. PDF
Ziyi Guan, et al, “An Isotropic Shift-Pointwise Network for Crossbar-Efficient Neural Network Design”, Design, Automation & Test in Europe Conference & Exhibition (DATE 2024 (CCF-B)), March 25, Valencia, 2024. PDF
Ziyi Guan,et al, “A Video-based Fall Detection Network by Spatio-temporal Joint-point Model on Edge Devices”, Design, Automation & Test in Europe Conference & Exhibition (DATE 2021 (CCF-B)). IEEE, 2021, pp. 422–427. pdf
Ziyi Guan, et al, “A Hardware-Aware Neural Architecture Search Pareto Front Exploration for In-Memory Computing.” in 2022 IEEE 16th International Conference on Solid-State Integrated Circuit Technology (ICSICT). IEEE, 2022, pp. 1–4. pdf
Other authors:
- Shuwei Li, Ziyi Guan, et al. “A Fall Detection Network by 2D/3D Spatio-temporal Joint Models with Tensor Compression on Edge.” in ACM Transactions on Embedded Computing Systems (TECS) vol. 21, no. 6, pp. 1–19, 2022 PDF
(Last updated on Aug. 13, 2026)
