Portrait of Jiayu Zhao

Biography

I am currently a third-year master’s student at the School of Microelectronics, University of Science and Technology of China (USTC), supervised by Assoc. Prof. Song Chen since Sep. 2024. Prior to that, I obtained my B.S. degree in Physics from the School of Physical Sciences at USTC in 2024.

My research lies at the intersection of AI systems and electronic design automation (EDA). I work on efficient inference and compression for large language models, particularly Mixture-of-Experts (MoE) LLMs, and on structure-aware LLM pipelines for hardware code generation.


Recent News

  • Sep/2026: VeriGRAG is accepted by ASPDAC 2027.

  • May/2026: BitsMoE is available on arXiv, focusing on mixed-precision quantization for MoE LLMs.


Research Interests

My research interests include:

  • Efficient MoE LLMs: quantization, model compression, and efficient inference.
  • LLMs for RTL Generation and Verification: structure-aware code generation and verification-aware reasoning.

Selected Research Directions

MoE LLMs Efficient Inference

Efficient MoE deployment requires preserving expert-specialized capacity while reducing resident memory and inference overhead. My work studies quantization and allocation strategies that make sparse LLMs practical under strict deployment budgets.

BitsMoE overview

BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization. arXiv preprint, 2026. (Paper) (Code)

Combines shared-basis spectral decomposition with factorized quantization cost modeling to allocate bits across expert-specific spectral components under a fixed memory budget, enabling accurate low-bit MoE quantization and efficient GPU inference.

Keywords: MoE LLMs; mixed-precision quantization; spectral decomposition; cost-aware bit allocation; efficient inference.


Education


Experience

Parallel and Distributed Computing Lab (PDCL), SCSE, Nanyang Technological University (NTU), Singapore

– visitors