Yuannuo Feng
Beihang University · Efficient LLM Inference · Compute-in-Memory
Beihang University
Beijing, China
I am a researcher at Beihang University, working on AI infrastructure — efficient large language model (LLM) inference and compute-in-memory (CiM) architectures.
Research
Current — AI infrastructure for efficient inference. I work on making LLM inference faster and cheaper. My recent project ASD (arXiv:2608.03447) introduces a bounded-regret approximate acceptance policy for speculative decoding. I am currently working on efficient inference optimization for Diffusion Language Models (DLMs).
Past — software–hardware co-optimization for compute-in-memory. I studied how to deploy neural networks robustly on analog CiM, a new AI acceleration substrate, addressing hardware noise through co-design: noise-aware training and straight-through estimation (ASICON’25a), hybrid projection decomposition for state space models (ASICON’25b), noise-aware sampling for diffusion models (DATE’26, ASP-DAC’27), and noise-resilient LLM inference on CiM (ROMER, KV cache protection).
Beyond research
I enjoy photography. You can find my photographic works on Xiaohongshu and Douyin.
Feel free to reach out via yanruo.f@gmail.com (or ynfeng@buaa.edu.cn) — I’m always happy to discuss inference systems and hardware–algorithm co-design.
latest posts
selected publications
- DATENoise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-MemoryIn 2026 Design, Automation & Test in Europe Conference (DATE), 2026
- ASICONHPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CiM HardwareIn 2025 IEEE 16th International Conference on ASIC (ASICON), 2025
- ASICONExtending Straight-Through Estimation for Robust Neural Networks on Analog CiM HardwareIn 2025 IEEE 16th International Conference on ASIC (ASICON), 2025