publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- DATENoise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-MemoryYuannuo Feng, W. Zhou, Y. Lv, and 5 more authorsIn 2026 Design, Automation & Test in Europe Conference (DATE), 2026
@inproceedings{feng2026nas, title = {Noise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-Memory}, author = {Feng, Yuannuo and Zhou, W. and Lv, Y. and Liu, H. and Wang, G. and Liu, Z. and Wong, N. and Kang, W.}, booktitle = {2026 Design, Automation \& Test in Europe Conference (DATE)}, pages = {1--3}, year = {2026}, } - arXivROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory SystemsYuannuo Feng, W. Zhou, Y. Chen, and 6 more authors2026
Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per token, yet their frequent expert switching creates memory bandwidth bottlenecks that compute-in-memory (CIM) architectures are well-suited to mitigate. However, analog CIM systems suffer from inherent hardware imperfections that perturb stored weights, and its negative impact on MoE-based LLMs in noisy CIM environments remains unexplored. In this work, we present the first systematic investigation of MoE-based LLMs under noise model calibrated with real chip measurements, revealing that hardware noise critically disrupts expert load balance and renders clean-trained routing decisions consistently suboptimal. Based on these findings, we propose ROMER, a post-training calibration framework that (1) replaces underactivated experts with high-frequency ones to restore load balance, and (2) recalibrates router logits via percentile-based normalization to stabilize routing under noise. Extensive experiments across multiple benchmarks demonstrate that ROMER achieves up to 58.6%, 58.8%, and 59.8% reduction in perplexity under real-chip noise conditions for DeepSeek-MoE, Qwen-MoE, and OLMoE, respectively, establishing its effectiveness and generalizability across diverse MoE architectures.
@misc{feng2026romer, title = {ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems}, author = {Feng, Yuannuo and Zhou, W. and Chen, Y. and Wu, T. and Xu, W. and Qi, W. and Liu, Z. and Kang, W. and Wong, N.}, year = {2026}, archiveprefix = {arXiv}, } - arXivSelective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-in-Memory SystemsYuannuo Feng, W. Zhou, Y. Ma, and 5 more authors2026
Analog compute-in-memory (CIM) arrays have emerged as a promising substrate for energy-efficient LLM inference, particularly for weight-stationary computations in linear layers. However, extending analog CIM to attention mechanisms introduces a fundamental challenge: KV cache operations demand repeated in-situ weight updates, and the resulting mismatch with the weight-stationary paradigm exposes dynamic computations to significant hardware noise, a critical problem that remains largely unexplored. In this paper, we present the first systematic study of dynamic attention computation on analog CIM arrays, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise. Motivated by this token-level insight, we propose a hierarchical token protection strategy that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM. A co-designed scheduler combines analog programming, ownership transition, and bulk-MVM tile formation to bound digital overhead. Evaluations on nine LLMs show that our approach lowers average perplexity under analog noise from 33.91 to 11.95, close to the clean baseline of 11.06, while improving dynamic-KV programming-row utilization from 23.1% to 91.2%.
@misc{feng2026kvcache, title = {Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-in-Memory Systems}, author = {Feng, Yuannuo and Zhou, W. and Ma, Y. and Chen, Y. and Yao, W. and Xie, Y. and Wong, N. and Kang, W.}, year = {2026}, archiveprefix = {arXiv}, } - arXivApproximate Speculative DecodingYuannuo Feng, Z. Peng, Y. Xie, and 5 more authors2026
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target-greedy under the realized prefix. In this paper, we introduce Approximate Speculative Decoding (ASD), a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection. ASD accepts selected mismatches subject to a local target-logit regret gate, a per-block exception cap, and a persistent request-level regret budget, then reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes. ASD requires neither a new draft model nor fine-tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed-workload throughput by 3.05%-15.26% over matched strict verification and averages a 7.78% gain across seven Qwen3-14B + DSpark-14B tasks. On DeepSeek-V4-Flash (284B) with DSpark it also raises verifier-side acceptance by roughly 10%-16% on GSM8K and MATH-500 in an FP4-to-FP8 compatibility setting.
@misc{feng2026asd, title = {Approximate Speculative Decoding}, author = {Feng, Yuannuo and Peng, Z. and Xie, Y. and Ye, Y. and Chen, Y. and Yao, W. and Zhou, W. and Kang, W.}, year = {2026}, archiveprefix = {arXiv}, } - ASP-DACASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory HardwareYuannuo Feng, Yizhe Chen, Wenshuai Yao, and 4 more authors2026Accepted by ASP-DAC 2027
Diffusion models achieve strong image generation quality but incur high iterative denoising costs. Analog compute-in-memory (CIM) can accelerate matrix-vector multiplications, yet spatial memory variations perturb weights and accumulate during sampling. Unlike conventional neural networks, diffusion models’ temporal sensitivity to hardware noise remains underexplored. We investigate diffusion inference using a noise model calibrated and validated against measurements collected from multiple physical CIM chips. Our results show that the early, high-noise denoising stage is substantially more vulnerable than the final refinement stage. A first-order trajectory analysis attributes this behavior to the repeated propagation of correlated prediction errors induced by a fixed hardware mapping. Based on this observation, we propose ASSERT, a training-free sampler that uses higher stochasticity early and smoothly transitions to deterministic denoising. The injected stochasticity changes subsequent activation trajectories and thereby reduces their alignment with persistent spatial errors. Across the evaluated settings, ASSERT achieves up to 2.58x lower FID than deterministic DDIM on high-resolution datasets and 7.68x lower FID in the CIFAR-10 step-count study, without changing model parameters or the number of network evaluations.
@misc{feng2026assert, title = {ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware}, author = {Feng, Yuannuo and Chen, Yizhe and Yao, Wenshuai and Xie, Yuxin and Wong, Ngai and Zhou, Wenyong and Kang, Wang}, year = {2026}, archiveprefix = {arXiv}, note = {Accepted by ASP-DAC 2027}, }
2025
- ASICONHPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CiM HardwareYuannuo Feng, W. Zhou, Y. Lyu, and 4 more authorsIn 2025 IEEE 16th International Conference on ASIC (ASICON), 2025
@inproceedings{feng2025hpd, title = {HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CiM Hardware}, author = {Feng, Yuannuo and Zhou, W. and Lyu, Y. and Liu, H. and Liu, Z. and Wong, N. and Kang, W.}, booktitle = {2025 IEEE 16th International Conference on ASIC (ASICON)}, pages = {1--4}, year = {2025}, } - ASICONExtending Straight-Through Estimation for Robust Neural Networks on Analog CiM HardwareYuannuo Feng, W. Zhou, Y. Lyu, and 4 more authorsIn 2025 IEEE 16th International Conference on ASIC (ASICON), 2025
@inproceedings{feng2025ste, title = {Extending Straight-Through Estimation for Robust Neural Networks on Analog CiM Hardware}, author = {Feng, Yuannuo and Zhou, W. and Lyu, Y. and Zhang, Y. and Liu, Z. and Wong, N. and Kang, W.}, booktitle = {2025 IEEE 16th International Conference on ASIC (ASICON)}, pages = {1--4}, year = {2025}, }