2026

ICML '26
HyPER: Bridging Exploration and Exploitation for Scalable LLM Reasoning with Hypothesis Path Expansion and Reduction
Shengxuan Qiu*, Haochen Huang*, Shuzhang Zhong, Pengfei Zuo, Meng Li
43rd International Conference on Machine Learning (ICML), 2026
OSDI '26
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
Shuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo, Runsheng Wang, Meng Li
20th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2026
TCAD '26
HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing
Haochen Huang, Shuzhang Zhong, Shengxuan Qiu, Zhe Zhang, Shuangchen Li, Cong Li, Dimin Niu, Hongzhong Zheng, Guangyu Sun, Runsheng Wang, Meng Li
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), 2026
arXiv '26
S²-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
Haochen Huang, Shengxuan Qiu, Meng Li
arXiv preprint, 2026

Ongoing work such as Tetris is presented on the Projects page until a public archival/preprint version is available.