HyPER

Dynamic test-time scaling. A closed-loop controller for multi-path LLM reasoning that reallocates compute between exploration and exploitation under a fixed budget, combining branch expansion, path reduction, and multi-granularity scheduling. PDF · Code

SPEX

Speculative Tree-of-Thought serving. An SGLang-based system that removes reward-dependency bottlenecks through speculative branch expansion, cross-query budget coordination, and adaptive early termination. PDF · Code

MASBench

Multi-agent LLM serving benchmark. Composable collaboration patterns and causal execution traces for studying synchronization barriers, context movement, tool stalls, queueing, TTFT/TPOT, and KV-cache behavior across GPU and Ascend platforms.

HDA-MoE

MoE on 3D near-memory processing. Hybrid parallelism and dynamic adaptive scheduling for efficient MoE inference with memory-centric architectures. PDF · Code

S²-MoE

Self-speculative MoE inference. Efficient self-speculative decoding for Mixture-of-Experts models on resource-constrained edge devices. PDF · Code

Tetris

Composable test-time scaling on GPU-PIM. Ongoing work on adaptive mapping, control, and memory management for emerging memory-centric acceleration of reasoning workloads.