PhD candidate at Harbin Institute of Technology, working on efficient inference for large language and vision-language models.
- Quantization and mixed-precision inference
- Adaptive and inference-time computation
- Multimodal model efficiency
- Model-system co-design
- Efficient inference algorithms for LLMs and VLMs
- Reproducible evaluation of accuracy, memory, latency, and throughput
- Open-source contributions to modern LLM inference systems
Python · PyTorch · Transformers · vLLM · Linux · Git
Building public contributions around multimodal inference, KV-cache efficiency, and reliable performance evaluation.