About Me 🚀
I am Tao Zhang, a Ph.D. candidate at the University of Science and Technology of China (USTC), working with the Future Network Laboratory. My research focuses on efficient LLM inference serving, AI infrastructure, and multi-agent systems.
I am currently a research intern at Huawei 2012 Laboratories, working on PTO optimization for Simpler-side graph construction, graph solving, and scheduling.
News 📰
- Joined Huawei 2012 Laboratories as a research intern on PTO optimization.
- SpecCache accepted by ACL 2026 as an Oral paper.
- HAWK accepted by CVPR 2026 as a Poster paper.
- SAVP, LatCom, and GSTEP accepted by EMNLP 2026 and ACM Multimedia 2026.
- FAESR published in IEEE TCCN.
Education 🎓
2023.09 - Present
Ph.D. Candidate, University of Science and Technology of China Institute of Advanced Technology · Future Network Laboratory
2019.09 - 2023.06
B.Eng., Chongqing University of Posts and Telecommunications School of Communication and Information Engineering
Internship Experience 💼
Huawei 2012 Laboratories
Research Intern · PTO Optimization / AI Infrastructure
2026 - Present
Focus on PTO optimization for dynamic and static graph construction on the Simpler side, efficient computational graph construction and solving, and scheduling.
Publications 📚
DisHelis: Optimizing Deployment of Disaggregated LLMs Inference Serving over Heterogeneous Environments via Hierarchical Max-Flow
SpecCache: Speculative KV Cache Reuse for Efficient RAG Serving
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
SAVP: Scene-Aware Vision Token Pruning for Efficient Video Large Language Models
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
LatCom: Latent Compression for Efficient Multi-Agent Collaboration
Multi-Timescale Joint Optimization for Disaggregated LLM Serving
FAESR: Fine-Grained Rate Adaptation for Energy-Aware Super Resolution in Mobile Panoramic Video Streaming
Research Interests 🔍
Efficient LLM Inference Serving Scheduling, cache reuse, and deployment planning for scalable model serving.
AI Infrastructure Graph construction, graph solving, and execution scheduling on heterogeneous clusters.
Multi-Agent Systems Efficient collaboration and system-aware coordination for agent workloads.
Multimodal Efficiency Visual token pruning and serving optimization for multimodal workloads.
Honors and Awards 🏆
2025
National Scholarship University of Science and Technology of China
2023 / 2024
Graduate Academic First-Class Scholarship University of Science and Technology of China
2023
Outstanding Graduate of Chongqing Municipality Chongqing, China
2022
MathorCup National Undergraduate Mathematical Modeling Competition, National First Prize National competition award
