Linli Yao 姚林丽
I am a PhD student in the Language Computing and Machine Learning Group (Lanco) at Peking University, advised by Prof. Xu Sun.
Previously, I received my master's and bachelor's degrees from Renmin University of China, advised by Prof. Qin Jin at the AI·M3 Lab.
I expect to graduate in June 2027 and am exploring full-time opportunities in academia and industry.
I welcome discussions and collaborations on long-horizon multimodal intelligence, multimodal agents, and embodied intelligence. Feel free to get in touch!
Research Interests
I study efficient and time-aware video understanding with multimodal large language models. My research connects two complementary directions:
- Efficient video understanding. Visual token compression, adaptive frame sampling, and efficient processing of long and streaming videos.
- Time-aware video-language modeling. Temporal grounding, temporal reasoning, and fine-grained, structured audio-visual captioning.
News
| 2026 | Claw-Eval accepted to NeurIPS 2026 ED Track as a poster. |
|---|---|
| 2026 | Released MiMo-V2.6 — honored to contribute as a Core Contributor. |
| 2026 | Two papers accepted to Findings of EMNLP 2026: AdaC-GRPO and Quality, Not Just Outcome. |
| 2026 | DiaDem, on dialogue descriptions in audio-visual video captioning, accepted to ECCV 2026. |
| 2026 | TimeChat-Captioner and ReaForest accepted to ICML 2026. |
| 2026 | Received the ICML 2026 Silver Reviewer Award. |
| 2026 | AVoCaDO accepted to ICLR 2026. |
| 2026 | Conan accepted to CVPR 2026. |
| 2026 | RICo, on instruction-tuning data selection, accepted to AAAI 2026. |
| 2026 | Trajectory-Enhanced Camera Motion Understanding accepted to ICASSP 2026. |
| 2026 | Released our survey on multimodal token compression and its open paper collection. |
| 2026 | Joined Xiaomi's MiMo LLM-Core team through the “顶尖人才计划”. |
| 2025 | RICO, on image recaptioning, accepted to EMNLP 2025. |
| 2025 | TimeChat-Online accepted to ACM Multimedia 2025. |
Education
PhD in Computer Software and Theory · Advisor: Xu Sun
Master's in Computer Application Technology · Advisor: Qin Jin
Bachelor's in Computer Science and Technology
Selected Publications (Full List citations—)
* Equal contribution.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
MLLM Token Compression Survey
Generative Frame Sampler for Long Video Understanding
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Edit As You Wish: Video Caption Editing with Multi-grained User Control
CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge
Image Difference Captioning with Pre-training and Contrastive Learning
More publications & collaborationsHide additional publications
-
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
NeurIPS 2026 ED Track · PosterAn open evaluation suite for autonomous agents across real-world tasks. (contribute to the multimodal evaluation component) -
AdaC-GRPO: Towards Mastering Long-Horizon Visual Sudoku Puzzles via Adaptive Curriculum Learning
Findings of EMNLP 2026 -
Quality, Not Just Outcome: A Structural Scoring Framework for Selecting Coding-Agent Trajectories for Supervised Fine-Tuning
Findings of EMNLP 2026 -
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
ECCV 2026 -
ReaForest: Fostering Generative Video Reasoning for Spatial Planning
ICML 2026 -
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
ICLR 2026 -
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
CVPR 2026 -
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
AAAI 2026 -
Trajectory-Enhanced Camera Motion Understanding for Multimodal Large Language Models
ICASSP 2026 -
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
EMNLP 2025 -
Temporal Reasoning Transfer from Text to Video
ICLR 2025 -
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
ICMR 2024 -
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-text Generation?
NAACL 2024 -
Rethinking Benchmarks for Cross-modal Image-text Retrieval
SIGIR 2023, long paper -
基于语言描述的细粒度美妆图片排序
计算机科学 · 2020
Industry Experience
Honors & Awards
- Silver Reviewer Award · ICML · 2026
- ACM SIGMM Student Travel Grant · ACM Multimedia · 2024 & 2025
- National Scholarship · Ministry of Education of China · 2022
- Outstanding Graduate · Renmin University of China · 2023 & 2020
- 1st Class Grade Scholarship · Renmin University of China · 2022 & 2021
- Merit Student · Renmin University of China · 2021 & 2018
- 1st Prize of China Undergraduate Mathematical Contest in Modeling (Beijing) · Beijing · 2019
- Meritorious Winner of American Mathematical Contest In Modeling · U.S. · 2018
Academic Service & Teaching
Workshop & Challenge Organization
The MTVG and MDVC challenges attracted 40 teams worldwide.
Reviewing
- Conferences:
- CVPR (2024–2026), ICLR (2025–2026), ICML (2026), ECCV (2026), NeurIPS (2024–2026), AAAI (2023–2024), and ACM Multimedia (2024–2026).
- Journals:
- IEEE Transactions on Image Processing (TIP) and IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
Teaching Assistant
- Human Language and Artificial Intelligence · Peking University · 2024, 2026
- Academic Criterion and Writing · Renmin University of China · 2022
- Spoken Language Processing · Renmin University of China · 2020
- Multimedia Application Technology · Renmin University of China · 2020
Visitor map & total pageviews · MapMyVisitors









