Linli Yao 姚林丽

I am a PhD student in the Language Computing and Machine Learning Group (Lanco) at Peking University, advised by Prof. Xu Sun.

Previously, I received my master's and bachelor's degrees from Renmin University of China, advised by Prof. Qin Jin at the AI·M3 Lab.

I expect to graduate in June 2027 and am exploring full-time opportunities in academia and industry.

I welcome discussions and collaborations on long-horizon multimodal intelligence, multimodal agents, and embodied intelligence. Feel free to get in touch!

Linli Yao hiking on the Seceda Ridge Trail, with mountain peaks in the background
山就在那里。The mountain is there.

Research Interests

I study efficient and time-aware video understanding with multimodal large language models. My research connects two complementary directions:

  • Efficient video understanding. Visual token compression, adaptive frame sampling, and efficient processing of long and streaming videos.
  • Time-aware video-language modeling. Temporal grounding, temporal reasoning, and fine-grained, structured audio-visual captioning.

News

2026Claw-Eval accepted to NeurIPS 2026 ED Track as a poster.
2026Released MiMo-V2.6 — honored to contribute as a Core Contributor.
2026Two papers accepted to Findings of EMNLP 2026: AdaC-GRPO and Quality, Not Just Outcome.
2026DiaDem, on dialogue descriptions in audio-visual video captioning, accepted to ECCV 2026.
2026TimeChat-Captioner and ReaForest accepted to ICML 2026.
2026Received the ICML 2026 Silver Reviewer Award.
2026AVoCaDO accepted to ICLR 2026.
2026Conan accepted to CVPR 2026.
2026RICo, on instruction-tuning data selection, accepted to AAAI 2026.
2026Trajectory-Enhanced Camera Motion Understanding accepted to ICASSP 2026.
2026Released our survey on multimodal token compression and its open paper collection.
2026Joined Xiaomi's MiMo LLM-Core team through the “顶尖人才计划”.
2025RICO, on image recaptioning, accepted to EMNLP 2025.
2025TimeChat-Online accepted to ACM Multimedia 2025.

Education

Peking University

PhD in Computer Software and Theory · Advisor: Xu Sun

Sep 2023 – Jun 2027 (expected)
Renmin University of China

Master's in Computer Application Technology · Advisor: Qin Jin

Sep 2020 – Jun 2023
Renmin University of China

Bachelor's in Computer Science and Technology

Sep 2016 – Jun 2020

Selected Publications (Full List citations—)

* Equal contribution.

Xiaomi MiMo mountain illustration

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

LLM-Core Xiaomi (As a Core Contributor)
At Xiaomi, I work on multimodal SFT data, video-agent training data, webdev evaluation and RL grader design. I contributed to MiMo-V2-Omni, MiMo-V2.5 and MiMo-V2.6.
timechat-captioner research overview

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Linli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li, Xinlong Chen, Feifan Song, et al.
ICML 2026 CCF-A
A task, benchmark, dataset, and model for timestamped, structured audio-visual descriptions of multi-scene videos.
timechat-online research overview

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Linli Yao*, Yicheng Li*, Yuancheng Wei*, Lei Li, Shuhuai Ren, Yuanxin Liu, et al.
ACM MM 2025 CCF-A
Efficient streaming video understanding through temporal redundancy reduction, connecting video perception with visual token compression.
timechat research overview

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Shuhuai Ren*, Linli Yao*, Shicheng Li, Xu Sun, Lu Hou
CVPR 2024 CCF-A
A time-sensitive multimodal large language model for long-video understanding that unifies timestamp-based tasks, including dense video captioning, temporal video grounding, and video highlight detection.
Multimodal token compression survey overview

MLLM Token Compression Survey

Towards Efficient Multimodal Large Language Models: A Survey on Token Compression
Linli Yao*, Long Xing*, Yang Shi*, Sida Li, Yuanxin Liu, Yuhao Dong, Yi-Fan Zhang, Lei Li, Qingxiu Dong, et al.
TechRxiv preprint · 2026
gens research overview

Generative Frame Sampler for Long Video Understanding

Linli Yao, Haoning Wu, Kun Ouyang, Yuanxing Zhang, Caiming Xiong, Bei Chen, Xu Sun, Junnan Li
Findings of ACL 2025
A generative frame sampler that selects question-relevant frames for efficient long-video understanding.
deco research overview

DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Linli Yao, Lei Li, Shuhuai Ren, Lean Wang, Yuanxin Liu, Xu Sun, Lu Hou
Preprint · 2024
Decoupling visual token compression from semantic abstraction to simplify and improve multimodal representations.
vdedit research overview

Edit As You Wish: Video Caption Editing with Multi-grained User Control

Linli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou, Tiezheng Ge, Yuning Jiang, Xu Sun, Qin Jin
ACM MM 2024 CCF-A
Fine-grained, user-controlled video caption editing at the entity, attribute, and sentence levels.
capenrich research overview

CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge

Linli Yao, Weijing Chen, Qin Jin
The Web Conference (WWW) 2023 CCF-A
Enriching web-image captions with cross-modal pretrained knowledge.
idc research overview

Image Difference Captioning with Pre-training and Contrastive Learning

Linli Yao, Weiying Wang, Qin Jin
AAAI 2022 CCF-A
Describing fine-grained differences between images with pretraining and contrastive learning.
More publications & collaborationsHide additional publications

Industry Experience

2026.01 - Present
Research Intern · “顶尖人才计划”
MiMo LLM-Core, Xiaomi.
2024.12 - 2025.12
Research Intern
Kling Team, Kuaishou Technology, Advised by Yuanxing Zhang.
2024.08 - 2024.11
Research Intern
Multimodal Group @ 01.AI, Advised by Bei Chen and Junnan Li.
2022.10 - 2023.07
Research Intern
Alimama CV&NLP Group @ Alibaba, Advised by Tiezheng Ge.

Honors & Awards

  • Silver Reviewer Award · ICML · 2026
  • ACM SIGMM Student Travel Grant · ACM Multimedia · 2024 & 2025
  • National Scholarship · Ministry of Education of China · 2022
  • Outstanding Graduate · Renmin University of China · 2023 & 2020
  • 1st Class Grade Scholarship · Renmin University of China · 2022 & 2021
  • Merit Student · Renmin University of China · 2021 & 2018
  • 1st Prize of China Undergraduate Mathematical Contest in Modeling (Beijing) · Beijing · 2019
  • Meritorious Winner of American Mathematical Contest In Modeling · U.S. · 2018

Academic Service & Teaching

Workshop & Challenge Organization

Person in Context Workshop
2022.04 - 2022.10
Organizer / Workshop Chair
Person in Context (PIC) Workshop @ ACM MM 2022

The MTVG and MDVC challenges attracted 40 teams worldwide.

YouMakeup Video Challenge
2020.04 - 2020.07
Organizer
YouMakeup Video Challenge @ CVPR LVVU Workshop 2020

Reviewing

Conferences:
CVPR (2024–2026), ICLR (2025–2026), ICML (2026), ECCV (2026), NeurIPS (2024–2026), AAAI (2023–2024), and ACM Multimedia (2024–2026).
Journals:
IEEE Transactions on Image Processing (TIP) and IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).

Teaching Assistant

  • Human Language and Artificial Intelligence · Peking University · 2024, 2026
  • Academic Criterion and Writing · Renmin University of China · 2022
  • Spoken Language Processing · Renmin University of China · 2020
  • Multimedia Application Technology · Renmin University of China · 2020

Visitor map & total pageviews · MapMyVisitors