VDC-Agent
When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- 1 State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University
- 2 Kuaishou Technology
- 3 Shenzhen University of Advanced Technology
- 4 Guangdong Provincial Key Laboratory of Computility Microelectronics
- 5 Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
ECCV 2026
Highlights
- Self-evolving captioner — closed loop of generation, principle-guided scoring, and prompt refinement; no human labels, no larger teacher models.
- VDC-Agent-19K — 18,886 preference pairs built automatically from unlabeled Cockatiel-4K videos.
- Curriculum DPO — train from large to small score gaps (easy → hard).
- VDC-Agent-7B — 49.08% Acc / 2.50 score on VDC (+5.13 Acc / +0.27 over Qwen2.5-VL-7B) at the same inference cost.
Method
Module A runs agentic self-reflection on unlabeled videos; Module B turns caption–score trajectories into preference pairs; Module C fine-tunes with curriculum DPO.
Results
VDCscore on the VDC benchmark (Accuracy / Score). Higher is better.
| Model | Camera | Short | Background | Main Object | Detailed | Average |
|---|---|---|---|---|---|---|
| Cockatiel-8B | 42.25 / 2.19 | 44.01 / 2.27 | 43.89 / 2.26 | 43.85 / 2.26 | 44.00 / 2.27 | 43.60 / 2.25 |
| AVC-DPO-7B | 50.40 / 2.66 | 39.00 / 2.03 | 49.90 / 2.57 | 50.50 / 2.58 | 48.90 / 2.54 | 47.70 / 2.47 |
| Qwen2.5-VL-7B | 42.61 / 2.18 | 37.76 / 1.91 | 44.10 / 2.22 | 49.00 / 2.47 | 46.25 / 2.35 | 43.95 / 2.23 |
| VDC-Agent-7B | 50.52 / 2.67 | 39.49 / 1.99 | 51.93 / 2.62 | 53.23 / 2.65 | 50.21 / 2.55 | 49.08 / 2.50 |
Citation
@article{vdcagent,
title={VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection},
author={Wang, Qiang and Gao, Xinyuan and He, Yuhang and Han, Jizhou and Li, Jiangyang and Dong, Songlin and Ma, Zhiheng and Gong, Yihong},
journal={arXiv preprint arXiv:2511.19436},
year={2025}
}