VDC-Agent

When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

Qiang Wang1, Xinyuan Gao2, Yuhang He1, Jizhou Han1, Jiangyang Li1, Songlin Dong3,4, Zhiheng Ma3,4,5, Yihong Gong1,3

ECCV 2026

Highlights

Method

Module A runs agentic self-reflection on unlabeled videos; Module B turns caption–score trajectories into preference pairs; Module C fine-tunes with curriculum DPO.

Overview of VDC-Agent framework
Figure 1. Overview of VDC-Agent: agentic self-reflection, dataset construction, and curriculum DPO.

Results

VDCscore on the VDC benchmark (Accuracy / Score). Higher is better.

Model Camera Short Background Main Object Detailed Average
Cockatiel-8B 42.25 / 2.19 44.01 / 2.27 43.89 / 2.26 43.85 / 2.26 44.00 / 2.27 43.60 / 2.25
AVC-DPO-7B 50.40 / 2.66 39.00 / 2.03 49.90 / 2.57 50.50 / 2.58 48.90 / 2.54 47.70 / 2.47
Qwen2.5-VL-7B 42.61 / 2.18 37.76 / 1.91 44.10 / 2.22 49.00 / 2.47 46.25 / 2.35 43.95 / 2.23
VDC-Agent-7B 50.52 / 2.67 39.49 / 1.99 51.93 / 2.62 53.23 / 2.65 50.21 / 2.55 49.08 / 2.50

Citation

@article{vdcagent,
  title={VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection},
  author={Wang, Qiang and Gao, Xinyuan and He, Yuhang and Han, Jizhou and Li, Jiangyang and Dong, Songlin and Ma, Zhiheng and Gong, Yihong},
  journal={arXiv preprint arXiv:2511.19436},
  year={2025}
}