‹ 返回 2026-06-17

自我演化视觉提问器

Self-Evolving Visual Questioner

▲ 9 💬 1 2026-06-17

Yijun Liang, Hengguang Zhou, Ming Li, Lichen Li, Cho-Jui Hsieh, Tianyi Zhou

摘要

视觉语言模型通常被训练成被动的提问者,而它们主动提出各种非平凡、以视觉为中心且与现实场景相关的问题的能力却尚未得到充分探索。现有的视觉提问模型的性能受到高质量训练数据的可用性或数据收集成本的限制。我们证明,无需外部监督,视觉语言模型就能持续提升自身的提问能力。我们提出了一种自进化的框架:利用视觉语言模型本身作为提问者和过滤器,从而生成更困难、更具信息量的、以视觉为中心的问题。同时,该框架还能保持问题的多样性,避免训练过程中出现停滞现象。这些问题随后被用于训练视觉语言模型,使其同时具备提问和回答的能力。为了评估这种提问模型的效果,我们引入了一种基于代理的评估机制,该机制可以从感知、推理以及问题多样性等多个维度来评估问题质量。通过对多种主流视觉语言模型进行实验验证,我们发现我们的方法显著提升了问题的质量,并大幅拓宽了自主生成问题的难度范围。在相同的预算下,这种自监督方式比基于静态训练数据的传统方法更为有效。此外,这种自进化的提问模型在回答问题时也表现得相当出色,甚至优于传统的提问模型。

English Abstract

Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high-quality training data or the cost of curating them. We show that a VLM can continuously improve itself as a visual questioner without any external supervision. We propose a self-evolving framework that uses a VLM itself as both a proposer and a filter to produce harder, more informative, and visual-centric questions, while maintaining their exploration diversity to avoid training collapse. These questions are then used to train the VLM in both questioner and answerer modes. To evaluate the questioner, we introduce an agentic protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across various backbone VLMs show that our method substantially enhances the quality and substantially expands the difficulty boundary of autonomous question generation. Under the same budget, our self-supervision is more effective than training on the static source data. Moreover, the self-evolving questioner remains a competitive or even better answerer.