Robusto-2:在利马和纽约市对自动驾驶领域的人类与虚拟智能体进行基准测试
Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
摘要
随着自动驾驶汽车在国际上不断普及,这些汽车开始使用VLM等多模态系统作为其动作模型的认知基础。那么,这些系统在新的环境中能够发挥多大的作用呢?尤其是在那些不属于常规分布范围的新地区中,它们又如何表现呢?在本文中,我们通过对利马的人类驾驶员、纽约市的人类驾驶员以及VLM进行全面的分析,并展示从利马和纽约市收集的行车记录仪视频内容,同时以视觉问答模式向他们提出各种问题。我们选择了这两个城市作为研究对象,因为这两个地方都是极具挑战性的驾驶环境,目前还没有任何自动驾驶汽车公司在此地开展业务。我们提出了四个类别的问题:事实性问题、评分问题、反事实问题和推理问题。我们发现,人类驾驶员和VLM的回答存在差异——不过这种差异受到所提问题的类型影响;而人类的回答则不受他们来自哪个地方的影响(利马或纽约)。令人惊讶的是,我们并没有发现由于地理位置而产生的明显差异,这可能是因为这些场景属于高挑战性的非常规分布场景。我们的数据集可以在以下链接获取:https://hf-mirror.com/datasets/Artificio/robusto-2
English Abstract
As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://hf-mirror.com/datasets/Artificio/robusto-2