Robusto-2: Avaliação Comparativa de Humanos e VLMs para Condução Autônoma em Lima e Nova York
Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
June 18, 2026
Autores: Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma, Victor Flores-Benites, David Ortega, Arturo Deza
cs.AI
Resumo
À medida que os Carros Autônomos continuam a se expandir internacionalmente e a utilizar sistemas multimodais, como os VLMs, como espinha dorsal cognitiva para seus Modelos de Ação, qual será a capacidade desses sistemas de generalizar em novos cenários, especialmente em situações extremas fora da distribuição (OOD) em novas geografias? Neste artigo, investigamos essa questão em aberto por meio de uma análise fatorial completa envolvendo motoristas humanos de Lima, motoristas humanos da Cidade de Nova York e VLMs, exibindo a eles filmagens de câmeras de painel coletadas em Lima e Nova York, e os instigando com uma variedade de perguntas dentro de um paradigma de Perguntas e Respostas Visuais (VQA). Em particular, escolhemos essas duas cidades por serem locais de direção altamente desafiadores, onde nenhuma empresa de Carros Autônomos opera atualmente, e formulamos perguntas que abrangem quatro categorias: Factual, Avaliações, Contrafactual e Raciocínio. Constatamos que humanos e VLMs divergem em suas respostas – embora isso seja modulado pelo tipo de perguntas feitas – e que os humanos respondem de forma semelhante, independentemente de sua origem (Lima/NYC). Para nossa surpresa, não encontramos uma forte diferença nas respostas (humanos ou VLMs) modulada pela geografia, provavelmente devido à sua natureza altamente fora da distribuição. Nosso conjunto de dados está disponível em: https://huggingface.co/datasets/Artificio/robusto-2
English
As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://huggingface.co/datasets/Artificio/robusto-2