ChatPaper.aiChatPaper

OpenBioRQ: Questões de Pesquisa Biomédica Não Resolvidas para Agentes

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

June 20, 2026
Autores: Minbyul Jeong
cs.AI

Resumo

Uma citação funcional parece prova -- mas o fato de um link ser resolvido não significa que o artigo citado apoia a afirmação. Descubro que os modelos agentes atuais raramente fabricam citações (mais de 99% são resolvidas), mas cerca de 15,9% delas apontam para o artigo errado. Os benchmarks existentes ignoram esse modo de falha: quando uma pergunta tem um gabarito fixo, um modelo pode reproduzir a fonte esperada a partir desse gabarito, em vez de verificar independentemente se a fonte apoia a afirmação. Apresento o \openbiorq{}, um benchmark agente fundamentado em recuperação com 12.553 perguntas não resolvidas de pesquisa biomédica em 12 domínios, que trata perguntas abertas como uma sonda de fidelidade e abstenção. Até onde sei, este é o primeiro benchmark biomédico a combinar um cenário agente -- onde o modelo deve realizar múltiplas chamadas de ferramentas -- com perguntas não resolvidas que não possuem gabarito. A abertura é verificada contra evidências reais de acompanhamento, em vez do conhecimento paramétrico do modelo. A dificuldade é empírica: fundamento-a em perguntas que três modelos de referência de pesos abertos não conseguem responder, em vez de rótulos subjetivos de dificuldade. Nesse subconjunto mais difícil, modelos retidos da mesma linhagem dos âncoras de dificuldade resolvem apenas ~17%, enquanto três agentes de fronteira independentes (Gemini-3-Pro, Opus-4.7, GPT-5.5) abrangem uma ampla faixa de 29% a 60%. O benchmark é, portanto, difícil, não saturante (o melhor agente ainda deixa ~33-40% não resolvidos) e discriminativo entre níveis de capacidade. Além da dificuldade, observo colapso agente nas perguntas mais difíceis, onde os agentes param de usar suas ferramentas. Para o modelo mais propenso a colapso, bloquear completamente o acesso às ferramentas mal altera sua pontuação -- portanto, as ferramentas deixam de ser úteis exatamente onde são mais necessárias. Uma lista de verificação fixa por pergunta eleva a concordância entre avaliadores de Spearman 0,35 para 0,82.
English
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over 99% resolve), yet roughly 15.9% link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source supports the claim. I introduce \openbiorq{}, a retrieval-grounded agentic benchmark of 12{,}553 unsolved biomedical research questions across 12 domains that treats open questions as a faithfulness-and-abstention probe. To my knowledge, this is the first biomedical benchmark to combine an agentic setting -- where the model must issue multiple tool calls -- with unsolved questions that have no answer key. Openness is verified against real follow-up evidence rather than a model's parametric knowledge. Difficulty is empirical: I anchor it on questions that three open-weight reference models fail to answer, rather than on subjective hardness labels. On this hardest subset, held-out models from the same lineage as the difficulty anchors solve only ~17%, while three independent frontier agents (Gemini-3-Pro, Opus-4.7, GPT-5.5) span a wide 29-60% range. The benchmark is thus hard, non-saturating (the best agent still leaves ~33-40\% unsolved), and discriminating across capability tiers. Beyond difficulty, I observe agentic collapse on the hardest questions, where agents stop using their tools. For the most collapse-prone model, blocking tool access entirely barely changes its score -- so tools stop paying off exactly where they are needed most. A frozen per-question checklist raises inter-judge agreement from Spearman 0.35 to 0.82.