Harder Data, Same Agent: How Far Each Model Fell
Databricks made OfficeQA harder, from OfficeQA Pro to OfficeQA Pro V2: 2× the corpus, 3.7× the documents cited per question, about 9× the text to read. Energent kept its model-agnostic agent fixed and compared seven frontier models on both editions. Fable 5.1 and GPT-5.6 Sol held within 3 points; the other five fell by 6.5 to 28.8, and the ranking changed.
Read article
