Energent scored ahead of the cited best-published result on four of five public engineering-document benchmarks. Every score comes from the benchmark's own scorer, executed unmodified.
TL;DR
- Energent leads on DrawingVQA, BlueprintSymVL, CircuitSense, and CADBench.
- On CircuitSense, Energent scores 69.6% versus 12.3%. On BlueprintSymVL, it scores 74.0% versus 50.5%.
- The results evaluate Energent's complete agentic workflow, including its ability to crop, re-inspect, and execute code.
Context
Engineering-document understanding spans several distinct skills. Construction drawings, P&IDs, circuit diagrams, and CAD programs each require different outputs and use purpose-built scoring methods.
We ran Energent against five public benchmarks using each benchmark's original scoring implementation. Each result is presented separately so it retains the meaning of its native metric.
Each benchmark uses its own metric on a 0–100 scale, so each row is a self-contained comparison.
Results
| Domain | Benchmark and metric | Energent | Best published |
|---|---|---|---|
| Construction drawings | DrawingVQA — visual QA accuracy | 89.1% | 77.2% (Gemini-3-pro) |
| Process and piping | BlueprintSymVL — P&ID symbol exact match | 74.0% | 50.5% |
| Analog electronics | CircuitSense — verified symbolic derivation | 69.6% | 12.3% |
| Mechanical CAD | CADBench — Aligned IoU | 55.3% | 45.1% (Claude Opus 4.7) |
| Mechanical CAD | BenchCAD — part QA score | 52.6% | 58.7% |
DrawingVQA: 89.1% vs. 77.2%
DrawingVQA evaluates visual question answering on construction drawings. Energent scores 89.1%, compared with 77.2% for Gemini-3-pro on the current dataset leaderboard.
The paper reports human baselines of 78.2% for a young professional and 94.9% for an experienced professional. The Energent result uses a majority vote of three runs.
BlueprintSymVL: 74.0% vs. 50.5%
BlueprintSymVL measures symbol recognition in engineering blueprints and P&IDs. Energent reaches 74.0% exact match, compared with the paper's best single-turn visual-language-model result of 50.5% across Tables 2–5.
The dataset is publicly available.
CircuitSense: 69.6% vs. 12.3%
CircuitSense tests whether a system can move from a circuit diagram to a verified symbolic derivation. Energent scores 69.6%, compared with the published 12.3% reference from Table 6, reweighted to match our level composition.
The dataset is available under the MIT license.
CADBench: 55.3% vs. 45.1%
CADBench evaluates program generation for mechanical CAD. Energent reaches 55.3% Aligned IoU, compared with a 45.1% published anchor from Claude Opus 4.7.
This result covers four of CADBench's six families. We reweight the published per-stratum values from Table 5 across those same four families, keeping the evaluation field consistent for Energent and the published reference.
The dataset is available under the MIT license.
BenchCAD: 52.6% vs. 58.7%
BenchCAD evaluates visual question answering on mechanical CAD parts. Energent scores 52.6% on the Vision-QA field, alongside the published leaderboard result of 58.7% on the repository's leaderboard.
The dataset is publicly available.
Methodology
For every benchmark, we vendored the scorer at a pinned commit and executed it unmodified. Extraction is ours; scoring is theirs.
The evaluation measures the complete Energent agentic system. Energent can crop a drawing, inspect a region again, and execute code during its reasoning loop; the cited published figures provide the model baselines for each benchmark.
Conclusion
Across construction drawings, P&IDs, circuit diagrams, and mechanical CAD, Energent leads the cited best-published result on four of five benchmarks. The results demonstrate the value of an agentic workflow for understanding complex engineering documents.
