CaseVBench: a benchmark of 11 multimodal AI models on 119 architectural drawing pages, September 2026. Models are re-run as new versions become available.
GPT-6 Astra scores 92%, outperforming all other models
We tested eleven multimodal AI models on 119 pages of architectural drawings. Each model received the same task: find every cabinet, countertop, floor plan, interior elevation, and callout on the page and draw a box around each one.
The reference set contained 1,430 objects marked by a human annotator, giving us a consistent way to compare the model results.
GPT-6 Astra scored 92% overall. Claude Fable 5.1 and Gemini 3.8 Flash followed at 70%. No other model scored above 65%.
That is a substantial change from our August test, when the highest score among the models we tested was 62%.
Large drawing elements are easier for most models
Floor plans and interior elevations were the easiest categories for most models.
Astra scored 99% on floor plans and 94% on elevations. Most of the other models also performed well on these categories, with floor-plan scores of at least 89% for all but one model and elevation scores of 86% or higher for eight of the eleven models.
These elements are generally larger and occupy more of the drawing sheet, which makes them easier to identify.
The larger differences appear with smaller objects.
Countertops
Astra scored 84% on countertops. The next highest model scored 34%.
Countertops can appear as thin lines along cabinet runs, making them more difficult to identify and outline on a full drawing sheet.
Callouts
Astra scored 96%, compared with 68% for the next highest model.
Callouts can be very small when an entire drawing sheet is provided to a model as a single image. At lower image resolutions, a callout may occupy only a few pixels.
Cabinets
Astra scored 82%. The next highest model scored 61%.
These results show that the difference between models is not simply about whether they can understand a drawing. Large drawing elements are handled reasonably well by many of the models in this test. The bigger differences appear when the objects are smaller or require more precise identification.
[coxit_compare ids=”1362,1363″ labels=”Ground truth|GPT-6 Astra” caption=”The expert marked 32 objects on this interior elevations sheet — 16 cabinets, 6 countertops, 9 elevations and 1 callout. GPT-6 Astra returned the same counts, scoring F1 1.00.”]
How the models compare
| Model | Cabinets | Countertops | Floor plans | Elevations | Callouts | All objects |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 82% | 84% | 99% | 94% | 96% | 92% |
| Claude Fable 5.1 | 61% | 34% | 98% | 87% | 64% | 70% |
| Gemini 3.8 Flash | 59% | 20% | 97% | 86% | 68% | 70% |
| GPT-5.6 Sol | 53% | 24% | 98% | 88% | 57% | 65% |
| Qwen3.8-Max | 60% | 32% | 98% | 90% | 47% | 62% |
| GPT-5.6 Terra | 51% | 23% | 94% | 87% | 38% | 56% |
| Gemini 3.5 Flash | 45% | 16% | 96% | 90% | 33% | 53% |
| Gemini 3.1 Pro Preview | 39% | 10% | 89% | 78% | 37% | 51% |
| Claude Opus 5 | 51% | 19% | 96% | 89% | 15% | 40% |
| Grok 4.6 | 5% | 7% | 91% | 69% | 0% | 30% |
| Claude Sonnet 5 | 10% | 0% | 73% | 47% | 4% | 21% |
Here we see two patterns stand out:
First, the models are generally stronger at finding large drawing elements than smaller ones. The rankings are largely determined by performance on cabinets, countertops, and callouts.
Second, newer models within the same family generally performed better than earlier versions in this test.
Gemini 3.8 Flash scored 17 points higher than Gemini 3.5 Flash. Claude Fable 5.1 scored 30 points higher than Claude Opus 5, with much of the difference coming from callouts: 64% compared with 15%.
GPT-6 Astra scored 27 points higher than GPT-5.6 Sol.
Fable shows a large improvement over Opus 5
Claude Fable 5.1 scored 70% overall, compared with 40% for Opus 5. Most of the improvement came from smaller objects: callouts increased from 15% to 64%, countertops from 19% to 34%, and cabinets from 51% to 61%.
Fable also produced fewer incorrect extra boxes. Opus often labelled circles, tags, and other symbols as callouts, while Fable was better at distinguishing them. Fable produced about 440 extra boxes across the test, compared with about 1,500 for Opus.
The broader pattern is similar across the models: large drawing elements are already handled well, while newer generations are making more progress on small symbols, thin lines, and other details on crowded pages.
Gemini 3.8 Flash improves mainly on callouts
Gemini 3.8 Flash also scored 70% overall, but its results differ from Fable’s. It found slightly fewer objects, while a higher share of its reported objects were correct. In practice, this means less deleting of incorrect boxes but more checking for objects it missed.
Most of its improvement over Gemini 3.5 Flash came from callouts, which increased from 33% to 68%. The trade-off is processing time: Gemini 3.5 Flash took about 20 seconds per page, while version 3.8 took about 104 seconds.
Cost and speed
Accuracy is only one consideration when choosing a model. We also recorded processing time and cost for each page.
| Model | Cost per page | Time per page | All objects |
|---|---|---|---|
| Gemini 3.5 Flash | $0.035 | 20 s | 53% |
| Gemini 3.1 Pro Preview | $0.040 | 23 s | 51% |
| GPT-5.6 Terra | $0.056 | 32 s | 56% |
| Qwen3.8-Max | $0.057 | 161 s | 62% |
| Claude Sonnet 5 | $0.073 | 60 s | 21% |
| GPT-5.6 Sol | $0.120 | 49 s | 65% |
| Claude Opus 5 | $0.134 | 48 s | 40% |
| Grok 4.6 | $0.145 | 362 s | 30% |
| Gemini 3.8 Flash | $0.150 | 104 s | 70% |
| Claude Fable 5.1 | $0.164 | 26 s | 70% |
| GPT-6 Astra | $0.220 | 39 s | 92% |
Astra was the most expensive model in the test at $0.22 per page. It also had the highest overall score.
Cost and accuracy did not follow the same pattern across the other models. Some of the more expensive models scored below less expensive alternatives.
Processing speed also varied independently of cost and accuracy.
Astra processed a page in about 39 seconds. Fable was the fastest of the three highest-scoring models at 26 seconds per page. Gemini 3.8 Flash was less expensive than both but took about 104 seconds per page.
Does this mean AI can do a takeoff?
This is where the benchmark needs some context. The test measures one part of a takeoff: finding predefined objects on a drawing page and drawing a box around each one.
A complete takeoff involves several additional steps. An estimator may need to read dimensions and schedules, distinguish between different cabinet types, match elevations to the rooms they belong to, calculate linear feet of countertop rather than simply count countertop views, and make decisions based on information that may be implied rather than explicitly shown.
None of those tasks are measured here.
What the results show is that AI has become more capable at locating objects within architectural drawings.
With Astra, 92% of the reference objects were correctly identified at the benchmark’s main scoring threshold. It scored 96% on callouts and 94% on elevations.
That makes its output useful as a starting point for navigating a drawing set and locating objects. But the remaining errors still require review, and performance varies between projects.
Astra also produced more additional boxes than missed objects. In practical terms, this means a reviewer may spend more time removing incorrect boxes than searching for objects the model failed to find. That can make the output easier to review, but it does not remove the need for review.
For the other ten models, the results suggest a greater level of checking is needed. At 70% overall or below, the output should be treated as an initial reference rather than a finished result. Countertops and callouts remain particularly difficult.
AI can assist with takeoffs, but object detection is only one step
AI can help locate and organize information in architectural drawings, but these results do not show that AI can complete a construction takeoff without human review.
Object detection is becoming more capable, particularly for large drawing elements and, with the best-performing model in this test, for smaller elements as well. A takeoff, however, requires more than locating objects.
We re-run this benchmark as new models ship; the running results live on CaseVBench.

