Yes, to a degree. Models detect many cabinets, but they can merge adjacent units into one object or miss units entirely. Counting accuracy depends heavily on how the drawing is processed before the model sees it.
How accurate is AI in reading the drawings?
We measure how accurately AI models read AEC drawings and detect objects: cabinets, countertops, elevations, and callouts.
Accuracy by model and object type
We gave each model a set of casework drawings and asked it to find every cabinet, countertop, elevation, and callout. This table shows how accurate the models are when they read a full sheet directly, the way it comes out of the drawing set.
| Model | Cabinets | Countertops | Elevations | Callouts |
|---|---|---|---|---|
| [Model A] | — | — | — | — |
| [Model B] | — | — | — | — |
| [Model C] | — | — | — | — |
| [Model D] | — | — | — | — |
With the COXIT booster
The models in this table are the same. What changes is how the drawing is prepared before the model reads it. Our processing layer divides each sheet into smaller views, similar to how an estimator scans a page in parts. The last row shows what this step adds.
| Model | Cabinets | Countertops | Elevations | Callouts |
|---|---|---|---|---|
| [Model A] | — | — | — | — |
| [Model B] | — | — | — | — |
| [Model C] | — | — | — | — |
| [Model D] | — | — | — | — |
| [Best model + COXIT booster] | — | — | — | — |
Results with the COXIT booster
The COXIT booster is a processing layer that prepares a drawing before an AI model reads it.
How the benchmarks were measured
Accuracy numbers are only useful when it is clear how they were measured. Here is how the benchmark works.
Drawings and markup
The benchmark is built on [N] casework projects from construction sets. An expert marks every cabinet, countertop, elevation, and callout by hand, [N] objects in total. This markup serves as the answer key.
One test for every model
Each model receives the same sheets and the same task. It has to find the objects and indicate where they are. A detection counts as correct when it matches the expert’s box.
Open results
We publish the results, the drawings, and the scoring method, including the cases where models fail. When a new model is released, we run it through the benchmark and update the tables.
Try on your own drawings
The tables above show what to expect in general. To find out what accuracy looks like on your drawings, send us a few sheets. We will run them through the same test and return the report.
//What happens next
- [1] You send a few sheets, in whatever state they are in.
- [2] We run them through the same benchmark, unchanged.
- [3] You get the marked-up drawings back, with your numbers next to the models'.
Frequently Asked Questions
-
-
The most common reasons are image resolution limits, small and dense symbols, and objects that look similar to construction lines. Much of the miscounting starts before inference, when the drawing is compressed to fit the model’s input size.
-
GPT models can read drawings and detect objects on them. In our benchmark, we test GPT alongside seven other models on the same pages and with the same evaluation rules, so part 2 will show how it compares.
-
We will publish the comparison in part 2. Our early runs suggest the workflow around the model affects results more than the choice between leading models.
-
Not on the basis of object recognition alone. It can automate parts of quantity takeoff, but estimating also involves scope, specifications, and judgment. See the section above.