How accurate is AI in reading the drawings?

We measure how accurately AI models read AEC drawings and detect objects: cabinets, countertops, elevations, and callouts.

Accuracy by model and object type

We gave each model a set of casework drawings and asked it to find every cabinet, countertop, elevation, and callout. This table shows how accurate the models are when they read a full sheet directly, the way it comes out of the drawing set.

Accuracy by model and object type
Model Cabinets Countertops Elevations Callouts
[Model A]
[Model B]
[Model C]
[Model D]

With the COXIT booster

The models in this table are the same. What changes is how the drawing is prepared before the model reads it. Our processing layer divides each sheet into smaller views, similar to how an estimator scans a page in parts. The last row shows what this step adds.

With the COXIT booster
Model Cabinets Countertops Elevations Callouts
[Model A]
[Model B]
[Model C]
[Model D]
[Best model + COXIT booster]

Results with the COXIT booster

The COXIT booster is a processing layer that prepares a drawing before an AI model reads it.

[Model A]

[ screenshot — Model A on sheet A-201 ]
Cabinets boxed where the model placed them, misses left in.

[Model B]

[ screenshot — Model B on sheet A-201 ]
The same sheet, read straight through with no preparation.

[Model C]

[ screenshot — Model C on sheet A-201 ]
Detections on the same sheet, including the false ones.

[Model D]

[ screenshot — Model D on sheet A-201 ]
Adjacent units merged into single boxes in the run of base cabinets.

[Best model + COXIT booster]

[ screenshot — Best model + COXIT booster on sheet A-201 ]
The same sheet after the booster splits it into smaller views first.

How the benchmarks were measured

Accuracy numbers are only useful when it is clear how they were measured. Here is how the benchmark works.

[1]

Drawings and markup

The benchmark is built on [N] casework projects from construction sets. An expert marks every cabinet, countertop, elevation, and callout by hand, [N] objects in total. This markup serves as the answer key.

[2]

One test for every model

Each model receives the same sheets and the same task. It has to find the objects and indicate where they are. A detection counts as correct when it matches the expert’s box.

[3]

Open results

We publish the results, the drawings, and the scoring method, including the cases where models fail. When a new model is released, we run it through the benchmark and update the tables.

How well can AI read your drawings? Read the full methodology

Try on your own drawings

The tables above show what to expect in general. To find out what accuracy looks like on your drawings, send us a few sheets. We will run them through the same test and return the report.

A ground floor plan open on a laptop, at the scale an estimator reads it

//What happens next

  • [1] You send a few sheets, in whatever state they are in.
  • [2] We run them through the same benchmark, unchanged.
  • [3] You get the marked-up drawings back, with your numbers next to the models'.

Frequently Asked Questions

  • Yes, to a degree. Models detect many cabinets, but they can merge adjacent units into one object or miss units entirely. Counting accuracy depends heavily on how the drawing is processed before the model sees it.

  • The most common reasons are image resolution limits, small and dense symbols, and objects that look similar to construction lines. Much of the miscounting starts before inference, when the drawing is compressed to fit the model’s input size.

  • GPT models can read drawings and detect objects on them. In our benchmark, we test GPT alongside seven other models on the same pages and with the same evaluation rules, so part 2 will show how it compares.

  • We will publish the comparison in part 2. Our early runs suggest the workflow around the model affects results more than the choice between leading models.

  • Not on the basis of object recognition alone. It can automate parts of quantity takeoff, but estimating also involves scope, specifications, and judgment. See the section above.