Understanding F1 Scores in AI Drawing Benchmarks

An estimator scratching their head over an architectural floor plan marked with red question marks, with a calculator and a scale ruler on the desk.

When we publish CaseVBench results, one of the main numbers we report is called F1.

CaseVBench is a benchmark for measuring how accurately AI models can locate objects on architectural drawings, such as cabinets, countertops, floor plans, elevations and callouts.

What is F1?

F1 is a single score that combines two questions about a model’s output: how much of what it found was correct (precision), and how much of what exists it found (recall). It always lands closer to the weaker of the two, so a model can’t score well by being good at only one.

F1 is a standard metric in machine learning, but the name does not tell you much about what it measures.

Precision tells you how many of the model’s predictions are correct

Imagine you hand someone a kitchen drawing and ask them to circle every cabinet. There are 6 cabinets on the page.

They circle 5 things. When you check, 4 of those are cabinets. The fifth is a fridge they mistook for a cabinet.

A kitchen drawing with 6 cabinets, a fridge and a window. The person circled 5 things: 4 cabinets, marked correct, and the fridge, marked wrong.

Of the 5 things they identified, 4 were correct. That gives a precision of 80%.

Recall tells you how many of the actual objects the model found

There are 6 cabinets on the drawing and the person found 4 of them. That gives a recall of 67%.

Recall: of the 6 cabinets on the page, 4 were found and 2 were missed, shown with a dashed outline, so recall is 4 / 6 = 67%.

The two measurements describe different types of errors. Precision is affected by incorrect predictions. Recall is affected by objects that were missed.

F1 combines precision and recall so that both types of error affect the score

Each measurement can look good on its own.

A cautious person might circle one cabinet they are completely sure about. Precision would be 100%, but recall would be only 17%.

Another person might circle everything that looks even slightly like a cabinet, say 12 things. All 6 cabinets are in there, so recall is 100%, but half of what they circled is wrong, so precision is 50%.

F1 combines precision and recall into one number, and it always lands closer to the weaker of the two. This is what stops a lopsided score from hiding behind a good one. For the person with 80% precision and 67% recall, F1 is approximately 73%, about the same as a plain average. For the person with 100% recall and 50% precision, it is 67%, where a plain average would give 75%.

The formula is F1 = 2 × (precision × recall) / (precision + recall).

The gap is widest when precision and recall are far apart. With 100% precision and 17% recall, for example, F1 is approximately 29%, rather than the 58% that a simple average would produce.

Three examples compared. Cautious: precision 100%, recall 17%, plain average 58%, F1 29%. Our example: precision 80%, recall 67%, plain average 73%, F1 73%. Reckless: precision 50%, recall 100%, plain average 75%, F1 67%.

A 70% F1 doesn’t mean 70% of the drawing is correct

An F1 score of 70% does not mean that the model understands 70% of the drawing.

In CaseVBench, a model is asked to find specific objects and draw a rectangle around each one. That rectangle is called a bounding box. Each prediction is compared with a rectangle a person drew around the same object, and the two count as a match only if they overlap enough.

F1 measures detection performance under a specific evaluation method. It is not a general measure of how well an AI model understands an architectural drawing.

How close the box has to be to count

Finding the right cabinet is only half of a match. The model also has to draw its box in roughly the right place. The rule for “roughly” is called IoU.

What is IoU?

Lay the model’s box and the person’s box on top of each other. IoU is the area they share, divided by the total area the two boxes cover together. Boxes sitting exactly on top of each other score 1; boxes that don’t touch score 0. The name stands for Intersection over Union, which is the maths way of saying overlap over combined area.

CaseVBench counts a prediction as a match when IoU is at least 0.50, so at least half of the combined area has to be shared. A box that is a little off still passes. A box that slid halfway off the cabinet, or one drawn twice too big, does not, even though the model clearly saw the cabinet.

One cabinet and three model boxes. A little off: IoU 0.78, counts. Slid half off: IoU 0.33, does not count. Twice too big: IoU 0.36, does not count.

The threshold changes F1 too. The same set of predictions scores higher at IoU 0.20 and lower at 0.70, which is why every F1 in the results says which threshold it was measured at. The primary CaseVBench results use 0.50.

F1 can be very different for different types of objects

The same model can perform well on one type of architectural object and poorly on another.

In the initial CaseVBench results, Qwen3.8-Max scored 98% F1 on floor plans, 90% on elevations, 60% on cabinets, 47% on callouts and 32% on countertops. Its overall F1 was 62%.

Large objects such as floor plans and elevations are easier to locate than small callouts. Countertops can also be difficult because they may appear as thin lines within an elevation rather than as clearly defined shapes.

This is why an overall F1 score needs to be read alongside the object type being measured.

Two models with the same F1 can make different kinds of mistakes

F1 does not show whether a model’s errors mainly come from missed objects or incorrect predictions.

One model might have relatively high precision but lower recall. It finds fewer objects, but most of its predictions are correct.

Another might have higher recall but lower precision. It finds more objects, but also produces more incorrect detections.

The two models could have similar F1 scores while requiring different amounts of human review. That is why precision and recall are useful alongside F1 when comparing models.

An F1 score needs to be interpreted in the context of the task

There is no universal F1 score that defines whether a model is good enough for architectural drawings. The useful level of performance depends on the task, the object type and how the output will be used.

When reading a CaseVBench result, look at F1, precision, recall, IoU threshold and object type together. F1 gives you the combined result, while the other measurements provide the context needed to understand it.

Follow the benchmark

A ground floor plan open on a laptop, at the scale an estimator reads it

We re-run this benchmark whenever a notable new model comes out. Leave your email and the updated results land in your inbox.

Frequently Asked Questions

  • F1 is a machine learning metric that combines precision and recall into a single score.

  • IoU, or Intersection over Union, measures the overlap between a predicted bounding box and a human-labelled bounding box.

  • It means the model has a particular balance of precision and recall under the evaluation criteria used. It does not mean that 70% of the drawing is correct.

  • Different objects have different visual characteristics. In CaseVBench, models show substantially different F1 scores across object types, with floor plans and elevations generally scoring higher than smaller or more difficult objects.

  • Two models can have the same F1 score while making different types of errors. Precision and recall show whether those errors are mainly incorrect predictions or missed objects.

Related articles

An architect's workstation with a residential floor plan open in CAD software on a large monitor, a rolled drawing set and a notebook of hand sketches on the desk. AI for Drawings

GPT-6 Astra Scores 92% on Architectural Drawings: What Does It Mean for AI Takeoffs?

woodworking-operations-ai AI for Drawings

From 6 Hours to 10 Minutes: How AI Transformed Estimating at the Largest U.S. Casework Factory

takeoff-woodworking-ai AI for Drawings

6 Questions About AI Takeoffs for Casework Shops

Let's collaborate

Tell us a bit about your project or challenge, and we'll get back to you shortly.

Volodymyr Hresko Volodymyr Hresko Co-Founder & COO

Reach out directly

[email protected]
This field is for validation purposes and should be left unchanged.
Full name
By submitting the form, you agree to Coxit’s Privacy Policy.