Abstract
Architectural millwork and casework drawing sets — cabinet/casework elevations, floor plans, and reflected ceiling plans — are dense, low-redundancy technical documents that general-purpose vision-language models (VLMs) receive no domain-specific training for. We study whether an off-the-shelf, non-fine-tuned VLM can nonetheless localize the objects on such a sheet from a single instruction prompt and a single full-page image per request ("One-Stage" detection), and at what accuracy, latency, and dollar cost.
We introduce CaseV-Bench (Casework Vision Benchmark), a benchmark built around five object categories — elevation, floor_plan, cabinet, countertop, and callout — hand-annotated on real construction-document PDFs, and evaluate fourteen current VLMs from five providers (Google, Anthropic, OpenAI, xAI, and Alibaba/Qwen) under one fixed single-pass prompting protocol, scored by greedy IoU matching against ground truth at IoU ≥ 0.5 with real, provider-reported per-request dollar cost.
Across nine annotated sheet sets (119 pages, 1,430 ground-truth objects), aggregate micro-averaged F1 is 51.0%, ranging from 20.0% to 91.6% across models — a 4.6× spread wide enough that model choice dominates prompt design entirely. One model, gpt-6-astra, stands apart from the rest of the field: it is the only model tested that keeps a high F1 (>80%) on all three small, densely packed object types (cabinet, countertop, callout), but it is also the most expensive model per page in the evaluation. A second model, claude-opus-5.5, reaches 78.8% F1 at roughly one fifth of that cost and the lowest latency of any model tested, closing part — but not all — of the small-object gap.
An IoU-threshold sweep (0.10–0.90) shows that the model ranking is stable across thresholds, and suggests that the small-object gap has two different causes depending on the model: loosely placed boxes for mid-table models, and objects that go unreported altogether for the weakest ones. We report per-model, per-document, and per-object-type results and conclude that single-pass VLM prompting is, for at least one current model, a plausible unattended first pass for this document class, while remaining an assisted-review signal for most of the field.
This paper reports the Sep 2026 run, 14 models. The CaseVBench results page is updated as new models are tested.
Key findings
- F1 ranges from 20.0% to 91.6% across fourteen models on the same prompt, documents and threshold — model choice is the largest lever measured.
- GPT-6 Astra reaches 91.6% F1 and is the only model above 80% on all three small object types (cabinets, countertops, callouts). It is also the most expensive per page.
- Claude Opus 5.5 reaches 78.8% F1 at about one fifth of GPT-6 Astra’s cost, with the lowest latency of any model tested.
- A newer version is not reliably better: Grok 4.7 and GPT-6 Sol score below the models they succeed.
- The ranking holds across IoU thresholds from 0.10 to 0.90.
Cite this paper
Chumak, A., Mykytyn, I., Kozynets, A., Mykhailov, V., Didyk, Y., Mykytyn, Y., Hresko, V. (2026). Single-Pass Vision-Language Model Prompting for Object Localization in Architectural Millwork Drawings. COXIT technical report. https://coxit.co/ai-drawing-benchmark/paper/
@techreport{coxit2026casevbench,
title = {Single-Pass Vision-Language Model Prompting for Object Localization in Architectural Millwork Drawings},
author = {Andrii Chumak and Iryna Mykytyn and Andrian Kozynets and Victor Mykhailov and Yurii Didyk and Yelysaveta Mykytyn and Volodymyr Hresko},
institution = {COXIT},
year = {2026},
month = sep,
url = {https://coxit.co/ai-drawing-benchmark/paper/},
note = {Code and data sample: https://github.com/COXIT-CO/CaseV-bench}
}