Logical Anomaly Detection: Turning Inspection Rules Into Questions
LogicQA turns inspection rules into visual questions. Explore the connector example, benchmark tradeoffs and a practical way to test generated checklists.
7
Select a figure to open it at full size.
A product can look undamaged and still be wrong. A connector assembly may have the wrong number of clamps, a package may contain an extra item, or individually valid parts may be arranged incorrectly. These are logical defects: the problem is a violated requirement rather than an unusual texture.
LogicQA, developed by Yejin Kwon, Daeun Moon, Youngje Oh and Hyunsoo Yoon, investigates whether a vision-language model can detect such violations by answering a generated checklist. For manufacturing and inspection teams, the attraction is a system that can explain which requirement failed without training a new model for every product class.
The research supports a specific approach to constructing and testing those questions. It does not show that a general-purpose visual model can reliably inspect any production line from a few photographs.
The Core Insight: Define Normal Before Asking What Is Wrong
A generic question such as “Is this image defective?” leaves the model to invent the inspection standard. LogicQA instead begins with a provided definition of normality and three normal images. The definition supplies constraints such as required objects, counts and relationships; the photographs show how those constraints appear.
The model describes each normal image, summarizes their shared characteristics, and turns that summary into yes/no questions. A separate set of normal examples helps remove questions that are too specific to the initial photographs: questions answered correctly on fewer than 80% of the normal validation cases are excluded.
That filtering step matters. Three examples might all happen to have orange levers or a particular cable shape. A generated question can mistakenly turn that coincidence into a universal requirement. LogicQA does not eliminate this risk by generating more fluent descriptions; it tests whether its questions accept other valid examples.
The original pipeline makes the supplied rules, image descriptions and generated questions visible as separate stages.
Original Figure 2, Kwon and colleagues. LogicQA converts a definition of normality and a few images into testable questions, then checks new images against them. Source workflow.
Walk Through the Connector Example
In the paper’s example, two splicing connectors should be joined by one cable, with matching numbers of clamps. The generated checklist asks about the number of connectors and the configuration of their clamps.
For a test image, the model answers five differently worded versions of each main question. The majority answer determines whether that requirement passes. If any required condition fails, the system marks the image anomalous and identifies the failing question.
The illustrated assembly passes the two-connector check but fails the clamp check: one connector has only two clamps. The output therefore offers something more useful than a red warning light. It directs the operator toward the part of the assembly that violates the expectation.
There is a caution inside the example itself. One illustrated question specifies five clamps, while the broader rule is that the two connectors must match. A five-clamp requirement is appropriate only if the product specification actually requires it. For another valid connector variant, that question could reject a correct assembly. A useful implementation needs a product owner to approve the requirements, not merely accept the model’s checklist.
Repeated questions can reduce sensitivity to wording, but five answers from the same model are not five independent inspectors. A model that consistently miscounts can repeat the same mistake in every paraphrase.
What the Research Actually Shows
The main public evaluation uses the logical-anomaly portion of MVTec LOCO AD. Its five categories include breakfast boxes, juice bottles, pushpins, screw bags and splicing connectors. The tests use three normal reference images, require no model fine-tuning, and average LogicQA’s results over three runs.
The original comparison is useful because it exposes the uneven results across products.
Original Table 1, Kwon and colleagues. AUROC measures separation between normal and anomalous images across score thresholds; F1-max is the best combined precision/recall score across thresholds. Neither is the percentage of production defects caught. Source comparison and conditions.
LogicQA’s average AUROC is 87.6, compared with 86.0 for LogicAD. The connector result is substantially stronger: 92.4 versus 73.4. But the screw-bag result is weaker, at 71.5 versus 83.8. The practical conclusion is that the question-based approach helps some forms of structured inspection more than others; an average improvement cannot decide whether it fits your product.
The expanded appendix comparison includes methods using additional in-house annotations that score higher overall. LogicQA’s contribution is the combination of limited examples, generated questions and explanations—not universal superiority over every inspection approach.
A second evaluation uses a private semiconductor dataset with spot and bridge defects. It reports an average AUROC of 90.3 for LogicQA with GPT-4o, versus 79.2 for PatchCore. This is useful evidence from another setting, but the actual industrial images are not public; the appendix’s sample illustrations should not be mistaken for released test data.
The Images Still Have to Be Understandable
The checklist does not bypass perception. In the public benchmark, the researchers use preprocessing to reduce distracting background and isolate repeated components. Background masking focuses attention on the object; a language-guided segmentation pipeline separates relevant parts so the model can inspect them more effectively.
This is consequential for implementation. A model may understand the rule “two matching connectors” while failing to see the clamps at the supplied resolution. Better wording cannot recover details that the image never captured clearly.
The paper also evaluates the explanations with two human annotators. Their agreement with the model is high for normal images and lower for anomalous images—around 85–86% for the latter. That supports the usefulness of the explanations while leaving meaningful room for incorrect rationales. An explanation should tell the operator where to look; it should not be treated as proof that the defect exists.
Real-World Applications: Build a Checklist You Can Audit
A sensible first application is a product family with explicit, visible requirements: required components, allowable counts and permitted arrangements. Avoid beginning with a loosely defined standard such as “looks well assembled.”
Build a normal reference set containing legitimate variation: approved colors, orientations, suppliers and camera conditions. Keep some normal variants out of question generation and use them to test the checklist. Include known defect examples for final evaluation, even if the model does not need them for fine-tuning.
Review failures at the question level. If correct products repeatedly fail a color question, the rule may be wrong. If a count question fails only under glare, the imaging process may be wrong. If the same question is unstable across paraphrases, the model may be unsuitable for that requirement.
Measure missed defects and unnecessary rejections at the chosen operating rule, along with total model calls and inspection latency. A production decision needs these quantities, not a threshold-swept benchmark score alone. For cases where the main concern is a classifier relying on irrelevant visual clues, explanation-based inspection review addresses a different, complementary problem.
Implementation Frameworks
The simplest implementation is a versioned question-generation and evaluation pipeline around a vision-language model. Store the approved normality definition, reference images, generated questions, rejected questions and model version. Keep generation separate from production inspection so a new image cannot silently change the standard it is being judged against.
Use the LogicQA paper’s appendix as the starting specification for prompts and evaluation, rather than reproducing the approach from a short description. Recreate one category, inspect each intermediate question, and confirm that your result parsing handles incomplete or ambiguous answers.
For small repeated components, the study’s language-guided segmentation approach is worth testing. Grounding DINO identifies objects from text descriptions, while segmentation provides more focused visual inputs. These components help the model see the relevant objects; they do not determine whether the assembly satisfies its requirements.
Compare the complete pipeline with a simpler count, geometry or established anomaly-detection baseline wherever that baseline can express the same rule. A language model earns its additional cost when it handles variation and explains failures usefully—not merely because it can produce a longer answer.
TechClarity’s View
LogicQA offers a practical way to make inspection assumptions visible. The generated questions can become an artifact that engineering and quality teams inspect together, rather than an implicit standard hidden in a model.
The strongest adoption case is a narrow, well-specified inspection problem with enough normal variation to challenge the checklist. We would expand only after the questions, perception and operating threshold all survive product-specific testing. A fluent explanation is useful; a validated inspection rule is what makes it dependable.