Machine Vision

How Accurate Must Machine Vision Be for Defect Detection?

Publication Date

Oct 11, 2026

author

TSV Data Lab

A vision system can appear highly accurate on a test set and still be unsafe or uneconomical on a production line. A camera may find obvious scratches on a flat metal part, then miss a shallow crack when the part rotates slightly, the illumination ages, or oil residue changes the surface reflection. In another line, the system may catch every questionable feature but reject too many acceptable parts, creating rework queues and operator distrust.

How accurate must machine vision be for defect detection? It must be accurate enough to meet the risk of a missed defect, the allowable rate of false rejects, and the repeatability required under actual operating conditions. There is no universal pass mark such as 95% or 99%. A cosmetic inspection can tolerate a different error profile from an inspection intended to prevent a safety-critical assembly failure. The useful target is a validated combination of detection sensitivity, false-positive control, measurement capability, and stable performance across normal process variation.

Start with the consequence of being wrong

The first mistake in specifying vision accuracy is treating every defect as equally important. A missed label wrinkle, a missed contaminant on a food package seal, and a missed crack in a precision-machined component have very different consequences. The inspection requirement should follow the consequence of the miss, not the marketing specification of the camera or AI model.

Defects generally fall into three practical classes:

  • Critical defects: Conditions that may create a safety, functional, traceability, or containment failure. The system should be optimized for extremely high probability of detection, even when that increases review effort or false rejects.
  • Major defects: Conditions likely to affect fit, performance, reliability, or customer acceptance. Detection must be dependable, but a controlled secondary review process may be acceptable.
  • Minor or cosmetic defects: Conditions that do not affect function but may affect appearance or brand standards. The economic cost of unnecessary rejection often deserves more weight here.

A useful requirement statement does not simply say “detect scratches.” It defines the smallest relevant scratch, its location, contrast, orientation, surface finish, and whether it is critical. A 0.2 mm dark scratch on a matte surface may be easy to detect, while a similar feature on a curved, polished surface may be indistinguishable from a lighting artifact. Without this definition, accuracy claims have little engineering value.

Use more than one accuracy metric

Overall accuracy combines correct accepts and correct rejects into a single percentage. It can be misleading when defects are rare, which is typical in stable manufacturing. A system that passes nearly every part can report excellent overall accuracy while missing a substantial share of actual defects.

For defect detection, the core measures are usually:

Metric What it indicates Why it matters on the line
Recall / sensitivity The share of real defects correctly identified Shows how often the system avoids false accepts or escapes
False-negative rate The share of real defects classified as acceptable Often the most important risk for critical defects
Precision The share of rejected parts that are genuinely defective Shows whether rejects create unnecessary inspection and scrap work
False-positive rate The share of good parts classified as defective Directly affects yield, throughput, and confidence in automation
Repeatability Whether the same part receives the same decision repeatedly Reveals sensitivity to lighting, focus, vibration, and positioning

For a critical condition, a team may set a minimum recall target and then determine whether the associated false-positive rate can be managed through process improvements or a review station. For a cosmetic classification task, the balance may shift: an aggressive threshold that floods the reject bin can cost more than a limited number of borderline features sent for manual judgment.

Precision and recall should also be reviewed by defect type, not only as combined averages. A model can perform well on common stains and still fail on the rare crack or edge chip that actually drove the project. A single blended score hides that weakness.

Translate the defect specification into image requirements

Machine vision cannot reliably classify information that the optical system does not capture. Before evaluating AI software or inspection logic, establish whether the defect is visible at a useful signal level in the acquired image.

Resolution is the starting point, but it is not the whole answer. Engineers often use a rule that the minimum feature should span multiple pixels rather than appear as a one-pixel event. The required pixel coverage depends on defect shape, contrast, motion, lens quality, and the decision being made. Measuring a dimensional edge requires a different image quality margin from detecting a high-contrast missing component.

Consider a defect that is physically large enough to occupy several pixels. It can still be poorly detected when the lens introduces distortion, depth of field is insufficient, or the part height changes between cycles. Likewise, a camera with more pixels does not solve a glare problem. On reflective materials, the defect may be visible only under a particular lighting angle, polarization arrangement, or diffuse illumination geometry.

The inspection design should therefore answer these questions before accuracy targets are finalized:

  • What is the smallest defect that must be detected, and how is its size defined?
  • Can the defect occur anywhere, or only within a controlled region of interest?
  • Does its appearance change with material batch, texture, coating, color, or orientation?
  • Will the feature be darker, brighter, raised, recessed, or visible only through texture disruption?
  • What part movement, vibration, speed, and working-distance variation occur during image capture?
  • Is a pass/fail decision enough, or must the system measure defect length, area, depth proxy, or location?

How Accurate Must Machine Vision Be for Defect Detection?

Accuracy must hold under production variation

A controlled demonstration often uses clean parts, stable lighting, and carefully positioned samples. Production introduces variation that is not noise from the system’s perspective; it is part of the real inspection task. Shift changes, lens contamination, ambient light intrusion, conveyor vibration, part-to-part placement variation, and changes in surface finish can all change the image distribution.

Validation should include the operational range rather than only nominal conditions. For example, inspect parts at expected position limits, with acceptable variation in orientation and height. Evaluate images across realistic exposure conditions, line speeds, and material appearances. Include acceptable parts with marks that resemble defects, such as grain patterns, tooling traces, print variation, dust, and harmless reflections. These are often the source of excessive false rejects.

Part presentation deserves special attention. A highly capable model cannot compensate indefinitely for uncontrolled rotation or a shifting field of view. When accuracy is unstable, it is often more productive to improve fixturing, triggering, lighting enclosure, or part separation before retraining the algorithm. Better image consistency reduces the burden placed on the classifier and produces more interpretable failure analysis.

Set separate acceptance criteria for escapes and false rejects

The most practical specification is usually a decision matrix rather than one accuracy number. It should define the required detection performance for each defect category and the maximum operational burden created by false alarms. It should also state how uncertain results are handled.

For example, a system can have three outcomes instead of two: pass, fail, and review. The review category is useful when a part falls near a confidence threshold, contains an image-quality warning, or presents a feature outside the trained appearance range. This approach is especially valuable when an immediate reject would be expensive but an unreviewed pass would be risky.

A sensible requirement may include:

  • Minimum detection rate for each critical and major defect class.
  • Maximum allowed false-reject rate for known-good production parts.
  • Maximum repeatability variation when the same part is inspected multiple times.
  • Allowed decision time, including image capture, processing, output signal, and part handling.
  • Rules for unclassifiable images, low confidence, camera faults, and missing triggers.
  • A defined human review method, when review is part of the control plan.

Do not set a false-reject target without considering the actual cost of a reject. A false reject may mean a quick visual confirmation, or it may stop a line, trigger destructive testing, and create traceability work. Similarly, do not treat every false negative as identical. The severity of an escaped defect should determine the level of redundancy, review, or process control required.

Build a validation set that reflects the decision boundary

Machine vision performance is only as credible as the samples used to test it. A validation set should contain both defective and non-defective parts that represent the range the system will encounter. It should not be made from nearly identical images collected in a single short run.

Defect examples need careful labeling. If inspection personnel disagree on whether a feature is acceptable, the model is being asked to learn an unstable rule. Establish a visual defect standard with clear examples, boundary conditions, measurement rules where applicable, and escalation criteria for ambiguous parts. The same standard should guide sample labeling, system tuning, manual review, and later audit decisions.

Rare defects create a particular problem. A model may not have enough representative examples to establish reliable behavior. In that situation, teams should avoid assuming that a high aggregate score proves rare-event detection. They may need controlled defect samples, additional process-based sensing, conservative review rules, or a staged deployment that collects more verified images before full reliance on automation.

Keep training, tuning, and final testing separate

Thresholds, model parameters, lighting adjustments, and region-of-interest settings are often refined during development. The final performance assessment should use samples that were not used to make those adjustments. Otherwise, the reported result may describe how well the system recognizes familiar images rather than how well it generalizes to normal production variation.

Testing should also preserve traceability. Store the image, timestamp, inspection result, relevant configuration version, and reference disposition where feasible. When a disputed part appears weeks later, this record makes it possible to determine whether the failure came from optics, classification logic, part handling, labeling, or a change in the manufacturing process.

Watch for signs that the target is poorly defined

Some commissioning problems are not tuning problems. They indicate that the acceptance requirement itself needs revision. Warning signs include a system that detects defects only after operators adjust lighting manually, a high false-reject rate concentrated in one material lot, inconsistent decisions on repeated images, or disagreement between quality staff about the pass/fail label.

Another warning sign is a requirement such as “detect all defects.” No inspection system can satisfy that phrase unless “defect” is bounded by detectable size, appearance, location, and operating conditions. The better formulation is to specify the minimum detectable condition and validate the probability of detection within the defined process window.

Choose the inspection architecture after defining the requirement

Rule-based vision is often effective when geometry, contrast, and part presentation are stable. It can be easier to validate for tasks such as presence checks, dimensional edge measurement, code reading, or well-defined surface features. AI-based classification can be useful where appearances vary and rules become fragile, but it requires representative data management, controlled updates, and monitoring for performance drift.

Some applications need more than a single 2D camera. A shallow dent may require directional lighting, photometric methods, laser profiling, or another sensing approach because its appearance changes too little in a conventional image. The correct question is not whether AI can compensate for missing information; it is whether the sensor and illumination reveal the physical condition reliably enough to support a decision.

Accuracy is therefore a system property. Camera resolution, lens selection, lighting, mechanical presentation, triggering, algorithm design, threshold setting, sample quality, and response workflow all contribute. A credible defect-detection target defines the defect risk first, proves that the image contains usable evidence, and then validates detection and false-reject performance across the conditions the line will actually produce.

Recommended News