Skip to content
Back to work

AI & Automation · Manufacturing

Image Classification System

For a manufacturer, a computer vision system on the production line that flags defects consistently across a full shift and hands uncertain items to an inspector.

Role
Product Owner & ML Solution Architect
Built on
PyTorch, OpenCV, containerized edge deployment
Industry
Manufacturing
Project type
AI & Automation
Primary service
AI & Automation
Scope
Vision model development, threshold and escalation design, edge deployment, monitoring.

Repository not shown as per company policy

  • 0%

    of automated decisions recorded and reviewable

  • 0

    outcomes only: confident classification, or escalate to a person

  • 0

    feedback loop returning every human review to training data

  • 0

    items pass without either model confidence or human sign-off

01Challenge

Manual inspection is consistent for the first hour of a shift.

Visual inspection is demanding, repetitive work, and consistency degrades across a shift in ways nobody intends and nobody records.

  • Defects are rare

    Class imbalance means a model can score extremely well by predicting "fine" every time, which is exactly the wrong behavior.

  • Training data barely exists

    The interesting cases are by definition uncommon, so the dataset starts thin precisely where it matters most.

  • Conditions vary

    Lighting, positioning and product variation change what the same defect looks like from one run to the next.

  • Trust has to be earned

    Inspectors will not defer to a system they cannot interrogate, and they are right not to.

02Approach

Optimize for recall, escalate uncertainty, and record everything.

  • Augment aggressively for scarce classes

    Targeted augmentation for defect classes, reflecting the real variation in lighting and positioning rather than generic transforms.

  • Make the threshold a business decision

    Where confidence stops and escalation starts is a cost trade-off the business owns, not a hyperparameter an engineer picks.

  • Close the loop from day one

    Every inspector review becomes labeled training data, so the system improves from normal operation rather than from a separate labeling project.

  • Deploy at the line

    Containerized at the edge, so a network problem slows nothing down and inspection does not depend on a round trip.

03Outcome

Built to assist inspectors, not to quietly replace their judgment.

Defects are rare, which makes them both hard to catch and hard to gather training data for. A system designed around the average case would be accurate on paper and useless on the line.

  • Design for the rare case

    Optimized for recall on defects, because the cost of a missed defect and the cost of a false alarm are nothing like equal.

  • Uncertainty escalates

    Anything the model is not confident about goes to an inspector rather than being forced into a decision.

  • Every decision is reviewable

    The image, the score and the outcome are recorded, so a disputed call can be examined rather than argued about.

What changed

  • Inspection applies the same standard in the last hour of a shift as in the first.
  • Items the model is unsure about are escalated to an inspector instead of being guessed.
  • Every inspector review returns to the training set, so the model learns most from the cases it found hardest.

How it works

From an image on the line to a recorded decision.

uncertainreviewed casespromotedriftLine imagingControlled lightingand positioningPre-processingNormalization and framingDecision recordOutcome, reviewer,timestampInspector queueEscalated items with theimage and the scoreReviewed case storeLabeled data fromnormal operationRetraining pipelineVersioned, evaluatedbefore promotionModel monitoringDrift and class balanceover timeInference at the edgeContainerized, runs without a network round trip1 ClassifierConvolutional model, recall-weighted2 Confidence thresholdBusiness-owned, not hard-coded3 Escalation routerUncertain items go to a person

One run, step by step.

01 / 05

Inside the system

What it is made of, and what keeps it safe.

What makes it safe to run unattended.

  • 1

    Augmentation strategy

    Targeted augmentation for scarce defect classes, modeled on the real variation seen on the line rather than generic transforms.

    GuardrailAugmentation is validated against held-out real defects, so it does not teach the model an artefact.

  • 2

    Confidence thresholds

    A configurable boundary between confident classification and escalation, owned by the business.

    GuardrailChanging the threshold is a recorded decision, because it directly trades false alarms against missed defects.

  • 3

    Escalation path

    Uncertain items are routed to an inspector with the image and the score, rather than being forced into a class.

    GuardrailThere is no silent third option. An item is either confidently classified or seen by a person.

  • 4

    Decision record

    Every classified item keeps its image, score, outcome and, where applicable, the reviewer.

    GuardrailA disputed decision can be reconstructed months later.

  • 5

    Feedback loop

    Reviewed cases flow back into the training set, with retraining evaluated before any promotion.

    GuardrailA new model has to beat the incumbent on defect recall before it replaces it.

Built with

What it runs on.

Modeling
PyTorchOpenCVPython
Serving
Flask REST inference APIDockerEdge deployment
Data
Reviewed case storeVersioned training sets
Operations
Retraining pipelineModel monitoringDecision audit trail

Roadmap

Where the capability goes next.

Model

  • Defect localization, so the inspector sees where as well as whether
  • Per-defect-class thresholds rather than one global boundary

Operations

  • Active learning, prioritizing the most informative cases for review
  • Automated evaluation and promotion gates in the retraining pipeline

Insight

  • Defect trend reporting back to production, to address causes not just symptoms
  • Shift and line comparison to surface systematic differences

Consistent inspection across a full shift, with a person on every uncertain call.

More work

Related projects.

AI & AutomationHealthcare

Autonomous AI Log Monitoring & Observability Platform

For a healthcare media and clinician engagement platform, a five-agent system that reads a production error, writes the fix and opens a reviewed pull request, with an engineer still deciding what ships.

AI & AutomationHealthcare

Drug Competitor Identification

A brand-intelligence tool that asks a language model who a drug competes with, then checks the answer against regulatory reference data before anyone is asked to trust it.

AI & AutomationHealthcare

AI Agents Platform

Seven agents that turn one upload into a recorded, print-ready batch of personalized posters, with a reviewer approving anything that carries commercial risk.

Discuss a similar project.

If something here is close to what you need, tell us about your situation and we will walk you through how we would approach it.

Intelligence → Innovation → Automation → Growth