For teams building computer-vision inspection

Generate the training set. Keep real data for the exam.

Realistic synthetic images are now easy to make. Labels that can be traced, checked and proven on real images are not. Synthetic Data Studio (SDS), by Aicadium, generates labelled training images for assembly inspection from a single correct example. That frees your scarce real data for the job it does best, which is measuring what the model has learned.

The numbers behind the problem

Most inspection models are limited by their data, and a synthetic image can hide it

63%
of organisations do not have, or are unsure whether they have, the right data management practices for AI

For inspection teams, that starts with how images are prepared and labelled.

60%
of AI projects unsupported by AI-ready data will be abandoned through 2026, Gartner predicted in 2025

AI-ready data means data that fits the use case. For inspection, that means labelled images covering the parts, layouts and errors the line will see.

0.25 → 0.89
mAP on real images (mean average precision, a standard detector accuracy measure from 0 to 1), before and after a sim-to-real adaptation step

A detector trained on 550 synthetic images was near-perfect in simulation and scored 0.2516 mAP on real images. An adaptation step using 50 paired real images, alongside 500 synthetic ones, lifted it to 0.8853.

What we believe

Three convictions shaping how we think about synthetic data for inspection.

CONVICTION I

The label is the hard part.

Generating a convincing image of a tray or a kit is no longer difficult. Knowing exactly what is in every frame, and being able to show how that was checked, is where inspection projects succeed or stall. A label that cannot be traced is a risk that travels into every model trained on it.

CONVICTION II

Real data is the exam. Synthetic data is the curriculum.

Real production images are scarce, slow to collect and often restricted. They are most valuable as a held-back test of whether a model works on the line. Synthetic data can carry the training load, supplying the volume, rare defects and variation that real production produces too slowly. A small, carefully chosen set of real images can help training as well. What matters is keeping a separate real set that never enters training.

CONVICTION III

Software does the heavy lifting. A person on your team makes the final check.

A training dataset can contain thousands of images, far too many to label by hand. So software draws the boxes and checks its own work. It gets most of them right, but not all. That is why a dataset should only count as finished once a person has looked through the results. For datasets Aicadium delivers, that review covers every annotated frame.

How it works

From one correct example to a labelled training set

Synthetic Data Studio starts from something an inspection team already has, which is a definition of a correctly assembled unit.

It plans every scene before generating it, so the class labels come from the plan rather than from guesswork. Each image arrives with its annotations, in whole-unit or per-compartment detail.

It runs entirely in the browser, with no server to set up, and exports in the formats your training pipeline already reads. Saving projects to a folder on your computer needs Chrome or Edge.

1
Golden sample
One photograph of a correctly assembled unit, a written description, or both.
2
Confirmed reading
An AI model describes what it sees. You confirm or correct it before anything is generated.
3
Scene plan
A dataset brief and scene-by-scene plan set out what each frame should contain.
4
Generation
Images are generated from the plan, and each one arrives with its annotations.
5
Labels and checks
Class labels come from the plan. General-purpose AI detectors tighten the boxes, and an automated pass checks the labels against each image.
6
Export
COCO JSON, YOLO and Pascal VOC, ready for the training pipeline you already use.
Five questions for any synthetic dataset

Do not just take the dataset. Ask how it was made.

Any synthetic data pipeline can hand over a folder of labelled images. These five questions show whether the labels will hold up when the model meets the real line.

01
Specification
“What was each image meant to contain?”

Every frame should trace back to a written plan that a person approved. SDS writes a dataset brief from your confirmed golden sample before generating anything.

02
Class labels
“Where did the class label come from?”

Labels taken from the scene plan can be traced back to it. Labels worked out afterwards by software looking at the image are only as reliable as that software. In SDS, class labels come from the plan.

03
Bounding boxes
“What drew the box, and where does it struggle?”

When software draws the boxes, it can fail on some object types. Ask for the detector’s known weak spots and how they are caught.

04
Review
“Who checked the output?”

An automated check is useful. A person reviewing the result is what catches the errors automation misses. Whoever generates the dataset should do that review before training. For datasets Aicadium delivers, a person reviews every annotated frame.

05
Evaluation
“Was performance measured on real images?”

Results on generated images say little about results on the line. Hold back a set of real frames and measure on those every time.

SDS in action

See a dataset take shape. From one photo to a labelled set.

From a single correct unit to a training set you can check.

  • 1Photograph one correct unit.A meal tray, a tool kit or a parts case, assembled exactly as it should be.
  • 2Confirm what the model read.Correct any misread item now, before it reaches every image in the dataset.
  • 3Review the brief and scene plan.See what each frame will contain, including which units are incorrect and why.
  • 4Generate and export.Images arrive with their annotations, in whole-unit or per-compartment detail, ready to export in standard formats.

Choose the level of detail your model needs to predict.

SDS offers two levels of label detail, so the labels match what your model needs to report.

  • 1Whole unit.One box per assembled unit, classed as correct or incorrect.
  • 2Per compartment.Each compartment boxed with its own error type, from complete and missing to substitution, arrangement, extra, portion, wrong_dish and occluded.
  • 3One source, three formats.COCO JSON, YOLO and Pascal VOC, all produced from the same labels and packaged with the images.
Questions buyers ask

What inspection teams ask before they commit

“

Will a model trained on synthetic images work with our real cameras?

Q
Not on its own, and independent research shows why. In that study, a small, carefully chosen set of real images used in training closed most of the gap. Keep a separate set of real photos from your own cameras for testing before you rely on the model.
“

Can it tell normal variation from a real error?

Q
Per-compartment labelling uses eight error types, including missing, substitution, arrangement, extra and portion. A model trained on those distinctions can report them separately.
“

Do we need new infrastructure?

Q
SDS runs in the browser, and saving projects to a folder on your computer needs Chrome or Edge. It exports COCO JSON, YOLO and Pascal VOC, formats existing training pipelines already read. Model calls go directly from your browser to each vendor, with no SDS server in the path.
Three predictions

What we expect to be obvious by 2028 and uncomfortable to admit today.

Within 12 months

Image generation becomes a default platform feature.

Major platforms already package synthetic defect generation, and manufacturers such as Wistron and Inventec report running it on inspection lines. The question buyers ask will move from “can you generate this?” to “how were the labels checked?”

Within 18 months

Real-image evaluation becomes a shortlisting requirement.

Vendor-reported gains are easy to publish and, with no standard quality metric across vendors, hard to compare. Buyers will ask to see results measured on their own real images before a pilot is approved.

Within 24 months

Every dataset comes with a record of how its labels were made.

As quality standards for synthetic data take shape, a dataset will be expected to carry a record of how each label was produced and reviewed. Teams that keep that record now will not need to rebuild it later.

Why this matters now

Inspection datasets have usually started with capture and labelling. They can start from a specification.

Without a specification-first pipeline
Capture first
training data follows whatever production happened to produce
  • Photograph every item, layout and error in live production
  • Wait for rare errors to happen often enough to capture
  • Label every box by hand, image by image
  • Rebuild the dataset when the specification changes
  • Measure on whatever real images are left over
→
With Synthetic Data Studio
Specification first
training data follows the standard your quality team signs off
  • Start from one correct unit and a confirmed reading
  • Plan which errors each scene contains
  • Receive class labels derived from the scene plan
  • Write a new brief when the specification changes
  • Keep a clean set of real images for evaluation

Ideas like this shape how we work at Aicadium.

Get updates about our latest AI Transformation projects.