ProcessCustom AI Model Development
Building a training dataset: curation, labeling and leakage control
A custom model is only as good as the examples it learns from and the test set that judges it. This page sets out the working sequence for turning raw proprietary records into a labeled, versioned dataset: scoping against a baseline, writing the label taxonomy, measuring annotator agreement, adjudicating disagreements, scaling labeling, splitting without leakage and documenting the result so it can be audited later.
On this page
- The decision and baseline that scope a labeling effort
- Seven stages from raw records to a versioned dataset
- What each stage of the annotation pipeline produces
- Choosing an agreement statistic for your annotation design
- Scaling labeling: experts, vendors or model-assisted selection
- Leakage and coverage checks before the splits are frozen
- A hypothetical survey dataset that leaked by vessel
- Versioning and datasheets that keep a dataset auditable
- Questions and answers
- Sources
The decision and baseline that scope a labeling effort
Labeling is usually the most expensive part of a custom model project, so scope it by the decision the model will support rather than by how much data happens to exist. Write down what the model's output will change: which queue a document goes to, which defect gets escalated, which record a reviewer checks first. That decision sets which distinctions the labels must capture.
Before annotation starts, score a simple baseline on a small hand-labeled sample: a rules-based filter, a keyword classifier or a general model with a careful prompt. ColdAI's approach to custom AI model development begins with that baseline rather than a training run1, because it tells you how much improvement the dataset has to buy.
The same sample exposes practical problems early: missing fields, documents that need text recognition, categories experts cannot tell apart and records you may not have the right to use. Rights have their own page on training data copyright and data protection.
Seven stages from raw records to a versioned dataset
- Scope and sample
Draw a representative sample and score the baseline to beat.
- Taxonomy and guidelines
Define each label with inclusion rules, exclusions and borderline cases.
- Pilot labeling round
Several annotators label the same batch independently.
- Agreement and adjudication
Measure agreement, resolve disputes and revise the guidelines.
- Scale labeling
Label at volume with experts, a vendor or model-assisted selection.
- Split and audit
Freeze the splits after leakage and coverage checks.
- Version and document
Record the version, its datasheet and the experiments using it.
What each stage of the annotation pipeline produces
Sample against the decision
Pull records from every source and period the model will see in use, not only the cleanest system. Hand-label a small sample and score the baseline on it.
Write the label taxonomy
Keep labels mutually exclusive where the decision demands it, and add an explicit "cannot tell" option so annotators are not forced to guess. Each label needs a definition, the evidence that justifies it and a borderline example.
Run a pilot round
Give the same batch to several annotators working independently, including someone who did not write the guidelines. Their questions show where the guidelines are ambiguous.
Measure agreement and adjudicate
Compute a chance-corrected agreement statistic per label, then review each disagreement with the guideline author: a genuinely ambiguous case, a guideline gap or an annotator error. Revise the guidelines before the next round.
Scale with quality controls
Once agreement is stable, label at volume. Seed batches with gold items whose answers are known, keep a share double-labeled to watch agreement over time, and route disputes to the adjudicator.
Split, audit and release
Assign related records to splits by entity or time, remove near-duplicates across splits and confirm rare classes reach the test set. Give the release an immutable version and record the guideline version behind every label.
Choosing an agreement statistic for your annotation design
| Criterion | Cohen's kappa | Fleiss' kappa | Krippendorff's alpha |
|---|---|---|---|
| Annotators per item | Exactly two, the same pair throughout | A fixed number per item; the raters may differ between items | Any number, and it may vary from item to item |
| Missing labels | Not handled: both raters label every item | Not handled: every item needs the same number of ratings | Handled: partly rated items still count |
| Label scales | Nominal; a weighted variant handles ordered labels | Nominal | Nominal, ordinal, interval or ratio, through its distance function |
| Best suited to | A pilot with two expert annotators | Vendor batches with fixed redundancy | Mixed teams, partial overlap and graded severity labels |
All three correct for agreement expected by chance, unlike raw percentage agreement. Report agreement per label as well as overall: a healthy average can hide one label nobody applies consistently.
Scaling labeling: experts, vendors or model-assisted selection
- If
Labels need professional judgment, such as clinical, legal or engineering assessment, or the data cannot leave your environment.
ThenLabel in-house with domain experts, on tooling inside your own infrastructure, and design the workflow around scarce expert time.
Vendor annotators cannot supply judgment they are not qualified to make, and moving restricted data adds legal and security work.
- If
Volume is high, the distinctions are learnable from the guidelines and the data may be shared under contract.
ThenUse a labeling vendor with gold items, double-labeling on a sample and an adjudicator on your side.
The quality controls, not the vendor's headcount, decide whether the labels are usable.
- If
A first model exists and most unlabeled records look alike.
ThenUse active learning: let the model nominate the records it is least certain about, or that differ most from what it has seen, and label those first.
Labeling records the model already handles well adds little information per hour.
- If
A large language model could pre-label the records.
ThenTreat its output as suggestions annotators confirm or correct, and measure its agreement with experts like any other annotator.
Unreviewed model labels teach the new model the old model's mistakes.
Leakage and coverage checks before the splits are frozen
Leakage makes a model look better in testing than it will be in use. A review of machine-learning-based science found leakage across many research fields and sorted it into recurring types2; most of them are caught by the checks below.
A hypothetical survey dataset that leaked by vessel
Versioning and datasheets that keep a dataset auditable
A result is reproducible only if you can say exactly which records, labels and guideline version produced it. Store each dataset release under an immutable identifier, such as a content hash or a tagged release in a data versioning tool, and log that identifier in the experiment tracker beside the code commit and configuration.
Document each release with a datasheet. The "Datasheets for Datasets" proposal by Gebru and colleagues poses questions about a dataset's motivation, composition, collection, preprocessing and labeling, intended uses, distribution and maintenance3. Answering them takes an afternoon while the dataset is fresh, and is far harder a year later.
If the model may become part of a high-risk system under the EU AI Act (Regulation (EU) 2024/1689), Article 10 sets data governance requirements for training, validation and testing data, including attention to relevance and representativeness and examination for possible biases4. A versioned dataset with a datasheet and a leakage report is the practical evidence for meeting them.
Questions and answers
How much labeled data does a custom model need?
There is no general number. It depends on how many classes the model must separate, how subtle the distinctions are and whether you start from a pretrained model. Draw a learning curve instead: train on growing fractions of the labeled pool and watch validation performance. If it is still climbing, more labels will help; if it has flattened, better labels or a different approach will.
Can synthetic data replace real labeled examples?
Synthetic data can add variety, balance rare classes or stand in for sensitive records, but it should not be the only evidence. Generated examples carry the assumptions of whatever produced them, so a model trained mostly on them can fail on real inputs. Keep validation and test sets made of real, adjudicated records, and check a generator's terms of use before training on its outputs.
What should happen when annotation guidelines change mid-project?
Record the guideline version against every label. When a definition changes, find the labels it affects, re-label those items or at least the test set, and publish the result as a new dataset version. Mixing labels made under different definitions quietly teaches the model contradictory rules and makes comparisons between experiments meaningless.
When should a production model's dataset be refreshed?
Refresh it when monitoring shows inputs or outcomes drifting away from the training data, when the decision itself changes, or when reviewers keep correcting the same kind of error. Sample recent production records, label them under the current guidelines and add them as a new version, while keeping an older test set so you can see whether familiar cases still work.
Sources
- Custom AI Model Development: approach, stages and deliverables — ColdAI
- Leakage and the Reproducibility Crisis in ML-based Science (Kapoor and Narayanan) — arXiv · checked 10 October 2026
- Datasheets for Datasets (Gebru et al.) — arXiv · checked 10 October 2026
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026