ProcessCustom AI Model Development

Building a training dataset: curation, labeling and leakage control

A custom model is only as good as the examples it learns from and the test set that judges it. This page sets out the working sequence for turning raw proprietary records into a labeled, versioned dataset: scoping against a baseline, writing the label taxonomy, measuring annotator agreement, adjudicating disagreements, scaling labeling, splitting without leakage and documenting the result so it can be audited later.

Reviewed 8 min read

On this page
  1. The decision and baseline that scope a labeling effort
  2. Seven stages from raw records to a versioned dataset
  3. What each stage of the annotation pipeline produces
  4. Choosing an agreement statistic for your annotation design
  5. Scaling labeling: experts, vendors or model-assisted selection
  6. Leakage and coverage checks before the splits are frozen
  7. A hypothetical survey dataset that leaked by vessel
  8. Versioning and datasheets that keep a dataset auditable
  9. Questions and answers
  10. Sources

The decision and baseline that scope a labeling effort

Labeling is usually the most expensive part of a custom model project, so scope it by the decision the model will support rather than by how much data happens to exist. Write down what the model's output will change: which queue a document goes to, which defect gets escalated, which record a reviewer checks first. That decision sets which distinctions the labels must capture.

Before annotation starts, score a simple baseline on a small hand-labeled sample: a rules-based filter, a keyword classifier or a general model with a careful prompt. ColdAI's approach to custom AI model development begins with that baseline rather than a training run1, because it tells you how much improvement the dataset has to buy.

The same sample exposes practical problems early: missing fields, documents that need text recognition, categories experts cannot tell apart and records you may not have the right to use. Rights have their own page on training data copyright and data protection.

Seven stages from raw records to a versioned dataset

01Scope and sample02Taxonomy andguidelines03Pilot labeling round04Agreement andadjudication05Scale labeling06Split and audit07Version and document
  1. Scope and sample

    Draw a representative sample and score the baseline to beat.

  2. Taxonomy and guidelines

    Define each label with inclusion rules, exclusions and borderline cases.

  3. Pilot labeling round

    Several annotators label the same batch independently.

  4. Agreement and adjudication

    Measure agreement, resolve disputes and revise the guidelines.

  5. Scale labeling

    Label at volume with experts, a vendor or model-assisted selection.

  6. Split and audit

    Freeze the splits after leakage and coverage checks.

  7. Version and document

    Record the version, its datasheet and the experiments using it.

Conceptual sequence for building a labeled dataset. The pilot and adjudication stages usually repeat before labeling scales; it is not a fixed timeline.

What each stage of the annotation pipeline produces

  1. Sample against the decision

    Pull records from every source and period the model will see in use, not only the cleanest system. Hand-label a small sample and score the baseline on it.

    Output
    Baseline score and sample notes
    Owner
    ML lead with a domain expert
  2. Write the label taxonomy

    Keep labels mutually exclusive where the decision demands it, and add an explicit "cannot tell" option so annotators are not forced to guess. Each label needs a definition, the evidence that justifies it and a borderline example.

    Output
    Taxonomy and annotation guidelines, first version
    Owner
    Domain expert
  3. Run a pilot round

    Give the same batch to several annotators working independently, including someone who did not write the guidelines. Their questions show where the guidelines are ambiguous.

    Output
    Pilot labels and question log
    Owner
    Annotation lead
  4. Measure agreement and adjudicate

    Compute a chance-corrected agreement statistic per label, then review each disagreement with the guideline author: a genuinely ambiguous case, a guideline gap or an annotator error. Revise the guidelines before the next round.

    Output
    Agreement report and revised guidelines
    Owner
    Guideline author and adjudicator
  5. Scale with quality controls

    Once agreement is stable, label at volume. Seed batches with gold items whose answers are known, keep a share double-labeled to watch agreement over time, and route disputes to the adjudicator.

    Output
    Labeled pool with an audit trail
    Owner
    Annotation lead
  6. Split, audit and release

    Assign related records to splits by entity or time, remove near-duplicates across splits and confirm rare classes reach the test set. Give the release an immutable version and record the guideline version behind every label.

    Output
    Frozen splits, leakage report and dataset release
    Owner
    ML engineer and data owner

Choosing an agreement statistic for your annotation design

CriterionCohen's kappaFleiss' kappaKrippendorff's alpha
Annotators per itemExactly two, the same pair throughoutA fixed number per item; the raters may differ between itemsAny number, and it may vary from item to item
Missing labelsNot handled: both raters label every itemNot handled: every item needs the same number of ratingsHandled: partly rated items still count
Label scalesNominal; a weighted variant handles ordered labelsNominalNominal, ordinal, interval or ratio, through its distance function
Best suited toA pilot with two expert annotatorsVendor batches with fixed redundancyMixed teams, partial overlap and graded severity labels

All three correct for agreement expected by chance, unlike raw percentage agreement. Report agreement per label as well as overall: a healthy average can hide one label nobody applies consistently.

Scaling labeling: experts, vendors or model-assisted selection

  • If

    Labels need professional judgment, such as clinical, legal or engineering assessment, or the data cannot leave your environment.

    Then

    Label in-house with domain experts, on tooling inside your own infrastructure, and design the workflow around scarce expert time.

    Vendor annotators cannot supply judgment they are not qualified to make, and moving restricted data adds legal and security work.

  • If

    Volume is high, the distinctions are learnable from the guidelines and the data may be shared under contract.

    Then

    Use a labeling vendor with gold items, double-labeling on a sample and an adjudicator on your side.

    The quality controls, not the vendor's headcount, decide whether the labels are usable.

  • If

    A first model exists and most unlabeled records look alike.

    Then

    Use active learning: let the model nominate the records it is least certain about, or that differ most from what it has seen, and label those first.

    Labeling records the model already handles well adds little information per hour.

  • If

    A large language model could pre-label the records.

    Then

    Treat its output as suggestions annotators confirm or correct, and measure its agreement with experts like any other annotator.

    Unreviewed model labels teach the new model the old model's mistakes.

Leakage and coverage checks before the splits are frozen

Leakage makes a model look better in testing than it will be in use. A review of machine-learning-based science found leakage across many research fields and sorted it into recurring types2; most of them are caught by the checks below.

0 of 8 checked

A hypothetical survey dataset that leaked by vessel

Versioning and datasheets that keep a dataset auditable

A result is reproducible only if you can say exactly which records, labels and guideline version produced it. Store each dataset release under an immutable identifier, such as a content hash or a tagged release in a data versioning tool, and log that identifier in the experiment tracker beside the code commit and configuration.

Document each release with a datasheet. The "Datasheets for Datasets" proposal by Gebru and colleagues poses questions about a dataset's motivation, composition, collection, preprocessing and labeling, intended uses, distribution and maintenance3. Answering them takes an afternoon while the dataset is fresh, and is far harder a year later.

If the model may become part of a high-risk system under the EU AI Act (Regulation (EU) 2024/1689), Article 10 sets data governance requirements for training, validation and testing data, including attention to relevance and representativeness and examination for possible biases4. A versioned dataset with a datasheet and a leakage report is the practical evidence for meeting them.

Questions and answers

How much labeled data does a custom model need?

There is no general number. It depends on how many classes the model must separate, how subtle the distinctions are and whether you start from a pretrained model. Draw a learning curve instead: train on growing fractions of the labeled pool and watch validation performance. If it is still climbing, more labels will help; if it has flattened, better labels or a different approach will.

Can synthetic data replace real labeled examples?

Synthetic data can add variety, balance rare classes or stand in for sensitive records, but it should not be the only evidence. Generated examples carry the assumptions of whatever produced them, so a model trained mostly on them can fail on real inputs. Keep validation and test sets made of real, adjudicated records, and check a generator's terms of use before training on its outputs.

What should happen when annotation guidelines change mid-project?

Record the guideline version against every label. When a definition changes, find the labels it affects, re-label those items or at least the test set, and publish the result as a new dataset version. Mixing labels made under different definitions quietly teaches the model contradictory rules and makes comparisons between experiments meaningless.

When should a production model's dataset be refreshed?

Refresh it when monitoring shows inputs or outcomes drifting away from the training data, when the decision itself changes, or when reviewers keep correcting the same kind of error. Sample recent production records, label them under the current guidelines and add them as a new version, while keeping an older test set so you can see whether familiar cases still work.

Sources

  1. Custom AI Model Development: approach, stages and deliverables — ColdAI
  2. Leakage and the Reproducibility Crisis in ML-based Science (Kapoor and Narayanan) — arXiv · checked 10 October 2026
  3. Datasheets for Datasets (Gebru et al.) — arXiv · checked 10 October 2026
  4. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026

More in Custom AI Model Development

Back to Custom AI Model Development

Next step

Send us a labeled sample and your draft label taxonomy

Share a small labeled sample, the taxonomy as it stands and the decision the model will support. We will reply with the baseline we would score first and where the labeling plan is most likely to need work.

Discuss your dataset plan