GuideChemicals

Using machine learning to reach a formulation target in fewer lab experiments

Formulation data is small, expensive and full of interacting ingredients, so the useful question is not whether a model can predict every property but whether it can choose better next experiments. This guide covers preparing ELN and LIMS records, representing a formulation, choosing between design of experiments, Bayesian optimization and active learning, handling cost and restricted substances, and running a loop in which the chemist keeps the final say.

Reviewed 7 min read

On this page
  1. Why formulation is a promising and stubborn machine learning problem
  2. Preparing ELN and LIMS records before any modeling
  3. How to describe a formulation to a model
  4. Design of experiments, Bayesian optimization and active learning compared
  5. Multiple objectives, raw-material cost and substances of concern
  6. The suggest, test and update loop
  7. Running a formulation campaign with the chemist in control
  8. Reformulating a coating to remove a restricted substance
  9. Measuring success by experiments to target, and protecting formulation IP
  10. Questions and answers
  11. Sources

Why formulation is a promising and stubborn machine learning problem

Formulation looks made for machine learning: a product is a recipe of ingredients and process steps, and performance is measured on defined tests. In practice the problem resists. A coating or adhesive may draw on dozens of candidate raw materials whose effects interact, several properties must be met at once, and each experiment costs days of lab and test time. Most teams hold tens or hundreds of relevant experiments, not the volumes deep learning expects.

That changes what a model is for. Its job is to choose the next few experiments so the team reaches the specification in fewer rounds at the bench. Methods built for small, expensive datasets do that well, provided the historical data is honest about what failed.

Preparing ELN and LIMS records before any modeling

0 of 8 checked

How to describe a formulation to a model

Composition vector
The fraction of each ingredient in the recipe. Transparent, but every new raw material adds a dimension the model has never seen.
Ingredient descriptors
Properties of each raw material, such as molecular weight, functionality, glass transition temperature or solubility parameters, so the model can generalize to a new ingredient with similar chemistry.
Process parameters
Conditions of making and applying the product, from mixing energy to cure temperature, held alongside composition.
Mixture constraint
Fractions must sum to one, which makes ingredients interdependent; mixture designs and suitable model forms account for this instead of treating each fraction as free.
Functional grouping
Pooling ingredients by role, such as binder, crosslinker, filler or surfactant, so a substitute inherits what is known about its group.

Design of experiments, Bayesian optimization and active learning compared

CriterionClassical DoEBayesian optimizationActive learning
Best suited toA few factors and a need to understand their effectsReaching a target with as few experiments as possibleA model that predicts well across the whole design space
Starting dataNone; the design is planned up frontA small designed set or relevant historyA small seed set
How experiments are chosenFixed in advance, such as factorial or mixture designsAn acquisition function balances promising and uncertain regionsWhere the model is most uncertain or most informative
Constraints and multiple objectivesBuilt into the design, awkward as factors growSupported through constrained and multi-objective acquisitionPossible, though the aim is coverage rather than a target
What you end withAn interpretable response-surface modelA candidate meeting the target and a working surrogateA surrogate reliable enough to screen future requests
Main weaknessExperiment count grows quickly with factorsCan exploit model errors without sensible boundsSpends experiments on regions nobody needs

Many campaigns combine them: a space-filling or mixture design to start, then Bayesian optimization once there is something to learn from.

Multiple objectives, raw-material cost and substances of concern

Real targets are rarely one number. A coating must meet adhesion, hardness, cure time and appearance at an acceptable cost, and these pull against each other. Multi-objective methods return a Pareto front, the set of formulations where no property improves without another getting worse, so the team chooses the trade-off openly instead of fixing weights before seeing the options.

Regulatory limits are hard constraints, not objectives. Exclude raw materials containing substances on the REACH Candidate List or restricted under Annex XVII before the model ever proposes them12. For a coating sold as a mixture, a Candidate List substance affects its safety data sheet; once it sits in a customer's article above 0.1% weight by weight, REACH Article 33 communication and SCIP notification duties follow23. The CLP hazard classes for endocrine disruptors and persistent substances may also change which ingredients customers accept4.

The suggest, test and update loop

01Target andconstraints02Model suggests abatch03Chemist screens04Lab or robot tests05Results captured06Model updated
  1. Target and constraints

    Specification, cost ceiling and excluded substances, agreed before the first suggestion.

  2. Model suggests a batch

    A handful of candidates chosen by the acquisition strategy.

  3. Chemist screens

    Removes impractical, unsafe or unstable candidates and records why.

  4. Lab or robot tests

    Manual benches or an automated platform prepare and test the batch.

  5. Results captured

    Stored with methods, replicates and failures.

  6. Model updated

    Refit on all data so the next suggestions reflect what was learned.

Conceptual loop for one formulation campaign, repeated until the target is met or the stopping rule ends it. It does not show measured results.

Running a formulation campaign with the chemist in control

  1. Agree the target and the stopping rule

    Write down the specification, cost ceiling, excluded substances and when the campaign stops, such as a test budget or no improvement over several rounds. ColdAI's research practice fixes how results will be evaluated before experiments run, for the same reason5.

    Output
    A one-page campaign brief
    Owner
    R&D lead
  2. Assemble and clean the history

    Apply the checklist above to past experiments in the product family and decide which old results remain comparable.

    Owner
    Formulation chemist with a data scientist
  3. Seed the model where history is thin

    Run a small space-filling or mixture design over the region of interest before the optimizer steers.

  4. Iterate in small batches

    Request a few suggestions at a time, matched to lab or robot capacity, and refit after every batch. The chemist can veto any suggestion; the reason is logged so both the model and the team learn from it.

    Owner
    Formulation chemist
  5. Confirm and hand over

    Replicate the best candidates, check stability and scale-up behavior, and pass the formulation to regulatory review with its full experimental trail.

    Output
    A confirmed formulation and its evidence

Reformulating a coating to remove a restricted substance

Measuring success by experiments to target, and protecting formulation IP

Judge the program by how many experiments and how much elapsed time it took to reach the specification, compared with similar campaigns run the traditional way. Model accuracy matters only insofar as it moves that number. A simple log per campaign of rounds, experiments, cost and outcome is enough.

Formulations are often a company's most valuable trade secrets, so where models run and who owns them matter. Keep data and models in your own environment or a private deployment, and settle ownership of models, code and results before work starts. ColdAI's research practice settles publication and IP terms at the outset5, and our custom model development work is designed to maintain data sovereignty6.

Questions and answers

How much historical data do we need before machine learning helps?

Less than most teams expect, because Bayesian optimization is designed for small datasets. Relevance and honesty matter more than volume: a few dozen well-recorded experiments in the same product family, including failures, beat thousands of loosely related records. Where history is thin, start with a small designed experiment and let the optimizer take over once it has something to learn from.

Can large language models design formulations?

Not as the optimizer. Language models help around the loop: searching ELN notes and literature, extracting conditions from old reports into structured data and drafting experiment plans for a chemist to check. Property predictions should come from models trained on your measured data and confirmed by experiment, because a fluent answer about a mixture's performance is not evidence.

Does machine learning for formulation require a robotic lab?

No. Most campaigns start on manual benches, with batch sizes matched to what the lab can run each week. Automated platforms make the loop faster and the data more consistent, and batch Bayesian optimization suits them, but the method and data discipline are the same. Prove the approach manually before investing in automation.

Who owns a model trained on our formulation data?

Settle it in the contract before any data is shared. The questions are who owns the trained model and code, whether the provider may reuse anything learned, where data is stored and how it is deleted afterwards. ColdAI's research engagements settle IP terms at the outset5; whoever you work with, avoid terms that let a vendor train shared models on your recipes.

Sources

  1. Candidate List of substances of very high concern for authorisation — European Chemicals Agency · checked 10 October 2026
  2. Regulation (EC) No 1907/2006 (REACH) — EUR-Lex · checked 10 October 2026
  3. SCIP database for substances of concern in articles — European Chemicals Agency · checked 10 October 2026
  4. Commission Delegated Regulation (EU) 2023/707 on hazard classes and criteria — EUR-Lex · checked 10 October 2026
  5. Frontier R&D: research questions, evaluation and IP terms — ColdAI
  6. Custom AI Model Development — ColdAI

More in Chemicals

Back to Chemicals

Next step

Test machine learning on one formulation campaign first

Tell us the product family, the target specification and how your ELN and LIMS records are kept. We will assess whether the history supports a campaign and propose how to run the first one.

Scope a formulation pilot