ArchitectureLife Sciences

A reference architecture for AI regulatory writing that traces every number

AI regulatory writing works when drafting is the last step of a data pipeline rather than a chat window. In this architecture, approved facts live as reusable components, each drafted sentence points to a specific table cell or protocol section, numbers are checked automatically, and medical writers, statisticians and regulatory leads remain the approvers. The result is a clinical study report or CTD summary draft in which no figure appears without a source.

Reviewed 7 min read

On this page
  1. Why submission writing is a data problem before it is a writing problem
  2. Six layers from source systems to the eCTD sequence
  3. Building blocks of a structured content layer for submissions
  4. How one efficacy sentence travels from table cell to approved text
  5. Extend the authoring platform, add a grounded drafting layer, or stay manual
  6. A mismatched figure caught before medical review
  7. Failure modes in AI-drafted submission content, and the control for each
  8. Publishing in eCTD v4, drafting controls and where ColdAI fits
  9. Questions and answers
  10. Sources

Why submission writing is a data problem before it is a writing problem

One trial's results get written up many times. The protocol and statistical analysis plan define the endpoints; the clinical study report presents results in the structure set out in ICH E31; the clinical overview and clinical summary in Module 2 of the Common Technical Document restate them for reviewers2; labeling and lay summaries restate them again for other readers. Each restatement is a chance for a number, a population or an endpoint name to drift.

Generative models speed that drift up when they draft from loose context. The fix is architectural: approved facts live once in a structured repository, drafts are assembled from them, and anything a model writes must point back to its source. EMA's reflection paper expects close human supervision and quality review of model-generated product information text before submission3, and the same discipline suits every submission document.

Six layers from source systems to the eCTD sequence

eCTD publishing01Review workflow02Drafting agents03Retrieval and checks04Structured content store05Source systems06
  1. eCTD publishing

    Approved documents assembled and validated in the existing publishing tool.

  2. Review workflow

    Medical writer, statistician and regulatory lead approve, with redlines and version history.

  3. Drafting agents

    Compose sections from components and retrieved sources, binding each sentence to its origin.

  4. Retrieval and checks

    Find the right table cell or protocol section and verify every number against it.

  5. Structured content store

    Approved, reusable components with owners, versions and controlled terminology.

  6. Source systems

    Protocol, analysis plan, tables, listings and figures, CTMS and safety data.

Conceptual layering of an AI-assisted writing stack, top to bottom. It shows responsibilities, not a specific product or a deployed system.

Building blocks of a structured content layer for submissions

Content component
A reusable, approved unit such as an endpoint definition, a study design paragraph or a population description, stored once with an owner and a version.
Source binding
A link from a sentence or value in a draft to the exact table cell, listing or protocol section it came from, carried with the document through review.
Tables, listings and figures (TLFs)
Statistical outputs produced under the analysis plan. Here they are the only permitted source for results numbers.
Numeric consistency check
An automated comparison of every number in a draft with its bound source, run on each draft and again after any TLF rerun.
Audience variant
The same component written for another reader, such as the summary for laypersons that the EU Clinical Trials Regulation requires alongside the results summary4.

How one efficacy sentence travels from table cell to approved text

Each number is checked twice: when it is retrieved, and again after drafting against the current table version.

Draft efficacy sectionFind primary endpointQuery bound tableCell values, versionGrounded sourcesDraft with bindingsRe-read each valuePass or mismatch listEdit and approve01Medical writer02Drafting agent03Retrieval layer04TLF store05Consistencycheck
  1. Medical writer

    Requests the section and approves the final text.

  2. Drafting agent

    Writes from components and retrieved sources only.

  3. Retrieval layer

    Resolves which table and cell answer each request.

  4. TLF store

    Holds versioned statistical outputs.

  5. Consistency check

    Re-reads every bound value before review.

  1. Medical writer to Drafting agentDraft efficacy section
  2. Drafting agent to Retrieval layerFind primary endpoint
  3. Retrieval layer to TLF storeQuery bound table
  4. TLF store to Retrieval layerCell values, version
  5. Retrieval layer to Drafting agentGrounded sources
  6. Drafting agent to Consistency checkDraft with bindings
  7. Consistency check to TLF storeRe-read each value
  8. Consistency check to Medical writerPass or mismatch list
  9. Medical writer to Medical writerEdit and approve
Conceptual message order for drafting one results paragraph. Actors are roles in the architecture, not named products.

Extend the authoring platform, add a grounded drafting layer, or stay manual

CriterionExtend existing RIM or authoring platformSeparate grounded drafting layerStructured content, manual drafting
Traceability to source cellsDepends on what the vendor exposes, often document-levelDesigned in: sentence-level bindings and checksComponent links, checked by people
Validation effortLower if the vendor supplies evidence for the featureHighest: you validate the pipeline and its checksLowest: existing authoring controls apply
Hosting of unpublished dataFollows the vendor's cloud and model choicesYour choice, including private deploymentUnchanged
Fit with eCTD publishingNative, inside one platformHands approved documents to the publishing toolNative
Changing models laterTied to the vendor's roadmapSwap models behind the same checksNot applicable
Best fitTeams standardized on one platform with modest volumeSponsors with many studies and strict traceability needsTeams building structured content before adding AI

The third column is often a first stage rather than a rival: drafting can be added once components and bindings exist. The table compares architectures, not named products.

A mismatched figure caught before medical review

Failure modes in AI-drafted submission content, and the control for each

Invented or misread numbers

Early signalA figure in the draft has no source binding.

MitigationBlock any number without a binding; the check fails the draft instead of warning.

Stale sources after a rerun

Early signalBindings point at a superseded table version.

MitigationPin bindings to versions and re-run the check whenever a TLF changes.

Unpublished data leaving your control

Early signalPrompts or source tables sent to a shared public endpoint.

MitigationHost models in a private deployment and log every retrieval.

Automation bias in review

Early signalReviewers approve sections without opening sources.

MitigationShow each sentence's source inline in the review view and sample-check bindings at every gate.

Lay text drifting into jargon or promotion

Early signalLay summaries reuse technical components unchanged.

MitigationKeep separate audience variants with their own reviewers and readability checks.

Publishing in eCTD v4, drafting controls and where ColdAI fits

Publishing stays in your validated eCTD tool. FDA has accepted new applications in eCTD v4.0 since September 20246, and the ICH M4 guideline has carried eCTD v4 tables since its 2016 revision2. Other agencies set their own acceptance and mandatory dates, so confirm them before changing publishing plans. Because eCTD v4 is designed to make document reuse easier, it rewards the component approach described here.

FDA's draft guidance on AI for regulatory decision-making leaves out drafting a submission when that use does not affect patient safety, drug quality or study reliability7, yet the sponsor still answers for every statement submitted, so drafting controls should match the stakes. ColdAI's life sciences work includes AI-assisted submission preparation with automated cross-referencing and compliance checking8, and its enterprise AI practice deploys models inside client infrastructure5.

Questions and answers

Will regulators accept submission text drafted with AI?

Agencies assess the content, and the sponsor is accountable for every statement however it was drafted. For product information texts, EMA's reflection paper asks for close human supervision and quality review so that model-generated text is factually and syntactically correct before submission3, a sensible bar for other documents too. A draft that traces each number to its source and has passed medical, statistical and regulatory review is assessed like any other document.

Who is accountable for AI-drafted content in a clinical study report?

The same people as before: the sponsor, through the medical writer, statistician and regulatory lead who approve it. The architecture helps them by showing each sentence's source and recording every edit and approval in the version history. Drafted text should never be able to reach the publishing tool without passing those named approvers.

Does FDA's draft AI guidance cover regulatory writing tools?

Generally not. Its scope excludes AI used for operational efficiencies and names drafting or writing a regulatory submission as an example, provided the use does not affect patient safety, drug quality or the reliability of study results7. That exclusion leaves the sponsor's responsibility for accuracy untouched, so review controls still matter.

Can we use a public model API with unpublished trial results?

Decide hosting before drafting starts, and check your data policy and contracts before uploading anything. Unpublished results and patient-level data are confidential, and a private deployment in your own cloud tenancy or data center keeps prompts, retrieved sources and outputs under your control while making access logging straightforward.

Sources

  1. ICH E3: Structure and Content of Clinical Study Reports — International Council for Harmonisation · checked 10 October 2026
  2. ICH M4(R4): Organisation of the Common Technical Document for the Registration of Pharmaceuticals for Human Use — International Council for Harmonisation · checked 10 October 2026
  3. Reflection paper on the use of Artificial Intelligence (AI) in the medicinal product lifecycle — European Medicines Agency · checked 10 October 2026
  4. Regulation (EU) No 536/2014 on clinical trials on medicinal products for human use (Article 37 and Annex V) — EUR-Lex · checked 10 October 2026
  5. Enterprise AI: on-premise and private-cloud deployments — ColdAI
  6. Electronic Common Technical Document (eCTD) v4.0 — U.S. Food and Drug Administration · checked 10 October 2026
  7. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (draft guidance) — U.S. Food and Drug Administration · checked 10 October 2026
  8. Life Sciences: regulatory submission automation — ColdAI

More in Life Sciences

Back to Life Sciences

Next step

Test traceable drafting on one completed study report

Tell us the document type you want to start with and how your TLFs and protocols are stored. We will reply with an outline content model, the checks it needs and where it would sit beside your current authoring and publishing tools.

Discuss AI regulatory writing