GuideFrontier R&D

How to write an R&D brief that a result can actually answer

A research brief decides whether an experiment can ever change your mind. It names the decision the work informs, records what prior work already settles, states hypotheses that a measurement could refute, fixes the baseline to beat, and agrees in advance which result counts as negative and when to stop. This guide walks through each part, with weak and strong hypotheses side by side and a one-page template you can reuse.

Reviewed 8 min read

On this page
  1. Anchor the brief to a decision with a date and an owner
  2. Prior-art checks that can end the project early
  3. Weak and falsifiable hypotheses side by side
  4. Turning a question into an evaluation plan
  5. Checkpoint decisions: continue, narrow or stop
  6. A one-page R&D brief template
  7. Hypothetical brief: agent memory across sessions
  8. Habits that make research unfalsifiable
  9. Questions and answers
  10. Sources

Anchor the brief to a decision with a date and an owner

Research questions drift when nobody says what they are for. Before drafting a single hypothesis, write one sentence that names the decision the results will feed, who makes it and roughly when. "Should we redesign our agent memory layer or change model provider before next year's platform budget is set?" is a decision. "Explore memory in agent systems" is a topic, and topics absorb any amount of effort without ever producing an answer.

The decision also sets the precision you need. A choice between two designs of similar cost needs only a rough answer; one that commits several teams for years needs the evidence a sceptical reviewer would ask for.

ColdAI's research engagements open the same way, with a one-page question statement and a review of published work, standards and open-source implementations before any hypothesis is signed off1.

Prior-art checks that can end the project early

A short, honest search often shows that the question is answered, partly answered or answerable without new experiments. Each check below either narrows the brief or removes the need for it.

0 of 5 checked

Weak and falsifiable hypotheses side by side

A hypothesis is ready when you can describe the measurement that would prove it wrong. Each row rewrites a typical first draft from one of the research areas we work in.

AreaFirst draftWhy it cannot failFalsifiable rewrite
Agent memoryA vector store will improve our agent's memory.No task, metric or baseline is named, so any outcome can be read as improvement.On our long-horizon task suite, adding vector retrieval raises correct recall of earlier facts above the no-retrieval baseline across repeated runs.
ConsensusProtocol B is faster than protocol A.Faster at what load, network and fault conditions? One favourable run confirms it.Under our recorded transaction mix and emulated inter-region latency, protocol B has lower tail time-to-finality than A at equal committee size, with and without a crashed replica.
Post-quantum signaturesPost-quantum signatures are too big for our ledger."Too big" has no threshold, so the claim survives any measurement.Replacing Ed25519 with ML-DSA keeps peak-hour gossip bandwidth and daily state growth within the headroom in our capacity plan.
Architecture searchArchitecture search will find a better model.Better on which metric, against which baseline and at what search cost?With the same training budget, a searched architecture matches the hand-designed baseline's held-out accuracy while using less inference compute on our target device.

Every rewrite names a task, a metric, a baseline and the conditions. That is what lets a negative result mean something.

Turning a question into an evaluation plan

  1. Choose metrics the decision depends on

    Pick the one or two measurements that would change the decision, plus guard-rail metrics that must not get worse. A memory redesign judged only on recall could hide a rise in latency or running cost.

    Output
    Primary and guard-rail metrics
    Owner
    Sponsor and principal researcher
  2. Fix the baseline before building anything

    The baseline is the strongest simple alternative you would otherwise choose: the current system, a published method or a well-tuned off-the-shelf option. Beating a weak one proves little.

    Output
    Named baseline and its configuration
  3. Secure data and a test harness

    Decide which datasets, recorded traffic or task suites the experiments use, who owns them and whether they may leave your environment. Build the harness so every candidate runs under identical conditions.

    Output
    Dataset register and harness specification
  4. Agree the analysis in advance

    State how many repeated runs each configuration gets, how variance is reported and which comparison answers each hypothesis. Changing the analysis after seeing results is how false positives are born.

    Output
    Pre-registered analysis plan
  5. Write the negative result down

    Describe the outcome that would show the hypothesis is false, in the same units as the metric. A sponsor who has signed this is far less likely to argue for one more experiment later.

    Output
    Negative-result statement
  6. Set checkpoints and stopping rules

    Place a checkpoint after each phase with a rule for continuing, narrowing or stopping. Common triggers to stop are a baseline that cannot be beaten in principle, data that cannot be obtained, or an effect too small to matter for the decision.

    Output
    Checkpoint schedule
    Owner
    Sponsor

Checkpoint decisions: continue, narrow or stop

At each checkpoint the readout is compared with the signed brief, not with what the team hoped to see. Three outcomes cover almost every case.

  • If

    Interim results move in the predicted direction and the harness has survived review.

    Then

    Continue to the next phase as planned, without widening the scope.

    Expanding a study that is working is the most common way good research overruns.

  • If

    The effect appears only under some conditions, such as one workload or one model family.

    Then

    Narrow the hypothesis to those conditions and re-agree the remaining phases.

    A precise, limited answer can still drive the decision; a vague broad one cannot.

  • If

    The baseline cannot be beaten, the data is unavailable or the effect is too small to change the decision.

    Then

    Stop, write up what was found and close the question.

    Stopping on a rule agreed in advance turns sunk cost into a documented answer.

A one-page R&D brief template

Every field fits on one page. If a field cannot be filled in yet, that gap is the first thing to resolve.

Decision
The choice the results feed, who makes it and by when.
Question
One sentence that a measurement can answer.
Prior work
What papers, standards and existing implementations already show, and the gap that remains for your conditions.
Hypotheses
Two or three falsifiable statements, each naming a task, a metric, a baseline and the conditions.
Evaluation plan
Metrics, guard rails, datasets, harness, repetitions and the comparison that tests each hypothesis.
Negative result
The outcome that would show each hypothesis is false.
Checkpoints and stopping rules
When sponsor and researchers decide to continue, narrow or stop, and on what evidence.
Terms
Data handling, ownership and publication, agreed now rather than once results exist. Our page on IP and publication terms compares the options.

Hypothetical brief: agent memory across sessions

Habits that make research unfalsifiable

Success is defined after the results arrive

Early signalThe readout introduces metrics that were never in the brief.

MitigationFreeze metrics and the analysis plan at sign-off, and label anything added later as exploratory.

The baseline moves

Early signalThe comparison quietly switches to an older or weaker system when the strong one wins.

MitigationName the baseline and its configuration in the brief and keep it fixed for the whole study.

Nobody owns the decision

Early signalResults arrive and no one is ready or able to act on them.

MitigationPut the decision owner on the first line of the brief and invite them to every checkpoint.

Questions and answers

How long should an R&D brief be?

One page for the core: decision, question, hypotheses, evaluation plan, negative result and stopping rules. The prior-art review, dataset descriptions and harness design can sit in appendices. A short core forces precision, and it lets the sponsor, the researchers and IP counsel sign the same document rather than separate summaries of it.

Who should write the brief, the sponsor or the researchers?

Both. The sponsor owns the decision, the constraints and what counts as useful; the researchers own whether the hypotheses are testable and the evaluation is sound. We draft the brief together with the sponsor, and nothing runs until both sides have signed it. A brief written by one side alone tends to be either commercially vague or technically narrow.

What if the prior-art check answers the question?

Then the brief has done its job early. You receive the review with notes on how far each source applies to your conditions, and the work moves to implementation instead of experiments. This happens often where a standard or a maintained open-source project already exists, and it costs far less than finding the same thing out after months of experimental work.

Can hypotheses change once experiments start?

Yes, but only through a recorded checkpoint decision. If early results show a hypothesis was framed badly, the sponsor and researchers agree a revised version and the readout notes the change and the reason. What must never happen is silent rewording to fit the results, which turns a test into a description and removes the reason for doing research at all.

Sources

  1. Frontier R&D: how a research engagement is framed — ColdAI

More in Frontier R&D

Back to Frontier R&D

Next step

Send us the question and the decision behind it

Share the question you cannot yet answer and the decision that depends on it. We will tell you whether a prior-art check might settle it and what a first brief would test.

Propose a research question