Deep diveExecutive Recruitment
Executive assessment by work sample: designing fair tests for technical leaders
An executive assessment work sample asks a shortlisted candidate to do a compressed, realistic piece of the job, such as reviewing an architecture decision or drafting a board update mid-incident, and scores it against anchors agreed in advance. Paired with a structured interview and structured referencing, it gives a panel comparable evidence rather than impressions. This page covers each instrument, role-specific exercises, fairness and how to combine the evidence.
On this page
- Why conversation-only executive interviews mislead panels
- Vocabulary for designing an executive assessment
- Turning a search brief into questions and anchors
- Anchors for one criterion: explaining technical risk to a board
- Role-specific work samples and the limits of each
- How evidence travels from the brief to an offer decision
- Keeping exercises proportionate and transparent for senior candidates
- Panel biases that structured assessment is built to catch
- Combining interview, work sample and references into one decision
- Questions and answers
- Sources
Why conversation-only executive interviews mislead panels
Senior candidates are practised interviewees. They have presented to boards and pitched investors many times, and an unstructured conversation rewards that fluency: the panel leaves with a vivid sense of the person and little evidence about how they make the decisions the role requires.
Three things tend to go wrong. Each assessor asks different questions, so candidates are compared on different material. Impressions form early and colour later answers. And the discussion converges on the most senior or most confident voice, with dissent rarely written down.
Structure does not remove judgement. It moves judgement to the design stage, where the panel agrees what good looks like before anyone has a favourite.
Vocabulary for designing an executive assessment
- Criterion
- One capability the panel will score, phrased so that evidence can confirm or contradict it, for example "makes reversible decisions quickly and irreversible ones deliberately".
- Behaviourally anchored rating scale
- A scoring scale where each level is described by observable behaviour, so that a middling score means the same thing to every panel member.
- Past-behaviour question
- A question about what the candidate actually did in a specific situation, followed by probes for their own actions, the options they rejected and what happened next.
- Work sample
- A compressed, realistic task drawn from the role, completed by the candidate and scored against anchors. It shows how someone works rather than how they describe their work.
- Integration meeting
- The closing panel session where results from every instrument sit side by side and a decision is recorded with its reasons.
Turning a search brief into questions and anchors
Extract four to six criteria from the brief
Take each problem the hire must solve and ask what capability it demands. "Turn strong individual contributors into a managed function" yields criteria about building management layers and holding people accountable, not technical brilliance.
Write past-behaviour questions with prepared probes
Ask for a specific occasion: a programme they built, a control they chose not to implement, a time they told a board something unwelcome. Prepare probes such as "what did you decide personally?" and "what did you argue against?", because executives naturally say "we" and the probes separate leadership from proximity to a good team.
Draft anchors for weak, adequate and strong evidence
Describe what each level sounds like for every criterion before interviews begin. Anchors written after meeting candidates tend to describe the favourite.
Anchors for one criterion: explaining technical risk to a board
| Evidence point | Weak | Adequate | Strong |
|---|---|---|---|
| Framing | Describes the technology in detail; the decision needed is unclear | States the risk and its likely impact in business terms | Opens with the decision required, the options and what each costs |
| Uncertainty | Presents estimates as facts or avoids estimates altogether | Flags what is unknown | Separates known, assumed and unknown, and says what would change the advice |
| Example given | Generic or about the team rather than themselves | A real occasion, described at a high level | A specific occasion with their own recommendation, the board's response and the outcome |
| Reflection | Nothing they would change | A minor improvement | Names what they misjudged and how they now brief differently |
Illustrative only: each search writes its own anchors from the brief before the first interview.
Role-specific work samples and the limits of each
| Role | The task | Materials given | What strong work shows | What it cannot tell you |
|---|---|---|---|---|
| CTO | Review a past architecture decision, present the trade-offs and what to change | A sanitised decision record, system diagram and the constraints at the time | Separates what was knowable then from hindsight; sequences a realistic change | How they lead a large organisation day to day |
| CISO | Walk through an incident timeline and draft the board update it requires | A fictional timeline with gaps and conflicting signals, plus notification duties in scope | Distinguishes facts from assumptions; sets notification decision points; writes a short note | Depth of hands-on forensic skill |
| CDO | Critique a data strategy and propose what to stop, keep and start | A short strategy paper, an excerpt of the data inventory and business priorities | Ties data work to decisions the business makes; candid about ownership and quality gaps | Ability to deliver a large platform migration |
| Head of AI | Critique a model-evaluation report attached to a launch request | An evaluation write-up with selective metrics and a description of the test data | Spots missing failure analysis, leakage risk and untested user groups; sets launch conditions | Originality as a researcher |
All material is fictional or sanitised; interview questions cover what each exercise cannot.
How evidence travels from the brief to an offer decision
- Approved search brief
Problems to solve and the criteria derived from them.
- Questions and anchors
Past-behaviour questions, probes and anchored scales fixed before candidates are met.
- Structured interview
Same questions for every shortlisted candidate, scored alone by each assessor.
- Work sample
A realistic role task, reviewed by a practitioner against its own anchors.
- Structured references
Managers, peers and people led, asked the same prompts mapped to the criteria.
- Integration meeting
Scores side by side, disagreements explained, decision and reasons recorded.
Keeping exercises proportionate and transparent for senior candidates
Panel biases that structured assessment is built to catch
Halo from pedigree
Early signalDiscussion keeps returning to a former employer, title or university rather than the evidence.
MitigationRequire an evidence note for every score, and where practical let the work-sample reviewer see the output without the CV.
Similarity disguised as fit
Early signalAssessors praise "cultural fit" without naming a behaviour from the brief.
MitigationReplace fit with the brief's criteria; a score that cannot point to an anchor does not count.
The loudest voice decides
Early signalScores move towards the chair's view once discussion starts.
MitigationCollect written scores first, let the most senior assessor speak last and record dissent in the decision note.
An exercise that screens out a group for reasons unrelated to the job
Early signalOne group scores consistently lower on a single exercise while doing well elsewhere.
MitigationReview the task for job relevance and redesign it. In the United States, the Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607) describe how procedures with adverse impact should be shown to relate to the job2.
Combining interview, work sample and references into one decision
Agree the weight of each criterion before the first interview, ideally when the brief is approved. Then, in the integration meeting, add up the evidence first and debate second. If the panel wants to overrule the combined result, it can, but the reason goes in writing.
Disagreement between instruments is information, not noise. A candidate who interviews brilliantly but produces a weak work sample may be better at narrating past success than at doing the work fresh; the reverse can point to someone undersold by modesty. References often settle it, because they describe behaviour over years rather than hours.
For an external yardstick, ISO 10667-2:2020 sets requirements for providers delivering work-related assessments, including how results are interpreted and reported and how personal data is handled1. The executive recruitment process produces the same trail: scored interview notes, work-sample results and reference reports measured against the brief4.
Questions and answers
How long should a work-sample exercise for an executive take?
Long enough to show judgement and short enough that a busy executive will agree to it. For most technical leadership roles that means a discussion or presentation of under an hour, with preparation measured in hours rather than days. Anything longer tends to test availability, and strong candidates with demanding current roles drop out for reasons unrelated to ability.
Should candidates be paid for completing a work sample?
Usually not, provided the task is fictional, bounded and disclosed in advance. Pay, or redesign the task, when it starts to resemble consulting work, uses the company's real problems or demands significant time. Either way, commit in writing never to use a candidate's output in the company's own work.
Can internal candidates sit the same assessment as external ones?
They should, so that the panel compares like with like. Adjust the materials to remove insider advantage: an internal CTO candidate should not review a decision they took part in. Agree in advance who sees internal candidates' scores, and offer development feedback whatever the outcome.
What if a strong candidate declines to complete a work sample?
Ask why first. The objection is often time or confidentiality, which a shorter format or a discussion of a past decision using the candidate's own anonymised material can address. If the panel substitutes something, score it against the same anchors and record the change rather than quietly lowering the standard for one person.
Sources
- ISO 10667-2:2020 Assessment service delivery: Procedures and methods to assess people in work and organizational settings, Part 2: Requirements for service providers — International Organization for Standardization · checked 10 October 2026
- 29 CFR Part 1607: Uniform Guidelines on Employee Selection Procedures — Legal Information Institute, Cornell Law School · checked 10 October 2026
- Equality Act 2010, section 20: Duty to make adjustments — legislation.gov.uk · checked 10 October 2026
- Executive Recruitment: assessment approach and search deliverables — ColdAI