A candidate evaluation form is one of the cheapest ways to improve hiring quality. It gives interviewers a shared structure, lets you compare people on evidence rather than instinct, and protects the candidate experience by keeping outcomes timely and transparent.
It is also one of the most commonly botched artefacts in recruitment. Most forms HR professionals inherit are a list of criteria with a 1 to 5 box beside each one and no definition of what any number means. That is not structure. It is unstructured judgement with a spreadsheet wrapped around it.
The templates below are built to fix that. Each one names the criteria, describes what each score looks like in observable behaviour, and leaves space for the candidate’s own words. Lift them straight into your ATS, Excel or Google Sheets.
A candidate evaluation form is a single, shared record of how interviewers assess skills, behaviours and motivation against a role profile. A complete one contains four things:
Used consistently, it improves objective assessment, reduces the influence of unconscious bias, sharpens debrief discussion, and speeds up offers without sacrificing fairness. The benefits compound: once you have a season of scored forms, you can start correlating criteria against actual on-the-job performance and cut the ones that predict nothing.
The most common failure is asking every stage to do every job. Decide what each step is genuinely for, then collect information accordingly.
| Stage | What it should establish | What it should not attempt |
| CV or application screen | Threshold requirements only: eligibility, licences, non-negotiable work experience | Ranking, culture assessment, potential |
| Structured first interview | Behavioural evidence against three to five core criteria | Deep technical depth |
| Work sample or task | Observable performance on realistic work | Personality inference |
| Panel or hiring manager interview | Depth on the criteria that carry most weight, plus scope and stakeholder context | Re-running questions already answered |
| Final review | Consolidation and the decision itself | New assessment |
Two rules worth holding to across all of them.
Four to six criteria, no more. Beyond that, interviewers stop discriminating between criteria and start scoring an overall impression six times. Go deeper with better prompts, not more boxes.
Be careful with educational background. Education is a legitimate criterion when a qualification is genuinely required to do the job or to be licensed for it. It is a bias vector when it is not. A degree, a university’s reputation or a specific course tells you a great deal about someone’s access to education and comparatively little about their ability to do the position you are hiring for. If you keep education on the form, write down the requirement it maps to. If you cannot, take it off.
Use for: high-volume first pass, before any interview. Scale: binary, not 1 to 5. A screen is a gate, not a ranking.
| Field | Entry |
| Candidate | |
| Position | |
| Reviewer | |
| Date | |
| Source |
| Threshold criterion | Met? | Evidence from application |
| Eligible to work in location | Yes / No | |
| Holds required licence or certification | Yes / No / N/A | |
| Minimum relevant work experience as defined in role profile | Yes / No | |
| Required qualification, where the position legally or technically demands one | Yes / No / N/A | |
| Availability matches the shift or contract pattern | Yes / No |
Notes on borderline cases Record why you are progressing or rejecting anyone who fails a threshold on a technicality. Adjacent work experience, career changers and non-linear paths are exactly where a rigid screen loses good people.
Outcome: Progress / Reject / Hold for another position
Do not score on this form. No 1 to 5, no ranking, no “seems strong”. A CV tells you what someone has had access to. It is a poor predictor of performance, which is why predictive assessment at the first mile outperforms it. Keep the screen mechanical and let the interview do the evaluating.
Use for: phone screens and first-round interviews. Scale: 1 to 5, anchored below.
| Field | Entry |
| Candidate | |
| Position | |
| Interviewer | |
| Stage | First interview |
| Date |
| Criterion | 1 – Limited | 3 – Proficient | 5 – Strong |
| Customer or stakeholder focus | Describes the situation but not the person’s need; no outcome | Identifies the need, acts on it, describes the result | Anticipates the need, balances it against a competing constraint, quantifies the outcome |
| Problem solving | Names the problem; the solution arrives without visible reasoning | Walks through diagnosis, options and choice; explains the trade-off | Handles ambiguity, revisits their own assumption, describes what they would do differently |
| Teamwork and communication | Uses “we” throughout; own contribution unclear | States their own role and how they worked with others | Describes managing disagreement or a difficult handover, and the effect on the team |
| Role-specific knowledge | Familiar with terms, cannot apply them | Applies knowledge correctly to a realistic scenario | Applies it, knows its limits, explains when it does not hold |
| Motivation and interest in the position | Generic reasons that would fit any employer | Specific interest in this role and business, with a reason | Connects the position to their own trajectory and can name what they want to learn |
Ask all four of every candidate, in the same order.
| Prompt | Criterion | Score (1–5) | Evidence – what you actually heard |
| 1 | Customer focus | ||
| 2 | Problem solving | ||
| 3 | Teamwork and communication | ||
| 4 | Motivation and interest | ||
| Throughout | Role knowledge |
Job-relevant concerns:
Adjustments provided during the interview:
Recommendation: Progress / Hold / Do not progress
One-sentence rationale, referencing evidence above:
Use for: any task, exercise or trial shift. Separate it from the interview form, because task performance and interview performance are different signals and blending them hides which one drove the decision.
| Field | Entry |
| Candidate | |
| Position | |
| Task issued | |
| Time allowed | |
| Assessor |
| What you are evaluating | 1 – Limited | 3 – Proficient | 5 – Strong | Score | Observed evidence |
| Output quality against the brief | Misses the core requirement | Meets the brief | Meets the brief and surfaces something you had not asked for | ||
| Reasoning made visible | No explanation offered | Explains the approach when asked | Explains it unprompted, including what they rejected | ||
| Response to constraint or new information | Ignores or stalls | Adapts and continues | Adapts and explains the cost of adapting | ||
| Craft and attention to detail | Errors that would reach a customer | Clean, checkable work | Clean work plus a self-check they ran |
Task conditions: (Note anything that differed from the standard brief: extra time granted, tooling unavailable, interruptions. A task score is only comparable if the conditions were.)
Recommendation on task alone: Progress / Hold / Do not progress
Use for: second-round or panel stages. Each panellist completes their own copy before any discussion. Weighting is set once, per position, before the first interview.
| Field | Entry |
| Candidate | |
| Position | |
| Panellist | |
| Prompts owned by this panellist | |
| Date |
| Criterion | Weight | Score (1–5) | Weighted (score × weight) | Evidence |
| Problem solving | 30% | |||
| Collaboration and influence | 25% | |||
| Role and domain knowledge | 25% | |||
| Values and motivation | 20% | |||
| Total | 100% | /5 |
Worked calculation. Scores of 4, 3, 4, 3 give (4 × 0.30) + (3 × 0.25) + (4 × 0.25) + (3 × 0.20) = 1.20 + 0.75 + 1.00 + 0.60 = 3.55.
Publish the threshold before you interview, not after. Deciding that 3.5 is the bar once you have seen the numbers is how weighting gets reverse-engineered to justify a favourite.
Assign prompt ownership. Write on each panellist’s copy which prompts they own. Without it, three interviewers ask the same question three times, the candidate repeats a rehearsed answer, and you record it as three independent data points.
Scope of role discussed with candidate:
Questions the candidate asked:
Recommendation: Hire / Hold / No hire
Rationale, one sentence:
One per candidate, completed by the hiring manager after individual forms are submitted. This is the only form on which a decision is recorded.
| Panellist | Stage | Weighted score | Recommendation |
| [Name] | |||
| [Name] | |||
| [Name] | |||
| Panel average | |||
| Score range (highest minus lowest) |
If the range is 1.5 or more, stop and reconcile. A wide spread usually means one of three things: panellists interpreted an anchor differently, they saw genuinely different evidence, or someone scored an overall impression rather than the criterion. Identify which before averaging, because averaging a disagreement you have not understood is how a panel launders a single strong opinion into a consensus.
Criterion-level disagreement to resolve:
Evidence that changed someone’s mind in the debrief:
Dissent recorded (name, criterion, position held):
Decision: Offer / Hold / Reject
Rationale, referencing the evidence:
Feedback to be provided to the candidate, and by when:
That last row is not administrative. Every candidate you evaluate has given you their time, and the completed form means you already have something specific to tell them. Committing to a date on the decision form is what makes it happen.
Use for: collecting feedback on the process itself, from everyone you interviewed, including those you rejected.
Send it after the outcome is communicated, not before. Keep responses anonymous and say so on the form, otherwise rejected candidates will tell you what they think you want to hear. Make clear that answering has no bearing on future applications.
Rate 1 to 5, where 1 is strongly disagree and 5 is strongly agree:
Then two open questions, which is where the qualitative feedback that actually changes anything lives:
How to use what comes back:
Score-only surveys drift towards a flat 4 and tell you nothing. Read the open answers as a set each month and look for the same phrase appearing in different people’s words. Repeated language is a process defect, not a mood. Track the six scores as a trend and publish the overall figure on your careers website once you are confident in it: candidates increasingly compare employers on process, and a real number beats a claim.
For teams starting from nothing, measuring candidate experience without dedicated tooling is a reasonable first step. As a reference point, Woolworths reached a 9.2/10 candidate satisfaction score after moving their first mile to structured chat interviews.
This is the part most guides skip, and the part interviewers actually need. Same prompt, same criterion, three real-shaped answers.
Prompt: Walk me through a recent problem you solved end to end. Criterion: Problem solving, 1 to 5.
Answer A – scored 1. “We had a backlog issue in the warehouse so we sorted it out and got back on track. It was a team effort, everyone pulled together.”
Why 1: the problem is named, the solution is asserted. No diagnosis, no options, no visible reasoning, no measurable outcome. Nothing here would let you predict how they would handle the next problem.
Answer B – scored 3. “Picking was running about 20% behind. I pulled the scan data and found most of the delay was on one aisle where fast-movers were on the top shelf. I moved them to waist height and picking times came back in line within a week.”
Why 3: diagnosis, evidence, action, measured result. This is a solid, proficient answer and 3 is a good score, not a polite one.
Answer C – scored 5. “Picking was 20% behind. My first assumption was staffing, so I asked for an extra picker, and it barely moved. That told me it was layout, not capacity. Scan data pointed at one aisle with fast-movers on the top shelf. I moved them down, times recovered within a week, and I set a monthly check because the mix shifts seasonally. If I did it again I would have checked the data before asking for headcount.”
Why 5: they tested and discarded their own assumption, used the failure as information, built in a control, and named their own mistake without being prompted.
What this shows interviewers:
The difference between B and C is whether the reasoning is visible and whether the person can evaluate their own judgement. Capture the phrases that carry that signal – “my first assumption was”, “if I did it again” – in the candidate’s own words. Adjectives in an evidence box are useless in a debrief three days later. Sentences are not.
Circulate two or three examples like this with the form. It is the fastest calibration exercise available and it costs an hour.
Rating scales are where most forms quietly fail.
Use 1 to 5 for behaviours. Wide enough to discriminate, narrow enough to anchor.
Use 1 to 4 for skills. An even number removes the safe middle, which matters when the honest answer is “I could not tell”.
Use binary for thresholds. Eligibility, licences, availability. There is no such thing as a 3 out of 5 for a right to work.
Avoid 1 to 10. Nobody can describe the difference between a 6 and a 7, so scores cluster between 6 and 8 and the scale stops doing any work.
Anchor at 1, 3 and 5 only. Anchoring all five levels produces a document nobody reads. Interviewers interpolate 2 and 4 accurately once they can see the shape.
Anchor in behaviour, never in adverbs. “Communicates effectively” is not an anchor. “States their own contribution rather than the team’s” is, because two people can look at the same transcript and agree.
Set the weighting before the first interview. Weight reflects what the position needs. It is not a dial you turn once you have met the candidates.
If you are standardising this across multiple panels, our interview score sheet templates go further on consistency mechanics, and the difference between structured and unstructured interviewing explains why the scale matters more than the questions.
A little process discipline goes a long way.
A structured form narrows the space in which affinity bias, halo effects and similarity bias operate. It does not eliminate them, and no honest vendor will tell you otherwise. What it does is make the reasoning inspectable, which is the precondition for fixing anything.
Three practices carry most of the weight.
Publish your adjustments process on your website. Candidates who need an adjustment and cannot find how to request one either disclose under pressure or drop out. Put it on the careers page, before the application form.
Keep the CV out of the interview. If your first-mile assessment is blind to name, educational background and career gaps, do not hand interviewers the CV before the structured interview. Sequencing is the whole intervention. LNER maintained 30% ethnic minority representation consistently across every hiring stage after moving from manual CV screening and video interviews, with representation of 32 to 37% across categories, while hiring time fell from seven weeks to three.
Audit selection rates by stage, not just at offer. Aggregate diversity figures at the end of the funnel hide where people are being lost. Break it down by stage and you can see which criterion, and often which interviewer, is doing it. This is the core of diversity hiring that survives scrutiny.
Collecting feedback from candidates is only half the loop. The other half is telling interviewers how their scoring compares to the panel’s.
A short quarterly note per interviewer, drawn from the forms:
| Metric | This quarter | Panel average |
| Mean score awarded | ||
| Score range used (highest minus lowest awarded) | ||
| Deviation from panel mean | ||
| Proportion of evidence boxes containing a verbatim quote | ||
| Proportion of recommendations that matched the final decision |
Two patterns matter.
Compression – an interviewer who only ever awards 3s and 4s is not evaluating, they are hedging.
Drift – an interviewer consistently a point above or below the panel needs recalibration on the anchors, not a reprimand.
Most interviewers have never been told how their scoring compares to anyone else’s. Showing them the numbers produces visible improvement within a quarter, and it is free.
Spreadsheet. Workable at low volume. Use data validation for the 1 to 5 fields, drop-downs for stage and recommendation, and conditional formatting to flag a panel range above 1.5. Keep the anchors on a locked tab so users can read them but not overwrite them. Version the file, or you will end up with four incompatible forms and no comparable data.
ATS. The right home once you manage multiple requisitions. Store templates centrally, pre-fill role criteria, and make evidence fields mandatory. The gain is not convenience, it is that scores become queryable, which is what lets you correlate criteria with performance later.
AI at the first mile. Neither of the above fixes the volume problem. If you are screening 400 to 600 applicants per position, the constraint is that nobody can apply a structured form to all of them, so the structure quietly reverts to CV skimming.
That is the specific gap Sapia.ai fills. A structured, mobile chat interview evaluates every applicant against your criteria and returns an explainable score with the language it was based on, so the output drops into the same evidence-and-score format your panel already uses. It integrates with interview scheduling and keeps candidates engaged with feedback rather than silence. For high-volume roles, it is the difference between a structured process and a structured intention.
Be clear about what it does not do. It does not replace the panel form, the task, or the decision. It handles the first mile so your interviewers spend their judgement on the shortlist instead of the pile. If your volumes are modest and your time-to-hire is fine, the templates above are enough on their own, and you should not buy software to fix a problem you do not have.
Keep the scoreboard small and look at it weekly.
| Metric | What it tells you | Act when |
| Average score by criterion, per stage | Whether a criterion discriminates at all | Everyone scores 4 – the criterion is decorative, cut or rewrite it |
| Rater variance across interviewers | Calibration health | Range above 1.5 recurring |
| Pass-through rate by stage | Where candidates are lost | A stage drops more than expected, or drops one group disproportionately |
| Time from final form to decision | Whether the process respects candidates | Beyond 5 working days |
| Correlation between score and 90-day performance | Whether your criteria predict anything | Any criterion with near-zero correlation after two quarters |
That last row is the one nobody runs and the only one that tells you whether the form works. If a criterion does not correlate with performance, it is costing you good candidates for nothing.
Copy everything from here to the end of the section into your ATS or document editor, then adapt the criteria to the position.
Candidate:
Position:
Interviewer:
Stage:
Date:
Prompts owned by this interviewer:
Set these before the first interview, not after. Weightings must total 100 per cent.
Criterion 1, weighting:
Criterion 2, weighting:
Criterion 3, weighting:
Criterion 4, weighting:
Criterion 5, weighting:
For each criterion, write one line describing what a 1, a 3 and a 5 look like in observable behaviour. Follow the pattern used in Template 2 above, for example:
Problem solving, 1: names the problem, but the solution arrives with no visible reasoning.
Problem solving, 3: walks through diagnosis, options and choice, and explains the trade-off.
Problem solving, 5: handles ambiguity, revisits their own assumption, describes what they would do differently.
Question 1:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:
Question 2:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:
Question 3:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:
Question 4:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:
Brief issued:
Anything that differed from the standard conditions:
Score, 1 to 5:
Observed evidence:
Weighted total, out of 5:
Job-relevant concerns: Adjustments provided during the interview:
Recommendation, delete as applicable: Hire / Hold / No hire
Rationale, one sentence that references the evidence recorded above:
Feedback to be sent to the candidate by:
A good candidate evaluation form makes hiring faster, fairer and easier to defend. The mechanics are not complicated: start from the outcomes that define success in the position, turn them into four to six criteria, anchor every rating scale in observable behaviour, and train interviewers to capture what they actually heard rather than how they felt about it.
The two habits that separate teams who do this well are unglamorous. They calibrate on real answers before interviewing, and they read what candidates tell them about the process afterwards. Everything else is a template, and the templates are above.
If you want to see how a structured, mobile-first first mile plugs into the forms you already use, book a Sapia.ai demo. Your people keep the decision. Candidates get a consistent process and an answer.
It standardises how interviewers score skills, behaviours, and values. You get consistent data that speeds the final decision and improves fairness.
Keep one core template but tweak prompts and weightings for CV screen, interview, and task review. Use a resume evaluation form for the screen, then a behavioural and task form for interviews.
Four to six. Go deeper with better prompts and a task rather than adding more checkboxes.
Use 1–5 for behaviours, and 1–4 for skills to reduce fence-sitting. Always include behavioural anchors.
Sapia.ai can run the structured first interview, generate explainable scores aligned to your rubric, and handle interview scheduling. You still review the evidence and make the decision.