Candidate evaluation forms: templates for consistent hiring decisions

TL;DR

  • A candidate evaluation form turns interviews into evidence, so hiring decisions are consistent, quick and defensible.
  • Six templates below cover the full journey: CV screen, first interview, work sample, panel interview, final decision and candidate experience survey.
  • Anchored rating scales are what make a form work. Naming a criterion without describing what a 3 looks like leaves interviewers scoring on instinct with extra paperwork.
  • Be deliberate about educational background. Scoring a degree or a course that the position does not genuinely require is one of the easiest ways to import socio-economic bias into an otherwise structured process.
  • Close the loop by collecting feedback in both directions: qualitative feedback from candidates on the process, and calibration feedback to interviewers on their scoring.
  • Sapia.ai can handle the first mile with mobile-first structured interviews, explainable scoring and interview scheduling, while hiring managers keep control of the decision.

A candidate evaluation form is one of the cheapest ways to improve hiring quality. It gives interviewers a shared structure, lets you compare people on evidence rather than instinct, and protects the candidate experience by keeping outcomes timely and transparent.

It is also one of the most commonly botched artefacts in recruitment. Most forms HR professionals inherit are a list of criteria with a 1 to 5 box beside each one and no definition of what any number means. That is not structure. It is unstructured judgement with a spreadsheet wrapped around it.

The templates below are built to fix that. Each one names the criteria, describes what each score looks like in observable behaviour, and leaves space for the candidate’s own words. Lift them straight into your ATS, Excel or Google Sheets.

What is a candidate evaluation form?

A candidate evaluation form is a single, shared record of how interviewers assess skills, behaviours and motivation against a role profile. A complete one contains four things:

  1. Criteria drawn from the outcomes that matter in the position, not a generic competency list.
  2. Anchored rating scales, so a 3 means the same thing to every assessor.
  3. Evidence boxes for verbatim snippets or observable actions.
  4. A recommendation that points back to the evidence rather than summarising a feeling.

Used consistently, it improves objective assessment, reduces the influence of unconscious bias, sharpens debrief discussion, and speeds up offers without sacrificing fairness. The benefits compound: once you have a season of scored forms, you can start correlating criteria against actual on-the-job performance and cut the ones that predict nothing.

Before you create the form: decide what information you need at each stage

The most common failure is asking every stage to do every job. Decide what each step is genuinely for, then collect information accordingly.

StageWhat it should establishWhat it should not attempt
CV or application screenThreshold requirements only: eligibility, licences, non-negotiable work experienceRanking, culture assessment, potential
Structured first interviewBehavioural evidence against three to five core criteriaDeep technical depth
Work sample or taskObservable performance on realistic workPersonality inference
Panel or hiring manager interviewDepth on the criteria that carry most weight, plus scope and stakeholder contextRe-running questions already answered
Final reviewConsolidation and the decision itselfNew assessment

Two rules worth holding to across all of them.

Four to six criteria, no more. Beyond that, interviewers stop discriminating between criteria and start scoring an overall impression six times. Go deeper with better prompts, not more boxes.

Be careful with educational background. Education is a legitimate criterion when a qualification is genuinely required to do the job or to be licensed for it. It is a bias vector when it is not. A degree, a university’s reputation or a specific course tells you a great deal about someone’s access to education and comparatively little about their ability to do the position you are hiring for. If you keep education on the form, write down the requirement it maps to. If you cannot, take it off.

Template 1: CV and application screening form

Use for: high-volume first pass, before any interview. Scale: binary, not 1 to 5. A screen is a gate, not a ranking.

FieldEntry
Candidate
Position
Reviewer
Date
Source
Threshold criterionMet?Evidence from application
Eligible to work in locationYes / No
Holds required licence or certificationYes / No / N/A
Minimum relevant work experience as defined in role profileYes / No
Required qualification, where the position legally or technically demands oneYes / No / N/A
Availability matches the shift or contract patternYes / No

Notes on borderline cases Record why you are progressing or rejecting anyone who fails a threshold on a technicality. Adjacent work experience, career changers and non-linear paths are exactly where a rigid screen loses good people.

Outcome: Progress / Reject / Hold for another position

Do not score on this form. No 1 to 5, no ranking, no “seems strong”. A CV tells you what someone has had access to. It is a poor predictor of performance, which is why predictive assessment at the first mile outperforms it. Keep the screen mechanical and let the interview do the evaluating.

Template 2: General candidate evaluation form, first interview

Use for: phone screens and first-round interviews. Scale: 1 to 5, anchored below.

FieldEntry
Candidate
Position
Interviewer
StageFirst interview
Date

Criteria and anchors

Criterion1 – Limited3 – Proficient5 – Strong
Customer or stakeholder focusDescribes the situation but not the person’s need; no outcomeIdentifies the need, acts on it, describes the resultAnticipates the need, balances it against a competing constraint, quantifies the outcome
Problem solvingNames the problem; the solution arrives without visible reasoningWalks through diagnosis, options and choice; explains the trade-offHandles ambiguity, revisits their own assumption, describes what they would do differently
Teamwork and communicationUses “we” throughout; own contribution unclearStates their own role and how they worked with othersDescribes managing disagreement or a difficult handover, and the effect on the team
Role-specific knowledgeFamiliar with terms, cannot apply themApplies knowledge correctly to a realistic scenarioApplies it, knows its limits, explains when it does not hold
Motivation and interest in the positionGeneric reasons that would fit any employerSpecific interest in this role and business, with a reasonConnects the position to their own trajectory and can name what they want to learn

Behavioural prompts

Ask all four of every candidate, in the same order.

  1. Tell me about a time you handled a difficult customer or stakeholder. What happened, and what did you do?
  2. Walk me through a recent problem you solved end to end.
  3. Give me an example of working with a team under time pressure.
  4. What attracted you to this position, and to this business?

Ratings and evidence

PromptCriterionScore (1–5)Evidence – what you actually heard
1Customer focus
2Problem solving
3Teamwork and communication
4Motivation and interest
ThroughoutRole knowledge

Job-relevant concerns:
Adjustments provided during the interview:
Recommendation: Progress / Hold / Do not progress
One-sentence rationale, referencing evidence above:

Template 3: Work sample evaluation form

Use for: any task, exercise or trial shift. Separate it from the interview form, because task performance and interview performance are different signals and blending them hides which one drove the decision.

FieldEntry
Candidate
Position
Task issued
Time allowed
Assessor
What you are evaluating1 – Limited3 – Proficient5 – StrongScoreObserved evidence
Output quality against the briefMisses the core requirementMeets the briefMeets the brief and surfaces something you had not asked for
Reasoning made visibleNo explanation offeredExplains the approach when askedExplains it unprompted, including what they rejected
Response to constraint or new informationIgnores or stallsAdapts and continuesAdapts and explains the cost of adapting
Craft and attention to detailErrors that would reach a customerClean, checkable workClean work plus a self-check they ran

Task conditions: (Note anything that differed from the standard brief: extra time granted, tooling unavailable, interruptions. A task score is only comparable if the conditions were.)

Recommendation on task alone: Progress / Hold / Do not progress

Template 4: Panel interview evaluation form, weighted

Use for: second-round or panel stages. Each panellist completes their own copy before any discussion. Weighting is set once, per position, before the first interview.

FieldEntry
Candidate
Position
Panellist
Prompts owned by this panellist
Date
CriterionWeightScore (1–5)Weighted (score × weight)Evidence
Problem solving30%
Collaboration and influence25%
Role and domain knowledge25%
Values and motivation20%
Total100%/5

Worked calculation. Scores of 4, 3, 4, 3 give (4 × 0.30) + (3 × 0.25) + (4 × 0.25) + (3 × 0.20) = 1.20 + 0.75 + 1.00 + 0.60 = 3.55.

Publish the threshold before you interview, not after. Deciding that 3.5 is the bar once you have seen the numbers is how weighting gets reverse-engineered to justify a favourite.

Assign prompt ownership. Write on each panellist’s copy which prompts they own. Without it, three interviewers ask the same question three times, the candidate repeats a rehearsed answer, and you record it as three independent data points.

Scope of role discussed with candidate:
Questions the candidate asked:
Recommendation: Hire / Hold / No hire
Rationale, one sentence:

Template 5: Final review and decision form

One per candidate, completed by the hiring manager after individual forms are submitted. This is the only form on which a decision is recorded.

PanellistStageWeighted scoreRecommendation
[Name]
[Name]
[Name]
Panel average
Score range (highest minus lowest)

If the range is 1.5 or more, stop and reconcile. A wide spread usually means one of three things: panellists interpreted an anchor differently, they saw genuinely different evidence, or someone scored an overall impression rather than the criterion. Identify which before averaging, because averaging a disagreement you have not understood is how a panel launders a single strong opinion into a consensus.

Criterion-level disagreement to resolve:
Evidence that changed someone’s mind in the debrief:
Dissent recorded (name, criterion, position held):

Decision: Offer / Hold / Reject
Rationale, referencing the evidence:
Feedback to be provided to the candidate, and by when:

That last row is not administrative. Every candidate you evaluate has given you their time, and the completed form means you already have something specific to tell them. Committing to a date on the decision form is what makes it happen.

Template 6: Candidate experience survey

Use for: collecting feedback on the process itself, from everyone you interviewed, including those you rejected.

Send it after the outcome is communicated, not before. Keep responses anonymous and say so on the form, otherwise rejected candidates will tell you what they think you want to hear. Make clear that answering has no bearing on future applications.

Rate 1 to 5, where 1 is strongly disagree and 5 is strongly agree:

  1. The process was explained clearly before I started.
  2. The questions I was asked felt relevant to the position.
  3. The time the process took was reasonable.
  4. I was treated with respect at every stage.
  5. I understand what happens next and when.
  6. I would apply here again.

Then two open questions, which is where the qualitative feedback that actually changes anything lives:

  • What was the single most frustrating part of the process?
  • What is one thing we could do to improve it for the next candidate?

How to use what comes back:

Score-only surveys drift towards a flat 4 and tell you nothing. Read the open answers as a set each month and look for the same phrase appearing in different people’s words. Repeated language is a process defect, not a mood. Track the six scores as a trend and publish the overall figure on your careers website once you are confident in it: candidates increasingly compare employers on process, and a real number beats a claim.

For teams starting from nothing, measuring candidate experience without dedicated tooling is a reasonable first step. As a reference point, Woolworths reached a 9.2/10 candidate satisfaction score after moving their first mile to structured chat interviews.

Worked example: one criterion, three answers

This is the part most guides skip, and the part interviewers actually need. Same prompt, same criterion, three real-shaped answers.

Prompt: Walk me through a recent problem you solved end to end. Criterion: Problem solving, 1 to 5.

Answer A – scored 1. “We had a backlog issue in the warehouse so we sorted it out and got back on track. It was a team effort, everyone pulled together.”

Why 1: the problem is named, the solution is asserted. No diagnosis, no options, no visible reasoning, no measurable outcome. Nothing here would let you predict how they would handle the next problem.

Answer B – scored 3. “Picking was running about 20% behind. I pulled the scan data and found most of the delay was on one aisle where fast-movers were on the top shelf. I moved them to waist height and picking times came back in line within a week.”

Why 3: diagnosis, evidence, action, measured result. This is a solid, proficient answer and 3 is a good score, not a polite one.

Answer C – scored 5. “Picking was 20% behind. My first assumption was staffing, so I asked for an extra picker, and it barely moved. That told me it was layout, not capacity. Scan data pointed at one aisle with fast-movers on the top shelf. I moved them down, times recovered within a week, and I set a monthly check because the mix shifts seasonally. If I did it again I would have checked the data before asking for headcount.”

Why 5: they tested and discarded their own assumption, used the failure as information, built in a control, and named their own mistake without being prompted.

What this shows interviewers:

The difference between B and C is whether the reasoning is visible and whether the person can evaluate their own judgement. Capture the phrases that carry that signal – “my first assumption was”, “if I did it again” – in the candidate’s own words. Adjectives in an evidence box are useless in a debrief three days later. Sentences are not.

Circulate two or three examples like this with the form. It is the fastest calibration exercise available and it costs an hour.

Rating scales: choosing and anchoring them

Rating scales are where most forms quietly fail.

Use 1 to 5 for behaviours. Wide enough to discriminate, narrow enough to anchor.

Use 1 to 4 for skills. An even number removes the safe middle, which matters when the honest answer is “I could not tell”.

Use binary for thresholds. Eligibility, licences, availability. There is no such thing as a 3 out of 5 for a right to work.

Avoid 1 to 10. Nobody can describe the difference between a 6 and a 7, so scores cluster between 6 and 8 and the scale stops doing any work.

Anchor at 1, 3 and 5 only. Anchoring all five levels produces a document nobody reads. Interviewers interpolate 2 and 4 accurately once they can see the shape.

Anchor in behaviour, never in adverbs. “Communicates effectively” is not an anchor. “States their own contribution rather than the team’s” is, because two people can look at the same transcript and agree.

Set the weighting before the first interview. Weight reflects what the position needs. It is not a dial you turn once you have met the candidates.

If you are standardising this across multiple panels, our interview score sheet templates go further on consistency mechanics, and the difference between structured and unstructured interviewing explains why the scale matters more than the questions.

How to use the forms in practice

A little process discipline goes a long way.

Before the interview

  • Share the role profile, criteria, and forms with interviewers.
  • Align on who owns which prompts to avoid overlap.
  • Brief interviewers on the rubric and the meaning of each score.

During the interview

  • Ask the same core questions of every candidate.
  • Capture short, verbatim notes linked to each prompt or task.
  • Use the rubric to score immediately after each answer, not from memory later.

After the interview

  • Submit individual forms first.
  • Hold a short debrief to compare scores and evidence.
  • Record a final decision with one-sentence rationale and next actions.

Fairness, bias and defensibility

A structured form narrows the space in which affinity bias, halo effects and similarity bias operate. It does not eliminate them, and no honest vendor will tell you otherwise. What it does is make the reasoning inspectable, which is the precondition for fixing anything.

Three practices carry most of the weight.

Publish your adjustments process on your website. Candidates who need an adjustment and cannot find how to request one either disclose under pressure or drop out. Put it on the careers page, before the application form.

Keep the CV out of the interview. If your first-mile assessment is blind to name, educational background and career gaps, do not hand interviewers the CV before the structured interview. Sequencing is the whole intervention. LNER maintained 30% ethnic minority representation consistently across every hiring stage after moving from manual CV screening and video interviews, with representation of 32 to 37% across categories, while hiring time fell from seven weeks to three.

Audit selection rates by stage, not just at offer. Aggregate diversity figures at the end of the funnel hide where people are being lost. Break it down by stage and you can see which criterion, and often which interviewer, is doing it. This is the core of diversity hiring that survives scrutiny.

Provide feedback to your interviewers, too

Collecting feedback from candidates is only half the loop. The other half is telling interviewers how their scoring compares to the panel’s.

A short quarterly note per interviewer, drawn from the forms:

MetricThis quarterPanel average
Mean score awarded
Score range used (highest minus lowest awarded)
Deviation from panel mean
Proportion of evidence boxes containing a verbatim quote
Proportion of recommendations that matched the final decision

Two patterns matter. 

Compression – an interviewer who only ever awards 3s and 4s is not evaluating, they are hedging. 

Drift – an interviewer consistently a point above or below the panel needs recalibration on the anchors, not a reprimand.

Most interviewers have never been told how their scoring compares to anyone else’s. Showing them the numbers produces visible improvement within a quarter, and it is free.

Running the forms in a spreadsheet, an ATS or an AI first mile

Spreadsheet. Workable at low volume. Use data validation for the 1 to 5 fields, drop-downs for stage and recommendation, and conditional formatting to flag a panel range above 1.5. Keep the anchors on a locked tab so users can read them but not overwrite them. Version the file, or you will end up with four incompatible forms and no comparable data.

ATS. The right home once you manage multiple requisitions. Store templates centrally, pre-fill role criteria, and make evidence fields mandatory. The gain is not convenience, it is that scores become queryable, which is what lets you correlate criteria with performance later.

AI at the first mile. Neither of the above fixes the volume problem. If you are screening 400 to 600 applicants per position, the constraint is that nobody can apply a structured form to all of them, so the structure quietly reverts to CV skimming.

That is the specific gap Sapia.ai fills. A structured, mobile chat interview evaluates every applicant against your criteria and returns an explainable score with the language it was based on, so the output drops into the same evidence-and-score format your panel already uses. It integrates with interview scheduling and keeps candidates engaged with feedback rather than silence. For high-volume roles, it is the difference between a structured process and a structured intention.

Be clear about what it does not do. It does not replace the panel form, the task, or the decision. It handles the first mile so your interviewers spend their judgement on the shortlist instead of the pile. If your volumes are modest and your time-to-hire is fine, the templates above are enough on their own, and you should not buy software to fix a problem you do not have.

What to measure from your forms

Keep the scoreboard small and look at it weekly.

MetricWhat it tells youAct when
Average score by criterion, per stageWhether a criterion discriminates at allEveryone scores 4 – the criterion is decorative, cut or rewrite it
Rater variance across interviewersCalibration healthRange above 1.5 recurring
Pass-through rate by stageWhere candidates are lostA stage drops more than expected, or drops one group disproportionately
Time from final form to decisionWhether the process respects candidatesBeyond 5 working days
Correlation between score and 90-day performanceWhether your criteria predict anythingAny criterion with near-zero correlation after two quarters

That last row is the one nobody runs and the only one that tells you whether the form works. If a criterion does not correlate with performance, it is costing you good candidates for nothing.

Download-ready master form

Copy everything from here to the end of the section into your ATS or document editor, then adapt the criteria to the position.

Header

Candidate:
Position:
Interviewer:
Stage:
Date:
Prompts owned by this interviewer:

Criteria and weightings

Set these before the first interview, not after. Weightings must total 100 per cent.

Criterion 1, weighting:
Criterion 2, weighting:
Criterion 3, weighting:
Criterion 4, weighting:
Criterion 5, weighting:

Anchors

For each criterion, write one line describing what a 1, a 3 and a 5 look like in observable behaviour. Follow the pattern used in Template 2 above, for example:

Problem solving, 1: names the problem, but the solution arrives with no visible reasoning.
Problem solving, 3: walks through diagnosis, options and choice, and explains the trade-off.
Problem solving, 5: handles ambiguity, revisits their own assumption, describes what they would do differently.

Prompts, scores and evidence

Question 1:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:

Question 2:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:

Question 3:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:

Question 4:
Criterion assessed:
Score, 1 to 5:
Evidence, in the candidate’s own words:

Task, where one was set

Brief issued:
Anything that differed from the standard conditions:
Score, 1 to 5:
Observed evidence:

Totals

Weighted total, out of 5:

Concerns and adjustments

Job-relevant concerns: Adjustments provided during the interview:

Recommendation

Recommendation, delete as applicable: Hire / Hold / No hire
Rationale, one sentence that references the evidence recorded above:
Feedback to be sent to the candidate by:

Conclusion

A good candidate evaluation form makes hiring faster, fairer and easier to defend. The mechanics are not complicated: start from the outcomes that define success in the position, turn them into four to six criteria, anchor every rating scale in observable behaviour, and train interviewers to capture what they actually heard rather than how they felt about it.

The two habits that separate teams who do this well are unglamorous. They calibrate on real answers before interviewing, and they read what candidates tell them about the process afterwards. Everything else is a template, and the templates are above.

If you want to see how a structured, mobile-first first mile plugs into the forms you already use, book a Sapia.ai demo. Your people keep the decision. Candidates get a consistent process and an answer.

FAQs about candidate evaluation forms

What is a candidate evaluation form interview template used for?

It standardises how interviewers score skills, behaviours, and values. You get consistent data that speeds the final decision and improves fairness.

Do we need different candidate evaluation forms by stage?

Keep one core template but tweak prompts and weightings for CV screen, interview, and task review. Use a resume evaluation form for the screen, then a behavioural and task form for interviews.

How many criteria should a candidate evaluation template include in 2026?

Four to six. Go deeper with better prompts and a task rather than adding more checkboxes.

What scale is best?

Use 1–5 for behaviours, and 1–4 for skills to reduce fence-sitting. Always include behavioural anchors.

Where does Sapia.ai fit with forms and scoring?

Sapia.ai can run the structured first interview, generate explainable scores aligned to your rubric, and handle interview scheduling. You still review the evidence and make the decision.

About Author

Get started with Sapia.ai today

Hire brilliant with the talent intelligence platform powered by ethical AI
Speak To Our Sales Team