Interview scorecard template

AI Trainer interview scorecard

Pre-screening scorecard for AI Trainer candidates.

See AI scoring
operations administrationai trainingdata annotationmodel evaluationrlhf
Complete evaluation framework

What to assess and how to score it

Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.

01
Evaluation factor

Execution and reliability

35% weight

Check the annotation or evaluation volume they have produced, the domains covered, and how their quality was measured.

Evidence to listen for

  • Describes a workload they owned and how they kept it from slipping
  • Names the tools and systems they ran day to day
  • Can talk about volume: tickets, inboxes, orders, uptime
  • Nothing quietly falls through when they are busy

Five-point scoring guide

1
Poor

Cannot describe their own workload; things slip without them noticing.

2
Needs Improvement

Handles routine volume; drops work under pressure.

3
Satisfactory

Reliable on steady-state work; struggles when volume spikes.

4
Very Good

Consistently reliable at real volume with a system for staying on top.

5
Excellent

Has produced real volume with measured quality scores, and knows where their own agreement rate was weakest.

02
Evaluation factor

Improving the process

25% weight

Test whether they improved the guidelines themselves rather than only applying them.

Evidence to listen for

  • Has changed a process rather than only following one
  • Can name what was slow or error-prone and what they did about it
  • Documents so the improvement survives them
  • Knows when a process is worth automating and when it is not

Five-point scoring guide

1
Poor

Follows instructions only; no sense that process can change.

2
Needs Improvement

Notices problems but escalates rather than solving.

3
Satisfactory

Makes small improvements; impact is local and undocumented.

4
Very Good

Has redesigned a real process with measurable effect.

5
Excellent

Has rewritten or sharpened annotation guidelines to resolve real ambiguity, and can cite the disagreement that prompted it.

03
Evaluation factor

Judgement and autonomy

25% weight

Assess how they handle an example the rubric does not cover, where any label is arguably defensible.

Evidence to listen for

  • Knows what to decide alone and what to escalate
  • Handles an exception without freezing or improvising recklessly
  • Protects confidentiality and access appropriately
  • Asks a clarifying question rather than guessing on something costly

Five-point scoring guide

1
Poor

Either escalates everything or acts recklessly on their own.

2
Needs Improvement

Needs frequent direction; uneasy with exceptions.

3
Satisfactory

Sound judgement on familiar decisions.

4
Very Good

Clear sense of their own authority; handles exceptions well.

5
Excellent

Handles rubric gaps by escalating the pattern rather than silently guessing, and is consistent once a rule is set.

04
Evaluation factor

Communication

15% weight

Judge writing quality and whether they can explain why a model output is wrong in a way an engineer can use.

Evidence to listen for

  • Writes clearly enough that people act without a follow-up
  • Manages expectations before a deadline slips, not after
  • Handles a frustrated colleague or customer calmly
  • Works across time zones or async where the role needs it

Five-point scoring guide

1
Poor

Unclear written communication; goes quiet when things slip.

2
Needs Improvement

Communication needs chasing; raises problems late.

3
Satisfactory

Clear enough day to day; proactive updates are inconsistent.

4
Very Good

Clear, proactive, and calm under pressure.

5
Excellent

Writes precisely about why an output is wrong, in terms specific enough for a model team to act on.

Put this rubric to work

Score every candidate against the same standard

Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.

Explore AI Scorecards