Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Method and rigour
35% weight
Check how they design playtests: sample sizing for small n, counterbalancing builds, think-aloud versus RITE iterations, survey scales like SUS or PENS, and mixing telemetry with lab sessions.
Evidence to listen for
Follows and can justify an established methodology
Understands contamination, bias, and chain of custody as they apply to the field
Knows the limits of their techniques and says so
Documents procedure so results are reproducible and defensible
Five-point scoring guide
1
Poor
Careless method; unaware of contamination, bias, or procedural integrity.
2
Needs Improvement
Knows procedures but applies them inconsistently; gaps in documentation.
3
Satisfactory
Sound standard practice; less certain outside familiar techniques.
4
Very Good
Rigorous and well documented; understands the limits of each method.
5
Excellent
Names specific method choices per question type, defends small-sample limits, and pairs behavioural telemetry with observed session data rather than relying on opinion polls.
02
Evaluation factor
Real casework
25% weight
Probe actual studies run on shipped or in-development titles: genre, build stage (vertical slice, alpha, soft launch), participant recruitment, and which design decisions changed as a result.
Evidence to listen for
Brings specific cases, sites, or projects rather than general description
States their own role and what they personally handled
Can describe an ambiguous or degraded case and how they proceeded
Knows what happened to the work afterwards
Five-point scoring guide
1
Poor
No hands-on casework; experience is entirely academic.
2
Needs Improvement
Limited exposure; cannot describe their own contribution clearly.
3
Satisfactory
Real casework with adequate detail; ownership sometimes vague.
4
Very Good
Specific cases with clear personal scope and outcomes.
5
Excellent
Cites named titles or projects with study counts, recruitment criteria, and concrete features cut, retuned, or redesigned because of their findings.
03
Evaluation factor
Interpretation and judgement
25% weight
Test how they separate tutorial comprehension failures from difficulty tuning or motivation drop-off, and how they read funnel drop, session length, and D1 or D7 retention signals.
Evidence to listen for
Separates what the evidence shows from what they infer
States confidence levels and what would change their conclusion
Comfortable saying the result is inconclusive
Handles pressure to reach a preferred conclusion without bending
Five-point scoring guide
1
Poor
Overstates findings; no separation of evidence from inference.
2
Needs Improvement
Reaches conclusions the evidence does not support; uneasy with uncertainty.
3
Satisfactory
Reasonable judgement; qualifies findings when prompted.
4
Very Good
Clearly separates evidence from inference and states confidence unprompted.
5
Excellent
Distinguishes signal from participant noise, states confidence levels honestly, and resists overclaiming from six sessions or a single A/B cohort.
04
Evaluation factor
Reporting and testimony
15% weight
Assess how they deliver findings to designers and producers: highlight reels, severity-ranked issue lists, one-page readouts, and holding a position when a creative director pushes back.
Evidence to listen for
Writes findings that a non-specialist can act on
Has presented or defended work to an external audience: court, client, review board, publication
Withstands challenge without overclaiming or retreating
Keeps records that hold up to scrutiny
Five-point scoring guide
1
Poor
Cannot communicate findings; records would not withstand review.
2
Needs Improvement
Reporting is unclear or incomplete; avoids external scrutiny.
3
Satisfactory
Adequate reports; limited experience defending work externally.
4
Very Good
Clear reporting and real experience presenting to an external audience.
5
Excellent
Delivers prioritised, actionable findings tied to design intent, uses clips as evidence, and negotiates scope with designers without softening uncomfortable results.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.