Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Technique and experimental design
35% weight
Probe how they design studies: within versus between subjects, counterbalancing, sample sizing, and instruments such as SUS, NASA-TLX, think-aloud protocols, eye tracking or Tobii heatmap setups.
Evidence to listen for
Runs the assays and instruments themselves rather than describing what a team does
Designs experiments with controls, replicates, and a stated hypothesis
Knows what each technique can and cannot resolve
Understands the science, not only the protocol
Five-point scoring guide
1
Poor
Protocol follower with no experimental design; cannot justify controls.
2
Needs Improvement
Runs standard assays; designs experiments poorly or not at all.
3
Satisfactory
Competent at the bench with sound routine design.
4
Very Good
Designs rigorous experiments and understands the limits of each technique.
5
Excellent
Names the study design and rationale, defends sample size and task order, and distinguishes formative usability work from controlled comparative experiments.
02
Evaluation factor
Results that went somewhere
25% weight
Ask which findings changed a shipped interface: navigation redesigns, error rate drops, task completion or time-on-task gains, WCAG 2.2 fixes that cleared an audit.
Evidence to listen for
Names projects where their results changed a decision, a process, or a product
States their own contribution rather than the group's
Has taken something from bench to a larger scale, a filing, or a publication
Knows what happened to the work after they handed it over
Five-point scoring guide
1
Poor
No results that went anywhere; work is entirely exploratory.
2
Needs Improvement
Contributed to projects but cannot say what their data changed.
3
Satisfactory
Real contributions; outcomes described loosely.
4
Very Good
Names results that changed a decision, with clear personal scope.
5
Excellent
Cites specific products and measured deltas, plus the recommendation engineers or designers actually implemented after the study.
03
Evaluation factor
Troubleshooting and reproducibility
25% weight
Test how they handle noisy or contradictory data: pilot failures, participant dropout, confounded prototypes, disagreement between behavioural logs and self-report scores.
Evidence to listen for
Treats a failed run as information rather than bad luck
Isolates reagent, instrument, operator, and biological causes systematically
Knows why a result failed to reproduce and can say when their own data was wrong
Keeps records good enough to diagnose from months later
Five-point scoring guide
1
Poor
Repeats failed runs unchanged; no diagnostic thinking.
2
Needs Improvement
Troubleshoots by substitution; cannot explain a reproducibility failure.
3
Satisfactory
Solid troubleshooting on familiar assays.
4
Very Good
Systematic isolation of causes, and honest about their own irreproducible results.
5
Excellent
Describes re-running pilots, triangulating telemetry against interviews, and openly reports where effects did not replicate or were underpowered.
04
Evaluation factor
Documentation and collaboration
15% weight
Look for study protocols, consent and ethics submissions (IRB or equivalent), tagged qualitative codebooks, and how they hand findings to product managers and front-end engineers.
Evidence to listen for
Keeps records to the standard the setting requires, whether that is GLP, GMP, or a defensible notebook
Writes up so someone else can repeat the work
Works with process, quality, or clinical colleagues rather than in a bench silo
Explains a result to a non-specialist without overclaiming
Five-point scoring guide
1
Poor
Records would not survive audit; work is not repeatable from them.
2
Needs Improvement
Documentation is thin; write-ups need heavy editing.
3
Satisfactory
Adequate records and write-ups; collaboration is limited.
4
Very Good
Audit-standard records and clear communication across functions.
5
Excellent
Keeps traceable protocols and coded transcripts, and turns them into prioritised, developer-legible recommendations rather than a slide deck of quotes.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.