Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Technique and experimental design
35% weight
Check hands-on sampling and identification methods: manta or neuston trawls, Niskin bottles, sediment cores, density separation in ZnCl2, FTIR or Raman polymer confirmation, and blank controls for airborne fibre contamination.
Evidence to listen for
Runs the assays and instruments themselves rather than describing what a team does
Designs experiments with controls, replicates, and a stated hypothesis
Knows what each technique can and cannot resolve
Understands the science, not only the protocol
Five-point scoring guide
1
Poor
Protocol follower with no experimental design; cannot justify controls.
2
Needs Improvement
Runs standard assays; designs experiments poorly or not at all.
3
Satisfactory
Competent at the bench with sound routine design.
4
Very Good
Designs rigorous experiments and understands the limits of each technique.
5
Excellent
Names mesh sizes, tow durations and separation densities used, and explains why procedural blanks and Nile Red staining were chosen over alternatives.
02
Evaluation factor
Results that went somewhere
25% weight
Probe where their data landed: peer-reviewed papers, MSFD or NOAA marine debris datasets, OSPAR beach litter reports, contributions to plastic treaty submissions or municipal waste policy.
Evidence to listen for
Names projects where their results changed a decision, a process, or a product
States their own contribution rather than the group's
Has taken something from bench to a larger scale, a filing, or a publication
Knows what happened to the work after they handed it over
Five-point scoring guide
1
Poor
No results that went anywhere; work is entirely exploratory.
2
Needs Improvement
Contributed to projects but cannot say what their data changed.
3
Satisfactory
Real contributions; outcomes described loosely.
4
Very Good
Names results that changed a decision, with clear personal scope.
5
Excellent
Points to specific published datasets, first or co-authored papers, and a decision, regulation or cleanup design that changed because of their findings.
03
Evaluation factor
Troubleshooting and reproducibility
25% weight
Test how they handle messy field and lab reality: fouled nets in heavy swell, spectral libraries misidentifying weathered polymers, low recovery rates, and inter-lab variability in particle counts.
Evidence to listen for
Treats a failed run as information rather than bad luck
Isolates reagent, instrument, operator, and biological causes systematically
Knows why a result failed to reproduce and can say when their own data was wrong
Keeps records good enough to diagnose from months later
Five-point scoring guide
1
Poor
Repeats failed runs unchanged; no diagnostic thinking.
2
Needs Improvement
Troubleshoots by substitution; cannot explain a reproducibility failure.
3
Satisfactory
Solid troubleshooting on familiar assays.
4
Very Good
Systematic isolation of causes, and honest about their own irreproducible results.
5
Excellent
Describes a specific failure (contaminated blanks, ambiguous spectra) plus the diagnostic steps and method revision that restored trustworthy counts.
04
Evaluation factor
Documentation and collaboration
15% weight
Assess record-keeping and teamwork: field logbooks, R or Python scripts on GitHub, metadata to Darwin Core or ERDDAP standards, plus work with ship crews, volunteers and coastal communities.
Evidence to listen for
Keeps records to the standard the setting requires, whether that is GLP, GMP, or a defensible notebook
Writes up so someone else can repeat the work
Works with process, quality, or clinical colleagues rather than in a bench silo
Explains a result to a non-specialist without overclaiming
Five-point scoring guide
1
Poor
Records would not survive audit; work is not repeatable from them.
2
Needs Improvement
Documentation is thin; write-ups need heavy editing.
3
Satisfactory
Adequate records and write-ups; collaboration is limited.
4
Very Good
Audit-standard records and clear communication across functions.
5
Excellent
Shares reproducible analysis code and archived metadata, and describes training volunteers or crew to sample consistently across a multi-site campaign.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.