Technical proficiency
Check hands-on depth in evaluation and safety tooling: adversarial robustness libraries, fairness metrics such as equalised odds, interpretability methods (SHAP, integrated gradients), and eval harness code they wrote.
Evidence to listen for
- Command of the languages, frameworks, and data tools the role actually uses
- Understands correctness, performance, and failure modes, not just syntax
- Has opinions on testing and can justify them
- Reads and reasons about code they did not write
Five-point scoring guide
Cannot work independently; fundamentals are missing.
Weak fundamentals; output needs heavy review.
Competent for the role; needs guidance on complex or unfamiliar work.
Strong practitioner; handles hard problems with little guidance.
Names specific metrics, attack methods and libraries used, and explains why each was chosen over cheaper alternatives for that model.