Technical proficiency
Check command of both sides: Meta Model and Milton Model language patterns, plus practical tooling such as prompt templating, Hugging Face datasets, Label Studio or an internal RLHF annotation stack.
Evidence to listen for
- Command of the languages, frameworks, and data tools the role actually uses
- Understands correctness, performance, and failure modes, not just syntax
- Has opinions on testing and can justify them
- Reads and reasons about code they did not write
Five-point scoring guide
Cannot work independently; fundamentals are missing.
Weak fundamentals; output needs heavy review.
Competent for the role; needs guidance on complex or unfamiliar work.
Strong practitioner; handles hard problems with little guidance.
Names specific linguistic patterns (presuppositions, embedded commands, reframes) and shows how each was encoded into prompts, rubrics or labelled turns.