Theoretical command
Probe command of PPO, SAC and Q-learning derivations: ask about bias variance in GAE, off-policy correction, entropy regularisation, and why value bootstrapping destabilises long horizons.
Evidence to listen for
- Explains the underlying theory at the level the role demands, and can go a layer deeper when pushed
- Knows which results are established and which are contested
- Distinguishes their own contribution from the field's
- Comfortable saying where the theory runs out
Five-point scoring guide
Recites terminology without understanding; cannot go one layer deeper.
Surface familiarity; conflates established results with speculation.
Solid grasp of the core theory; thin at the frontier.
Strong command; separates settled results from open questions.
Derives policy gradient variants from scratch, names failure modes of each, and cites specific papers behind their algorithm choices.