leadership strategyai strategyeu ai actmlops governancemodel risk
Complete evaluation framework
What to assess and how to score it
Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Record of outcomes
35% weight
Ask what AI systems reached production under their remit: model families, users served, inference cost per call, and revenue or cost impact tied to a P&L line.
Evidence to listen for
Names organisations, budgets, and headcount they were actually accountable for
Gives outcomes with numbers, not initiatives with adjectives
Separates what they drove from what the market or a predecessor did
Can describe something they owned that failed
Five-point scoring guide
1
Poor
Titles and initiatives only; no accountable outcomes.
2
Needs Improvement
Claims results the organisation would have had anyway; no scope clarity.
3
Satisfactory
Real accountability with some numbers; attribution occasionally generous.
4
Very Good
Clear scope and measured outcomes, including an honest failure.
5
Excellent
Names shipped systems (recommender, LLM assistant, forecasting) with adoption numbers, unit economics, and the business metric each moved.
02
Evaluation factor
Strategic judgement
25% weight
Probe how they chose build versus buy versus fine-tune, GPU capacity commitments, vendor lock-in on foundation models, and which AI initiatives they deliberately killed.
Evidence to listen for
Explains a decision where the options were genuinely close and the information incomplete
Can say what they chose not to do and why
Distinguishes a bet from a certainty
Changes course on evidence rather than defending a position past its life
Five-point scoring guide
1
Poor
Frameworks and slogans; no real decision they can walk through.
2
Needs Improvement
Describes decisions made elsewhere; cannot state the trade-off.
3
Satisfactory
Sound judgement on familiar decisions; less tested on ambiguous ones.
4
Very Good
Walks a genuinely close call, states what they gave up, and changed course on evidence.
5
Excellent
Explains a portfolio with clear kill criteria, defends a costly platform bet, and shows where classical methods beat generative approaches.
03
Evaluation factor
Building and leading teams
25% weight
Examine how they staffed research scientists, ML engineers, and data platform teams; ask about retention, leveling, and the split between central AI and embedded squads.
Evidence to listen for
Has hired, developed, and where necessary removed people
Names someone who grew under them and what they did to cause it
Handles an underperformer directly rather than waiting it out
Builds a team that functions when they are not in the room
Five-point scoring guide
1
Poor
No real people leadership; avoids difficult personnel decisions.
2
Needs Improvement
Managed a team but cannot describe developing or exiting anyone.
3
Satisfactory
Competent manager; development is informal.
4
Very Good
Demonstrable record of growing people and handling underperformance directly.
5
Excellent
Describes hiring senior ML talent against big-tech offers, a working operating model, and named people promoted into leadership.
04
Evaluation factor
Influence across the business
15% weight
Test how they handled the board, legal, and regulators: EU AI Act readiness, model risk documentation, and persuading skeptical business unit heads to adopt.
Evidence to listen for
Wins support from peers and boards without positional authority
Translates their function into terms the rest of the business cares about
Delivers unwelcome news early
Manages up without either capitulating or stonewalling
Five-point scoring guide
1
Poor
Relies entirely on authority; conceals bad news.
2
Needs Improvement
Struggles to influence peers; communicates in function-specific jargon.
3
Satisfactory
Works adequately with peers and leadership.
4
Very Good
Persuades peers and boards on merit and delivers bad news early.
5
Excellent
Cites board-level AI governance they authored, resolved a real compliance or safety escalation, and won over a resistant business owner.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.