Evaluate AI Governance Manager candidates across 4 weighted areas: technical depth, real incidents and findings, risk judgement, and getting things fixed. Technical depth leads at 35%, so check command of the EU AI Act risk tiers, ISO/IEC 42001, NIST AI RMF and how they map model cards, DPIAs and bias testing into. Use the rubric to compare role-specific evidence consistently.
security complianceai governanceeu ai actmodel risknist ai rmf
TL;DR
For technical depth, look for evidence the candidate cites specific clauses and controls, distinguishes high risk from limited risk systems, and knows where ISO 42001 overlaps ISO 27001. For real incidents and findings, look for evidence the candidate walks through concrete assessments with dates, systems, and findings, including a case where they blocked or conditioned a deployment.
Apply the written 1–5 anchors to every answer, record the evidence behind each rating, and use the factor weights to reach a consistent overall assessment.
Complete evaluation framework
What to assess and how to score it
Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Technical depth
35% weight
Check command of the EU AI Act risk tiers, ISO/IEC 42001, NIST AI RMF and how they map model cards, DPIAs and bias testing into an actual control set.
Evidence to listen for
Command of the specific attack surface, tooling, and controls the role covers
Understands how the underlying system works, not just how the tool reports on it
Can explain an attack or control chain end to end
Distinguishes what they found themselves from what a scanner flagged
Five-point scoring guide
1
Poor
Tool operator only; no understanding of the systems underneath.
2
Needs Improvement
Runs tooling but cannot explain findings or how the attack works.
3
Satisfactory
Solid working knowledge; depth thins outside familiar tooling.
4
Very Good
Strong command of the domain; explains attack and control chains clearly.
5
Excellent
Cites specific clauses and controls, distinguishes high risk from limited risk systems, and knows where ISO 42001 overlaps ISO 27001.
02
Evaluation factor
Real incidents and findings
30% weight
Probe named reviews they ran: which models or vendors they assessed, findings raised on training data provenance, drift, or explainability, and what evidence they demanded.
Evidence to listen for
Brings specific incidents, findings, or audits they personally worked
States their own role rather than the team's
Describes what was actually at risk and what changed afterwards
Can talk about a finding that turned out to be wrong
Five-point scoring guide
1
Poor
No hands-on work; knowledge is entirely certification or coursework.
2
Needs Improvement
Limited exposure; cannot describe their contribution to an incident.
3
Satisfactory
Real casework with adequate detail; ownership sometimes vague.
4
Very Good
Specific incidents with clear personal scope and what changed after.
5
Excellent
Walks through concrete assessments with dates, systems, and findings, including a case where they blocked or conditioned a deployment.
03
Evaluation factor
Risk judgement
20% weight
Test how they triage: ranking a customer-facing LLM against an internal HR screening tool, tolerating residual risk, and defending calls to legal, product, and the board.
Evidence to listen for
Prioritises by actual exploitability and business impact, not raw severity scores
Can argue for accepting a risk as well as fixing it
Knows the difference between a finding and a problem
Does not cry wolf or wave things through
Five-point scoring guide
1
Poor
Treats every finding as critical, or waves real risk through.
2
Needs Improvement
Follows severity scores mechanically; no business context.
3
Satisfactory
Reasonable prioritisation; less confident arguing for risk acceptance.
4
Very Good
Prioritises by exploitability and impact; can justify accepting a risk.
5
Excellent
Ranks by harm and exposure rather than novelty, states residual risk explicitly, and names who owns acceptance of it.
04
Evaluation factor
Getting things fixed
15% weight
Assess how they moved engineers and product owners to act: intake workflows, gate reviews, register tooling (OneTrust, Credo, internal), remediation deadlines and closure rates.
Evidence to listen for
Writes findings engineers can act on rather than a wall of output
Has persuaded a team to fix something they did not want to fix
Explains risk to executives in business terms
Works with the org rather than policing it
Five-point scoring guide
1
Poor
Adversarial with engineering; findings never get fixed.
2
Needs Improvement
Reports are unactionable; no influence beyond raising tickets.
3
Satisfactory
Adequate reporting; relies on mandate rather than persuasion.
4
Very Good
Actionable findings and a real record of getting fixes shipped.
5
Excellent
Describes a governance process teams actually used, with adoption numbers, and remediation items closed rather than logged indefinitely.
Evidence-led prompts
Interview questions for a AI Governance Manager
Use these prompts to surface evidence for the weighted factors above and compare candidates against the same role-specific criteria.
01
Which AI tools and technologies are you familiar with, and how deeply?
02
Do you have experience auditing AI systems for compliance?
03
How familiar are you with current laws and regulations relating to AI use?
04
Can you give an example where you identified a potential risk or policy violation?
05
Can you discuss a time you had to manage an AI-related incident?