Interview scorecard template

Responsible AI Advocate interview scorecard

Pre-screening scorecard for Responsible AI Advocate candidates.

See AI scoring
public sector communityai ethicsalgorithmic bias auditeu ai actnist ai rmf
Complete evaluation framework

What to assess and how to score it

Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.

01
Evaluation factor

Outcomes that landed

30% weight

Ask what changed because of their advocacy: a model blocked at review, a bias mitigation shipped, a model card standard adopted, or an internal AI use policy signed off.

Evidence to listen for

  • Names programmes or initiatives that were adopted, funded, or delivered
  • States their own role rather than the department's
  • Gives measured reach or impact
  • Distinguishes work that landed from work that stalled, and explains why

Five-point scoring guide

1
Poor

No delivered work; describes intent and process only.

2
Needs Improvement

Involved in initiatives but cannot say what resulted or what they owned.

3
Satisfactory

Real delivery with adequate ownership; impact described loosely.

4
Very Good

Named outcomes with clear personal scope and some measures.

5
Excellent

Names specific systems altered or halted, with dates, decision forums, and the residual risk accepted by named owners.

02
Evaluation factor

Stakeholder facilitation

25% weight

Probe how they worked ML engineers, legal counsel, procurement, and affected user groups through disagreement on a contested model or dataset without stalling the roadmap.

Evidence to listen for

  • Brings a real contested case, not a philosophy of engagement
  • Names the competing interests and the resolution method
  • Uses concrete engagement formats and can point to input that changed a decision
  • Treats every group as legitimate

Five-point scoring guide

1
Poor

Diplomacy-speak with no case attached, or contempt for one group.

2
Needs Improvement

Recalls conflict but no method; engagement is a box to check.

3
Satisfactory

Real case and workable approach; resolution thin on specifics.

4
Very Good

Names the tension and method; cites engagement that shaped the outcome.

5
Excellent

Describes running review boards or red-team workshops where engineers and counsel reached a documented decision both sides could defend.

03
Evaluation factor

Regulatory and policy command

25% weight

Test command of the EU AI Act risk tiers, NIST AI RMF, ISO/IEC 42001, GDPR Article 22, and sector rules such as NYC Local Law 144 or FDA SaMD guidance.

Evidence to listen for

  • Names the statutes, funding rules, and processes they have worked under
  • Explains how those requirements sequenced their work
  • Owns the compliance thinking rather than deferring it entirely
  • Knows where the discretion sits

Five-point scoring guide

1
Poor

Outsources all regulatory thinking; cannot name a framework.

2
Needs Improvement

Generalities about compliance; no sequencing or named rules.

3
Satisfactory

Knows the main frameworks; sequencing described loosely.

4
Very Good

Names relevant frameworks and how they shaped a timeline.

5
Excellent

Maps a given product to the correct risk tier and cites conformity, transparency, and human oversight duties without hedging.

04
Evaluation factor

Evidence and reporting

20% weight

Look for how they measured harm: disparate impact ratios, subgroup error rates, red-team findings, incident logs, and what those reports triggered downstream.

Evidence to listen for

  • Uses data to choose between options, not to justify a decision already made
  • Names the sources and methods behind their numbers
  • Reports to funders, councils, or the public in terms those audiences can use
  • Tracks whether the intervention worked

Five-point scoring guide

1
Poor

No use of evidence; decisions are assertion.

2
Needs Improvement

Cites data but cannot explain its source or limits.

3
Satisfactory

Uses evidence competently; evaluation after the fact is thin.

4
Very Good

Evidence drives choices and is reported clearly to lay audiences.

5
Excellent

Shows fairness metrics tied to a defined protected-attribute methodology, plus reports that drove remediation rather than sitting in a drive.

Put this rubric to work

Score every candidate against the same standard

Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.

Explore AI Scorecards