public sector communityaccessibilityai governancealgorithmic biaseu ai act
Complete evaluation framework
What to assess and how to score it
Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Outcomes that landed
30% weight
Check what changed because of their advocacy: a model card rewritten, a biased training set replaced, WCAG fixes shipped, or a launch paused pending a fairness audit.
Evidence to listen for
Names programmes or initiatives that were adopted, funded, or delivered
States their own role rather than the department's
Gives measured reach or impact
Distinguishes work that landed from work that stalled, and explains why
Five-point scoring guide
1
Poor
No delivered work; describes intent and process only.
2
Needs Improvement
Involved in initiatives but cannot say what resulted or what they owned.
3
Satisfactory
Real delivery with adequate ownership; impact described loosely.
4
Very Good
Named outcomes with clear personal scope and some measures.
5
Excellent
Names specific systems they influenced, the disparity metric before and after, and who signed off on the remediation.
02
Evaluation factor
Stakeholder facilitation
25% weight
Probe how they hold a room with ML engineers, product owners, disability groups and affected communities at once, including participatory design sessions or community review panels they ran.
Evidence to listen for
Brings a real contested case, not a philosophy of engagement
Names the competing interests and the resolution method
Uses concrete engagement formats and can point to input that changed a decision
Treats every group as legitimate
Five-point scoring guide
1
Poor
Diplomacy-speak with no case attached, or contempt for one group.
2
Needs Improvement
Recalls conflict but no method; engagement is a box to check.
3
Satisfactory
Real case and workable approach; resolution thin on specifics.
4
Very Good
Names the tension and method; cites engagement that shaped the outcome.
5
Excellent
Describes translating lived-experience testimony into concrete backlog items engineers accepted, naming the friction and how it resolved.
03
Evaluation factor
Regulatory and policy command
25% weight
Test working command of the EU AI Act risk tiers, NIST AI RMF, Section 508 and WCAG 2.2, plus emerging state rules on automated employment decision tools.
Evidence to listen for
Names the statutes, funding rules, and processes they have worked under
Explains how those requirements sequenced their work
Owns the compliance thinking rather than deferring it entirely
Knows where the discretion sits
Five-point scoring guide
1
Poor
Outsources all regulatory thinking; cannot name a framework.
2
Needs Improvement
Generalities about compliance; no sequencing or named rules.
3
Satisfactory
Knows the main frameworks; sequencing described loosely.
4
Very Good
Names relevant frameworks and how they shaped a timeline.
5
Excellent
Cites obligations by name, distinguishes binding law from voluntary framework, and knows which internal artefacts satisfy each.
04
Evaluation factor
Evidence and reporting
20% weight
Assess how they evidence harm: disaggregated performance metrics, demographic parity or equalised odds testing, red-team findings, incident logs, and how results reached executives or regulators.
Evidence to listen for
Uses data to choose between options, not to justify a decision already made
Names the sources and methods behind their numbers
Reports to funders, councils, or the public in terms those audiences can use
Tracks whether the intervention worked
Five-point scoring guide
1
Poor
No use of evidence; decisions are assertion.
2
Needs Improvement
Cites data but cannot explain its source or limits.
3
Satisfactory
Uses evidence competently; evaluation after the fact is thin.
4
Very Good
Evidence drives choices and is reported clearly to lay audiences.
5
Excellent
Shows a real fairness assessment they authored, explains metric choice and limits, and reports uncomfortable findings unfiltered.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.