Why pre-screen resilience testing governance roles before the interview
Deliberately failing a production system is a legitimate engineering practice and it uses real customers as the test population. That makes the governance question concrete: how small is the blast radius, who could be harmed, and what stops an experiment that is going further than intended. Officers worth hiring have halted one. A short screen asks about that, and about what came after.
What actually matters when screening Chaos Engineering Ethics Officer candidates
- 01
Technical depth
Check fluency with fault injection tooling (Gremlin, AWS FIS, Chaos Mesh, Litmus), blast radius scoping, abort criteria, steady-state hypotheses, and how error budgets gate experiment approval.
- 02
Real incidents and findings
Probe game days they reviewed or vetoed: what experiment leaked to real customers, which region or tenant was affected, and what the post-incident ethics review changed.
- 03
Risk judgement
Test how they weigh learning value against user harm: latency injection on payment paths, experiments touching PII stores, healthcare or safety-adjacent traffic, and consent or notification thresholds.
- 04
Getting things fixed
Assess how they got engineers to adopt review gates: pre-registration templates, on-call sign-off, dry runs in staging, and turnaround times that avoided becoming a bottleneck.
Pre-screening questions to ask Chaos Engineering Ethics Officer candidates
12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.
Experiments they governed
3 questions01What is your understanding of the ethical questions raised by resilience testing?
Listen forConcrete issues named such as unconsented customer impact and effects on staff during an experiment.
Answered as general principles, or customer impact not identified as the central question.
02What policies would you establish to govern these practices?
Listen forApproval thresholds tied to blast radius, with defined limits on time, traffic share and services affected.
Policy described as principles, or approval required for everything regardless of scale.
03Can you give an example where poor judgement in this area caused harm?
Listen forA specific incident with the mechanism understood, whether from their own work or the public record.
Hypothetical examples only, or no awareness of incidents where testing caused real outages.
Blast radius limited
3 questions04How do you balance resilience testing against the risk to customers?
Listen forBlast radius quantified as a proportion of traffic and duration, with automatic stop conditions defined.
Blast radius described qualitatively, or no automatic termination when error rates rise.
05What steps would you take to reduce risk during these experiments?
Listen forStaged rollout from non-production, with a rollback tested before the experiment runs anywhere real.
Experiments run directly in production, or rollback assumed to work without being exercised.
06How do you ensure experiments do not disproportionately affect vulnerable users?
Listen forTraffic selection examined for who it actually includes, with critical user paths excluded deliberately.
Random traffic selection assumed fair, or safety-critical user journeys not excluded.
Impact assessed first
3 questions07How do you ensure transparency about these experiments?
Listen forInternal teams and support notified in advance, with a route to report unexpected effects quickly.
Experiments run without notifying support, or incidents investigated as real by unaware teams.
08What is your position on consent for these experiments?
Listen forAn honest position that customer consent is impractical, with that treated as a reason for tighter limits.
Consent claimed through terms of service, or the absence of consent not treated as significant.
09How do you handle sensitive data during these experiments?
Listen forData handling unchanged by the experiment, with no relaxation of controls to make testing easier.
Controls relaxed during experiments, or production data copied for experiment analysis.
When harm happens
3 questions10How would you handle an experiment that caused unintended harm?
Listen forImmediate stop and remediation, with affected customers informed and the incident reviewed openly.
Harm treated as an acceptable cost, or customer impact not disclosed to those affected.
11How would you respond to concerns raised by employees about an experiment?
Listen forConcerns treated as information with the experiment paused while they are assessed.
Concerns overruled by the engineering team, or no route for someone to stop an experiment.
12How do you ensure accountability for these practices?
Listen forA named owner per experiment with decisions recorded, so responsibility is traceable afterwards.
Accountability described collectively, or no record of who approved a given experiment.
How to score responses
Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.
Technical depth
35%5Names specific injection primitives, defines steady-state metrics and automatic halt conditions, and ties experiment scope to error budget headroom.
Real incidents and findings
30%5Recounts named experiments including one that breached blast radius, with the customer impact numbers and the guardrail added afterwards.
Risk judgement
20%5Separates recoverable degradation from irreversible harm, states when they refuse an experiment outright, and justifies notification thresholds concretely.
Getting things fixed
15%5Describes a review workflow teams actually used, with adoption or cycle-time figures and evidence of guardrails landing in pipeline code.
Deliberately failing production uses real customers who did not agree to it. A one-way video screen asks how the reach is limited.
Try it on HirevireScreening FAQ
Process basics
How long should a pre-screening round for this role take?
Fifteen minutes across eight to ten questions, answered async. Enough to establish experiments they governed, test their blast radius thinking, and hear how harm was handled.
How technical does this role need to be?
Technical enough to assess an experiment design and challenge a blast radius claim. Someone working from policy alone will approve experiments whose actual reach they cannot evaluate.
Evaluating answers
What is the strongest signal when screening this role?
An experiment that went further than planned. Officers with real experience have one and describe the stop and the review. Anyone whose experiments always stayed contained has governed very few.
How do I judge their impact assessment?
Ask who could be harmed by an experiment. Real answers name specific user groups and dependent services. Anyone answering in general terms has approved experiments without assessing reach.
























