Pre-Screening Interview Questions to Ask a Chaos Engineering Ethics Officer

Last updated on

Deliberately breaking production tests resilience using real customers who did not agree to it. These questions test blast radius control and what happens when harm occurs.

TL;DR, what to screen for

The best pre-screening questions for a chaos engineering ethics officer test four things: experiments they governed rather than principles they hold, whether blast radius is genuinely limited, whether customer and employee impact is assessed before running, and what happens when an experiment causes harm. Ask about one that went further than planned.

  • Experiments they governed
  • Blast radius limited
  • Impact assessed first
  • When harm happens

Why pre-screen resilience testing governance roles before the interview

Deliberately failing a production system is a legitimate engineering practice and it uses real customers as the test population. That makes the governance question concrete: how small is the blast radius, who could be harmed, and what stops an experiment that is going further than intended. Officers worth hiring have halted one. A short screen asks about that, and about what came after.

What actually matters when screening Chaos Engineering Ethics Officer candidates

  1. 01

    Technical depth

    Check fluency with fault injection tooling (Gremlin, AWS FIS, Chaos Mesh, Litmus), blast radius scoping, abort criteria, steady-state hypotheses, and how error budgets gate experiment approval.

  2. 02

    Real incidents and findings

    Probe game days they reviewed or vetoed: what experiment leaked to real customers, which region or tenant was affected, and what the post-incident ethics review changed.

  3. 03

    Risk judgement

    Test how they weigh learning value against user harm: latency injection on payment paths, experiments touching PII stores, healthcare or safety-adjacent traffic, and consent or notification thresholds.

  4. 04

    Getting things fixed

    Assess how they got engineers to adopt review gates: pre-registration templates, on-call sign-off, dry runs in staging, and turnaround times that avoided becoming a bottleneck.

Pre-screening questions to ask Chaos Engineering Ethics Officer candidates

12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.

Experiments they governed

3 questions
  1. 01What is your understanding of the ethical questions raised by resilience testing?

    Listen for

    Concrete issues named such as unconsented customer impact and effects on staff during an experiment.

    Answered as general principles, or customer impact not identified as the central question.

  2. 02What policies would you establish to govern these practices?

    Listen for

    Approval thresholds tied to blast radius, with defined limits on time, traffic share and services affected.

    Policy described as principles, or approval required for everything regardless of scale.

  3. 03Can you give an example where poor judgement in this area caused harm?

    Listen for

    A specific incident with the mechanism understood, whether from their own work or the public record.

    Hypothetical examples only, or no awareness of incidents where testing caused real outages.

Blast radius limited

3 questions
  1. 04How do you balance resilience testing against the risk to customers?

    Listen for

    Blast radius quantified as a proportion of traffic and duration, with automatic stop conditions defined.

    Blast radius described qualitatively, or no automatic termination when error rates rise.

  2. 05What steps would you take to reduce risk during these experiments?

    Listen for

    Staged rollout from non-production, with a rollback tested before the experiment runs anywhere real.

    Experiments run directly in production, or rollback assumed to work without being exercised.

  3. 06How do you ensure experiments do not disproportionately affect vulnerable users?

    Listen for

    Traffic selection examined for who it actually includes, with critical user paths excluded deliberately.

    Random traffic selection assumed fair, or safety-critical user journeys not excluded.

Impact assessed first

3 questions
  1. 07How do you ensure transparency about these experiments?

    Listen for

    Internal teams and support notified in advance, with a route to report unexpected effects quickly.

    Experiments run without notifying support, or incidents investigated as real by unaware teams.

  2. 08What is your position on consent for these experiments?

    Listen for

    An honest position that customer consent is impractical, with that treated as a reason for tighter limits.

    Consent claimed through terms of service, or the absence of consent not treated as significant.

  3. 09How do you handle sensitive data during these experiments?

    Listen for

    Data handling unchanged by the experiment, with no relaxation of controls to make testing easier.

    Controls relaxed during experiments, or production data copied for experiment analysis.

When harm happens

3 questions
  1. 10How would you handle an experiment that caused unintended harm?

    Listen for

    Immediate stop and remediation, with affected customers informed and the incident reviewed openly.

    Harm treated as an acceptable cost, or customer impact not disclosed to those affected.

  2. 11How would you respond to concerns raised by employees about an experiment?

    Listen for

    Concerns treated as information with the experiment paused while they are assessed.

    Concerns overruled by the engineering team, or no route for someone to stop an experiment.

  3. 12How do you ensure accountability for these practices?

    Listen for

    A named owner per experiment with decisions recorded, so responsibility is traceable afterwards.

    Accountability described collectively, or no record of who approved a given experiment.

How to score responses

Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.

  1. Technical depth

    35%

    5Names specific injection primitives, defines steady-state metrics and automatic halt conditions, and ties experiment scope to error budget headroom.

  2. Real incidents and findings

    30%

    5Recounts named experiments including one that breached blast radius, with the customer impact numbers and the guardrail added afterwards.

  3. Risk judgement

    20%

    5Separates recoverable degradation from irreversible harm, states when they refuse an experiment outright, and justifies notification thresholds concretely.

  4. Getting things fixed

    15%

    5Describes a review workflow teams actually used, with adoption or cycle-time figures and evidence of guardrails landing in pipeline code.

Deliberately failing production uses real customers who did not agree to it. A one-way video screen asks how the reach is limited.

Try it on Hirevire

Screening FAQ

Process basics

How long should a pre-screening round for this role take?

Fifteen minutes across eight to ten questions, answered async. Enough to establish experiments they governed, test their blast radius thinking, and hear how harm was handled.

How technical does this role need to be?

Technical enough to assess an experiment design and challenge a blast radius claim. Someone working from policy alone will approve experiments whose actual reach they cannot evaluate.

Evaluating answers

What is the strongest signal when screening this role?

An experiment that went further than planned. Officers with real experience have one and describe the stop and the review. Anyone whose experiments always stayed contained has governed very few.

How do I judge their impact assessment?

Ask who could be harmed by an experiment. Real answers name specific user groups and dependent services. Anyone answering in general terms has approved experiments without assessing reach.

Go deeper on this role

Sanat Hegde
Sanat Hegde
Founder, Hirevire

Sanat has been hiring since 2012 and watching the recruitment industry change up close ever since, and turned that screening process into Hirevire's video screening platform. LinkedIn

Trusted by 500+ Companies

Screen Chaos Engineering Ethics Officer candidates on Hirevire

Turn this question list into an async video screen in minutes. Every applicant answers the same blast radius, impact and harm questions on camera, so you compare governance rather than principles.