Why pre-screen IT resilience managers before the interview
The standard failure is a plan that has never been run. It names systems that were decommissioned, contacts who left, and recovery times nobody has measured, and it looks entirely credible until the day it is needed. Managers worth hiring exercise plans, publish what the exercise found, and fix it. A short screen asks what the last test revealed, which is a question that separates real practice from document maintenance immediately.
What actually matters when screening IT Resilience Manager candidates
- 01
Technical depth
Check command of RTO and RPO setting, BIA methodology, ISO 22301 or NIST SP 800-34, failover architectures, backup immutability, and dependency mapping across cloud and on-prem estates.
- 02
Real incidents and findings
Probe actual invocations: ransomware recovery, datacentre or region outage, failed failover. Ask about the DR test schedule they ran, tabletop exercises, and post-incident actions tracked to closure.
- 03
Risk judgement
Assess how they rank single points of failure, third party and SaaS concentration risk, and where they accepted residual risk rather than funding hot standby.
- 04
Getting things fixed
Look for evidence of moving application owners, infrastructure teams and vendors to update runbooks, close audit findings, and meet regulator expectations such as DORA or FCA operational resilience.
Pre-screening questions to ask IT Resilience Manager candidates
12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.
Plans they exercised
4 questions01Can you tell us about your previous IT resilience planning experience?
Listen forScope in systems and services covered, with what they owned rather than what a committee approved.
Experience described by framework used, or a remit that produced documents and no exercises.
02Describe a situation where you successfully implemented a disaster recovery plan.
Listen forA plan built with the technical teams and tested, including recovery times measured rather than estimated.
Recovery times taken from engineering estimates with no test, or a plan implemented but never run.
03How do you test the resilience of an IT system?
Listen forExercises that go beyond a walkthrough, with a failure actually induced and what the test revealed.
Testing limited to tabletop discussion, or exercises that have never surfaced a problem.
04What strategies have you used to ensure service continuity following a disruption?
Listen forContinuity arrangements that were used in practice, including manual workarounds agreed with the business.
Continuity described entirely as technical failover, with no manual process for the interim.
Real disruptions
2 questions05Tell us about a time when you managed the impact of a critical system failure.
Listen forA real failure with their own decisions, covering communication, prioritisation and what changed afterwards.
Failures described from a plan rather than experience, or no structural change after a serious event.
06Can you explain your experience coordinating major incident response activities?
Listen forA defined role during incidents with stakeholder updates on a stated cadence while the team works.
Coordination that competes with the technical response, or updates given only once resolved.
Supplier dependencies
3 questions07What do you consider important when evaluating the resilience of a third-party supplier?
Listen forConcentration and dependency mapped, with contractual recovery commitments tested rather than accepted.
Supplier resilience accepted from a questionnaire, or no view on what happens if a key supplier fails.
08Can you talk about your experience with cloud technologies and their role in resilience?
Listen forAwareness that cloud availability is not the same as their own recoverability, with region and account risk considered.
Cloud treated as inherently resilient, or backups held in the same account as production.
09How familiar are you with risk management, business continuity and disaster recovery?
Listen forThe distinctions applied practically, with business impact analysis driving technical priorities rather than the reverse.
Terms used interchangeably, or recovery priorities set by technical convenience rather than business impact.
Agreeing what it costs
3 questions10How would you handle conflicting opinions from stakeholders during recovery planning?
Listen forPriorities resolved with business impact evidence, with a decision recorded and an owner named.
Every service classed as critical, or conflicts escalated with no recommendation attached.
11What is your approach to documenting and communicating resilience plans?
Listen forDocumentation accessible during an outage, including offline copies and contacts verified on a schedule.
Plans stored only in a system that would be unavailable during the incident they cover.
12How do you continuously monitor, review and improve resilience strategies?
Listen forA review cadence tied to change, with a specific gap found through review and closed afterwards.
Annual review as the only mechanism, or plans that did not change when the estate did.
How to score responses
Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.
Technical depth
35%5Explains how they derived tiered RTO/RPO from a business impact analysis and matched them to real replication and backup designs.
Real incidents and findings
30%5Recounts a live invocation or full failover test with timings, what missed target, and the remediation items they drove afterwards.
Risk judgement
20%5Ties resilience investment to quantified impact per hour of downtime and defends deliberate acceptance of specific residual risks.
Getting things fixed
15%5Names owners, deadlines and closure rates for resilience gaps, showing runbooks and impact tolerances that were actually kept current.
A plan nobody has run names decommissioned systems and people who left, and reads perfectly. A one-way video screen asks what the last exercise found.
Try it on HirevireScreening FAQ
Process basics
How long should a pre-screening round for this role take?
Fifteen minutes across eight to ten questions, answered async. Enough to establish what they exercised, hear one real disruption they managed, and check how they assess supplier and cloud dependencies.
How much technical depth does this role need?
Enough to challenge a recovery time estimate. A manager who accepts engineering estimates without testing them will publish objectives the organisation cannot meet, which is worse than having none.
Evaluating answers
What is the strongest signal when screening this role?
What their last exercise found. Real tests always reveal something: a missing dependency, an out-of-date contact, a recovery that took three times the estimate. Anyone whose exercises pass cleanly is running a walkthrough.
How do I judge their handling of cost conversations?
Ask about a recovery objective the business would not fund. Good answers document the accepted risk with the decision-maker named. Anyone who simply published the objective anyway has created a false assurance.
























