Why pre-screen data center managers before the site walkthrough and panel interview
Pre-screening data center managers protects the most expensive hour on your calendar: the site walkthrough. Applicants arrive from colocation floors, enterprise server rooms, NOC shifts, and facilities contracting, and the resume flattens all of that into "managed data center operations". It will not tell you whether they ran 40 racks or 4,000, whether they owned the UPS and chiller maintenance contracts, or whether they have ever stood in front of a tenant after a failed transfer switch. Ten minutes of recorded answers settles it.
What actually matters when screening Data Center Manager candidates
- 01
Execution and reliability
Check the scale they have run: racks, power draw, uptime commitments, and what happened during their worst outage.
- 02
Improving the process
Test what they changed about capacity planning, power and cooling efficiency, or maintenance scheduling.
- 03
Judgement and autonomy
Assess how they decide during an incident when information is incomplete and every option has downtime.
- 04
Communication
Judge how they brief tenants, vendors, and executives during and after an outage.
Pre-screening questions to ask Data Center Manager candidates
12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.
Scale and uptime
4 questions01Walk us through the scale you have run: rack count, total power draw, and the uptime commitment you were held to.
Listen forConcrete figures (racks, kW per rack, total load, SLA tier or uptime percentage) plus what they personally owned versus what a landlord or vendor owned.
Vague scale claims with no numbers, or they cannot say who owned power, cooling, and the SLA.
02How do you ensure high availability and reliability in a data center, in terms of the redundancy you actually design and test?
Listen forSpecifics on N+1 or 2N topology, dual power paths, generator and UPS load bank testing, failover drills, and single points of failure they found and removed.
Talks about redundancy on paper but has never load tested a generator or run a live failover drill.
03How do you monitor the health and performance of a data center day to day?
Listen forNamed systems (BMS, DCIM, environmental sensors, SNMP traps, alerting thresholds), plus how alarms are triaged and who gets paged at what severity.
Relies on walking the floor or on tenant complaints as the primary way problems get discovered.
04Which data center infrastructure management tools have you used, and what did you actually use them for?
Listen forHands-on detail with tools such as Nlyte, Sunbird, EcoStruxure, or Netbox: asset records, rack elevations, circuit tracking, capacity dashboards they maintained.
Names tools they only viewed reports from, with no ownership of data accuracy or asset records.
Incident judgement
3 questions05Describe a time you had to handle a crisis at a data center, and take us through the timeline.
Listen forA clear sequence: detection, escalation, the call they made with incomplete information, the tradeoff accepted, restoration time, and the follow-up action.
Cannot recall a specific incident, or claims no serious incident ever happened on their watch.
06How do you prevent equipment failure, and what do you do when a critical unit fails anyway?
Listen forA real preventive maintenance calendar (UPS batteries, CRAC filters, chillers, transfer switches), vendor SLAs, spares held on site, and a failure they caught early.
Treats maintenance as deferrable, or has no spares strategy and no vendor response time commitments.
07What disaster recovery procedures have you implemented, and when were they last tested for real?
Listen forDefined RTO and RPO targets, documented runbooks, a secondary site or cloud failover, and dates of actual tests with findings that changed the plan.
A DR document that exists but has never been exercised, or no awareness of RTO and RPO.
Efficiency and capacity
4 questions08What strategies do you use to manage power and cooling, and what did they do to your PUE?
Listen forNamed interventions (hot and cold aisle containment, blanking panels, raised setpoints, variable speed fans, economiser hours) with before and after PUE or kWh figures.
Discusses cooling only as thermostat settings, with no PUE baseline or measured result.
09How would you handle a data center that is over capacity in power or under capacity in space?
Listen forCapacity modelling by circuit and rack, breaker headroom checks, phased decommissioning, densification limits, and when to trigger a build or colocation decision.
Answers only "add more racks" without checking power, cooling, or floor loading constraints.
10On camera, walk us through one project where you optimized a data center's efficiency: the baseline number, what you changed, and the number afterward.
Listen forA named project with a measured baseline, the specific change they drove, the resulting metric, and who they had to convince to approve the spend or downtime.
Describes an efficiency project with no baseline, no result, and no clarity on their own role in it.
11How do you measure whether a data center operation is running well, and which numbers do you report upward?
Listen forA short set of tracked metrics: unplanned downtime minutes, PUE, capacity utilisation, ticket resolution times, preventive maintenance completion rate, SLA credits issued.
No metrics at all, or reports only uptime while ignoring maintenance completion and capacity headroom.
Onboarding and logistics
1 question12If hired, what would your first initiative be as our data center manager, and what would you need from us in the first 30 days?
Listen forA grounded first step (asset audit, thermal survey, maintenance contract review, runbook refresh) plus questions about on-call rotation, maintenance windows, and vendor contracts.
Proposes a large capital rebuild before seeing the floor, or has no questions about on-call and shift coverage.
How to score responses
Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.
Execution and reliability
35%5Has run real facility scale to an uptime commitment, and can walk through their worst outage honestly.
Improving the process
25%5Has redesigned capacity, cooling, or maintenance practice with measured effect on cost or uptime.
Judgement and autonomy
25%5Decides clearly under incident pressure, knows their escalation authority, and documents the call afterwards.
Communication
15%5Communicates outages proactively and clearly to tenants, vendors, and executives alike.
Async video and audio let you hear how a candidate narrates an outage timeline under pressure: the pace, the ordering of decisions, and whether they stay calm and specific. That composure is what tenants and executives will hear at 3am.
Try it on HirevireScreening FAQ
Process basics
What should a data center manager screen cover before a site visit?
Cover four areas: scale and uptime commitments they have carried, monitoring and DCIM tooling they use daily, one real incident narrated end to end, and their efficiency or capacity work with numbers. Add availability for on-call rotation and after-hours maintenance windows. That combination tells you whether the walkthrough is worth booking, and it gives your facilities lead specific things to probe on site.
Should I screen facilities candidates or IT candidates for this role?
Screen both, then compare against your floor. A candidate from mechanical and electrical facilities knows UPS strings, generators, chillers, and CRAC units; a candidate from IT operations knows server platforms, patching, and NOC escalation. Ask each side about the other, and note which gap is coachable given your existing team, vendor contracts, and building management system.
Evaluating answers
How do I tell real data center experience from borrowed credit?
Real operators volunteer specifics without prompting: rack counts, kW per rack, design redundancy such as N+1 or 2N, alarm thresholds, and the vendor they called at 3am. Borrowed credit stays at the level of "we maintained high availability" and stalls when you ask who signed the method statement or who authorised the failover.
What answers should disqualify a data center manager candidate?
Disqualify anyone who reports zero outages across years of operations, blames every failure on a vendor or on "user error", or cannot name a single metric they tracked such as PUE, unplanned downtime minutes, or capacity utilisation. Also drop candidates who describe skipping preventive maintenance to keep uptime numbers looking clean.
























