Why pre-screen AGI researchers before the technical panel
Very little in this area is settled, which means confident timelines and strong positions are cheap and rigorous results are rare. Researchers worth hiring describe what their experiments established, what they did not, and where their own view could be wrong. A short screen asks what a result did not show, which distinguishes research from advocacy faster than any question about capability.
What actually matters when screening Artificial General Intelligence Researcher candidates
- 01
Theoretical command
Probe depth on transformer internals, scaling laws, RLHF and RL objectives, mechanistic interpretability; ask them to defend a position on emergence or sample efficiency with citations.
- 02
From theory to hardware or code
Check what they actually trained: model sizes, cluster and framework (JAX, PyTorch FSDP, Megatron), datasets curated, evals built, and open-sourced code or checkpoints.
- 03
Research judgement
Test how they pick problems: killed experiments, negative results, choosing between a scaling ablation and an architectural bet under fixed GPU-hours.
- 04
Explaining it to non-specialists
Assess how they brief policy staff, safety reviewers or funders on capability jumps and risk without hype or jargon; ask for a blog post or talk.
Pre-screening questions to ask Artificial General Intelligence Researcher candidates
12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.
Research they conducted
3 questions01Do you have hands-on experience with research projects in this area?
Listen forExperiments they designed and ran, with what was established stated separately from what was hoped.
Familiarity with the literature offered in place of research, or no experiments they personally ran.
02What types of systems have you worked with in your research?
Listen forSpecific systems and scales, with an honest account of what compute and data were available to them.
Systems described in general terms, or claims that do not match the resources they had.
03Do you have published research or contributions in this field?
Listen forPublications with their own contribution described, including work that did not produce a positive result.
Author lists offered with no personal contribution, or only successful results ever discussed.
Method and evaluation
3 questions04Can you explain the methodology of a study you have conducted?
Listen forHypothesis, controls and evaluation described clearly, with the limitations of the design acknowledged.
Method described as training a model, or evaluation designed after the results were seen.
05Can you describe your experience with machine learning methods?
Listen forReal depth in current methods, with an understanding of what they do and do not generalise across.
Methods described at a survey level, or generalisation claimed beyond the evaluated distribution.
06How proficient are you with the programming and tooling used in this research?
Listen forCode written and experiments run by them, with reproducibility handled through versioning and seeds.
Implementation delegated entirely, or experiments that cannot be reproduced from what was recorded.
Safety as technical work
2 questions07Can you discuss your understanding of safety research in this area?
Listen forSafety treated as concrete technical problems such as evaluation, oversight and specification.
Safety discussed only as a position, or dismissed as a distraction from capability work.
08What ethical considerations do you regard as most relevant to this research?
Listen forConcrete considerations raised, including release decisions, misuse potential and evaluation before any deployment.
Ethics answered abstractly, or release and misuse questions treated as somebody else's job.
Honest uncertainty
4 questions09How do you handle research that does not produce the results you anticipated?
Listen forNegative results treated as informative and reported, with hypotheses revised rather than reframed.
Results reframed to look positive, or failed directions never written up or shared.
10How would you assess the state of progress toward general capability?
Listen forDeep uncertainty acknowledged, with the assessment tied to specific measurable capabilities rather than dates.
Confident timelines given in either direction, or progress assessed from demonstrations alone.
11How familiar are you with current theories and models in this field?
Listen forCompeting accounts described fairly, including the strongest arguments against their own view.
One school of thought presented as settled, or opposing arguments not fairly represented.
12Do you collaborate with other researchers or organisations in your work?
Listen forReal collaborations with the contribution of each party clear, and disagreement handled productively.
Work done entirely in isolation, or collaborators described only as providing resources.
How to score responses
Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.
Theoretical command
35%5Argues precisely about optimisation dynamics, loss curves and inductive bias, citing specific papers and naming where the theory breaks down.
From theory to hardware or code
30%5Names runs they owned end to end, from data pipeline to eval harness, with parameter counts, compute budgets and released artefacts.
Research judgement
20%5Describes a research direction they abandoned early with the evidence that triggered it, and how they reallocated compute.
Explaining it to non-specialists
15%5Explains a result like grokking or deceptive alignment plainly, separating measured findings from speculation, and states uncertainty ranges.
Confident timelines are cheap in this field and rigorous results are rare. A one-way video screen asks what a result did not show.
Try it on HirevireScreening FAQ
Process basics
How long should a pre-screening round for this role take?
Fifteen minutes across eight to ten questions, answered async. Enough to establish research they conducted, test their methodology, and hear how they handle uncertainty and safety.
How should I weigh publications?
As evidence of contribution rather than a count. Ask what they personally did on a paper and what the result actually established. Author lists tell you far less than that account.
Evaluating answers
What is the strongest signal when screening this role?
Stating what a result did not show. Researchers with discipline draw that line clearly. Anyone whose findings support broad claims about general capability is overreaching from narrow evidence.
What should worry me in an answer?
Confident timelines for general capability, in either direction. The honest position is that nobody knows, and someone who is certain will make claims your organisation has to defend.
























