Why pre-screen genomic data curators before the interview
The failures in curation are silent. A sample swap, a reference build mismatch, a batch effect that looks like biology, metadata that says one thing while the file says another. Nothing errors; the analysis downstream is simply wrong. Curators worth hiring have found one of these in data everyone had already trusted. A short screen asks for that, along with the checks that would have caught it earlier.
What actually matters when screening Genomic Data Curator candidates
- 01
Method and rigour
Check command of ACMG/AMP classification criteria, HGVS nomenclature, and ontology mapping with HPO or MONDO; ask which reference builds, transcripts, and gnomAD population filters they apply.
- 02
Real casework
Probe volume and type of curated content: variants per week, gene-disease validity assessments, ClinVar or ClinGen submissions, panel content reviews, or literature triage backlogs they cleared.
- 03
Interpretation and judgement
Test how they resolve conflicting evidence: discordant ClinVar submissions, weak functional assays, VUS reclassification triggers, or segregation data that contradicts computational predictors.
- 04
Reporting and testimony
Assess written curation records: evidence summaries, internal review notes, discrepancy escalations to variant scientists, and how they document decisions for audit or reanalysis cycles.
Pre-screening questions to ask Genomic Data Curator candidates
12 questions grouped by what they test. Ask the same set in every screen and score answers on a consistent scale, or send them as an async video screen and compare answers side by side.
Datasets they curated
3 questions01Can you describe your experience with genomic data annotation and curation?
Listen forDatasets they curated with scale and data types named, and what they were curated for.
Analysis experience presented as curation, or no dataset they were responsible for.
02What experience do you have with sequencing data from current platforms?
Listen forPlatform-specific artefacts understood, with the processing steps that were applied before curation.
Data treated as platform-agnostic, or no awareness of platform-specific error profiles.
03Can you discuss a challenging project involving genomic data?
Listen forA specific difficulty such as inconsistent sample identifiers or a mixed reference build.
Challenges described as data volume, or no problem that required investigation.
Quality control that catches
3 questions04Can you provide examples of quality control measures you implement?
Listen forChecks that would catch a sample swap or contamination, run before release rather than on request.
Quality control limited to summary statistics, or checks run only when something looks wrong.
05How do you ensure data integrity and accuracy in your curation process?
Listen forChecksums, provenance and version control applied, with changes to a dataset tracked and reversible.
Files edited in place, or no record of what changed between dataset versions.
06What is your experience with variant calling and annotation?
Listen forAwareness that annotation depends on the reference and tool version, with both recorded.
Annotations treated as fixed truth, or reference build not recorded with the data.
Metadata that survives
3 questions07What experience do you have with metadata, and why does it matter here?
Listen forMetadata recorded well beyond the minimum, including protocol, batch and full processing history.
Only repository-required fields captured, or metadata reconstructed after the fact.
08Can you explain your approach to normalisation and standardisation?
Listen forControlled vocabularies and ontologies applied, so datasets can be combined without manual mapping.
Free-text fields left unstandardised, or each dataset standardised in its own way.
09How familiar are you with the major public genomic databases?
Listen forSubmission experience with their requirements, including what each expects and where they disagree.
Databases used for lookup only, or no experience preparing a submission.
Consent respected
3 questions10How do you address ethical considerations and data privacy in your work?
Listen forConsent scope checked before any sharing, with access controls and re-identification risk understood.
Consent treated as a completed step, or genomic data shared as though it were anonymous.
11How do you collaborate with researchers and other stakeholders on data projects?
Listen forRequirements gathered from the people who will use the data, with pushback on unusable requests.
Requests fulfilled without question, or curation decisions made with no user input.
12How do you handle large datasets and ensure efficient processing?
Listen forPipelines that scale with checkpoints, so a failure does not require reprocessing everything.
Processing done manually at scale, or pipelines that restart from the beginning after any failure.
How to score responses
Score every candidate on the same four criteria immediately after the screen. At this stage you are shortlisting for panel interviews, not making the final call.
Method and rigour
35%5Cites specific evidence codes (PS3, PM2, BP4), explains transcript selection and build liftover, and names population frequency thresholds used.
Real casework
25%5Quantifies curated variants or genes, names the submitting lab or ClinGen expert panel, and describes concordance rates with reviewers.
Interpretation and judgement
25%5Walks through a reclassification they drove, weighting each evidence line explicitly and stating what would have changed the call.
Reporting and testimony
15%5Produces traceable curation notes citing PMIDs and criteria applied, and defends calls calmly in expert panel or sign-out review.
A sample swap or a build mismatch does not error; the analysis is simply wrong. A one-way video screen asks what they caught.
Try it on HirevireScreening FAQ
Process basics
How long should a pre-screening round for this role take?
Fifteen minutes across eight to ten questions, answered async. Enough to establish datasets they curated, test their quality control, and check metadata and consent practice.
How much analysis skill should I expect?
Enough to know what downstream analysis needs and where it breaks. A curator who has never analysed data will preserve the wrong things and standardise fields nobody uses.
Evaluating answers
What is the strongest signal when screening this role?
A problem they found in trusted data. Curators with real experience have caught a sample swap or a build mismatch. Anyone whose datasets were always clean has not looked closely.
How do I judge their metadata practice?
Ask what they record beyond the minimum. Real answers cover protocol, batch and processing history. Anyone recording only what a repository requires produces data nobody can reuse.
























