healthcare clinicalcbt protocolsconversational aicrisis escalationdigital mental health
Complete evaluation framework
What to assess and how to score it
Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Clinical competence
35% weight
Probe their licensure (LPC, LCSW, LMFT or psychology registration) and which modalities they script or supervise: CBT, DBT skills, motivational interviewing, ACT, plus PHQ-9 and GAD-7 use.
Evidence to listen for
Command of the procedures, anatomy, and equipment the role requires
Holds current registration or certification
Knows normal from abnormal and what to do about each
Recognises when a case is outside their scope
Five-point scoring guide
1
Poor
Unsafe knowledge gaps; registration missing or lapsed.
2
Needs Improvement
Knowledge gaps that would affect patient care.
3
Satisfactory
Competent for standard cases; needs support on complex ones.
4
Very Good
Strong clinical knowledge; safe and reliable across the usual range.
5
Excellent
Holds an active clinical licence, names specific modalities and outcome measures, and shows how each maps to dialogue flows or model prompts.
02
Evaluation factor
Patient safety and protocol
30% weight
Test how they handle suicidal ideation, self-harm disclosure and psychosis in an automated channel: escalation triggers, human handoff SLAs, HIPAA or GDPR handling, and false-negative review.
Evidence to listen for
Follows identification, infection control, and documentation protocol without prompting
Can describe an error or near miss and what they did
Escalates deterioration early
Treats protocol as protection rather than bureaucracy
Five-point scoring guide
1
Poor
Casual about protocol; would not report an error.
2
Needs Improvement
Inconsistent protocol adherence; slow to escalate.
3
Satisfactory
Follows protocol reliably; documentation sometimes thin.
4
Very Good
Protocol is instinctive; escalates early and reports honestly.
5
Excellent
Describes concrete risk-detection triggers, warm handoff to a human clinician within a defined window, and audit of missed crisis flags.
03
Evaluation factor
Patient communication
20% weight
Assess how they write or evaluate empathic responses at scale: tone calibration, reflective listening in text, avoiding false rapport, and handling users who anthropomorphise the system.
Evidence to listen for
Explains a procedure to an anxious or confused patient
Handles distress, pain, or refusal without losing control of the interaction
Respects privacy and dignity in practice, not just in principle
Works with families and carers
Five-point scoring guide
1
Poor
Dismissive of patients; no bedside awareness.
2
Needs Improvement
Task-focused; struggles with distressed patients.
3
Satisfactory
Adequate rapport; less confident in difficult interactions.
4
Very Good
Calm, clear, and respectful with anxious or difficult patients.
5
Excellent
Shows sample dialogue they authored or red-teamed, explaining word-level choices and where the agent must state it is not human.
04
Evaluation factor
Working in a clinical team
15% weight
Look for collaboration with ML engineers, product and clinical safety boards: annotation guidelines, model evaluation rubrics, IRB or ethics review, and post-deployment incident debriefs.
Evidence to listen for
Hands over cleanly and completely
Challenges a colleague when patient safety requires it
Takes direction from clinicians without deferring blindly
Handles shift work and pressure without becoming difficult to work with
Five-point scoring guide
1
Poor
Poor handover; cannot work in a clinical team.
2
Needs Improvement
Handover gaps; avoids raising concerns about colleagues.
3
Satisfactory
Reliable team member; handover adequate.
4
Very Good
Clean handovers and willing to speak up on safety.
5
Excellent
Cites specific joint work with engineering on eval sets or guardrails, and names a review body they reported clinical findings to.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.