Review the evidence signals before interviewing. Then use the anchored descriptions—not instinct alone—to choose the score that best matches each answer.
01
Evaluation factor
Portfolio
35% weight
Ask to walk through shipped voice experiences: Alexa skills, Google Actions, IVR redesigns or in-car assistants. Look for sample dialogues, flow diagrams and persona documents they authored.
Evidence to listen for
Work exists and can be looked at, not just described
States what they made versus what the team or a template made
Shows range rather than one repeated style
Can walk through a piece from brief to final
Five-point scoring guide
1
Poor
No portfolio, or work that is unattributable or clearly templated.
2
Needs Improvement
Thin portfolio; unclear what they personally made.
3
Satisfactory
Real work with adequate range; contribution mostly clear.
4
Very Good
Strong varied portfolio with clear personal ownership.
5
Excellent
Shows named shipped assistants with dialogue flows, prompt scripts and persona docs, plus containment or task completion numbers before and after.
02
Evaluation factor
Craft and rationale
25% weight
Probe handling of error recovery, reprompts, barge-in, confirmation strategy and SSML tuning. Ask why they chose implicit over explicit confirmation in a specific turn.
Evidence to listen for
Explains why a layout, type choice, or colour decision serves the brief
Knows typography and hierarchy as craft, not decoration
Works to a brand system without either breaking it or hiding behind it
Names the tools they are genuinely fast in
Five-point scoring guide
1
Poor
Cannot explain any decision; work is arbitrary.
2
Needs Improvement
Talks in taste terms only; no link between choice and brief.
3
Satisfactory
Sound craft with some ability to justify decisions.
4
Very Good
Articulate about why each choice serves the brief.
5
Excellent
Explains reprompt ladders, no-match and no-input handling, and TTS prosody choices with reasoning tied to user intent and cognitive load.
03
Evaluation factor
Feedback and iteration
25% weight
Test how they use transcripts, utterance logs and intent confusion data to revise dialogue. Ask about a flow rewritten after Wizard of Oz or usability testing.
Evidence to listen for
Takes critique without treating it as an attack
Distinguishes a subjective preference from a real problem, and says so politely
Iterates fast rather than defending version one
Delivers files correctly and on time
Five-point scoring guide
1
Poor
Defensive about critique; will not revise.
2
Needs Improvement
Accepts feedback passively; iterations do not improve the work.
3
Satisfactory
Revises willingly; struggles to push back on weak feedback.
4
Very Good
Iterates quickly and can argue for the work when the feedback is wrong.
5
Excellent
Cites specific transcript findings, expanded utterance sets or rewritten prompts, and reports the drop in fallback or hang-up rates afterwards.
04
Evaluation factor
Working with the brief
15% weight
Check collaboration with NLU engineers, linguists and product owners: writing intent and slot specifications, scoping voice-only versus multimodal, and negotiating platform certification constraints.
Evidence to listen for
Asks about audience and goal before opening the design tool
Works with marketing, product, or clients rather than in isolation
Flags an impossible brief early
Hands over files and assets others can actually use
Five-point scoring guide
1
Poor
Designs in isolation; ignores the brief's purpose.
2
Needs Improvement
Starts designing before understanding the goal.
3
Satisfactory
Asks the right questions when prompted.
4
Very Good
Interrogates the brief up front and hands over cleanly.
5
Excellent
Translates ambiguous product goals into annotated flows, intent schemas and acceptance criteria engineers can build against without repeated clarification.
Put this rubric to work
Score every candidate against the same standard
Add these weighted factors to Hirevire and let AI evaluate recorded answers against your rubric.