Use cases / Hiring and talent
Interview transcripts scored against a rubric
Every interview scored against one rubric, with the quote behind each score.
- Unit
- One interview transcript plus the role’s rubric
- Ask
- Three competency scores. The evidence turn. Did they answer?
- Get
- Scores with distributions, and the one answer to read
Fireflies · Files · 2 reducers · 7 questions · updated Sep 21, 2026
The flow
Same frozen rubric for every candidate. Must-haves only for those above the bar. 343 requests for 212 interviews. Every score links to a rubric version and a turn.
Follow one record
- interviews
- Attach the rubric
- Score competencies
- Keep problem solving ≥ 2.0
- Attach the resume
- Must-haves
- Outputs
Interview I-0412, senior support engineer
snapshot v4, 42 minutes, 6 questions asked
- Q2Walk me through a time a customer's problem turned out to be different from what they reported.
- A2A customer said exports were slow. I reproduced it and the export was fine; their VPN was rewriting headers. I wrote up the header trace, sent it to their IT, and added a check to our diagnostics page so support could see it next time.
- Q3What would you have done if IT had pushed back?
- A3I had the packet capture, so I would have walked them through it live. If it still stalled I would escalate on our side with the evidence attached rather than argue.
Rubric for role R-17, version 2
left key role_id, right key role_id
- role
- Senior support engineer
- problem_solving
- Below the bar: guesses. Meets with gaps: reproduces but stops at the symptom. Meets: isolates the cause with evidence. Exceeds: also prevents recurrence.
- communication
- Below the bar: unclear. Meets with gaps: clear but incomplete. Meets: clear and complete. Exceeds: adapts to the audience.
- ownership
- Below the bar: defers. Meets with gaps: acts when asked. Meets: acts and follows through. Exceeds: changes the system.
Request 33 of 212
model jev-1.13.0, attempt 1, succeeded
- Below the bar0.02
- Meets with gaps0.28
- Meets0.68
- Exceeds0.02
confidence 0.66, provider reported
- Below the bar0.00
- Meets with gaps0.08
- Meets0.76
- Exceeds0.16
confidence 0.74, provider reported
- Below the bar0.06
- Meets with gaps0.64
- Meets0.30
- Exceeds0.00
confidence 0.62, provider reported
- A20.71
- A30.24
- A40.03
- A10.02
the answer the score rests on
- true0.93
answered the questions asked
{
"model": "jev-1.13.0",
"state": {
"interview_id": "I-0412",
"candidate_id": "CAND-2290",
"role_id": "R-17",
"rubric": {
"version": 2,
"problem_solving": "Below the bar: guesses. Meets with gaps: reproduces but stops at the symptom. Meets: isolates the cause with evidence. Exceeds: also prevents recurrence.",
"communication": "Below the bar: unclear. Meets with gaps: clear but incomplete. Meets: clear and complete. Exceeds: adapts to the audience.",
"ownership": "Below the bar: defers. Meets with gaps: acts when asked. Meets: acts and follows through. Exceeds: changes the system."
},
"transcript": [
{ "turn": "Q1", "speaker": "interviewer", "text": "Tell me about your current role." },
{ "turn": "A1", "speaker": "candidate", "text": "Tier 2 support at a logistics SaaS, about 40 tickets a week, on-call one week in four." },
{ "turn": "Q2", "speaker": "interviewer", "text": "Walk me through a time a customer's problem turned out to be different from what they reported." },
{ "turn": "A2", "speaker": "candidate", "text": "A customer said exports were slow. I reproduced it and the export was fine; their VPN was rewriting headers. I wrote up the header trace, sent it to their IT, and added a check to our diagnostics page so support could see it next time." },
{ "turn": "Q3", "speaker": "interviewer", "text": "What would you have done if IT had pushed back?" },
{ "turn": "A3", "speaker": "candidate", "text": "I had the packet capture, so I would have walked them through it live. If it still stalled I would escalate on our side with the evidence attached rather than argue." }
]
},
"questions": {
"problem_solving": { "type": "score", "instructions": "Using `rubric.problem_solving` as the scale, how does the candidate's reasoning in `transcript` rate?", "criteria": ["Below the bar", "Meets with gaps", "Meets", "Exceeds"] },
"communication": { "type": "score", "instructions": "Using `rubric.communication` as the scale, how clearly does the candidate explain their thinking in `transcript`?", "criteria": ["Below the bar", "Meets with gaps", "Meets", "Exceeds"] },
"ownership": { "type": "score", "instructions": "Using `rubric.ownership` as the scale, how much personal responsibility for outcomes does the candidate show in `transcript`?", "criteria": ["Below the bar", "Meets with gaps", "Meets", "Exceeds"] },
"evidence_turn": { "type": "choice", "instructions": "Which of the candidate's answers in `transcript` best evidences the problem-solving score?", "criteria": { "A1": "The candidate's first answer", "A2": "The candidate's second answer", "A3": "The candidate's third answer", "A4": "The candidate's fourth answer", "A5": "The candidate's fifth answer", "A6": "The candidate's sixth answer" } },
"answered": { "type": "noul", "instructions": "Did the candidate answer the questions actually asked in `transcript`, rather than a different question?" }
}
}
{
"model": "jev-1.13.0",
"answers": {
"problem_solving": { "expected": 2.70, "probabilities": [0.02, 0.28, 0.68, 0.02], "confidence": 0.66 },
"communication": { "expected": 3.08, "probabilities": [0.00, 0.08, 0.76, 0.16], "confidence": 0.74 },
"ownership": { "expected": 2.24, "probabilities": [0.06, 0.64, 0.30, 0.00], "confidence": 0.62 },
"evidence_turn": { "choice": "A2", "probabilities": { "A2": 0.71, "A3": 0.24, "A4": 0.03, "A1": 0.02, "A5": 0.00, "A6": 0.00 }, "confidence": 0.70 },
"answered": { "probability": 0.93 }
}
}
2.70 is at or above 2.0, kept
the scores stay with the row either way
- field
- answers.problem_solving.expected
- comparison
- gte 2.0
- value
- 2.70
- result
- matched collection
Resume joined on candidate_id
no model call
- 2023 to nowTier 2 Support Engineer, Ridgeline Logistics. Python tooling for log triage used by the whole team; on-call rotation.
- 2020 to 2023Support Specialist, Fernwood Health. Zendesk macros, escalations to engineering.
Request 21 of 131
model jev-1.13.0, attempt 1, succeeded
- true0.88
resume: team-wide Python tooling
- true0.34
handled incidents; leading one is not evidenced
{
"model": "jev-1.13.0",
"answers": {
"python_production": { "probability": 0.88 },
"led_incident": { "probability": 0.34 }
}
}
One row the hiring manager reads with the quote open
every score links to the rubric version and the evidence turn
- interview
- I-0412, senior support engineer
- problem solving
- 2.70, Meets (0.68)
- communication
- 3.08, Meets (0.76)
- ownership
- 2.24, Meets with gaps (0.64)
- evidence
- A2 (0.71): the VPN header trace answer
- must-haves
- Python in production yes 0.88; led an incident no 0.34
- trace
- 7 nodes, 2 attempts, rubric v2 frozen in the run
What comes out
| Interview | Problem solving | Communication | Ownership | Evidence | Python in prod | Led incident |
|---|---|---|---|---|---|---|
| I-0388 | 3.41 | 3.12 | 3.05 | A4 | 0.96 | 0.91 |
| I-0412 | 2.70 | 3.08 | 2.24 | A2 | 0.88 | 0.34 |
| I-0401 | 2.66 | 2.31 | 2.90 | A3 | 0.52 | 0.78 |
| I-0377 | 2.41 | 2.88 | 2.12 | A2 | 0.91 | 0.12 |
| I-0420 | 2.08 | 2.02 | 2.55 | A5 | 0.23 | 0.61 |
Synthetic rows in the app's Outputs layout. Scores are expected positions on the four-level rubric; the distributions are one click away.
Details, if you want them
whyWhy this shape
Why this shape
- One rubric, frozen. Every candidate scored against the same words.
- Ordinal scores. 2.7 means between two levels, and you can see how sure.
- The evidence turn. Read one answer, not the transcript.
- Must-haves only above the bar. The other 81 keep their scores.
What it replaces
- Six people remembering what was said.
- A 3 that means a 4 to someone else.
- Not the interviewer. People decide.
questionsThe 7 questions, verbatim
- Score competencies
problem_solvingordinal, score
Using `rubric.problem_solving` as the scale, how does the candidate’s reasoning in `transcript` rate?
Below the barMeets with gapsMeetsExceeds - Score competencies
communicationordinal, score
Using `rubric.communication` as the scale, how clearly does the candidate explain their thinking in `transcript`?
Below the barMeets with gapsMeetsExceeds - Score competencies
ownershipordinal, score
Using `rubric.ownership` as the scale, how much personal responsibility for outcomes does the candidate show in `transcript`?
Below the barMeets with gapsMeetsExceeds - Score competencies
evidence_turncategorical, choice
Which of the candidate’s answers in `transcript` best evidences the problem-solving score?
A1A2A3A4A5A6 - Score competencies
answeredbinary, noul
Did the candidate answer the questions actually asked in `transcript`, rather than a different question?
- Must-haves
python_productionbinary, noul
Do `transcript` and `resume` together show the candidate has written Python that ran in production?
- Must-haves
led_incidentbinary, noul
Do `transcript` and `resume` together show the candidate has led the response to a customer-facing incident?
buildHow each stage is built, as JSON
1 join Attach the rubric
role_id puts the rubric text into each interview’s state. The rubric is a snapshot, so every candidate in the run is scored against the same words; editing it later creates a new version and does not touch this run.stage definition
{
"id": "with_rubric",
"input": "source:interviews",
"kind": "join",
"left_key": [
"role_id"
],
"name": "Attach the rubric",
"right": "source:rubrics",
"right_key": [
"role_id"
]
}2 reducer Score competencies
evidence_turn picks the candidate’s answer that best supports the problem-solving score; answered catches the answer that dodged the question.stage definition
{
"id": "score",
"input": "with_rubric",
"kind": "evaluate",
"model": "jev",
"name": "Score competencies",
"questions": {
"answered": {
"instructions": "Did the candidate answer the questions actually asked in `transcript`, rather than a different question?",
"type": "noul"
},
"communication": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.communication` as the scale, how clearly does the candidate explain their thinking in `transcript`?",
"type": "score"
},
"evidence_turn": {
"criteria": {
"A1": "The candidate's first answer",
"A2": "The candidate's second answer",
"A3": "The candidate's third answer",
"A4": "The candidate's fourth answer",
"A5": "The candidate's fifth answer",
"A6": "The candidate's sixth answer"
},
"instructions": "Which of the candidate's answers in `transcript` best evidences the problem-solving score?",
"type": "choice"
},
"ownership": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.ownership` as the scale, how much personal responsibility for outcomes does the candidate show in `transcript`?",
"type": "score"
},
"problem_solving": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.problem_solving` as the scale, how does the candidate's reasoning in `transcript` rate?",
"type": "score"
}
},
"state_path": []
}3 filter Keep problem solving ≥ 2.0
stage definition
{
"comparison": "gte",
"field": [
"answers",
"problem_solving",
"expected"
],
"id": "above_bar",
"input": "score",
"kind": "filter",
"name": "Keep problem solving ≥ 2.0",
"value": 2
}4 join Attach the resume
candidate_id only for the candidates above the bar.stage definition
{
"id": "with_resume",
"input": "above_bar",
"kind": "join",
"left_key": [
"candidate_id"
],
"name": "Attach the resume",
"right": "source:resumes",
"right_key": [
"candidate_id"
]
}5 reducer Must-haves
stage definition
{
"id": "must_haves",
"input": "with_resume",
"kind": "evaluate",
"model": "jev",
"name": "Must-haves",
"questions": {
"led_incident": {
"instructions": "Do `transcript` and `resume` together show the candidate has led the response to a customer-facing incident?",
"type": "noul"
},
"python_production": {
"instructions": "Do `transcript` and `resume` together show the candidate has written Python that ran in production?",
"type": "noul"
}
},
"state_path": []
}6 definition The whole workflow, as saved
Saving publishes an immutable version. A run freezes this definition, the snapshot IDs and each reducer’s model profile, so the trace can always show exactly what ran. Edits create a new version; old runs keep theirs.
jevreduce.local.v1
{
"name": "Interview transcripts scored against a rubric",
"schema_version": "jevreduce.local.v1",
"sources": {
"interviews": "snap_fireflies_interviews_v4",
"resumes": "snap_resumes_v4",
"rubrics": "snap_rubrics_v2"
},
"stages": [
{
"id": "with_rubric",
"input": "source:interviews",
"kind": "join",
"left_key": [
"role_id"
],
"name": "Attach the rubric",
"right": "source:rubrics",
"right_key": [
"role_id"
]
},
{
"id": "score",
"input": "with_rubric",
"kind": "evaluate",
"model": "jev",
"name": "Score competencies",
"questions": {
"answered": {
"instructions": "Did the candidate answer the questions actually asked in `transcript`, rather than a different question?",
"type": "noul"
},
"communication": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.communication` as the scale, how clearly does the candidate explain their thinking in `transcript`?",
"type": "score"
},
"evidence_turn": {
"criteria": {
"A1": "The candidate's first answer",
"A2": "The candidate's second answer",
"A3": "The candidate's third answer",
"A4": "The candidate's fourth answer",
"A5": "The candidate's fifth answer",
"A6": "The candidate's sixth answer"
},
"instructions": "Which of the candidate's answers in `transcript` best evidences the problem-solving score?",
"type": "choice"
},
"ownership": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.ownership` as the scale, how much personal responsibility for outcomes does the candidate show in `transcript`?",
"type": "score"
},
"problem_solving": {
"criteria": [
"Below the bar",
"Meets with gaps",
"Meets",
"Exceeds"
],
"instructions": "Using `rubric.problem_solving` as the scale, how does the candidate's reasoning in `transcript` rate?",
"type": "score"
}
},
"state_path": []
},
{
"comparison": "gte",
"field": [
"answers",
"problem_solving",
"expected"
],
"id": "above_bar",
"input": "score",
"kind": "filter",
"name": "Keep problem solving ≥ 2.0",
"value": 2
},
{
"id": "with_resume",
"input": "above_bar",
"kind": "join",
"left_key": [
"candidate_id"
],
"name": "Attach the resume",
"right": "source:resumes",
"right_key": [
"candidate_id"
]
},
{
"id": "must_haves",
"input": "with_resume",
"kind": "evaluate",
"model": "jev",
"name": "Must-haves",
"questions": {
"led_incident": {
"instructions": "Do `transcript` and `resume` together show the candidate has led the response to a customer-facing incident?",
"type": "noul"
},
"python_production": {
"instructions": "Do `transcript` and `resume` together show the candidate has written Python that ran in production?",
"type": "noul"
}
},
"state_path": []
}
]
}limitsWhat it will not do
- A score is a position on your rubric, not a hiring decision.
- Rubric versions are frozen per run. Replay to rescore.
- Transcription errors are scored as said.
- Measure consistency on a labelled sample first.
- No written feedback. Jev returns scores and a turn.