Why impressions are a poor hiring instrument
After an interview what stays in your head is a verdict on the whole person: “liked them”, “something bothers me”. The problem isn't that it's subjective — subjectivity here is unavoidable. The problem is that it's incomparable: two interviewers who spoke to the same person come away with different impressions and cannot establish where exactly they diverged.
There's a second, less obvious effect. An impression forms in the first minutes and then defends itself: the rest of the conversation is spent looking for confirmation of a verdict already reached. So “I knew within five minutes” is usually an honest description of what happened, not a sign of experience.
It doesn't make assessment objective, and it doesn't save you from badly chosen criteria. If the requirements list says something other than what the job actually needs, the scale will simply measure the wrong thing precisely. Structure buys comparability, not truth.
Three foundations of a structured interview
- The same blocks for every candidate for the role. Not identical questions word for word — identical themes in the same order. Otherwise you're comparing conversations, not people.
- Criteria fixed before the first interview. A list assembled after the meetings always bends towards the candidate you already liked — quietly and very reliably.
- Assessment from evidence, not from feeling. Every score has to rest on something the candidate said or did — ideally a quotation from the transcript.
Evidence versus impression
The difference is settled by one question: could this be disputed by watching the recording? If yes, it's evidence. If no, it's an impression, and it has no place on the scorecard.
Weak on design. The experience seems more theoretical. Not sure they'd cope at our level.
Asked about sharding, named two approaches and explained why their project keyed on customer ID. On the follow-up about hot partitions gave no answer: “we never ran into that”.
The first note can't be discussed: there's nothing to grip, and you'd be arguing with a colleague's sensation. The second is checkable — open the transcript and decide whether that's enough for the role. And only the second survives three months: a quarter later nobody will recall what exactly “bothered” them about the candidate.
Scorecard: try scoring one
The scale deliberately has four points and no middle: the middle is a way of not deciding. The weight reflects how much the criterion matters for this particular role. Computed in your browser, sent nowhere.
| Criterion | Weight | Score |
|---|---|---|
| Depth in the core areaWorks through their own example rather than retelling an article | ×3 | |
| Solving a task from the roleGets to a working answer and sees its weak points | ×3 | |
| Handling uncertaintyAsks clarifying questions instead of guessing | ×2 | |
| Explaining complexityCan tell it so an adjacent specialist understands | ×2 | |
| Response to challengeShifts position on an argument, not under pressure | ×1 |
Scale: 1 — didn't show it, 2 — showed it partly, 3 — showed it confidently, 4 — showed more than expected. Work out the threshold in advance and write it down before the interview, or it will bend to fit the candidate.
Where the transcript earns its place
A scorecard demands evidence, and evidence demands precision. Reconstructing a candidate's exact wording from memory an hour later is impossible: what remains is a paraphrase, and a paraphrase already carries your interpretation. The transcript removes that step — the quotation is taken verbatim.
The second, less obvious gain: it frees your attention during the conversation. While you're writing you aren't listening; while you're catching up on notes you miss the next answer. An interviewer who isn't taking notes asks noticeably sharper follow-up questions.
Four mistakes when moving to structured interviews
- Too many criteria. Past six and the assessment turns intuitive again: nobody fills in twelve rows honestly.
- A scale with a middle. A five-point scale collects threes. An even number of points forces you to pick a side.
- Discussing scores before recording them. If interviewers talk first and fill in afterwards, you have one opinion in two copies. Independently first, reconciliation second.
- A threshold set afterwards. “Well, seventeen isn't bad” is the moment the whole construction stops working.
Checklist before the interview
What to do with divergences
A divergence in scores isn't a problem — it's the most valuable thing the method produces. It marks the place where two people saw different things in the same answer. The cause is usually one of three: a different reading of the criterion, a different weight given to the same fact, or one interviewer leaning on an impression and unable to produce evidence.
Only work through items that differ by two points or more. A one-point gap is ordinary noise; discussing it eats time and creates a false sense of the method's precision.
Whisperer keeps the interview transcript and assembles structured notes from it, so the quotation behind each score is seconds away. It doesn't make the hiring decision and shouldn't: criteria, weights and thresholds are what the team owns, not the tool.