Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  22

Article 22 · Management

I knew five minutes in (the line that survived every measurement against it)

"I can tell straight away" is the hiring policy of a good share of mid-sized companies, and one of the few management convictions you can actually measure. In 2013 three researchers did exactly that: they handed one side of the table a secret instruction the other side never knew existed. In 2022, another team redid the sums behind the table that ranks selection methods: almost every number came down, and what ended up on top was not what anybody expected. Four measurements, and what to do on Monday with the interview already in your calendar.

By 32sur · September 2026 · Reading time: 15 minutes

I knew five minutes in

Nine forty on a Tuesday morning, in the meeting room of a four-hundred-person company in Vicente López. Fourth interview for the Operations manager role; the calendar says forty-five minutes. It runs seventy.

There are three people inside: the industrial director, the head of Human Resources and the candidate. They talk about a boss they have in common, and about the traffic. At minute fifty the first two questions about the plant appear. Nobody takes notes. On the table, three coffees and a CV with a single handwritten mark: a circle around the name of one company.

The director comes out first and waits in the corridor to speak:

"This is the one. I knew five minutes in."

Nobody argues. The head of Human Resources had prepared a grid of questions and used none of them: the conversation went somewhere else and cutting it short felt rude. The candidate starts the following month and leaves seven months later; on the way out, the word "fit" appears.

None of that is a failure of the process: it is the process. The same company has a quality manual and requires two signatures for a salary advance; to choose the person who will run the plant, seventy minutes of conversation and a circle drawn by hand.

(The scene is a composite reconstruction. From here on, everything with a number in it has been measured and has a source.)

The line has a reasonable defence: twenty years of interviewing must have taught you something. But they only teach if somebody keeps score of how each one turned out, and nobody keeps it: deciding without measuring the outcome does not produce expertise, it produces confidence without accuracy.

A warning: none of the four measurements that follow was carried out on an Argentine owner choosing the manager for his plant. What travels is not the result but the mechanism.

In 2013, three researchers built the experiment that puts the line to the test. It is fairly cruel.

The answers picked at random

The task was concrete and verifiable: predict the grade point average another student would earn the following semester. The ones doing the predicting were students too, on the side of the table that concerns you: the interviewer's. They had that person's previous average —hard data from the file— and twenty minutes of conversation.

There were three versions: one with yes-or-no questions, another with open questions, and a third identical to the first, yes or no, except for something the interviewer never knew: the person across the table had been given a secret instruction.

The interview had a break at the ten-minute mark. From there on, in the third version, the interviewee stopped listening to the question and answered yes or no according to a mechanical rule about the letters of the last two words he had heard: a coin toss in disguise, executed in a second and leaving no pattern. (The literal rule, in reference 1.) You ask whether he sat a hard final exam, the other says yes, and you build meaning out of that: building meaning is what an interviewer does.

Does it show? The authors recorded sixteen of those interviews and showed them to sixty-four people from outside. Of the thirty-two times somebody watched a rigged one, nineteen judged it genuine: only thirteen caught the trick. Tossing a coin would have caught half; with full attention on it, fewer.

And confidence did not move: asked on a 1-to-4 scale how much they could infer about the person, the three groups came out equally convinced.

The number that carries the argument is still missing. First, ten seconds of translation, because this number comes back five times: when a selection method is measured, the result is a correlation — how closely what it predicted resembles what the person went on to do. It runs from 0, tossing a coin, to 1, predicting perfectly. In real jobs, no selection method reaches 0.60; in this laboratory experiment —predicting a grade when you already have the previous grade— accuracy starts higher, so the numbers that follow will clear that bar without contradicting it: they are playing an easier game.

Each participant predicted the average of several people: some he interviewed, others he assessed with the file alone. With the file alone, his accuracy was 0.61; with the file plus the twenty minutes of conversation, 0.31. Same data, same people, half the accuracy. And that 0.31 pools the three versions, sincere conversations included: nothing needs to be rigged.

One defence remains: that a person lost to a formula. The 0.61 against the 0.31 dismantles it — the same people, with and without the conversation. The 0.65 is only the bar: with the previous average used exactly as it stands, nobody in the room, accuracy would have been that.

Read it slowly. The ones who talked had the previous average on the same sheet, and a conversation on top of it. They ended up aiming worse than the sheet alone. This is called the dilution effect: the interview does not add noise alongside the good information, it eats its place.

The file alone beat the file plus the conversation The same people, the same data. The only difference is twenty minutes of conversation. accuracy — how closely the prediction matched the average the person actually earned 0.00 0.20 0.40 0.60 0.80 THE PREVIOUS AVERAGE USED AS IT STANDS, NOBODY IN THE ROOM the benchmark: not a group of people, the data used without interpretation 0.65 THE PARTICIPANTS, WITH THE FILE ALONE 0.61 the conversation did not add: it replaced. THE SAME PARTICIPANTS, WITH THE FILE PLUS TWENTY MINUTES OF CONVERSATION 0.31 Dana, J., Dawes, R. and Peterson, N., Belief in the Unstructured Interview: The Persistence of an Illusion, Judgment and Decision Making, 8(5), 2013. Students predicting other students’ grade point average for the next semester. The value with an interview pools the three versions of the experiment: the disaggregated correlations are too small as subsamples.
Figure 1 — The file alone beat the file plus the conversation. It is not that the free-form interview is useless: it is that it takes the place of what was working. Source: Dana, J., Dawes, R. and Peterson, N., Belief in the Unstructured Interview: The Persistence of an Illusion, Judgment and Decision Making, 8(5), 2013. Students predicting other students' grade point average for the next semester. The value with an interview pools the three versions of the experiment: the disaggregated correlations are too small as subsamples.

And then the punchline. Another group of a hundred and sixty-nine people had the whole experiment described to them and was asked to rank the options: 57% put the rigged interview above having none at all.

The people who hire for a living

Here somebody raises a hand: they were students, and what they were predicting was a grade. A manager with twenty years on the job does not interview like that.

It is the right objection, and there is a measurement that answers it. Three researchers gathered 132 people with real responsibility for hiring decisions, 39 years old on average, who rated themselves 4.6 out of 6 on experience. That self-rating is not a credential, and that is precisely why it works: it is the same certificate you hold about your own eye. (US decision-makers choosing between two files, not Argentine owners.)

The task was ten pairs of real applicants to an airline: pick which one would perform better. Tossing a coin gets you 50%. With the test scores in the file, they got 69% right. Adding a free-form interview —no fixed questions, no rubric, the one we all run—, 62%.

Confidence moved the other way: that is the finding. On the same scale as the hits: with the tests alone they said they were sure about 74 out of every 100 choices and got 69 right; adding the interview they said 78 and got 62 right. The gap between how sure and how right almost tripled: from 0.06 to 0.17 in the paper's finer decimals. The authors: unstructured interviews "not only fail to help selection decisions: they can hurt them".

Stop on that 69 against 62: those are hits out of a hundred, the only place where the numbers speak the language you decide in. Seven choices in a hundred flipping over sounds like little; put it on your own scale: with three manager searches a year, it is one extra bad hire every five years, and each one costs —between severance and the replacement's ramp-up— around six months of salary. The free-form conversation is not free: it has a price, and that is the calculation that turns this piece into money. Run it with your own numbers.

The table everyone cites went out of date in 2022

Why is free judgement still the default method? Scott Highhouse put it down in 2008 to comparing yourself against 100% rather than against the coin —seven points of accuracy feel like nothing— and to almost nobody keeping the list of what actually predicts up to date.

That list goes around on LinkedIn, in every HR course and in almost every assessment proposal. It ranks methods by how well each one anticipates later performance on the job, on the correlation scale from the section above. It is Schmidt and Hunter, 1998, and what stuck is that general intelligence rules, at 0.51 —the one everybody remembers, not the highest: that was work samples, at 0.54.

What they corrected is an arithmetic problem with a technical name, and it is worth telling first without the name. You only ever see how the people you hired perform, and the ones who got in resemble each other, so from the inside any method looks as if it predicts less than it does. To compensate, researchers add a correction. It is called the range restriction correction, and Sackett, Zhang, Berry and Lievens showed it had been applied too generously. They redid the sums: of the ten methods in the figure, nine came down.

In the redone list, the structured interview ended up on top.

Structured means three things, none of them expensive: the same questions for everyone, a rubric written before meeting anybody, and each interviewer scoring on their own. The end of this piece is how to build it.

Here the easy headline lies, so here is the full subtraction. The structured interview, which in 1998 tied with cognitive ability at 0.51, came down to 0.42. It is not that it improved: almost all of them came down, cognitive ability included, which collapsed to 0.31. It started high and it took the correction. (The tests that beat the free-form conversation in the Kausel study belong to this family: against the free-form one they win, against the structured one they lose.)

At the other end fell the unstructured interview, the free-form conversation from Vicente López: 0.19, an average across very different contexts. In the least favourable 10% —the edge of the range the authors report— it is −0.01: there it predicts nothing, or predicts backwards.

Almost all of them dropped. The interview dropped least, and came out first Ten selection methods, before and after the sums were redone. SCHMIDT AND HUNTER, 1998 the table still being cited SACKETT, ZHANG, BERRY AND LIEVENS, 2022 the sums redone 0.60 0.60 0.00 0.00 Work samples 0.54 Structured interview 0.51 General cognitive ability 0.51 Job knowledge 0.48 Unstructured interview 0.38 Biodata 0.35 Almost all of them dropped. This one dropped least, and came out first. 1 0.42 Structured interview 2 0.40 Job knowledge 3 0.38 Biodata 4 0.33 Work samples 5 0.31 General cognitive ability shares fifth place with integrity tests, also at 0.31 0.19 Unstructured interview +4 more methods, all of them down Both axes use the same scale, from 0.00 to 0.60, with ticks every 0.10. Values from Table 3 of Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F., Journal of Applied Psychology, 107(11), 2022. The 1998 column is from Schmidt and Hunter, Psychological Bulletin, 124(2), 1998. All ten values, in Reference 3.
Figure 2 — Almost all of them dropped. The interview dropped least, and came out first. The free-form interview ended at 0.19 — and in the least favourable 10% of contexts it predicts nothing, or predicts backwards. Values from Table 3 of Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F., Revisiting Meta-Analytic Estimates of Validity in Personnel Selection, Journal of Applied Psychology, 107(11), 2022, pp. 2040-2068. The 1998 column corresponds to Schmidt, F. L. and Hunter, J. E., Psychological Bulletin, 124(2), 1998. All ten values, in Reference 3.

And there is something else in that table, which orders everything that follows. "Show me how you do the work" beats "tell me what you are like": the four strongest predictors —structured interview, job knowledge, biodata (a scored record of past behaviour) and work samples— measure something specific to the job, not a general trait. Wernimont and Campbell called it, back in 1968, sign versus sample.

Go back to Vicente López. The circle around a company's name is a sign: it says where the person comes from. Nobody asked him for a sample —how he gets a plant moving on a Monday with two absences and a late order— which is exactly what the new table puts at the top.

What is disputed, and what nobody disputes. In 2023 another group challenged Sackett's criterion for not correcting when the test is taken by people who already work at the company. The fight is about cognitive ability: how far that 0.51 really fell. On the interview there is no dispute — neither camp argues that the structured one beats the free-form one. The big number is being fought over while the ranking that helps you on Monday is already beyond debate.

69% → 62%hit rate of real decision-makers choosing between two candidates: first with the tests in front of them, then adding the free-form conversation (Kausel, Culbertson and Madrid, 2016)
0.61 → 0.31the accuracy of the same people with the file alone, against their accuracy once they had also talked for twenty minutes (Dana, Dawes and Peterson, 2013)
0.42 → 0.19what the structured interview predicts, against what that same interview predicts when it is run free-form (Sackett, Zhang, Berry and Lievens, 2022)

Three subtractions, and always in the same direction. The first of the three is signed by Edgar Kausel from the Catholic University of Chile —the science that dismantles "I can tell straight away" is also written in the region— and it was repeated on a second independent group of 70 people, with the same pattern: 72% against 63%.

So far, what predicts better. The other half is still missing: what happens when the person deciding has the forecast on the table and chooses something else.

When the manager beats the system

The manager walks in with two files. "The system says this one is best, but I'm going with the other." The owner trusts him. Nobody keeps count of how often it happened or how it ended. Somebody did.

Three economists followed 15 service companies that rolled out a test —green, yellow or red for each applicant— site by site, not all at once: while some sites were already using it, others in the same company were not yet — and that difference is what makes it possible to isolate the effect of the test. That is 400,000 applicants and 91,000 hires in low-skill jobs, and quality is measured by how long the person stays, not by their performance.

When the test came in, tenure rose by 25%.

Then they measured the two-files scene: every time a manager hired a yellow when a green was available, or a red over a yellow. They gave that act —skipping the order the test proposes— a name and a count: an exception. With it they built two laboratory companies, one made of the managers who almost never skip the test and another of the ones who always do: in the first, people stay around 15% longer.

The classic defence is still missing: yes, they stay less, but while they are there they perform better. The authors went looking for it in daily output per hour and did not find it —absence of evidence in favour of the exception, not evidence that the person performs worse—. Gut feel had its chance and did not show up: the figure that follows is a measured cost beside an absent benefit.

They went looking for the gut-feel advantage in output per hour. It was not there. 15 service companies · 400,000 applicants · 91,000 hires · when the test came in, tenure rose by 25%. THE COST OF THE EXCEPTION ≈15% shorter tenure the company built from the managers who skip the test most, compared with the one built from those who almost never do. THE ADVANTAGE THAT WOULD JUSTIFY IT this is where the advantage that justifies the exception would be looked for in the daily output per hour of a subset of companies. It did not appear: the box stays empty. Hoffman, M., Kahn, L. B. and Li, D., Discretion in Hiring, The Quarterly Journal of Economics, 133(2), 2018, pp. 765-800. Low-skill service jobs; the study’s measure of quality is how long the person stays in the job. The two companies compared are built by the authors from the 10th and 90th percentiles of exception rate across 445 managers.
Figure 3 — Gut feel had its chance to justify itself with numbers and did not show up. Source: Hoffman, M., Kahn, L. B. and Li, D., Discretion in Hiring, The Quarterly Journal of Economics, 133(2), 2018, pp. 765-800. Low-skill service jobs; the study's measure of quality is how long the person stays in the job. The two companies on the left are a comparison built by the authors from the 10th and 90th percentiles of exception rate across 445 managers, not two observed companies. The search on the right was run on the daily output per hour of a subset of companies: it is absence of evidence in favour of the exception, not evidence that the person hired by exception performs worse.

In the mid-sized Argentine company, skipping the order is not exceptional: it is the procedure, because the owner does the interviewing and nobody reviews his verdict. The exception is not the enemy; the problem is that here it leaves no trace: in the study it could be measured because every one of them was logged, in your company it is not.

The minimum viable structure

None of this recommends you stop interviewing: the structured one came out first. What does not work is running it free-form. At 32sur we call it The Grid —six moves; the timings and thresholds are on the card at the end.

One: define. Before the job ad, the three or four things the person has to know how to do, in verbs, written by whoever will suffer if it goes wrong.

Two: write the questions. The same ones for everybody: if you asked one about his worst week and did not ask the other, you do not have two measurements, you have two conversations.

Three: ask for a sample. A real problem from the company, the same one for every finalist: "show me how you do the work" turned into a step.

Four: the rubric. Three levels with a written example, drafted before interviewing anybody. It is what turns a conversation into a measurement.

Five: score separately. Everyone hands in their sheet before the candidate is discussed, and the person who decides hands in theirs first: if they speak first, what comes after is not information, it is echoes. It is what was missing in Vicente López: the head of Human Resources had the questions written down and did not use them so as not to interrupt the director's conversation.

Six: deliberate at the end. The big differences get discussed: a 3 against a 1 on the same question is the conversation you need to have.

And if you interview alone. It is the most common case, and the one no measurement covers. Write your score before you leave the room: what it protects there is your own opinion from ten minutes ago. And add an interviewer who does not report to you. The ceiling comes from an internal Google analysis —four interviews, 86% confidence; a single company, no peer review: a rule of thumb—; the floor is ours, with no measurement behind it: two.

And that 0.42 is an average: across contexts it varies with a standard deviation of almost 0.20 —spread, not accuracy—. Which is why no method replaces measuring your own: question 6 in the closing list.

Against the dilution effect from the opening there is only one antidote: the interview adds to the file, it does not replace it. If the conversation erased what you knew before you walked in, the problem was not the candidate.

In other countries, not doing this is explained by a lack of data, of scale or of budget. Here it is not: the same questions for everybody, a sample of the real work, a one-page rubric and a scoring sheet do not cost money. They cost one meeting.

The interview is not eliminated: it is industrialised Six moves. Total cost: one meeting, one page and one scoring sheet. MOVE WHEN THE DETAIL 1 · DEFINE Before posting the job ad A 90-minute meeting with whoever will suffer if the person fails 2 · WRITE THE QUESTIONS Before the first interview Six to eight, the same for everyone, anchored to real situations ("tell me about the last time you…") 3 · ASK FOR A SAMPLE With the finalists The same case and the same time for everyone: a problem that happened here 4 · THE RUBRIC Written before interviewing anybody Three levels with a written example, for each question and for the sample 5 · SCORE SEPARATELY Before the candidate is discussed A complete sheet; whoever decides hands in theirs first 6 · DELIBERATE AT THE END With the scores on the table Only the big differences; decided with the full file in view WHAT DROPS OUT OF THE PROCESS the free-form opening chat that ends up deciding the outcome · the clever questions · the sixth round of interviews · the meeting where whoever decides speaks first · the "culture fit" nobody wrote down before meeting the candidate. Reference timeline: your next search. 32sur framework. The only third-party input is the criterion of job-specific measures (Sackett et al., Journal of Applied Psychology, 2022), which underpins move 3.
Figure 4 — The interview is not eliminated: it is industrialised. Six moves. Total cost: one meeting, one page and one scoring sheet. Reference timeline: your next search. 32sur framework. The only third-party input is the criterion of job-specific measures (Sackett et al., Journal of Applied Psychology, 2022), which underpins move 3. The discussion of how many interviewers is in the body of the article, with its caveats (Bock, Work Rules!, 2015).

The four lines that always come up

They will come up the moment you propose any of the six moves. All four are true up to a point, and that point is what keeps them alive.

The line that gets saidWhat was measuredWhat you are left with
"The interview is where you really get to know someone"The file alone predicted 0.65; the file plus the conversation, 0.31 (Dana et al., 2013)You do get to know them. What you get to know does not anticipate how they will work.
"I have an eye for this: I've been doing it for twenty years"69% hits with the tests; 62% adding the free-form conversation (Kausel et al., 2016)Accuracy does not improve with the years. Confidence does.
"So we should stop interviewing and use tests"0.42: the structured interview is today the best-ranked procedure (Sackett et al., 2022)The opposite: done properly, it is the best thing there is.
"I use the system, but if the candidate is worth it I skip it"People hired against the test's ranking stay less, with no evidence that they perform better (Hoffman et al., 2018)Skipping the test has a measured price. The advantage that would justify it never turned up.

It is the half that is true that keeps the line alive.

Questions for Monday

The first five need nothing opened. If getting started requires setting up a system, nobody gets started.

Five from memory, two that take time

  1. The questions from your last interview: could you write them down now? Which of them did you also ask the other finalist? If the answer is "none", you compared two people using different information.
  2. When whoever did the hiring gets it wrong, who pays the cost? Was that person in the room when it was decided?
  3. Of your last ten hires, in how many was the decision made before the first interview was over? Knowing whether it is two or eight is enough.
  4. For the role opening next month, write down on your own, right now, the three things the person has to know how to do at six months. In verbs. If nothing comes out in five minutes, you already know how the meeting will go.
  5. The last person you hired: how many interviewed them, and how many wrote down a score before talking to the others? If it is zero, that is not a reproach: it is everybody's starting point.

And two that cannot be answered today:

  1. Take the hires from the last two years and count how many are still there at twelve months. That is your baseline —no measurement in this piece was done on your company, this one is— and the only thing that leaves a trace where today nothing does: how many times you went against the file, and how it turned out.
  2. In your next three searches, ask everybody the same questions, give every finalist the same work sample, and have each person write down a score before the discussion. At six months, compare those three against the previous three: are they still there, and do they do the three things from question 4?

The Operations manager lasted seven months. "Fit" is the word used when nobody wrote down in advance what the person had to know how to do. The grid the head of Human Resources prepared that morning is still in a folder: writing the questions before walking in was, without anybody knowing it, the second move of The Grid — orphaned, because the first one, defining what had to be known, never happened. The most valuable thing on that table, worth more than the three coffees and a good deal more than the circle.

Nobody ever told the industrial director that he had got it wrong, and strictly speaking he had not: he liked the man, and he was right, he did like him. Nobody looked at the file of the man who left again. He opened the search for the replacement himself, a month later, and the first interview ran seventy minutes again. Nobody wrote anything down — and writing things down, not the coffee or the circle, was the only thing that was needed.

The line in the corridor has an honest and far more useful translation: five minutes in, I knew I liked him. That is real information, and confusing it with something else is paid for in months of salary.

The one thing it is not, is a prediction.

Opening a manager search next month?

You can do all of this on your own, and the person least able to is the one who decides: the key step is that they hand in their score first, and nobody inside the company is going to ask them for it. 32sur sits down for ninety minutes with the people who will do the interviewing, on the role you are opening, and out of it come the profile in verbs, the questions with their rubric, the sample and the scoring sheet. Let's talk.

Let's talk

References

  1. Dana, J., Dawes, R. and Peterson, N., "Belief in the Unstructured Interview: The Persistence of an Illusion", Judgment and Decision Making, 8(5), 2013, pp. 512-520 — the experiment that opens this article. Study 1: three conditions, all with a twenty-minute interview (two of them restricted to closed questions). The literal rule in the random condition, which the body of the article summarises as "a mechanical rule about the letters": from the ten-minute break onwards, the interviewee took the last two words of each question and answered yes if their initials fell in the same block of the alphabet, and no if they fell in different blocks; it was not applied from the start of the interview. Validity of the previous average used alone, r = 0.65; predictions with an interview, r = 0.31; predictions by the same participants without an interview, r = 0.61. Stated confidence on a 1-to-4 scale: 2.85 in the random condition, 2.72 in the closed-question condition with truthful answers and 2.80 in the open-question condition, with no significant differences. Study 2: 16 recorded interviews shown to 64 people who had not been in the room; 19 of the 32 viewings of a random interview were judged genuine, and in the other 13 the trick was caught. Study 3: 169 different people ranked the options as described, and 57% put the random interview above having none at all. The correlations the paper reports disaggregated by condition are not used here: they are too small as subsamples.
  2. Kausel, E. E., Culbertson, S. S. and Madrid, H. P., "Overconfidence in Personnel Selection: When and Why Unstructured Interview Information Can Hurt Hiring Decisions", Organizational Behavior and Human Decision Processes, 137, 2016, pp. 27-44 — the measurement run on people with real responsibility for hiring decisions, not students: 132 participants, 39 years old on average and 4.6 out of 6 in self-rated experience, of whom 21% declared themselves "extremely experienced". The task was a forced choice between pairs of real applicants to an airline, so the chance floor is 50%: 69% hits with tests against 62% adding an unstructured interview; confidence from 0.74 to 0.78; overconfidence bias from 0.06 to 0.17 (Table 4); and a second independent sample of 70 people with the same pattern (72% against 63%). The lead author signs from the Pontifical Catholic University of Chile and the University of Chile. Of the 132 participants, 42 were recruited at a job fair at a Midwestern US university and 90 by snowball sampling through 35 students on an MBA course; the authors found no significant differences between subsamples.
  3. Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F., "Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range", Journal of Applied Psychology, 107(11), 2022, pp. 2040-2068 — the backbone of the article: most of the best-ranked procedures lose between 0.10 and 0.20 points of mean validity once the overcorrection for range restriction is addressed. The ten methods in Figure 2, with their 1998 values (Schmidt and Hunter) and their 2022 values (Sackett et al.):
    Method19982022
    Structured interview0.510.42
    Job knowledge tests0.480.40
    Empirically keyed biodata0.350.38
    Work samples0.540.33
    General cognitive ability0.510.31
    Integrity tests0.410.31
    Assessment centres0.370.29
    Conscientiousness0.310.19
    Unstructured interview0.380.19
    Years of experience0.180.07
    Nine of the ten come down and empirically keyed biodata goes up. The last four —integrity, assessment centres, conscientiousness and years of experience— are the ones Figure 2 groups in grey with no individual label. The structured interview ends up as the best-ranked procedure (0.42, with a standard deviation of 0.19 across contexts) and the unstructured interview falls to 0.19, with a floor of −0.01 in the 80% credibility interval, that is, in the least favourable 10% of contexts. Fifth place on the new list is a tie between cognitive ability and integrity tests, both at 0.31, and the authors flag it explicitly. The average of the five best methods goes from 0.49 in 1998 to 0.37 in 2022. This is also the source of the pattern that organises the whole article: the four strongest predictors are job-specific measures, not general psychological traits.
  4. Schmidt, F. L. and Hunter, J. E., "The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings", Psychological Bulletin, 124(2), 1998, pp. 262-274 — the table still circulating in courses and presentations across the region, and the "before" column of Figure 2 in this article.
  5. Hoffman, M., Kahn, L. B. and Li, D., "Discretion in Hiring", The Quarterly Journal of Economics, 133(2), 2018, pp. 765-800 — what happens when the manager hires against the test's recommendation: 15 companies, 400,000 applicants, 91,000 hires and 445 managers. The test classified 48% of applicants as green and 32% as yellow, and it was rolled out in staggered fashion across each company's sites; that staggering is what identifies the effect. Completed job tenure rises by a little over 25% when it is introduced (0.23 log points), with +6 percentage points in the probability of reaching 6 months and +7.5 in that of reaching 12. A company made up of managers at the 10th percentile of exception rate would have tenures around 15% longer than one at the 90th percentile; these are two hypothetical companies constructed by the authors, not two observed companies. Examining daily output per hour, the authors find no evidence that exceptions are associated with higher productivity. These are low-skill service jobs and the measure of quality is how long the person stays.
  6. Highhouse, S., "Stubborn Reliance on Intuition and Subjectivity in Employee Selection", Industrial and Organizational Psychology, 1(3), 2008, pp. 333-342 — the two implicit beliefs that explain why a better method is not adopted: that near-perfect accuracy is achievable, and the myth that the ability to predict human behaviour improves with experience. In a footnote he cites three cases —Meyer (1956), Huse (1962) and Borneman and co-authors (2007)— where tests alone predicted better than tests combined with expert judgement, and warns that the comparison has rarely been made explicitly. The survey of HR professionals the article mentions is described by Highhouse himself as still unpublished: it is not cited here with figures.
  7. Oh, I.-S., Le, H. and Roth, P. L., "Revisiting Sackett et al.'s (2022) Rationale Behind Their Recommendation Against Correcting for Range Restriction in Concurrent Validation Studies", Journal of Applied Psychology, 108(8), 2023, pp. 1300-1310; the replies from the Sackett team the same year in that journal and in Industrial and Organizational Psychology, 16, pp. 371-377; and Oh, I.-S., Mendoza, J. L. and Le, H., "To correct or not to correct for range restriction, that is the question", Industrial and Organizational Psychology, 16(3), 2023, pp. 322-327 — the open controversy over how much to correct for range restriction. What the box calls "not correcting when the test is taken by people who already work at the company" is what the literature calls concurrent validation. The disagreement concentrates on cognitive ability; the relative order between structured and unstructured interviews is not in dispute between the parties. If somebody projects one of the two tables at you as the final word, the useful question is what the other one says.
  8. Bock, L., Work Rules!, Twelve, 2015 — the "rule of four" from the internal Google analysis attributed to Todd Carlisle: four interviews are enough to decide with 86% confidence whether to hire the person, and each additional interviewer adds around one point of predictive power. It is internal analysis from a single company, with no peer review or public replication: it is used here as a rule of thumb, not as evidence.
  9. Wernimont, P. F. and Campbell, J. P., "Signs, Samples, and Criteria", Journal of Applied Psychology, 52(5), 1968, pp. 372-376 — the distinction between sign and sample that Sackett and co-authors recover to interpret why the best predictors are the job-specific ones.