Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  09

Article 09 · High-performance teams

What actually moves the needle

Second of three parts. If performance is designed, what exactly do you design? Four named levers with evidence and effect sizes — and one finding that should change how you hand out incentives.

By 32sur · July 2026 · Reading time: 18 minutes

Two teams, the same company

Two development teams, in the same company, with comparable budgets and people of similar caliber. The first moves forward: people ask when they don't understand, flag problems early, argue hard and decide fast. The second drags: meetings are monologues by whoever talks most, mistakes surface late and expensive, and when something smells wrong no one says so until it blows up. Same profiles, same pay, same building. Opposite results.

In Part 1 we saw why: performance does not come from assembled talent but from the conditions around it. But "conditions" is a hollow word. This article fills it in. What, exactly, do you design when you design a team? The research of recent decades identified the concrete levers and measured how much each one weighs. Not all are worth the same. We go one by one, with the number alongside.

Lever 1: that people can speak up

Psychological safety is not a "nice" climate where no one is uncomfortable: it is the shared belief that one can take interpersonal risks —asking something obvious, admitting a mistake, disagreeing with the boss— without being left exposed. It is what makes productive conflict possible, not what avoids it. Amy Edmondson (1999) built the construct studying 51 teams: safety enables learning behavior, and that behavior is what mediates toward performance.

The figure that best captures it is dizzyingly counterintuitive. In nursing units, Edmondson found that those with the best leadership and climate reported more medication errors, not fewer. They did not make more: they reported them, because they felt safe. In the fearful units, the error was hidden —and a hidden error is not corrected, it is repeated—. A whole thesis on why the bad news that does not rise in time ends up being the most expensive.

How much does it weigh? The meta-analysis by Frazier and colleagues (2017) —136 samples, more than 22,000 people, close to 5,000 groups— found strong correlations: with learning behavior, 0.62; with information sharing, 0.52; with organizational citizenship, 0.32.

0 0.70 Learning behavior 0.62 Information sharing 0.52 Organizational citizenship 0.32 Creativity 0.13
Figure 1 — What psychological safety is associated with (corrected correlation ρ). Correlational relationships, not causal: safety enables the behaviors —learning, sharing— that in turn produce performance. Its relationship with creativity is weak. Source: Frazier et al., 2017.

Psychological safety makes no one brilliant; it makes the intelligence and knowledge already in the room come to light instead of staying silent. Almost always the problem is not that knowledge is missing, but that whoever knows does not speak.

Lever 2: a team's intelligence is not the sum of the IQs

In 2010, Anita Woolley, Thomas Malone and colleagues published in Science a study that looked in groups for what psychology had looked for in individuals: a general factor of intelligence. With 699 people in small groups solving very different tasks, they found a "collective intelligence" factor (factor c): a single factor explained more than 43% of the variance in performance across tasks, and predicted a final criterion task with a correlation of 0.52.

What does it depend on? Not on assembled talent: c was barely correlated with average individual intelligence (r ≈ 0.19, not significant) nor with that of the most brilliant member (r ≈ 0.27). What did predict it was the how: social sensitivity, equality in speaking turns and the proportion of women (mediated by social sensitivity).

0 Average individual IQ 0.19 · ns IQ of the most brilliant 0.27 · ns Social sensitivity 0.26 Equality in speaking turns 0.41 Proportion of women 0.23
Figure 2 — What predicts (and what does not) a team's intelligence. Assembled individual talent barely predicts (gray, not significant); how the group interacts does (teal). In a joint regression, only social sensitivity remains significant. Source: Woolley et al., Science, 2010.

The finding was replicated even in teams that collaborate only online, by text (factor c of 41–49% of the variance), and a 2021 reanalysis of 22 studies confirmed it.

The honest counterpoint. Not everyone buys the interpretation. Bates and Gupta (2017) found that, in their data, individual intelligence explained up to 80% of the differences in group intelligence, and they did not reproduce the gender or turn-taking effects. The factor c is well established; its independence from talent is still debated. The prudent reading: talent matters and is not enough; how the group interacts adds a layer of performance that no sum of IQs guarantees.

There is a bridge to the real world, built by Alex "Sandy" Pentland, of MIT. With sensors that measured how thousands of people communicated across 21 organizations, he concluded that communication patterns are the most important predictor of a team's success —as important as all the other factors combined—, and that close to 35% of the variation in performance is explained by the amount of face-to-face exchange. (It is popular science, not a peer-reviewed paper; it is worth taking as evidence convergent with Woolley, not as a laboratory coefficient. But the message is robust: how a team converses is not "the soft stuff"; it is the hard variable most open to intervention.)

Lever 3: goals that bite (and why the misplaced incentive destroys)

The goal-setting theory of Edwin Locke and Gary Latham —half a century, hundreds of studies— concludes that specific and difficult goals outperform vague ones of the "do your best" kind: they won in close to 90% of the studies, with effect sizes (d) of 0.42 to 0.80. In teams it also works (d ≈ 0.80), but with a trap: in interdependent work, giving each person "egoistic" individual goals destroys collective performance (an effect of −1.75). Only group-centric goals produce the improvement (+1.20).

A · Individual level (Cohen’s d) d 0.42 – 0.80 specific-hard vs. "do your best" B · In interdependent teams (Cohen’s d) 0 Group-centric goals +1.20 Selfish individual goals −1.75
Figure 3 — Goals: which help and which destroys. On a team, rewarding individual performance can do worse than setting no goals at all: it aligns each person against the whole. Sources: Locke & Latham, 2002; Kleingeld et al., 2011.

The consequence explains half the dysfunctional teams you see: a system of individual incentives mounted on interdependent work is a machine for manufacturing internal competition and sabotaging the shared result. The salesperson who hides the lead, the area that optimizes its number at the expense of the one next door: it is almost never malice, almost always a badly designed incentive pulling against the team.

The 80% myth (and why it still helps). A rule attributed to Noel Tichy circulates: 80% of conflicts stem from poorly defined goals; then roles; then processes; only 0.8% from relationships (80 / 16 / 3.2 / 0.8). The numbers are too tidy to be real: they are an artifact of applying Pareto in cascade, not a measurement. But the direction of the advice matches the evidence: when a team fights, the root is usually higher up —confused goals disguised as a clash of egos—. Use it as a compass, never as a statistic.

Lever 4: structure (size, roles, trust)

Size: the smallest that works. What saturates a team is not the number of people but the number of links, which grow by the formula n(n−1)/2: a team of 6 has 15 links; one of 7, 21; one of 12, 66. That is why Katzenbach and Smith observed that almost all effective teams had fewer than ten members, and Amazon institutionalized the "two-pizza rule". They are heuristics, but they point the same way: include the bare minimum, and not one person more.

Clear roles and diversity of the right kind. Diversity of knowledge and function helps —above all in complex tasks—, though with a small effect. Demographic diversity has, in the serious meta-analyses, an effect on performance close to zero on objective criteria; it turns negative only when the measure is a boss's subjective evaluation —which betrays evaluator bias—. The genuine mechanism: diverse groups process information more carefully, debate more facts and make fewer mistakes. (The "36% more profitability" figure comes from consultancy, is correlational and did not replicate; better not to rest anything important on it.)

Trust: the measurable lubricant. The meta-analysis by De Jong, Dirks and Gillespie (2016), over 112 studies and 7,763 teams, found that intra-team trust is associated with performance at a correlation of 0.30 —above the average of what gets measured in this field—, and weighs more the more interdependent the work is.

LeverWhat it measuresEvidenceSource
Psychological safetyBeing able to risk speaking upρ ≈ 0.62 with learning; 0.52 with information sharingFrazier et al., 2017
Collective intelligenceSocial sensitivity + even turnsfactor c: 43% of the variance; almost no relationship with the group's IQWoolley et al., 2010
Communication patternsHow (not what) the team converses~35% of the performance variation ‡Pentland, 2012 ‡
Specific and difficult goalsFocus and effortd ≈ 0.42–0.80 (individual); ≈ 0.80 (group)Locke and Latham, 2002
Intra-team trustMutual vulnerabilityρ ≈ 0.30 with performanceDe Jong et al., 2016
CohesionCommitment to the shared taskpositive, greater with interdependent workBeal et al., 2003

‡ popular science or practice, not peer-reviewed. The effect sizes are not comparable with one another (different metrics and designs).

0.62correlation between psychological safety and learning behavior (Frazier et al., 2017)
43%of group performance variance explained by a single collective intelligence factor (Woolley et al., 2010)
−1.75the effect of giving an interdependent team egoistic individual goals (Kleingeld et al., 2011)

Open questions

  1. Do people bring bad news in time and disagree with you to your face —or do you find out about problems once they are already expensive?
  2. In your meetings, is the floor shared —or do two people speak 80% of the time?
  3. Do you know who on your team reads others well —or do you build teams looking only at résumés?
  4. Are your goals specific and measurable —or elegant versions of "do your best"?
  5. Do your incentives reward the team's result —or individual performance in interdependent work, aligning each person against the whole?
  6. Does your team have the minimum size that works —or more links than it can sustain?
  7. Do your people trust each other enough to depend on one another?

If several answers are uncomfortable, you already know where to intervene. None of these levers requires changing the people; all of them require changing the conditions in which they work.

"High-performance team" is not a vague aspiration but a short list of intervenable variables. None is charisma or chemistry. But knowing which the levers are is not the same as knowing how to pull them. In Part 3 we close the circle: what the evidence shows about how to build, develop and sustain a team —and how much it is worth doing it well—.

Suspect your teams have talent to spare and design missing?

32sur diagnoses the conditions holding your teams back today, redesigns goals and incentives aligned to the collective result, and installs practices of psychological safety and good communication.

Let's talk

References

How to read the figures. An effect size (Cohen’s d) gauges how large a difference is (~0.2 small, ~0.5 moderate, ~0.8 large). A correlation (ρ) measures how much two things move together, from 0 (not at all) to 1 (perfectly); in team research, ~0.3 is already a meaningful relationship. ns means the result is not statistically significant (it could be due to chance).

  1. Edmondson, A. C., "Psychological Safety and Learning Behavior in Work Teams", Administrative Science Quarterly, 44(2), 1999 — the construct, in 51 teams.
  2. Edmondson, A. C., "Learning from Mistakes Is Easier Said than Done", Journal of Applied Behavioral Science, 32(1), 1996 — units with better climate report more errors (greater willingness to report).
  3. Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A. and Vracheva, V., "Psychological Safety: A Meta-Analytic Review and Extension", Personnel Psychology, 70(1), 2017 — 136 samples (>22,000 people); correlations with learning (0.62), information (0.52), citizenship (0.32) and creativity (0.13).
  4. Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N. and Malone, T. W., "Evidence for a Collective Intelligence Factor in the Performance of Human Groups", Science, 330(6004), 2010.
  5. Engel, D., Woolley, A. W. et al., "Reading the Mind in the Eyes or Reading between the Lines?", PLoS ONE, 9(12), 2014 — the factor c also in online teams.
  6. Bates, T. C. and Gupta, S., "Smart Groups of Smart People: Evidence for IQ as the Origin of Collective Intelligence", Intelligence, 60, 2017 — counterpoint: individual IQ explains much of group intelligence.
  7. Pentland, A., "The New Science of Building Great Teams", Harvard Business Review, April 2012 — communication patterns (popular science, not peer-reviewed).
  8. Locke, E. A. and Latham, G. P., "Building a Practically Useful Theory of Goal Setting and Task Motivation", American Psychologist, 57(9), 2002; Kleingeld, A., van Mierlo, H. and Arends, L., "The Effect of Goal Setting on Group Performance", Journal of Applied Psychology, 96(6), 2011.
  9. De Jong, B. A., Dirks, K. T. and Gillespie, N., "Trust and Team Performance: A Meta-Analysis", Journal of Applied Psychology, 101(8), 2016 — trust and performance ρ ≈ 0.30.
  10. van Dijk, H., van Engen, M. L. and van Knippenberg, D., Organizational Behavior and Human Decision Processes, 119(1), 2012; Bell, S. T. et al., Journal of Management, 37(3), 2011; Beal, D. J. et al., Journal of Applied Psychology, 88(6), 2003.