Two teams, the same company
Two development teams, in the same company, with comparable budgets and people of similar caliber. The first moves forward: people ask when they don't understand, flag problems early, argue hard and decide fast. The second drags: meetings are monologues by whoever talks most, mistakes surface late and expensive, and when something smells wrong no one says so until it blows up. Same profiles, same pay, same building. Opposite results.
In Part 1 we saw why: performance does not come from assembled talent but from the conditions around it. But "conditions" is a hollow word. This article fills it in. What, exactly, do you design when you design a team? The research of recent decades identified the concrete levers and measured how much each one weighs. Not all are worth the same. We go one by one, with the number alongside.
Lever 1: that people can speak up
Psychological safety is not a "nice" climate where no one is uncomfortable: it is the shared belief that one can take interpersonal risks —asking something obvious, admitting a mistake, disagreeing with the boss— without being left exposed. It is what makes productive conflict possible, not what avoids it. Amy Edmondson (1999) built the construct studying 51 teams: safety enables learning behavior, and that behavior is what mediates toward performance.
The figure that best captures it is dizzyingly counterintuitive. In nursing units, Edmondson found that those with the best leadership and climate reported more medication errors, not fewer. They did not make more: they reported them, because they felt safe. In the fearful units, the error was hidden —and a hidden error is not corrected, it is repeated—. A whole thesis on why the bad news that does not rise in time ends up being the most expensive.
How much does it weigh? The meta-analysis by Frazier and colleagues (2017) —136 samples, more than 22,000 people, close to 5,000 groups— found strong correlations: with learning behavior, 0.62; with information sharing, 0.52; with organizational citizenship, 0.32.
Psychological safety makes no one brilliant; it makes the intelligence and knowledge already in the room come to light instead of staying silent. Almost always the problem is not that knowledge is missing, but that whoever knows does not speak.
Lever 2: a team's intelligence is not the sum of the IQs
In 2010, Anita Woolley, Thomas Malone and colleagues published in Science a study that looked in groups for what psychology had looked for in individuals: a general factor of intelligence. With 699 people in small groups solving very different tasks, they found a "collective intelligence" factor (factor c): a single factor explained more than 43% of the variance in performance across tasks, and predicted a final criterion task with a correlation of 0.52.
What does it depend on? Not on assembled talent: c was barely correlated with average individual intelligence (r ≈ 0.19, not significant) nor with that of the most brilliant member (r ≈ 0.27). What did predict it was the how: social sensitivity, equality in speaking turns and the proportion of women (mediated by social sensitivity).
The finding was replicated even in teams that collaborate only online, by text (factor c of 41–49% of the variance), and a 2021 reanalysis of 22 studies confirmed it.
The honest counterpoint. Not everyone buys the interpretation. Bates and Gupta (2017) found that, in their data, individual intelligence explained up to 80% of the differences in group intelligence, and they did not reproduce the gender or turn-taking effects. The factor c is well established; its independence from talent is still debated. The prudent reading: talent matters and is not enough; how the group interacts adds a layer of performance that no sum of IQs guarantees.
There is a bridge to the real world, built by Alex "Sandy" Pentland, of MIT. With sensors that measured how thousands of people communicated across 21 organizations, he concluded that communication patterns are the most important predictor of a team's success —as important as all the other factors combined—, and that close to 35% of the variation in performance is explained by the amount of face-to-face exchange. (It is popular science, not a peer-reviewed paper; it is worth taking as evidence convergent with Woolley, not as a laboratory coefficient. But the message is robust: how a team converses is not "the soft stuff"; it is the hard variable most open to intervention.)
Lever 3: goals that bite (and why the misplaced incentive destroys)
The goal-setting theory of Edwin Locke and Gary Latham —half a century, hundreds of studies— concludes that specific and difficult goals outperform vague ones of the "do your best" kind: they won in close to 90% of the studies, with effect sizes (d) of 0.42 to 0.80. In teams it also works (d ≈ 0.80), but with a trap: in interdependent work, giving each person "egoistic" individual goals destroys collective performance (an effect of −1.75). Only group-centric goals produce the improvement (+1.20).
The consequence explains half the dysfunctional teams you see: a system of individual incentives mounted on interdependent work is a machine for manufacturing internal competition and sabotaging the shared result. The salesperson who hides the lead, the area that optimizes its number at the expense of the one next door: it is almost never malice, almost always a badly designed incentive pulling against the team.
The 80% myth (and why it still helps). A rule attributed to Noel Tichy circulates: 80% of conflicts stem from poorly defined goals; then roles; then processes; only 0.8% from relationships (80 / 16 / 3.2 / 0.8). The numbers are too tidy to be real: they are an artifact of applying Pareto in cascade, not a measurement. But the direction of the advice matches the evidence: when a team fights, the root is usually higher up —confused goals disguised as a clash of egos—. Use it as a compass, never as a statistic.
Lever 4: structure (size, roles, trust)
Size: the smallest that works. What saturates a team is not the number of people but the number of links, which grow by the formula n(n−1)/2: a team of 6 has 15 links; one of 7, 21; one of 12, 66. That is why Katzenbach and Smith observed that almost all effective teams had fewer than ten members, and Amazon institutionalized the "two-pizza rule". They are heuristics, but they point the same way: include the bare minimum, and not one person more.
Clear roles and diversity of the right kind. Diversity of knowledge and function helps —above all in complex tasks—, though with a small effect. Demographic diversity has, in the serious meta-analyses, an effect on performance close to zero on objective criteria; it turns negative only when the measure is a boss's subjective evaluation —which betrays evaluator bias—. The genuine mechanism: diverse groups process information more carefully, debate more facts and make fewer mistakes. (The "36% more profitability" figure comes from consultancy, is correlational and did not replicate; better not to rest anything important on it.)
Trust: the measurable lubricant. The meta-analysis by De Jong, Dirks and Gillespie (2016), over 112 studies and 7,763 teams, found that intra-team trust is associated with performance at a correlation of 0.30 —above the average of what gets measured in this field—, and weighs more the more interdependent the work is.
| Lever | What it measures | Evidence | Source |
|---|---|---|---|
| Psychological safety | Being able to risk speaking up | ρ ≈ 0.62 with learning; 0.52 with information sharing | Frazier et al., 2017 |
| Collective intelligence | Social sensitivity + even turns | factor c: 43% of the variance; almost no relationship with the group's IQ | Woolley et al., 2010 |
| Communication patterns | How (not what) the team converses | ~35% of the performance variation ‡ | Pentland, 2012 ‡ |
| Specific and difficult goals | Focus and effort | d ≈ 0.42–0.80 (individual); ≈ 0.80 (group) | Locke and Latham, 2002 |
| Intra-team trust | Mutual vulnerability | ρ ≈ 0.30 with performance | De Jong et al., 2016 |
| Cohesion | Commitment to the shared task | positive, greater with interdependent work | Beal et al., 2003 |
‡ popular science or practice, not peer-reviewed. The effect sizes are not comparable with one another (different metrics and designs).
Open questions
- Do people bring bad news in time and disagree with you to your face —or do you find out about problems once they are already expensive?
- In your meetings, is the floor shared —or do two people speak 80% of the time?
- Do you know who on your team reads others well —or do you build teams looking only at résumés?
- Are your goals specific and measurable —or elegant versions of "do your best"?
- Do your incentives reward the team's result —or individual performance in interdependent work, aligning each person against the whole?
- Does your team have the minimum size that works —or more links than it can sustain?
- Do your people trust each other enough to depend on one another?
If several answers are uncomfortable, you already know where to intervene. None of these levers requires changing the people; all of them require changing the conditions in which they work.
"High-performance team" is not a vague aspiration but a short list of intervenable variables. None is charisma or chemistry. But knowing which the levers are is not the same as knowing how to pull them. In Part 3 we close the circle: what the evidence shows about how to build, develop and sustain a team —and how much it is worth doing it well—.
Suspect your teams have talent to spare and design missing?
32sur diagnoses the conditions holding your teams back today, redesigns goals and incentives aligned to the collective result, and installs practices of psychological safety and good communication.
References
How to read the figures. An effect size (Cohen’s d) gauges how large a difference is (~0.2 small, ~0.5 moderate, ~0.8 large). A correlation (ρ) measures how much two things move together, from 0 (not at all) to 1 (perfectly); in team research, ~0.3 is already a meaningful relationship. ns means the result is not statistically significant (it could be due to chance).
- Edmondson, A. C., "Psychological Safety and Learning Behavior in Work Teams", Administrative Science Quarterly, 44(2), 1999 — the construct, in 51 teams.
- Edmondson, A. C., "Learning from Mistakes Is Easier Said than Done", Journal of Applied Behavioral Science, 32(1), 1996 — units with better climate report more errors (greater willingness to report).
- Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A. and Vracheva, V., "Psychological Safety: A Meta-Analytic Review and Extension", Personnel Psychology, 70(1), 2017 — 136 samples (>22,000 people); correlations with learning (0.62), information (0.52), citizenship (0.32) and creativity (0.13).
- Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N. and Malone, T. W., "Evidence for a Collective Intelligence Factor in the Performance of Human Groups", Science, 330(6004), 2010.
- Engel, D., Woolley, A. W. et al., "Reading the Mind in the Eyes or Reading between the Lines?", PLoS ONE, 9(12), 2014 — the factor c also in online teams.
- Bates, T. C. and Gupta, S., "Smart Groups of Smart People: Evidence for IQ as the Origin of Collective Intelligence", Intelligence, 60, 2017 — counterpoint: individual IQ explains much of group intelligence.
- Pentland, A., "The New Science of Building Great Teams", Harvard Business Review, April 2012 — communication patterns (popular science, not peer-reviewed).
- Locke, E. A. and Latham, G. P., "Building a Practically Useful Theory of Goal Setting and Task Motivation", American Psychologist, 57(9), 2002; Kleingeld, A., van Mierlo, H. and Arends, L., "The Effect of Goal Setting on Group Performance", Journal of Applied Psychology, 96(6), 2011.
- De Jong, B. A., Dirks, K. T. and Gillespie, N., "Trust and Team Performance: A Meta-Analysis", Journal of Applied Psychology, 101(8), 2016 — trust and performance ρ ≈ 0.30.
- van Dijk, H., van Engen, M. L. and van Knippenberg, D., Organizational Behavior and Human Decision Processes, 119(1), 2012; Bell, S. T. et al., Journal of Management, 37(3), 2011; Beal, D. J. et al., Journal of Applied Psychology, 88(6), 2003.