Chapter 5
When is a team more — or less — than the sum of its members?
Above every workstation on Toyota's assembly line in Georgetown, Kentucky, runs a cord. Any worker who spots anything wrong — a misaligned seal, a missing bolt, a doubt — is expected to pull it. Lights flash, a chime sounds, and within seconds a team leader arrives: not to assign blame, but to help solve the problem before the car moves on. If it can't be solved within the work cycle, that segment of the line stops. Nobody is punished for pulling; pulling is the job.56 In 2007, a BBC report captured the astonishing arithmetic: Georgetown workers were pulling the cord roughly 2,000 times a week. At a new Ford truck plant in Dearborn that had installed the identical hardware, workers pulled it about twice a week. Nobody believes Ford had a thousand times fewer problems. The difference was not the cord. It was what it felt safe to say.4
In February 2024, a panel of experts convened by the Federal Aviation Administration reported on what it felt safe to say inside Boeing. Congress had ordered the study after two crashes of the 737 MAX — Lion Air in 2018 and Ethiopian Airlines in 2019, 346 people dead — and the panel delivered its findings weeks after a door plug blew out of an Alaska Airlines MAX whose retaining bolts, investigators found, were missing after factory work.1 After more than 250 interviews and 4,000 pages of documents, the panel described a "disconnect" between senior management and everyone else on safety culture; employees who hesitated to report problems "for fear of retaliation," since the managers who investigated safety reports often also controlled the reporters' salaries and promotions; and widespread distrust of the anonymity of "Speak Up," the very reporting channel created after the crashes. Many employees, the panel found, did not know how to raise a safety issue at all.123 Boeing had installed the cord. It had not made it safe to pull.
Here is what makes this genuinely puzzling rather than a simple morality tale. Boeing employed some of the best aerospace engineers alive; Toyota's Georgetown line was staffed by ordinary workers, many with no manufacturing background. If team performance is about the talent on the team, Boeing should win every comparison. If it is about something in the room between the people — what can be said, who speaks, what happens next — then a line of average workers who pull a cord 2,000 times a week can outperform a company of brilliant engineers who stay silent. Thoughtful people resist this conclusion, and for a defensible reason: silence is cheap to condemn in hindsight, and every organization must also ship products, make schedules, and stop the line only when it matters. When does voice create quality, and when is it noise?
The central question: what makes a group of people more than the sum of its members — and what makes it less? The tools: how groups form and actually develop; the process losses (social loafing, coordination costs) that shrink teams below their potential; psychological safety, the best-supported team-level performance factor we have; groupthink, the classic account of smart teams deciding badly; collective intelligence and what actually predicts it; and communication structure — media richness and what large-scale remote-work evidence now shows. We close with Boeing and Toyota as the two ends of one design variable: the price of speaking up.
A group is people who interact and see themselves as a unit; a team is a group with a shared goal, mutual accountability, and interdependent work — the difference between people who ride the same elevator and people who build the same engine. The most famous account of how teams develop is Bruce Tuckman's 1965 sequence: forming, storming, norming, performing (with "adjourning" added later).9
Tuckman's stages are a useful expectation-setter — early conflict is normal, not a sign of failure — but treat them as a memory aid, not a law: the model came from reviewing therapy and training groups, and field studies of real work teams often fail to find the tidy sequence. The most important correction is Connie Gersick's punctuated equilibrium finding: project teams she observed did not climb stages; they locked into a pattern in their very first meeting, coasted on it, and then — almost precisely at the midpoint of their allotted time — underwent a burst of alarm and wholesale reorganization before racing to the deadline.10 Two managerial implications follow. First meetings matter enormously, because the patterns set there persist; and the midpoint is a natural, predictable window for intervention, when teams are briefly open to changing how they work. The broader lesson: team development is shaped less by internal psychology than by deadlines, context, and first impressions — which is why the design choices in the rest of this chapter matter more than waiting for "performing" to arrive.
A team's actual output equals its potential output minus process losses — the productivity destroyed by working together. The oldest documented loss is social loafing: individuals exert less effort in groups than alone, a phenomenon measured as early as Ringelmann's rope-pulling studies and established experimentally by Bibb Latané, Kipling Williams, and Stephen Harkins, whose participants shouted and clapped with progressively less effort as (they believed) group size grew.11
Loafing grows with group size and shrinks under three conditions the experiments isolate: identifiability (individual contributions can be seen), meaningfulness (the task matters to the person), and indispensability (my effort actually changes the outcome). Coordination costs compound the effort losses: as teams grow, communication links multiply geometrically, which is why adding people to a late project so often makes it later. The design translation is direct — keep teams as small as the task allows, make individual contributions visible, and connect each member's work to consequences. The boundary conditions matter too: loafing reverses into social facilitation (working harder in others' presence) on simple, well-learned tasks, and cohesive teams with strong norms can largely suppress it. Loafing, in other words, is not a fact about lazy people; it is a fact about invisible contributions — fix the visibility and you mostly fix the loafing.
Psychological safety is a team's shared belief that interpersonal risk-taking is safe — that admitting a mistake, asking a question, flagging a problem, or disagreeing will not be punished or humiliated. Amy Edmondson introduced and validated the construct in a 1999 field study of 51 manufacturing work teams, showing that psychological safety predicted team learning behavior, which in turn predicted team performance.7 Her path to the idea is itself instructive: studying hospital teams, she found the better teams reporting more errors — not because they made more, but because they were the teams where errors could be surfaced.
The construct's fame rests on replication at scale. Google's Project Aristotle studied 180 of its own teams expecting to find that who was on the team — stars, seniority, skill mix — would predict performance; instead, the strongest differentiator of its best teams was psychological safety, followed by dependability, structure and clarity, meaning, and impact. Group norms — especially roughly equal conversational turn-taking and members' sensitivity to one another — mattered more than the roster.8 A caution worth teaching: most of this evidence is correlational at the team level, and psychological safety is not niceness or comfort — Edmondson's own framing pairs it with high standards. Safety without standards produces a pleasant, mediocre team; standards without safety produce Boeing's silence.
Psychological safety is the mechanism that converts a team's private knowledge into collective performance: every error caught, risk flagged, and half-formed idea contributed passes through the same gate — the speaker's split-second forecast of what speaking will cost. That is why it is the load-bearing wall for every safety-critical and innovation-critical team, and why Toyota's cord and team-leader response are best understood as psychological safety engineered into hardware and roles decades before the term existed.56 The limits: safety can be faked in surveys while fear governs behavior (measure what people do, as the andon pull-counts do); and it is built or destroyed locally — one leader's reaction to one bad-news message teaches the whole team the real price list.
Groupthink is Irving Janis's name for the deterioration of judgment in cohesive groups under pressure for unanimity: members self-censor doubts, pressure dissenters, assume silence means agreement, and construct an illusion of invulnerability — his case studies ran from the Bay of Pigs to Pearl Harbor.12 Janis's antidotes remain the standard toolkit: assign a devil's advocate, have the leader withhold their view until others speak, break into independent subgroups, and revisit decisions in a "second-chance" meeting. Honest appraisal requires the caveat: groupthink was built from selected historical cases, and controlled research finds its full syndrome less often than its fame suggests — cohesion alone does not reliably degrade decisions; directive leadership plus insulation from outside views are the more dangerous ingredients. Read it as a checklist of failure modes to design against, not a diagnosis to throw at any group that agrees with itself. Its deepest overlap with this chapter: every groupthink symptom is a psychological-safety failure wearing a different name — self-censorship is simply the individual-level act that Section 2.3 measures the price of.
Is there such a thing as a smart team, over and above smart members? Anita Williams Woolley and colleagues' study in Science found evidence for a general collective intelligence factor ("c"): a team's performance across one set of diverse tasks predicted its performance on others, just as individual IQ does. The striking part was what predicted c — not the average or maximum member intelligence, but three interaction properties: members' social sensitivity (reading others accurately), equality of conversational turn-taking (c fell when a few voices dominated), and the proportion of women on the team, an effect statistically carried by women's higher average social-sensitivity scores.13 The convergence with Project Aristotle — conducted independently, years later, inside a company — is the most reassuring pattern in this literature: how a team talks beats who is on it.813 Cautions for honest use: subsequent replications find the c-factor and the turn-taking result more consistently than the exact size of other predictors, and none of this implies member ability is irrelevant — it implies ability is a ceiling that interaction quality determines whether you reach. For diversity more broadly, the fairest summary of a large literature is conditional: diverse teams have higher potential (more information, less redundancy) and higher process costs (more conflict, slower trust), so diversity's payoff depends on exactly the variables in Sections 2.2–2.3.
Communication succeeds when meaning survives the trip from one head to another. Richard Daft and Robert Lengel's media richness theory ranks channels by capacity — cues carried, feedback speed, personalization: face-to-face at the top, then video, phone, chat, email, formal documents — and prescribes matching richness to ambiguity: rich media for equivocal, emotional, or novel matters; lean media for routine ones.14
The theory earned new relevance when the pandemic ran the largest communication experiment in history. Yang and colleagues analyzed the emails, calendars, messages, and calls of 61,182 Microsoft employees as the firm went fully remote — a natural experiment, since some employees were remote already. The causal results: collaboration networks became more static and siloed, with about 25 percent less of collaboration time spent on cross-group ties; workers added new collaborators more slowly; and communication shifted from synchronous, rich channels toward asynchronous, lean ones — precisely the channels least suited to ambiguous, novel information.15 The pattern's significance for teams: strong ties inside a group survive distance well; the weak ties and bridges that carry fresh information across an organization are what atrophy — a network-level echo of groupthink's insulation ingredient. The design responses follow the theory: deliberately schedule rich, synchronous contact for the ambiguous work (kickoffs, conflict, bad news), engineer cross-group collisions on purpose, and stop conducting emotionally loaded conversations over the leanest channel in the building. Limits: media richness is a matching rule, not a moral ranking — teams also drown in synchronous meetings, and hybrid arrangements were not what Yang's data tested.
These cases hold industry logic constant — both are safety-critical, high-volume manufacturers where a missed defect can kill — and vary one design variable: the organizational price of speaking up. If team performance flowed from member talent, Boeing's engineering workforce should dominate. If it flows from voice — psychological safety expressed through structure — then the company that pays people to stop the line should produce quality the silent one cannot. The cases test that proposition from opposite directions.
The facts: after the 2018–2019 MAX crashes killed 346 people, Congress ordered an independent review of Boeing's safety culture; the expert panel — FAA officials plus airline, labor, and aerospace-safety representatives — worked from March 2023 through February 2024, conducting more than 250 interviews and reviewing over 4,000 pages.12 Its findings map with eerie precision onto this chapter's concepts. Psychological safety: employees hesitated to report "for fear of retaliation," because managers who investigated safety reports could also control the reporters' evaluations, salaries, and promotions — the interpersonal risk calculation tilted, structurally, toward silence.23 Communication structure: reporting channels were inconsistent and confusing; employees distrusted the anonymity of the post-crash "Speak Up" program and preferred telling their managers — the one channel the incentive problem contaminated; many did not learn the outcomes of reports they did file.3 Groupthink's real ingredients: a "disconnect between Boeing's senior management and other members of the organization on safety culture" — leadership insulated from the floor's information.2 The panel issued 50 recommendations; weeks earlier, the Alaska Airlines door plug had blown out with its retaining bolts missing after factory work — a defect of exactly the kind a functioning voice system exists to catch.1 The honest caveats: the panel documented culture, not causation for any specific incident, and Boeing has contested none of the findings while claiming progress; silence here was not employee cowardice but a rational response to the price structure employees accurately perceived.
The facts: the andon system, a pillar of the Toyota Production System since the mid-century work of Taiichi Ohno, gives every line worker a cord (now often a button) whose pull summons a team leader within seconds; most problems are resolved within the work cycle, and only unresolved ones stop a line segment.56 Steven Spear and H. Kent Bowen's classic analysis of the system's "DNA" identified the deeper design: every worker is a tester of hypotheses about the work, and every problem surfaced is treated as information, not accusation.6 The 2007 Georgetown-versus-Dearborn comparison quantifies the cultural variable this chapter has been circling: identical hardware, 2,000 pulls a week versus two — three orders of magnitude of difference in what it felt safe to say, with team-leader response capacity and non-punishment as the mechanisms.4 The NUMMI joint venture supplies the natural experiment: a workforce GM had written off as its worst became, under the same voice system, among its best within a year.5 Honest caveats here too: Toyota is not a utopia — it has had its own recalls and crises, high pull-rates carry real coordination costs that the system's buffers and responders exist to absorb, and the cord works only because an entire management system (stable teams, trained responders, no-blame norms) answers it. Ford installed the cord and got two pulls a week; hardware without the response system is theater.45
| Boeing (per 2024 FAA panel) | Toyota (andon system) | |
| Price of speaking up | Perceived retaliation risk; evaluator and investigator often the same manager | Zero — pulling is the job; no punishment for 'unnecessary' pulls |
| Response to voice | Unclear channels; outcomes often unreported to the reporter | Team leader arrives in seconds; problem solved or line segment stops |
| Information flow to top | 'Disconnect' between senior management and workforce | Problems surface at the station, thousands of times a week |
| Observed behavior | Hesitation, distrust of Speak Up, preference for informal reporting | ~2,000 pulls/week (Georgetown) vs. 2 at a cord-equipped Ford plant |
| Chapter concept made visible | Psychological safety failure + contaminated channels + insulated leadership | Psychological safety engineered into hardware, roles, and response time |
Three conclusions, held with appropriate care. First, the pairing demonstrates that voice is a system property, not a personality trait: the same human beings behave as silent Boeing inspectors or prolific Toyota pullers depending on the price list and the response their organizations attach to speaking — Ford's two-pulls-a-week cord is the controlled comparison that proves hardware alone changes nothing.4 Second, measure behavior, not slogans: Boeing had a program literally named "Speak Up" and a workforce that didn't trust it; Toyota rarely uses the phrase "psychological safety" and generates thousands of acts of it weekly — pull-counts, error reports, and question rates are the real dashboard.35 Third, resist the halo in both directions: Boeing's culture was produced by identifiable structural choices (who investigates whom, what metrics rule) that any company under margin pressure can drift into, and Toyota's system costs real money in responders, buffers, and stopped segments — it survives because the company treats those costs as the price of quality rather than a variance to eliminate. The question for any team you join is the one the two factories answer differently: when someone sees a problem, what happens to them next?
Return to the central question: what makes a team more than the sum of its members, and what makes it less? Four conclusions carry.
Teams are designed, not assembled. Development follows first meetings and midpoints more than mystical stages;10 process losses respond to size, visibility, and indispensability;11 and interaction quality — turn-taking, social sensitivity — predicts collective intelligence better than the roster does.13 The manager's leverage is in the design variables, not the draft picks.
Make speaking cheaper than silence. Psychological safety is the gate through which every error, risk, and idea must pass,78 and it is set by structure — who investigates whom, what happens within seconds of a pull — far more than by posters. Pair it with standards: safety without standards is comfort; standards without safety is Boeing's silence.2
Guard the information diet. Groupthink's active ingredients are directive leadership and insulation;12 remote work's measured cost is the quiet atrophy of cross-group bridges.15 Both are cured the same way: deliberately import dissent and distance-spanning ties, and match rich channels to ambiguous conversations.14
Audit behavior, not vocabulary. Count what the cord-pull count counts: questions asked in meetings, errors self-reported, bad news that reaches you while it is still cheap. A team's true communication culture is visible in those numbers long before it is visible in an FAA report — and by then, the price has been paid by someone else.14
Notes appear as superscript numbers in the text and correspond to the numbered sources above. DOIs are provided where available; classic books are cited to their original publishers.