What Makes a High-Performing Team?
What a decade of research across 39,000 professionals tells us about building teams that actually deliver
Over the years, I’ve been asked countless times what a high-performing team actually looks like and how to build or coach one. As a consultant who has helped many organisations embrace a better way of working, I have a pretty good idea of what the hallmarks of a high-performing team are — both cultural and technical. But personal experience, however hard-won, only takes you so far. There is also a vast body of research around this topic, and I wanted to surface it all in one place.
Over the past decade, rigorous evidence — spanning Google’s Project Aristotle, the DORA research programme, academic studies, and hard-won lessons from companies like Spotify, Amazon, and GitLab — has converged on a surprisingly consistent set of themes. These studies don’t all measure the same things or use the same methods, and they don’t agree on everything. What follows is not a unified theory of team performance — it’s a curated synthesis of the strongest evidence available. Some of the conclusions are intuitive. Others are deeply counterintuitive. All of them are grounded in data, though the strength of that grounding varies, and I’ll try to be honest about where.
It’s Not About the Stars — It’s About the Dynamics
One of the most influential studies in this space is Google’s Project Aristotle, a two-year internal study of 180 teams examining hundreds of attributes. The single most critical factor in team success was not individual talent, technical skill, or seniority — it was psychological safety, defined by Harvard professor Amy Edmondson as “a shared belief held by members of a team that the team is safe for interpersonal risk taking.”
Google identified five key dynamics, in decreasing order of importance: psychological safety, dependability, structure and clarity, meaning, and impact. What did not matter in that study, with those measures was equally revealing: colocation, consensus decision-making, individual extroversion, individual performance, workload size, seniority, team size, and tenure showed no significant correlation with team effectiveness. That doesn’t mean these factors never matter — anyone who’s managed a distributed team or scaled a 30-person department knows better. But within Google’s context, they weren’t the differentiators.
This finding challenges the persistent myth of the “10x engineer.” The evidence says that who is on the team matters far less than how the team works together. Edmondson’s foundational 1999 study — one of the most cited papers in organisational psychology — established that psychological safety enables learning behaviour, which in turn drives performance improvement. The DORA research programme validated the connection: generative cultures that include psychological safety predict both software delivery performance and organisational performance.
An important nuance: psychological safety enables performance, but it doesn’t cause it on its own. Plenty of teams feel safe and still don’t ship anything useful. Safety without standards, direction, or healthy pressure can produce very comfortable mediocrity. That’s why Google’s research found it was the foundation — the other four dynamics (dependability, structure, meaning, impact) stack on top. Safety alone is not the goal. Safety in service of high expectations is.
The practical implication is profound. Hiring alone won’t get you there. You have to build the conditions for high performance, not just recruit for it.
Culture Is Not a Perk — It’s a Performance Predictor
Sociologist Ron Westrum identified three types of organisational culture: pathological (power-oriented, fear-based), bureaucratic (rule-oriented, turf-protecting), and generative (performance-oriented, mission-focused). The DORA research demonstrated that generative culture predicts software delivery performance, organisational performance, and job satisfaction.
Generative cultures are characterised by good information flow, high cooperation and trust, bridging between teams, and conscious inquiry. But here’s what makes this finding actionable: the research shows that adopting technical practices — continuous integration, test automation, trunk-based development — actively shifts culture over time. Organisations do not need to “fix culture first.” The relationship is bidirectional: culture enables technical practices, and technical practices shape culture.
A word of caution, though. This bidirectional claim is easier to write than to live. In messy organisations — the ones with blame-heavy post-mortems, territorial middle managers, and approval chains that strangle momentum — trying to introduce trunk-based development or continuous delivery can backfire. The practice gets adopted in name only, or it surfaces conflicts that the culture isn’t ready to handle. The research says the path exists; it doesn’t say the path is clean. In reality, it’s often political, uncomfortable, and slow. But the evidence that it can work, even if it doesn’t always, is worth something.
Size Matters — But Not the Way You Think
Amazon’s two-pizza teams, Hackman and Vidmar’s research, Scrum guidance, and Dunbar’s number all point in the same direction: the optimal team size is somewhere around 5 to 9 members. These sources come from different domains — organisational psychology, military anthropology, practitioner heuristics, corporate operating models — and they aren’t studying the same phenomenon in the same way. But the convergence is suggestive. The underlying mechanism is well-understood: communication channels grow quadratically — a 10-person team has three times the channels of a 6-person team. Beyond that threshold, productivity per person declines as coordination costs overwhelm individual contributions.
And when a project is already late? Fred Brooks’s 1975 observation remains validated by five decades of experience: “adding manpower to a late software project makes it later.” The ramp-up time for new members, combined with the explosion in communication overhead and the non-divisibility of many software tasks, means that the intuitive response — throw more people at the problem — is usually the wrong one.
This applies to team composition as well — but with an essential caveat. Cognitive diversity drives innovation only when combined with psychological safety and inclusive practices. Without those conditions, diverse teams may actually underperform homogeneous ones. When the conditions are right, however, the effects are striking: research published by the National Institutes of Health found that cognitively diverse teams solve complex tasks significantly faster and produce more innovative solutions. The implication is that hiring for diversity must be accompanied by investing in inclusive team dynamics — one without the other can backfire.
Your Software Architecture Is Your Org Chart (and Vice Versa)
In 1967, Melvin Conway observed that organisations design systems that mirror their communication structures. Nearly sixty years later, MIT and Harvard Business School researchers found “strong evidence to support the mirroring hypothesis” — products built by loosely-coupled organisations are significantly more modular than those from tightly-coupled ones.
This has moved from observation to strategic tool. Martin Fowler endorses the “Inverse Conway Manoeuvre” — deliberately structuring teams to produce the desired architecture. Matthew Skelton and Manuel Pais built the entire Team Topologies framework on this principle, defining four team types: stream-aligned (delivering direct customer value), platform (reducing cognitive load via self-service), enabling (coaching other teams temporarily), and complicated-subsystem (handling specialist technical components).
The key insight from Team Topologies is that the primary benefit of a platform team is not efficiency — it’s reducing the cognitive load on stream-aligned teams. When cognitive load is managed, teams can focus on their core mission. When it isn’t, every team becomes a bottleneck for every other team.
Measure What Matters — And Only What Matters
The DORA research programme originally identified four key metrics that predict software delivery performance: deployment frequency, lead time for changes, change failure rate, and time to restore service. A fifth metric — reliability — was added in subsequent years, and the 2024 report further refined the framework by replacing “mean time to recovery” with failed deployment recovery time (moving it from stability to throughput) and introducing rework rate as a new stability measure.
From 2018 through 2024, DORA grouped teams into four performance clusters — elite, high, medium, and low — derived each year via cluster analysis of survey respondents. The table below, based on the original Accelerate benchmarks, illustrates the scale of the gap between top and bottom performers:
ClusterDeploy FrequencyLead TimeFailure RateRecovery TimeEliteOn demand< 1 hour~5%< 1 hourHighDaily to weekly1 day - 1 week~10%< 1 dayMediumWeekly to monthly1 week - 1 month~15%1 day - 1 weekLowMonthly+1 month+~64%1 week+
In 2025, DORA retired these tiers entirely and replaced them with seven team archetypes — such as “Harmonious High Achiever,” “Stable and Methodical,” and “Legacy Bottleneck” — that blend delivery metrics with human factors like burnout and friction. The shift reflects a recognition that a single linear ranking oversimplifies how teams actually perform: a team can have high throughput but crushing burnout, or modest deployment frequency but exceptional stability. For practitioners, the implication is significant: “high-performing” is no longer a single target to aim for. It means understanding which archetype your team resembles and addressing its specific weaknesses, rather than chasing a universal ideal. The field is moving from “are we elite?” toward “what kind of team are we, and what’s holding us back?”
Regardless of how the clusters are drawn, one finding has remained consistent across every year of the research: throughput and stability are not trade-offs — the highest-performing teams excel at both. This demolishes the common excuse that moving fast requires sacrificing quality. If you’re deploying monthly, aim for weekly first — not “on demand.”
But beware the wrong metrics. McKinsey’s 2023 attempt to measure individual developer productivity generated fierce backlash from the engineering community, with critics including Gergely Orosz and Kent Beck arguing the framework would “do far more harm than good.” Among engineering leaders, the prevailing view is that individual-level productivity metrics are counterproductive and damage trust. Measure teams and systems, not individuals. The SPACE framework — Satisfaction, Performance, Activity, Communication, and Efficiency — formalises the principle that developer productivity cannot be reduced to a single dimension or metric. At least three of the five dimensions should be measured simultaneously to avoid misleading conclusions.
The Practices That Separate the Best From the Rest
Technical practices are not just implementation details — they are culture shapers. The Accelerate research found that teams practising trunk-based development deploy over two hundred times more frequently with over a hundred times faster lead times than the lowest-performing teams. Teams that work off trunk or branches lasting less than a day have significantly higher performance.
A necessary caveat: most of this evidence is correlational, not causal. High-performing teams practice trunk-based development, but high-performing teams also tend to be the ones capable of adopting these practices well. The arrow of causation isn’t always clear. Treat these practices as strongly associated with high performance rather than as guaranteed recipes for it. And note that these practices done badly can actively hurt: trunk-based development without solid CI and feature flags leads to broken mainlines; test automation without discipline produces brittle suites that slow everything down and erode trust in the tests themselves.
Code review velocity matters too. Google’s code review practices demand a maximum one-business-day response time, and the overwhelming majority of their engineers report satisfaction with the process. The primary purpose is not gatekeeping — it’s “making sure that the overall code health of Google’s code base is improving over time.” Speed and quality in code review are not trade-offs. Fast turnaround prevents context-switching costs and unblocks developers, while the review itself maintains code quality and spreads knowledge.
Test automation is another proven driver. When Google Web Server hit a crisis in 2005 — 80% of pushes causing user-impacting bugs — mandatory test automation and the “Test Certified” programme resolved the crisis within two years. Mike Cohn’s test pyramid (many unit tests, fewer integration tests, few end-to-end tests) remains the most validated testing strategy.
Pair programming, by contrast, shows mixed results. A meta-analysis found a small positive effect on quality and a medium positive effect on duration, but a medium negative effect on effort. Quality benefits are strongest for complex tasks. The evidence supports pair programming as a tool to deploy selectively — for complex tasks, knowledge transfer, and onboarding — not as a mandatory practice.
Autonomy With Alignment — The Balancing Act
High-performing teams need both freedom and direction. OKRs (Objectives and Key Results) provide the mechanism: at Google, the majority of OKRs are developed bottom-up and then aligned to company goals. Fully top-down OKRs reduce autonomy; fully bottom-up OKRs risk misalignment. This balance enables what the DORA research calls “loosely coupled architecture with tightly aligned goals.”
Spotify’s organisational model — squads, tribes, chapters, guilds — offers a cautionary tale about the limits of structure. Spotify themselves have publicly acknowledged the model never fully worked as described. Tribes became siloed, cross-tribe coordination proved difficult, guilds lost effectiveness at scale, and high autonomy without collaboration processes led to duplicated effort. Companies that adopted the structure without the underlying culture of trust and psychological safety consistently failed.
The lesson: organisational structure is necessary but insufficient. It works only when underlaid by generative culture. You can’t copy another company’s org chart and expect their results.
Developer Experience Is a Strategic Lever
The DevEx framework, published in ACM Queue in 2023, distils developer experience to three core dimensions: feedback loops (speed and quality of responses to developer actions), cognitive load (mental processing required for tasks), and flow state (energised focus, full involvement, and enjoyment).
All three are actionable. Feedback loops can be measured through build times, CI/CD pipeline duration, and code review turnaround. Cognitive load can be reduced through better documentation, simpler architecture, and fewer context switches. Flow state can be enabled by reducing interruptions and providing blocks of uninterrupted time.
Asynchronous communication plays a role here. GitLab, a fully remote company with 2,000+ employees, operates on an “async-first” basis. Async communication minimises distractions and enables deeper focus, but success requires deliberate investment in documentation and clear communication guidelines. The trade-off is real: async can slow decision-making and reduce personal connection if not managed carefully.
Invest in Beginnings and Belonging
According to widely-cited research, organisations with strong onboarding processes see dramatically higher new hire retention and productivity — though these figures originate from a single industry study and should be treated as directional rather than precise. More robustly sourced: Google found that new hires paired with buddies reached full efficiency significantly faster, and one company reduced time-to-productivity from six weeks to ten days through structured onboarding with automation.
The most effective onboarding programmes share common elements: hardware and access configured before day one, a buddy or mentor pairing, a 30/60/90-day milestone framework, a first meaningful commit within the first week, and architecture documentation that explains why, not just how.
Mentorship extends beyond onboarding. Deloitte found that millennials with mentors are roughly twice as likely to stay at their organisation long-term, and employees involved in mentoring programmes have markedly higher retention rates. The evidence broadly supports mentorship as beneficial, though specific, evidence-based guidance for implementing it in software engineering contexts remains a real gap. Most of the research is cross-industry, and what works for management consulting or sales may not transfer directly to a codebase with steep domain complexity. Teams should treat mentorship as a high-probability bet worth making, while measuring retention and ramp-up outcomes locally rather than assuming the general statistics apply.
The Anti-Patterns That Destroy Performance
Knowing what to avoid is as important as knowing what to do.
Burnout is not an individual weakness — it’s an organisational design problem. The Accelerate research identifies six organisational risk factors: work overload, lack of control, insufficient rewards, breakdown of community, absence of fairness, and value conflicts. The vast majority of developers report experiencing work-related burnout. The 2024 DORA report found a specific, actionable intervention: teams with stable priorities face significantly less of it.
Measuring developers by lines of code, commit count, or individual output metrics creates perverse incentives and damages team culture. Organisations with generative culture — which avoids such metrics — consistently outperform those that don’t. The SPACE framework was explicitly designed as an alternative to these reductive metrics.
Platform engineering can backfire. The 2024 DORA report found measurable decreases in both throughput and change stability when teams were required to exclusively use internal platforms. The benefit comes from voluntary adoption of well-designed platforms, not mandated use.
And AI? The early data is more nuanced than the hype suggests. The 2024 DORA report found that as AI adoption increases, delivery stability drops measurably — even as developers feel more productive. One plausible interpretation — though not yet proven — is that AI-assisted coding tends to produce larger changesets, and larger changesets introduce more risk. Without AI-specific code review practices and batch size discipline, teams ship more code but break more things. This is not an argument against AI adoption — it’s an argument for treating it as a practice that requires the same measurement rigour as any other. Monitor your change failure rate and batch sizes before and after AI rollout. If stability drops, slow down and adjust.
The Bottom Line
Before listing what the research points to, it’s worth naming what it largely leaves out: the external forces that often dominate outcomes more than any internal practice. Funding models, leadership churn, regulatory pressure, market dynamics, reorgs — these shape what teams can actually do far more than whether they’ve adopted trunk-based development. The research tends to study teams in relatively stable contexts. If your organisation is in the middle of a merger, a layoff, or a pivot, the advice here still applies in principle, but the path to applying it is much harder and much more political than the research alone suggests.
With that caveat, there are clear patterns across these studies, even if the details vary. High-performing software teams are not built by assembling the most talented individuals. They are built by creating the conditions in which ordinary professionals can do extraordinary work together.
Those conditions are:
Psychological safety as the foundation — people must feel safe to take risks, ask questions, and admit mistakes
Generative culture — mission-focused, high-trust, with good information flow
Small, cognitively diverse teams — 5 to 9 members with managed cognitive load
Technical excellence — CI/CD, trunk-based development, test automation, fast code reviews
Autonomy with alignment — bottom-up goals connected to organisational mission
Great developer experience — fast feedback loops, low cognitive load, protected flow state
Investment in people — structured onboarding, mentorship, inclusive practices
The right metrics — team-level outcomes, not individual activity
The relationship between these elements is not linear. Technical practices shape culture. Culture enables technical practices. Structure constrains architecture. Architecture constrains structure. The organisations that succeed are the ones that treat all of these as a system, not a checklist.
If I had to distil it even further, the research points to what I think of as T*D — three practices that, taken together, capture the essence of high-performing teams. Trunk-based development: integrate continuously, keep branches short-lived, and ship in small increments. Test-driven development: build quality in from the start rather than inspecting it in at the end. And team-focused development — what some call social programming — pairing, mobbing, and ensemble work that spreads knowledge, reduces bus-factor risk, and turns code review from a bottleneck into a conversation. Each of these reinforces the others. Trunk-based development demands good tests. Good tests demand shared understanding. Shared understanding comes from working together. T*D is where technical excellence and team culture meet.
The good news? You don’t have to fix everything at once. Pick one practice. Implement it well. Measure the result. The evidence says that even small changes — adopting trunk-based development, speeding up code reviews, running genuine retrospectives with follow-through — create virtuous cycles that build momentum over time.
The evidence is strong, but it’s suggestive rather than prescriptive. Context matters enormously. The hard part was never knowing the practices — it’s navigating the trade-offs in your specific environment, with your specific constraints, and your specific people. The research won’t make those decisions for you. But it can tell you which directions have worked for others, and which ones haven’t. That’s worth paying attention to.
References
AWS Executive Insights. “Amazon’s Two Pizza Teams.” aws.amazon.com.
Martin Fowler. “Two Pizza Team.” martinfowler.com.
Mountain Goat Software. “The Just-Right Size for Agile Teams.” mountaingoatsoftware.com.
Robin Dunbar. “Dunbar’s Number.” Referenced via psychsafety.com and Wikipedia.
Fred Brooks. “The Mythical Man-Month: Essays on Software Engineering.” Addison-Wesley, 1975.
Alan MacCormack, John Rusnak, Carliss Baldwin. “Exploring the Duality between Product and Organizational Architectures.” Harvard Business School.
Nagappan, Murphy, Basili. “The Influence of Organizational Structure on Software Quality.” University of Maryland.
“Conway’s Law Revisited: The Evidence for a Task-Based Perspective.” IEEE Software.
Matthew Skelton, Manuel Pais. “Team Topologies: Organizing Business and Technology Teams for Fast Flow.” IT Revolution, 2019.
Nicole Forsgren, Jez Humble, Gene Kim. “Accelerate: The Science of Lean Software and DevOps.” IT Revolution, 2018.
Amy Edmondson. “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly, 44(2), 1999.
Google re:Work. “Guide: Understand team effectiveness.” rework.withgoogle.com.
DORA. “Capabilities: Generative organizational culture.” dora.dev.
Ron Westrum. “A typology of organisational cultures.” BMJ Quality & Safety, 2004.
IT Revolution. “Westrum’s Organizational Model in Technology Organizations.” itrevolution.com.
Boston Consulting Group. “How Diverse Leadership Teams Boost Innovation.”
PMC/NIH. “When and how is team cognitive diversity beneficial?” pmc.ncbi.nlm.nih.gov.
ACM Digital Library. “Diversity and Teamwork in Student Software Teams.” dl.acm.org.
DORA. “DORA’s software delivery performance metrics.” dora.dev.
DORA. “Accelerate State of DevOps Report 2024.” dora.dev.
Google Cloud Blog. “Use Four Keys metrics to measure your DevOps performance.” cloud.google.com.
Forsgren, Storey, Maddila, Zimmermann, Houck, Butler. “The SPACE of Developer Productivity.” ACM Queue, Vol 19(1), 2021.
Gergely Orosz. “Measuring developer productivity? A response to McKinsey.” newsletter.pragmaticengineer.com.
McKinsey. “Developer Velocity: How software excellence fuels business performance.” mckinsey.com.
InfoQ. “Trunk Based Development as a Cornerstone for Continuous Delivery.” infoq.com.
Google Engineering Practices. “The Standard of Code Review.” google.github.io.
Google. “Code Review - Software Engineering at Google.” abseil.io.
Hannay, Dyba, Arisholm. “The effectiveness of pair programming: A meta-analysis.” Information and Software Technology, 2009.
Mike Cohn. “Succeeding with Agile: Software Development Using Scrum.” Addison-Wesley, 2009.
Google re:Work. “Guide: Set goals with OKRs.” rework.withgoogle.com.
John Doerr. “Measure What Matters.” Portfolio/Penguin, 2018.
Henrik Kniberg, Anders Ivarsson. “Scaling Agile @ Spotify.” Crisp, 2012.
Noda, Storey, Forsgren, Greiler. “DevEx: What Actually Drives Productivity.” ACM Queue, 2023.
GitLab Handbook. “How to embrace asynchronous communication for remote work.” handbook.gitlab.com.
Cortex. “Developer Onboarding Guide.” cortex.io.
Deloitte. Mentorship and retention research. Referenced via guider-ai.com.
OSTI. “An Exploration of the Mentorship Needs of Research Software Engineers.” osti.gov.
ScienceDirect. “Burnout in software engineering: A systematic mapping study.” 2022.
Ellahi, Rehman, Javed, Sultan, Rehman. “Impact of Servant Leadership on Project Success.” SAGE Open, 2022.
Spotify Engineering. “Squad Health Check model.” engineering.atspotify.com, 2014.
Patrick Lencioni. “The Five Dysfunctions of a Team.” Jossey-Bass, 2002.
ADR GitHub. “Architectural Decision Records.” adr.github.io.
Spotify Engineering. “When Should I Write an Architecture Decision Record.” 2020.
PMC. “Perceived diversity in software engineering: a systematic literature review.” 2021.
Martin Fowler. “Team Topologies.” martinfowler.com.
Inc. “Google Spent 2 Years Studying 180 Teams.” inc.com.
LeaderFactor. “Project Aristotle Psychological Safety.” leaderfactor.com.
LaunchDarkly. “Elite Performance with Trunk-based Development.” launchdarkly.com.
Dr. Michaela Greiler. “Code Reviews at Google are lightweight and fast.” michaelagreiler.com.
LeadDev. “What McKinsey got wrong about developer productivity.” leaddev.com.
GetDX. “Highlights from the 2024 DORA State of DevOps Report.” getdx.com.
ScienceDirect. “Burnout in software engineering: A systematic mapping study.” 2022.
DORA. “State of AI-assisted Software Development 2025.” dora.dev.
Kent Beck. “Measuring developer productivity? A response to McKinsey.” Substack, 2023.

