Updated 30 June 2026

Added Matrix to the software-factory section as an example of the orchestration-platform, or agent-company, shape.

There is a scene at the start of McKinsey's spring report on agentic delivery that the whole industry seems to want to be true. A product owner logs in at nine in the morning. Overnight, while nobody was at a desk, a feature has moved from a pile of requirements to tested code; the edge cases are flagged, the architecture dependencies are validated, and the trade-offs are written up for her to read over coffee. No one worked late. The did .

It is a good scene. It is also a brochure. Most teams I see are not running a clean night shift handing finished work to a calm morning review; they are running something messier, where a person kicks off three agents, two come back useful and one comes back confidently wrong, and the afternoon goes on figuring out which was which. The near-term reality is not agents replacing the team, and it is not a team that keeps agents at arm's length. It is a mixed team, where humans and agents share the work, and the hard design questions are about the seams: who hands what to whom, at which point a human has to look, and whose name is on the result.

The open questions about hybrid teams are not about raw model capability, which improves on its own schedule. They are about division of labour, review, accountability, and the shape of the team, and those are management problems rather than model problems.

In this article:

  • The shape that's actually emerging. Humans directing and reviewing, agents drafting and executing, with the design work sitting in the handoffs.
  • are older than the agents. A 57-year-old idea, why it kept failing, and what is different this time.
  • What this means if you build software for a living. The part the buyer-side reports skip: what happens to the people who deliver software for other people.
  • What it does to the team. Smaller, flatter teams, fewer hands on boilerplate, and a hole where the apprenticeship used to be.
  • Build more, keep less. Cheaper features mean less rationing up front, and more shipping, A/B testing, and quietly retiring what loses.
  • Where the line between human and agent sits. Which work goes where, and how that line moves as the scaffolding improves.
  • Review doesn't compress. The one part of the job that gets more important, not less.
  • The consensus, and the gap underneath it. Every major firm agrees; their own surveys say most clients are still in pilots.

The shape that's actually emerging

The McKinsey report's own diagrams split the work in two. Humans hold the daytime work: reviewing output, resolving ambiguity, strengthening the structure the agents build against, aligning stakeholders. Agents do the overnight execution: enriching requirements, checking the architecture is in place, generating and testing a first version, packaging it for review. In the same report, humans sit at defined review gates rather than relaying work between every step .

So the interesting unit of design is the gate, not the agent. An agent that can write a service is not the constraint; the constraint is deciding what "done" means precisely enough that the agent can aim at it, and catching the cases where it aimed well and missed anyway. Bain, describing the same shift, puts the developer's new position as moving from writing code to architecting and orchestrating the agents that write it . The verbs that survive are the ones that were always the harder part of the job.

An org-chart of a mixed team: some nodes drawn as solid filled circles, others as open outlines, joined by thin arrows, several of which loop back on themselves.
People Direction and review.
Agents Drafting and execution.
Some nodes are filled, some are outlines; the arrows that matter are the ones that loop back. Author's illustration.

Software factories are older than the agents

The current word for the industrialised version of this is the software factory. The term is 57 years old. Hitachi applied the factory label to software with its Hitachi Software Works programme at the end of the 1960s [R14a]. Through the 1980s the approach was formalised across Hitachi, Toshiba, NEC and Fujitsu, with process control, quality metrics, and reuse libraries, all documented in Michael Cusumano's study of the period . In 2004 Jack Greenfield and Keith Short gave it a software-engineering form at Microsoft: product lines, domain-specific languages, model-driven generation, the assembly of applications from configured components . Then DevOps and CI/CD inherited the name, and the US Air Force ran literal ones.

Every wave promised the same thing, and every wave hit the same wall. Cusumano found that as of 1986, around 94% of the software sold in Japan was wholly or partly customised . Even at the high point of the factory metaphor, the work was overwhelmingly bespoke, and bespoke is exactly what an assembly line is bad at. The factory could industrialise the repeatable parts and never quite swallow the custom ones, which is where the cost always lived.

The 2026 claim is different in one specific way. Factory, the company, describes a continuous loop: signals from the outside world (bug reports, feedback, requirements) get triaged into planned changes, which are built, tested, reviewed, secured, shipped, and monitored, and the monitoring produces the next signals . The pitch is that agents let you standardise the pipeline and customise the output at the same time, so the assembly line finally handles the bespoke work it used to choke on. Factory's own line for what this does to engineers is that they stop being the people who build the software and become the people who build the factories that build it . 8090, Chamath Palihapitiya's version, frames today's tooling as single-player AI dropped into a multiplayer business, where one engineer's agent ships a change that contradicts last quarter's architecture and nobody notices; its Software Factory ties requirements, architecture, work orders and tests together through a knowledge graph, and EY has rebuilt part of its own delivery on it [R4b].

A retro computing magazine-style illustration of an automated factory line with empty workstations, where each product coming off the conveyor is a different shape and colour.
No hands Every station unattended.
No two alike Each unit customised, not stamped.
A line with nobody at the stations, and not one identical unit coming off it. Author's illustration.

The factory is old; the customisation is the new part

A software factory has meant "industrialise software" since 1969. What is new in 2026 is the claim that the line can finally absorb the custom work, which is the only work that ever paid the bill.

Two things get conflated here constantly. A software factory is a standardised requirements-to-code pipeline. An is the thing that coordinates agents as if they were an organisation, and the clearest example is Paperclip, an open-source project that gives agents an org chart, budgets, and approval gates, launched in March 2026 and past tens of thousands of stars within weeks . Matrix, a macOS app that surfaced in mid-2026, pushes that shape to its limit. It casts the whole thing as an "agent company runtime": a CEO office, departments, persistent leads that own a domain, and that execute scoped tasks beneath them, with the human left as the owner who sets objectives and approves results . It claims to run the marketing, the payments and the distribution as well as the engineering, a zero-person company that earns. Treat that with care, because the people behind it are not named, the headline is self-reported, and there is no independent account of it yet. The framing is the part that matters: it is the orchestration platform taken all the way up to the company itself. A factory is the production line; an orchestration layer is the management structure. You can have either without the other, and a lot of confusion comes from pitching one and shipping the other.

What this means if you build software for a living

The reports written for buyers tend to skip one question: what happens to the people whose living is building and customising software for other people. The factory vendors do not skip it. They are aiming at exactly that market.

BCG put numbers on it in February, from a survey of more than 115 enterprise executives and more than 75 at service providers . The headline most people quote is the $200 billion of new value, but the body is more interesting. BCG's own read is that agentic AI is deflationary in parts of the and expansionary overall, at least in the short to medium term . Providers expect the delivery pyramid to shrink by 10% to 20% over the following two years as agents absorb the routine work, while total headcount still grows, with a different mix of skills . The body-shop arithmetic, bill person-years times bodies, does not survive that. You cannot bill the hours when the hours evaporate.

The model that replaces it is showing up in the contract terms. More than 70% of the enterprise buyers BCG surveyed want output- or outcome-linked commercial models, while roughly 60% of providers still write time-and-materials or fixed-price contracts . Gartner expects the cost-to-value gap in process-centric service contracts to close by at least half by 2027 through agentic reinvention, with the impact landing squarely on IT services, consulting, and outsourcing, and contracts shifting from hours to outcomes [R8b]. Selling effort is becoming a worse business than selling results, and that is a different business to be in.

Value does not vanish in that shift; it splits by tier. The commodity middle, standard enterprise-app builds, on a deadline, offshore maintenance of well-understood systems, is the part that thins fastest, because it was already a race to the bottom on rate and it is the easiest thing for a factory to absorb. What holds its value is the work the factory cannot do without you: the domain knowledge, the institutional context that becomes the knowledge graph, the last-mile fit of software to a business that does not behave the way the requirements said it would, and the governance and accountability around all of it. BCG names this last one as a recurring revenue pool in its own right, the continuous monitoring, drift detection, human-in-the-loop escalation, audit trails and compliance reporting that agents in regulated workflows demand . Watching the agents becomes a line item.

There is a version of this that is good news for a small shop. One experienced person with agents now looks a lot like a small delivery team, which is roughly the experience behind a side project of mine that became a real product without a team behind it.

Sovereignty matters more here than the productivity multiplier. Factory sells the ability to run the whole loop inside your own walls, from fully hosted down to air-gapped with no external network , and a client who can stand up their own governed factory has less reason to rent yours. The consultancy that survives sells the factory and the accountability, not the people. EY rebuilding its delivery on a factory platform is the cleanest example of a firm doing this to its own model before someone does it to them .

I am not sure which way this goes. Customisation has beaten every software factory for 57 years. If the 2026 factories really do absorb the bespoke work, the supplier model changes in a deep way and the commodity tier goes quickly. If they hit the same last-mile wall as every prior wave, the value concentrates in the people who can do the fit, which is, again, the experienced domain people. Both roads thin the same middle. I do not think anyone selling you a platform knows which road we are on, and I am suspicious of the confidence on display.

What it does to the team

The team-shape change is the easiest part to picture and the part with the nastiest second-order problem. McKinsey's illustrative numbers have teams of eight to twelve giving way to pods of three or four, with total effort on a programme roughly halved and average team size down about 60% . Fewer hands on boilerplate, more on direction and review. The named cases point the same way: Nubank used 's to migrate a six-million-line monolith that had been scoped as a multi-year effort across more than a thousand engineers, reporting a roughly twelvefold gain in engineering hours and around twenty-fold cost savings, with humans defining the migration patterns up front and approving the output . A data migration with a human-shaped hole in the middle is what a hybrid team looks like, not the empty night shift.

Smaller usually means flatter. A pod of three or four supervising agents needs less of the coordination scaffolding a team of twelve required, and a good deal of middle management exists to coordinate hands that are now agents. McKinsey's own framing moves the human role away from manual coordination toward architecture and supervision . A layer or two of the org chart, the part that mostly existed to route work between people, has less to do once the work routes itself. I doubt those layers go quietly, but the pressure on them is real, and the more autonomous the pod, the fewer people it needs sitting above it.

And the second-order problem is the one without a fix yet: the rung juniors used to climb, the small bugs and the boilerplate and the first drafts, is the rung agents now occupy, and nobody has built the replacement.

Build more, keep less

If a feature costs an afternoon instead of a sprint, the logic that made teams choosy up front starts to weaken. A lot of product discipline was rationing in disguise: features were expensive and slow, so you argued hard about which few to build, because a wrong bet cost a quarter you would not get back. Lower the cost of the bet and you can afford to make more of them.

My guess is that teams build far more speculatively than they do now and grow less precious about pulling something back out. Online news already runs this way: a desk will test several headlines and front-page images against each other, keep the one that performs, and quietly drop the rest, or serve different versions to different readers based on what it knows about them. Features can move the same way once building three versions costs about what building one used to. Ship them, measure them, keep the one that works, retire the others.

Cheap to build is not the same as free to keep. Every shipped feature is still something to maintain, secure, and carry in your head, and a thing you ship and later remove has still cost users a little trust and the team a little attention. The choosiness does not vanish; it moves from before the build to after it, and turns into a measurement problem. Bain makes the point that raw velocity can mislead, since a team can ship faster while piling up defects and risk underneath, which is why it argues for measuring outcomes rather than output . Once the constraint stops being how fast you can build and becomes how well you can tell what to keep, the work that matters is choosing what survives.

Where the line between human and agent sits

The division of labour is not fixed; it moves as the scaffolding around the model improves. The , the skills, the context handling, the orchestration layer, all of it shifts work from the human column to the agent column over time, which is why "what should the agent do" is a question with a different answer every quarter.

The tooling that has shipped this year is converging on the same answer about where the human belongs. GitHub's app, announced at Build in June, is built to supervise several agents at once, each in its own isolated worktree, with shared surfaces for a human to redirect the work and an Agent Merge step that carries a change through review and checks under conditions the human sets . The persistent-teammate pattern is everywhere too: background tasks in the major assistants, , and the always-on personal agents like Microsoft's Scout, built on the open-source OpenClaw project . The common thread across all of it is that the agent does more and the human keeps the decision about what ships.

rejected

accepted

Human: intent, spec,
acceptance criteria

Agents: draft,
build, test

Human
review gate

Merge and ship

Human: refine rules,
reprioritise

How far the line can move before the loop falls apart is the open question.

Review doesn't compress

If execution gets cheap, the bottleneck moves to the one thing that does not get cheaper: deciding whether what came back is any good. A team's output is only as trustworthy as its review discipline.

A team's output is only as trustworthy as its review discipline, and a hybrid team's output is mostly review.

The tooling knows this. code review now runs automatically on pull requests, on GitHub's own Actions infrastructure, pulling in repository skills and context and routing harder changes to a higher-reasoning tier . Review is becoming an agent in its own right, which is useful and also raises the obvious problem of who reviews the reviewer. 's own coding research puts a number on the limit: developers use AI for roughly 60% of their work but report being able to fully hand off only 0% to 20% of tasks, a gap Anthropic calls the . The other 40% to 80% is a human staying in the loop because the output cannot yet be trusted unattended.

The failure mode is a confident wrong answer

A broken build announces itself. An agent that misread the spec ships something that compiles, passes the tests it wrote, and is wrong in a way only a reviewer who understands the domain will catch. The review is not a formality at the end; it is where the trust comes from.

Accountability does not transfer to the agent. Whoever merged the change owns it, and the BCG buyers asking for outcome-linked contracts are, in effect, asking someone to own the outcome . That is a human signing their name, and it is the part of the job that gets more valuable as the typing gets automated.

The consensus, and the gap underneath it

Every major firm is saying the same thing at once. Bain calls it the shift from AI-assisted to AI-led development, with hybrid human-agent teams it projects at five to ten times current productivity . Accenture is running AI reinvention across more than 2,000 engagements and selling the "10x bank", where one person leads a team of AI co-workers . BCG, Capgemini and EY have each built a branded factory or platform and a delivery practice around it . When parties with different incentives describe the same shape, the shape is probably real.

Their own survey data tells a smaller story. The big multipliers are expectations or single case studies; the measured numbers are modest. Bain is explicit that most companies are still seeing single-digit efficiency gains while expectations run to five and ten times, and that the multiplier is what executives now anticipate, not what they have banked . BCG found enterprises expecting 30% to 40% productivity while providers commit to 6% to 15%, and nearly 60% of enterprises reporting no measurable change in total cost of ownership yet on deals that include agentic AI . McKinsey's global survey of nearly 2,000 organisations found 62% experimenting with agents but only about a third scaling AI across the enterprise at all, and 39% reporting enterprise-level profit impact . 's benchmark of 1,050 IT leaders found an average of twelve agents per company, half of them operating in isolation rather than as a coordinated system . Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 over cost, unclear value, and weak controls, and says outright that fully autonomous agents are not ready for most enterprise use cases .

Experimenting with agents Scaling AI enterprise-wide Tasks fully handed off Projects cancelled by 2027
62% ~1 in 3 0–20% >40%
McKinsey, n≈2,000 McKinsey Anthropic Gartner forecast

In financial services, the vertical I work in, the gap is wider still: Capgemini's research found only about 10% of firms running agents at scale while 80% are in ideation or pilot, and roughly half of banks and insurers are creating new roles specifically to supervise the agents . That last detail is the one I keep coming back to, because a job whose entire description is "watch the agents" is the hybrid team made into an org chart.

The team is the unit now

The unit of software delivery is shrinking from the team to the human-agent team, and in its industrialised form to the factory, and that smaller unit changes both how you compose a team and how you sell the work a team does. The capability questions answer themselves on the model-makers' schedule. The questions that are yours are where to put the handoffs, how to keep the review discipline that makes the output trustworthy, who signs their name to what shipped, and which tier of the old business you are still in once the commodity middle thins out.

The firms selling this are right that it is real and right that it is early. They are quieter about the fact that, by their own numbers, most of their clients are still in the pilot, the multipliers are mostly still expectations, and roughly half the agents already deployed do not talk to each other.

Final thought

Composing a team out of people and agents, and running it so the results can be trusted, is the management skill of the next few years. And the firms with the loudest decks are not the ones most likely to be good at it.


References20
  1. 1McKinsey & Company, "Rewiring software delivery for the agentic era" (Jared Moon, Rory Walsh, Vito Di Leo, Adam Thelwall), May 2026. mckinsey.com ↗
  2. 2Factory, "Factory 2.0: From coding agents to software factories", 2026. factory.ai ↗ · x.com ↗ Accessed 2026-06-17
  3. 3Bain & Company, "The Rise of the AI Development Life Cycle", June 4, 2026. Surveys: Bain AI Deployment Perspectives Survey, August 2025 (n=205); Bain Tech & Engineering Survey, March 2026 (n=293). bain.com ↗
  4. 4Boston Consulting Group, "The $200 Billion Agentic AI Opportunity for Tech Service Providers" (Vikash Jain, Sudhanshu Chawla, et al.), February 20, 2026. Surveys: 115+ enterprise executives, 75+ service-provider executives. bcg.com ↗
  5. 5Cognition, "Nubank" customer story (Devin). devin.ai ↗ Accessed 2026-06-17
  6. 6McKinsey & Company / QuantumBlack, "The state of AI in 2025: Agents, innovation, and transformation", November 5, 2025. Survey fielded June 25–July 29, 2025; n=1,993. mckinsey.com ↗
  7. 7Anthropic, "2026 Agentic Coding Trends Report". resources.anthropic.com ↗
  8. 8Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", June 25, 2025. gartner.com ↗
  9. 9Salesforce / MuleSoft, "Connectivity Benchmark Report 2026" (11th annual; Vanson Bourne with Deloitte Digital), February 2026; n=1,050 IT leaders, fielded October–November 2025. salesforce.com ↗
  10. 10Capgemini Research Institute, "World Cloud Report — Financial Services 2026", March 2026. capgemini.com ↗
  11. 11Accenture, "Agentic AI and the future of work in financial services" (Banking Top Trends 2026), February 2026. bankingblog.accenture.com ↗
  12. 12Gartner, "Hype Cycle for Agentic AI, 2026", April 2026. gartner.com ↗
  13. 13EY (Ernst & Young LLP) and 8090, "Ernst & Young LLP and 8090 launch AI-native EY.ai Product Development Lifecycle (PDLC)", March 18, 2026. ey.com ↗
  14. 14Michael A. Cusumano, "Japan's Software Factories: A Challenge to U.S. Management", Oxford University Press, 1991 (the ~94%-customised figure cited to Cusumano, 1988). global.oup.com ↗
  15. 15Jack Greenfield and Keith Short, "Software Factories: Assembling Applications with Patterns, Models, Frameworks, and Tools", Wiley, 2004. wiley.com ↗
  16. 16GitHub, "GitHub Copilot app: the agent-native desktop experience", GitHub Blog, June 2, 2026. github.blog ↗
  17. 17GitHub Docs, "About GitHub Copilot code review". docs.github.com ↗ Accessed 2026-06-17
  18. 18Computerworld, "Microsoft unveils Scout, an autonomous AI agent built on OpenClaw", June 2026. computerworld.com ↗
  19. 19Paperclip, open-source agent orchestration platform (paperclipai/paperclip), launched March 2026. paperclip.ing ↗ · github.com ↗ Accessed 2026-06-17
  20. 20Matrix (matrix.build), product site and "Matrix Guide": an "agent company runtime" with owner/lead/worker/system roles, a CEO Office and OKR memory, departments and structured handoffs; a downloadable macOS app ("web app coming soon"); tagline "Launch a 0-Person Company that actually earns". The GDPval-Bench score is self-reported, and the operating company, founders and funding are undisclosed as of access. matrix.build ↗ · matrix.build ↗ Accessed 2026-06-30