AI now writes half the code—but humans still clean up the mess

Syntax

The gist

AI is now writing half the code at major companies, but humans are left cleaning up a bigger, messier pile—review times are soaring and error rates are on the rise.

What to know

  • AI generates up to 50% of new code at firms like Coinbase and Zapier, but code review times have surged by as much as 91%.
  • Non-traditional developers now make up 63% of new tool users, accelerating prototyping but increasing software errors by 41%.
  • Despite widespread adoption, 82% of organizations don’t track AI’s true impact—leaving teams to restructure and hire blindly.

Productivity Gains, New Bottlenecks

AI supercharges code generation but shifts the real work—and slowdowns—to review, validation, and maintenance, exposing the limits of traditional productivity metrics.

AI-powered coding tools have delivered undeniable acceleration in code generation, with organizations like Coinbase reporting that AI now writes nearly half their code and Zapier leveraging agent automation to boost output per engineer. However, this surge in productivity has not translated into unqualified gains; new bottlenecks have emerged downstream, particularly in code review, validation, and quality assurance. Studies show that while individual developers may merge up to 98% more pull requests and save 10+ hours per week on coding, review times have ballooned by as much as 91%, and maintenance burdens have grown, shifting the productivity constraint rather than eliminating it.

The measurement paradox remains a persistent challenge: despite widespread adoption—nearly 90% of developers now use AI coding assistants daily—most organizations lack robust frameworks to accurately assess AI's true impact. Traditional metrics like lines of code or pull request counts are increasingly seen as misleading, as they fail to capture nuances such as code quality, maintainability, and the hidden costs of rework or error correction. Companies like Microsoft, Dropbox, and GitHub are pioneering more sophisticated, blended measurement approaches, tracking everything from 'bad developer days' and change failure rates to prompting efficiency and trust calibration, yet even these efforts are hampered by data access issues, subjective reporting, and the evolving nature of both tools and workflows.

AI's impact on productivity is highly variable and context-dependent, amplifying both organizational strengths and weaknesses. While some teams report transformational gains—up to 25% or even tenfold increases in individual productivity—others see negligible improvements or even negative ROI due to increased rework, erratic code quality, and persistent process bottlenecks. Factors such as developer expertise, workflow integration, prompt engineering skill, and organizational practices play a decisive role, with experienced developers benefiting most when they maintain control and oversight, while less experienced users may inadvertently generate unreviewable or insecure code, compounding downstream challenges.

The evolving landscape has prompted a shift in how productivity is conceptualized and measured: developer roles are moving from pure code writing to orchestrating, validating, and synthesizing AI-generated outputs, with frameworks like DORA and SPACE adapting to capture these new dimensions. As Nicole Forsgren and others argue, developer experience—encompassing flow state, cognitive load, and feedback loops—has become central to understanding productivity in the AI era. Ultimately, the organizations realizing the greatest gains are those that invest in robust measurement infrastructure, continuous training, and systemic workflow improvements, rather than relying on superficial adoption metrics or chasing raw output volume.

Sources
SyntaxThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The PitchTechTalksPractical Engineering ManagementThe Pragmatic EngineerOne Useful Thing

Rise of the AI Orchestrator

As AI blurs the lines between developer, designer, and product manager, technical judgment and cross-disciplinary fluency now outshine raw coding speed.

AI is fundamentally redrawing the boundaries between developers, product managers, and designers, giving rise to new archetypes like orchestrators, editors, and AI-powered full stack builders. By early 2026, organizations like LinkedIn have replaced traditional Associate Product Manager programs with 'Full Stack Builder' tracks, empowering individuals from any background to take products from idea to launch by blending coding, design, and product management skills—all amplified by AI. This shift is not just about tooling: it demands a mindset change, as the most successful builders are those who embrace AI as a collaborative partner, focusing on judgment, vision, and cross-disciplinary fluency rather than manual coding alone.

The traditional 'Coding Machine' archetype—valued for rapid, high-volume code output—has collapsed as AI democratizes code generation, shifting the bottleneck from creation to evaluation and judgment. As Copilot, Claude Code, and similar tools allow even mid-level engineers to produce volumes of code that once set senior developers apart, the new premium is on the ability to review, validate, and select the best AI-generated solutions. As the CEO of Warp notes, 'developers became orchestrators of AI agents—a role that demands the same technical judgment, critical thinking, and adaptability they’ve always had,' with data showing PRs are now 18% larger and incidents per PR up 24%, underscoring the rising importance of human oversight.

AI is accelerating the convergence of roles, enabling designers and product managers to build and iterate on software directly, while developers increasingly focus on higher-level orchestration, architectural oversight, and outcome validation. Tools like OpenAI’s GPT, Cursor, and Claude Code now let designers turn sketches into live websites and PMs to ship MVPs without writing code, fostering hybrid roles such as the 'AI Product Engineer.' However, this democratization also raises the bar for technical judgment and system-level thinking, as even non-technical builders must now guide, critique, and refine AI-driven outputs to ensure quality and strategic alignment.

While AI automates much of the technical heavy lifting, the enduring value of human roles lies in accountability, judgment, and the ability to synthesize context that AI cannot access. As software creation becomes more spec-centric—with natural language specifications treated as durable, versioned source code—developers, PMs, and designers are increasingly evaluated on their ability to define intent, make trade-offs, and maintain trust. The rise of orchestrators and editors reflects this shift: success now depends less on artifact production and more on the quality of decisions, context synthesis, and collaborative problem-solving that keep AI-generated work aligned with business and user needs.

Sources
Lenny's PodcastLenny's NewsletterDesign + AI@jasmine’s substackThe Pragmatic EngineerPaired Ends

Non-Engineers Take the Helm

AI tools flood software creation with non-traditional builders, but fragile code and new skill gaps mean oversight and onboarding matter more than ever.

AI-powered tools like Claude Code, Cursor, and Replit have dramatically lowered the technical barriers to software creation, enabling a surge of non-traditional developers—designers, marketers, product managers, and even CEOs—to build and ship applications without deep coding expertise. By late 2025, platforms such as Lovable and Vercel’s v0 were already empowering non-coders to prototype and validate ideas in hours, with data showing that 63% of 'vibe coders' were non-developers. This shift is not just anecdotal: developer tool signups have soared from 3,000 to 16,000 per day, with many new users identifying as non-engineers, and Netlify CEO Matt Biilmann notes that AI agents are fundamentally reshaping the user base and economics of software development.

While AI democratizes access, it also introduces new challenges around onboarding, skill development, and quality assurance, especially as non-experts begin shipping production code. Studies like METR’s found that experienced developers were 19% slower with AI due to time spent cleaning up AI-generated code, and the rise of 'vibe coding'—accepting code without thorough review—risks creating fragile software and developers ill-equipped to debug or maintain it. Companies like Egnyte and practitioners such as Zevi Arnovitz at Meta emphasize that human oversight, robust code review, and foundational engineering skills remain essential, even as AI accelerates productivity and enables junior professionals to take on full-lifecycle responsibilities earlier.

The evolution of AI coding tools is shifting the bottleneck in software development from technical syntax to decision-making, context management, and prompt engineering, requiring users—regardless of background—to develop new mental models and workflows. As Alexandre Pesant and others argue, successful adoption hinges on expressive interfaces, context engineering, and practical evaluation workflows that help non-technical users monitor and improve AI outputs. This new landscape is fostering hybrid roles—product engineers, design engineers, AI product creators—who blend design, engineering, and product management, but also highlights the need for continuous skill development and accessible onboarding to ensure that democratization does not simply amplify existing gaps.

Despite the promise of 'anyone can build,' onboarding hurdles and the lack of 'software vision'—the ability to recognize and frame problems as software solutions—remain significant barriers for the truly uninitiated. Firsthand accounts reveal that even with tools like Claude Code, non-engineers often struggle with initial setup, permissions, and understanding workflows, likening the experience to 'cooking with ingredients from a stranger’s fridge.' As a result, effective democratization requires not just accessible tools but also targeted training, practical evaluation methods, and a cultural shift in how software problems are identified and approached.

Sources
Super Data Science: ML & AI Podcast with Jon KrohnBartek PucekDesign + AIAI + a16zGrowth AlgorithmBen's Bites

Speed vs. Software Quality

AI-fueled rapid prototyping unleashes more bugs and bigger pull requests, making human critical review the last line of defense against disaster.

The surge in AI-driven 'vibe coding' has democratized rapid prototyping, enabling non-engineers to quickly test and iterate on ideas, but this speed comes at a cost: studies show AI coding tools can introduce up to 41% more bugs, and prototypes often lack the robustness required for production. As a result, the need for human review and engineering expertise has only intensified, with developers increasingly tasked with transforming these AI-generated drafts into secure, production-ready applications. This dynamic not only underscores the enduring value of critical human judgment in software development, but also highlights how AI's promise of accessibility paradoxically increases the demand for skilled developers to ensure quality and maintainability.

While AI tools have undeniably accelerated code generation and increased output—Coinbase, for example, reports AI now writes nearly half its code—the resulting code is often more error-prone, with enterprise AI commits showing a 30% higher rate of logic errors and PRs growing 18% larger on average. This rapid output has shifted the bottleneck from code creation to the need for robust human judgment, as incidents per PR have risen by 24% and change failure rates by 30%, making critical review and collaborative oversight more vital than ever to prevent fragile, unmaintainable software from reaching production. As James Gosling bluntly put it, 'vibe coding produces software that mostly works. That gap—between mostly and always—is where disasters hide.'

Across organizations, the impact of AI on software quality is highly variable, with some teams experiencing significant improvements while others face notable declines in maintainability and change confidence. Industry averages can be misleading, masking the volatility and risk that come with AI adoption; for example, while average improvements in metrics like Change Confidence and Code Maintainability hover around 2 points, individual companies can see swings of over 40 points in either direction. This variability underscores the necessity of organization-specific measurement, rigorous engineering hygiene, and the irreplaceable role of human oversight to ensure that AI-driven productivity gains do not come at the expense of long-term codebase health.

The evolving landscape of AI-assisted development has redefined the developer's role from code producer to orchestrator and critical evaluator, with leading organizations like MongoDB and ProductBoard emphasizing that core product code and strategic decisions remain firmly in human hands. Anthropic research and industry leaders agree that while AI can handle implementation and routine tasks, decisions requiring high-level thinking, organizational context, and 'taste'—the nuanced judgment that aligns code with business needs and user realities—are not delegable to machines. As the bottleneck shifts from coding speed to decision-making quality, the enduring value of human judgment, critical thinking, and collaborative review becomes the true gatekeeper of software quality and trust.

Sources
Super Data Science: ML & AI Podcast with Jon KrohnSyntaxSiliconANGLE theCUBEPath to Staff EngineerGrowth AlgorithmEngineering Enablement

Teams Restructure in the Dark

Organizations overhaul roles and hiring for the AI era, yet most lack the measurement to know what’s actually working or where new risks lie.

Organizational adaptation to agentic AI is fundamentally reshaping team structures, talent strategies, and the measurement of impact—often before companies are fully equipped to understand or optimize these changes. Despite 82% of organizations not tracking the effectiveness of AI tools and 60% citing the lack of metrics as a key challenge, nearly half (48%) report significant shifts in team structures, and 54% expect a decrease in hiring junior engineers. This disconnect between rapid structural evolution and lagging measurement frameworks leaves many organizations flying blind, risking suboptimal deployment of AI and missed opportunities to cultivate the right mix of skills and roles for the future.

Sources
Engineering Leadership

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.