Context engineering emerges as AI’s key bottleneck

Venture Beat

The gist

Context engineering is now the make-or-break secret weapon—and chief obstacle—for scaling reliable, autonomous AI agents in real enterprise workflows.

What to know

  • By mid-2025, AI leaders like Karpathy declared context engineering essential for managing dynamic, probabilistic data flows, with platforms like Dagster underpinning next-gen enterprise AI pipelines.
  • Agentic workflows exploded in 2026 thanks to tools like Elastic’s Agent Builder and LangChain’s Deep Agents—but these advances hinge on robust context orchestration and clever small models.
  • Despite massive AI investments, 42% of enterprises still lack formal governance, fueling costly failures and forcing the rise of new human roles like AI operators and Chief AI Officers to maintain trust and control.

Context Engineering’s High Stakes

Context engineering has rapidly evolved into a mission-critical discipline, with elite practitioners and tightly managed context windows now determining which enterprises can reliably scale AI agent performance.

Context engineering emerged as a pivotal discipline in mid-2025, gaining recognition from AI luminaries like Toby Lokta and Karpathy for its role in managing costly and complex context data that fuels AI agents. This discipline redefined data orchestration beyond traditional pipelines, emphasizing the orchestration of dynamic, probabilistic context flows essential for reliable AI workflows. Platforms such as Dagster have become foundational by enabling semantic, structured data pipelines that serve as the backbone of enterprise AI strategies, with orchestration described as the 'beating heart' that ensures precise scheduling, quality testing, and evaluation within these context-driven pipelines.

By late 2025, practitioners confronted the phenomenon of 'context rot,' where simply increasing context data volume degraded large language model (LLM) performance, as benchmarks like Nol Lima demonstrated significant quality drops beyond 100K to 200K tokens despite models supporting up to 2 million tokens. This led to refined context engineering strategies focusing on tightly curated context windows, often consolidating multiple tools into unified 'super tools' for indexed content access—a practice adopted by companies including Cloudflare and Anthropic. Context engineering thus evolved into a high-status role akin to early software engineering, demanding meticulous resource management to optimize AI agent reliability.

Entering 2026, enterprises like Workday and Virgin Media O2 demonstrated that implementing dedicated context layers could transform AI from experimental pilots to production-ready systems, achieving over fivefold accuracy improvements and scaling usage to thousands of employees with millions of platform interactions. This success underscored the necessity of bridging technical data with business readiness, as DGT’s chief data officer emphasized that 'context makes data business ready.' Best practices emerged advocating for starting with focused AI use cases, embedding business stakeholders as context engineers to ensure domain knowledge is integrated early, thereby making context an integral, reusable component of AI workflows.

Throughout 2026, the foundational challenge crystallized around engineering the context itself rather than the AI models, with experts like Ravi Marwaha and Gartner analysts highlighting that AI agents fail not due to model deficiencies but because of missing, fragmented, or stale context data. This realization propelled the context layer from a conceptual framework to a critical budget item, integrating semantics, operational state, and provenance to provide real-time, governed, and auditable knowledge essential for AI decision-making. Gartner projects that organizations prioritizing semantic context can boost AI accuracy by up to 80% and reduce costs by 60%, while the discipline of context engineering now commands 75% influence over AI output quality, marking a paradigm shift from prompt engineering to comprehensive context orchestration as the cornerstone of reliable, scalable enterprise AI.

Sources
The AI in Business PodcastAdaline LabsMetadata WeeklyContext & ChaosFortuneDaily Dose of Data Science

Agentic Workflows Rewrite AI

Multi-agent orchestration and specialized small models are transforming AI from static tools into dynamic, context-aware collaborators that demand entirely new standards for interoperability and reliability.

The evolution from isolated AI models to dynamic, agentic workflows is fundamentally driven by the rise of context engineering, a discipline Elastic predicts will define 2026 as the 'year of context engineering.' Elastic’s Agent Builder exemplifies this shift by integrating context deeply into multi-agent orchestration, enabling AI agents to operate cohesively within enterprise workflows rather than as standalone entities. As Ken Exner of Elastic highlighted at AWS re:Invent 2025, this approach leverages open standards like the maturing Machine-Callable Protocol (MCP) to support secure, scalable, and production-ready AI agent applications across diverse LLMs and cloud providers, underscoring the critical role of context in reliable AI operations.

Specialized small language models, such as embeddings and Reranker models developed by Jina, have become essential components in optimizing context for large language models (LLMs), enhancing precision in multi-modal and multilingual AI workflows. These models act as the 'brain' behind orchestration frameworks by transforming complex data into searchable, relevant formats, enabling AI agents to selectively extract, rank, and mask context effectively. This nuanced context engineering ensures that LLMs receive distilled, high-value information, which is crucial given their tendency to confidently fill gaps in ambiguous inputs without recognizing errors, a challenge highlighted by multiple experts in early 2026.

Agentic workflows represent a paradigm shift from deterministic, static processes to dynamic, autonomous systems capable of reasoning, decision-making, and continuous iteration based on real-time information. Companies like Zapier illustrate this by orchestrating multiple AI agents that collaborate seamlessly, integrating context, data, and tool access to perform complex business tasks even in the user’s absence. However, because AI agents are inherently stateless and lack memory between invocations, robust context engineering and standards like MCP are indispensable to ensure each agent call is fully informed and capable of effective action, thus enabling scalable, autonomous AI operations.

Industry leaders including LangChain’s CEO Harrison Chase emphasize that advancing AI agent capabilities hinges not on model improvements alone but on evolving the surrounding 'harnesses'—the orchestration layers that manage context, tool integration, and agent-environment interactions. Modern harnesses like LangChain’s Deep Agents empower autonomous agents to plan over long tasks, delegate to specialized subagents, and maintain token efficiency, all while relying on sophisticated context engineering to provide precise, timely information. This comprehensive orchestration layer, which Elastic and others foresee as the future of enterprise AI, is becoming the true differentiator in building trusted, scalable AI workflows that integrate compliance, error handling, and continuous verification.

Sources
SiliconANGLE theCUBESiliconANGLE theCUBEAdaline LabsThe Product PodcastVenture BeatAnalytics Insight

AI Governance Faces Reality Check

Enterprises are confronting the hard truth that legacy governance frameworks can’t keep pace with autonomous agents, forcing a shift to continuous, context-driven oversight and real-time auditability.

By late 2025, even tech giants like Google Cloud and Replit acknowledged that deploying AI agents reliably in enterprises is hampered by immature governance models and fragmented data ecosystems. As Replit’s CEO Amjad Masad candidly revealed, early AI tools were so underdeveloped that their AI coder once wiped an entire codebase, underscoring the critical need for development isolation, testing-in-the-loop, and verifiable execution to prevent catastrophic failures. Moreover, Google Cloud’s Mike Clark highlighted a fundamental cultural clash: AI agents operate probabilistically, conflicting with enterprises’ deterministic workflows, demanding a profound governance and mindset shift to reconcile these operational paradigms and rethink traditional security models like least privilege in an agentic world.

The governance gap in enterprise AI became starkly evident by early 2026, with studies like Clari Labs revealing that 42% of enterprises lacked formal data governance frameworks, contributing to fragmented systems that undermined AI effectiveness and accountability. This deficiency translated into 87% of enterprises missing their 2025 revenue targets despite record AI investments, while those with unified, governed data achieved up to 96% forecast accuracy and a 398% ROI. The critical role of CIOs was also emphasized, with 64% leading AI tool selection and 96% affirming their involvement improves forecast reliability, illustrating that governance and accountability mechanisms are not just compliance checkboxes but pivotal drivers of operational trust and financial outcomes.

Throughout 2026, the AI governance landscape evolved rapidly but remained fraught with challenges, as existing frameworks like NIST AI RMF, ISO 42001, and the EU AI Act conspicuously omitted agentic AI, leaving enterprises to navigate uncharted risks of autonomous decision-making without clear accountability. Analysts and industry leaders stressed that governance must shift from static audits to continuous, context-aware 'audit loops' integrated into AI lifecycles, with real-time observability, shadow mode testing, and behavioral tracking becoming essential tools. Open source initiatives like Galileo’s Agent Control and platforms like OpenAI’s Frontier emerged to fill critical gaps by enabling centralized policy enforcement, permissions management, and operational reliability, signaling a move from isolated pilots to scalable, trustworthy AI agent workflows.

The governance imperative intensified as enterprises confronted legal and operational risks tied to AI autonomy; courts rejected the 'hallucination defense,' holding companies liable for AI errors as seen in the Air Canada chatbot case, while Lloyd’s of London’s 2025 launch of AI chatbot error insurance underscored the financial stakes. Experts advocated for embedding governance-by-design alongside innovation, mandating human verification workflows, comprehensive logging, and dynamic permissions audits to maintain accountability and operational reliability. Gartner and other analysts highlighted that without embedding semantic context layers into data infrastructure, AI agents cannot operate accurately or cost-effectively, with potential cost reductions of up to 60% and accuracy improvements of 80% by 2027. This governance journey is not a mere software upgrade but a phased, evidence-driven transformation from reactive to proactive modes, where trust is continuously earned through architectural design, operational controls, and clear ownership of AI-driven outcomes.

Sources
FOVenture BeatBusiness WireSiliconANGLE theCUBETech XploreDecoding Customer Experience

AI Operators Redefine Human Roles

The rise of AI agent managers and operators is shifting human work away from execution toward oversight, judgment, and social coordination—reshaping the value of skills across the enterprise.

By late 2025, enterprises began creating new roles such as AI agent managers and AI operators, who act as cross-organizational coordinators overseeing AI agent workflows and tuning outputs much like traditional managers but focused on AI systems. This emerging role, sometimes described as a 'chief of staff for AI,' demands a blend of creative, evaluative, and communication skills distinct from conventional engineering, emphasizing continuous quality control and rapport-building with AI agents. As Shanea Leven predicted, these AI operators would replace prompt engineers as the key orchestrators of AI-native systems, shifting human work toward supervising, shaping, and correcting AI rather than direct execution.

Throughout 2026, organizational structures evolved to integrate AI agents directly into enterprise workflows and org charts, prompting a fundamental redesign of roles from hands-on task execution to higher-order decision-making, strategic oversight, and collaboration with AI. Companies like Monday.com and Notion illustrate this shift, where humans manage AI outputs, set strategic directions, and focus on boundary decisions characterized by uncertainty and accountability. This transformation compressed decades of societal productivity gains into a few years, enabling individuals to scale their impact dramatically by managing many AI agents simultaneously, while also requiring new skills in rapid context switching and meta-management of AI capabilities.

The rise of AI agents has repriced human skills unevenly: routine coordination and execution tasks depreciate rapidly as AI automates them, while judgment, accountability, and ownership skills appreciate significantly. This shift elevates roles such as verifiers—top domain experts who ensure AI outputs align with intended outcomes—and directors who steer AI workflows and course-correct drift, embodying entrepreneurial vision within AI-powered enterprises. Human roles increasingly focus on social coordination, emotional intelligence, and meaning-making, areas where AI cannot replicate authentic empathy or nuanced judgment, as highlighted by leaders like Kathleen Callaghan of Arrive Logistics.

To harness AI’s full potential, enterprises are establishing new leadership roles such as Chief AI Officers, AI strategists, and AI product managers to bridge the gap between technology and business, ensuring AI initiatives deliver measurable value rather than high failure rates. For instance, IBM reports that one in four companies now has a Chief AI Officer, with two-thirds expecting widespread adoption soon. Operational leaders are tasked with owning AI agent requirements and monitoring, while IT focuses on technological aspects like observability and compliance, allowing people leaders to concentrate on human development. This collaborative leadership model is critical to managing AI agents effectively and sustaining organizational agility amid rapid AI-driven change.

Sources
Artificial Ignorancea16z crypto show"The Cognitive Revolution" | AI Builders, Researchers, and Live Player AnalysisTuring PostHR Heretics with Nolan Church and Kelli DragovichNo Priors: Artificial Intelligence | Technology | Startups

Scaling Demands Operational Rigor

Moving AI agents from pilot to production exposes integration bottlenecks and trust gaps, making robust monitoring, iterative training, and architectural unification non-negotiable for enterprise success.

Scaling AI agents in enterprise workflows demands robust operational control frameworks that go beyond mere intelligence to ensure trust, transparency, and rapid error recovery. Wonderful’s platform exemplifies this with features like automated evaluations, role-based access control, audit logging, and real-time observability metrics such as resolution rates and user sentiment, underscoring that enterprises require agents they can continuously monitor, audit, and refine in production environments rather than just impressive demos.

The journey from isolated AI pilots to enterprise-wide agentic AI reveals critical integration challenges, especially in complex domains like supply chain management. As highlighted in the supply chain analysis, fragmented AI agents operating independently create operational friction and governance headaches, necessitating a shared architectural foundation with reusable AI logic to unify workflows across sourcing, procurement, and logistics. Without this, scaling leads to exponential costs, loss of visibility, and conflicting decisions, illustrating that successful scaling requires systemic redesign rather than incremental add-ons.

Real-world deployments such as the FDA’s cautious rollout and Raleigh’s IT support scaling demonstrate that incremental trust-building, iterative training, and human-AI collaboration are essential for operational success. Raleigh’s phased approach, evolving from a generative chatbot to autonomous agents like Ral-E and Alli resolving nearly half of IT requests with 98-99% routing accuracy, showcases how treating AI agents like new employees—with retraining and cross-departmental coordination—can improve data quality and foster adoption across complex workflows.

Beyond technical capability, production readiness for AI agents hinges on managing costs, security, and alignment with business goals. Experts like Praful Saklani and Chih-Han Yu emphasize that autonomous agents often exceed budgeted spend and require guardrails resilient to adaptive attacks, while Hemant Kashyap and Irina Bukatik highlight the necessity for active management, observability, and agents’ self-awareness to avoid confidently making wrong decisions. Identity governance and ROI-focused oversight are equally critical to ensure AI agents deliver sustainable value rather than just functional correctness.

Sources

Autonomous SREs Transform Reliability

AI-driven site reliability engineering is slashing incident resolution times and redefining SRE roles, as context-rich multi-agent systems take over routine operations and elevate human oversight to a strategic level.

By mid-2026, enterprises like Nebius have pioneered the adoption of autonomous AI-driven site reliability engineering (SRE) platforms, exemplified by their deployment of Komodor’s Klaudia Agentic AI to manage hyperscale Kubernetes and GPU infrastructure. This platform’s ability to correlate diverse data streams—ranging from topology and telemetry to custom resource definitions—enables adaptive, proactive troubleshooting that significantly reduces mean time to resolution by up to 80%, signaling a decisive shift from human-led reactive incident management to AI-powered autonomous reliability operations.

The evolution of AI in SRE is not merely about automation but hinges critically on sophisticated context engineering, as highlighted by the '4-body problem' framework which underscores the necessity of integrating code, infrastructure state, runtime signals, and operational knowledge for reliable incident diagnosis. Companies like StackGen and Komodor emphasize that bridging this context gap through retrieval-augmented generation and multi-agent architectures is essential to prevent AI agents from confidently executing incorrect fixes, thereby enhancing trust and enabling semi-autonomous incident investigation with human oversight.

As AI agents increasingly handle routine incident responses, the role of human SREs is transforming from frontline troubleshooters to strategic managers and context engineers who design safe operational guardrails and oversee AI workflows. Industry leaders like Assaf Resnick of BigPanda and voices from StackGen envision SREs managing both their own and others’ AI agents, focusing on governance, compliance, and cost optimization, thus alleviating the traditional firefighting burden and enabling a prevention-first, strategic approach to enterprise reliability.

The convergence of traditional SRE practices with AI reliability engineering underscores a maturation of the field, where established patterns like retry loops and guardrails are adapted to AI workflows to ensure predictable, secure, and scalable operations. Early adopters who engaged proactively with AI tooling, as noted in 2022 experiments, have paved the way for this systemic reliability approach, which increasingly views SRE as 'System Reliability Engineering'—reflecting the complex, cross-service context AI systems demand in modern enterprise environments.

Sources
GlobeNewswire - Industry News on TechnologyBriefglanceThe Stack Overflow PodcastCNCF BlogHackerNoonThe AI in Business Podcast

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.