AI steps up as science’s co-pilot: from solving erdős mysteries to designing experiments

Air Street Press

The gist

AI has graduated from assistant to scientific co-pilot, autonomously cracking decades-old mathematical mysteries and turbocharging discovery across disciplines.

What to know

  • In early 2026, AI systems like GPT-5.2 Pro and Google Althia independently solved long-standing Erdős problems, with formal verification by experts such as Terence Tao.
  • Multi-agent AI platforms—including DeepMind’s Co-Scientist and FutureHouse’s Robin—now generate hypotheses, design experiments, and refine research, all under human oversight.
  • Breakthroughs in AI hardware and domain-specific frameworks (like Mercury 2, Nemotron 3 Super, and Berkeley Lab’s MatterChat) are enabling persistent, scalable agents to drive high-throughput, explainable scientific workflows.

AI’s New Scientific Role

AI models now generate, test, and refine scientific hypotheses across disciplines, acting as tireless collaborators that identify overlooked patterns and solve long-standing problems beyond human capacity.

By late 2025, AI had decisively moved beyond philosophical speculation to empirically generating new scientific knowledge, actively shaping hypotheses, experiments, and interpretations. Models like GPT-5 demonstrated reliable reasoning capabilities, producing intermediate scientific steps accepted by experts and enabling a novel division of labor where AI explores vast hypothesis spaces while humans frame problems and validate results. This foundational shift was exemplified by GPT-5’s contributions to four previously unsolved mathematical problems and its cross-disciplinary impact in fields ranging from plasma physics to materials science, where it identified symmetries, corrected subtle errors, and optimized computational methods.

A landmark milestone arrived in early 2026 when GPT-5.2 Pro, integrated with Harmonic's Aristotle proof verification system, autonomously solved three decades-old Erdős problems (#728, #729, and #397), with proofs formalized in Lean and verified by Terence Tao. While these problems were considered low-hanging fruit, Tao emphasized that AI’s ability to systematically tackle such neglected conjectures represents a foundational proof of concept, challenging prior skepticism about AI’s capacity to generate novel scientific knowledge. This achievement underscored AI’s emerging role in complementing human mathematicians by exhaustively addressing the ‘long tail’ of easier, obscure problems that often go overlooked.

Beyond mathematics, early AI systems had already begun to participate actively in the scientific method by proposing underlying rules and hypotheses rather than merely analyzing data. Notably, neural networks in 2020 rediscovered fundamental physical quantities like conserved momentum and energy from raw observational data, while DeepMind’s 2023 GNoME system vastly expanded the search for stable crystal structures, predicting over two million possible new materials—an order of magnitude beyond existing databases. These advances, alongside symbolic regression systems recovering physical laws from experimental data, illustrate AI’s unique ability to explore vast hypothesis spaces without human intuition, uncovering patterns and laws that might otherwise remain hidden.

System-level AI architectures such as FutureHouse’s Kosmos exemplify the automation of large portions of research workflows by ingesting thousands of scientific papers and generating tens of thousands of lines of code to produce structured scientific artifacts that link claims directly to supporting evidence. This integration points toward a future where large-scale AI models trained across diverse domains evolve into general-purpose discovery tools, expanding the scope of scientific inquiry rather than replacing scientists. Early successes across disciplines, including theoretical physics where AI uncovered previously assumed impossible particle interactions, reinforce the foundational potential of AI-driven discovery to accelerate incremental and groundbreaking advances alike.

Sources
Air Street PressLiberty’s HighlightsNoahpinionExploring ChatGPT

Agentic AI Breaks Barriers

Autonomous AI agents like Google’s Althia are independently publishing research and verifying their own work, but their rise is tempered by infrastructure gaps and persistent developer skepticism.

By early 2026, Google's Althia AI agent had set a new benchmark in autonomous scientific research by independently selecting and solving open mathematical problems, culminating in the submission of publishable papers without human intervention. Achieving a level 2 classification in Google's novel AI research contribution ranking, Althia not only solved four out of 700 unsolved Erdos conjecture problems but also inspired subsequent human-led research expansions, demonstrating a self-verification loop that mimics human iterative refinement to ensure correctness and reliability.

Despite the excitement around AI agents in 2025, the initial hype gave way to a more pragmatic recognition of significant infrastructural and trust-related challenges that slowed their seamless integration into scientific workflows. Key obstacles included the absence of AI-ready data centers, multi-node architectures for collaborative agent operation, and the difficulty of accessing and transforming legacy enterprise data, compounded by a pervasive developer trust deficit—nearly half of developers remained skeptical about AI-assisted coding despite widespread interest.

Autonomous AI agents distinguish themselves from traditional large language models by operating within an agentic loop that enables reasoning, tool use, error recovery, and multi-step problem solving over extended time horizons, necessitating sophisticated evaluation frameworks especially as these agents enter high-stakes domains like medicine and coding. This agentic loop represents a pivotal innovation, allowing AI to autonomously assess intermediate results and dynamically correct errors, thereby enhancing their effectiveness in complex scientific research workflows.

The emergence of multi-agent autonomous AI systems such as Google DeepMind's Co-Scientist and FutureHouse's Robin in 2026 marked a transformative leap in AI-assisted scientific discovery by orchestrating specialized AI agents to collaboratively generate hypotheses, design experiments, interpret data, and iteratively refine research directions. Validated in biomedical contexts—Co-Scientist proposing novel drug candidates for acute myeloid leukemia and Robin identifying potential treatments for dry age-related macular degeneration—these systems underscore a new paradigm of AI-human collaboration, ensuring scientists remain integral to validating AI-generated insights, as highlighted in their landmark Nature publication.

Sources
TheAIGRIDThe Stack Overflow PodcastDeep (Learning) FocusTech Xplore

Hardware Fuels AI Discovery

Breakthroughs in memory architecture, model throughput, and agent orchestration are enabling AI systems to process massive scientific datasets and sustain complex, persistent workflows at unprecedented speed and scale.

By early 2026, scaling AI in scientific research has increasingly hinged on overcoming critical hardware bottlenecks, particularly memory capacity rather than just compute or bandwidth. Research highlights from February emphasize the necessity of disaggregated, heterogeneous compute architectures with specialized accelerators and large-capacity memory disaggregation to support AI agent inference workloads with context lengths surpassing one million tokens, especially in coding agents. This hardware evolution underpins tools like OpenScholar, which integrates retrieval-augmented language models with iterative self-feedback loops to synthesize literature from 45 million open-access papers, outperforming GPT-4o by 6% in accuracy and achieving citation reliability comparable to human experts, thereby exemplifying scalable AI tool development for complex scientific workflows.

Advances in AI model architectures and inference systems have dramatically boosted throughput and efficiency, enabling more robust and scalable scientific workflows. Inception's Mercury 2 model leverages diffusion-based generation to achieve over 1,000 tokens per second on NVIDIA Blackwell GPUs—more than five times faster than typical speed-optimized models—while supporting a 128K context window and native tool use. Complementing this, the DualPath inference system tackles storage I/O bottlenecks by introducing a dedicated data path for key-value cache loading, nearly doubling offline and online throughput. NVIDIA’s Nemotron 3 Super further pushes the envelope with a 120B parameter open model optimized for agentic workloads, featuring hybrid architectures and native multi-token prediction that enable up to 2.2x faster inference compared to peers, all critical for sustaining the demands of persistent, long-context AI agents in scientific research.

The maturation of AI agent frameworks and orchestration tools is reshaping scientific workflows into persistent, observable, and controllable agentic systems akin to a 'bigger IDE.' Platforms like Perplexity’s Personal Computer, Replit Agent 4, and Base44 Superagents exemplify this shift by integrating multiple specialized models and tools to support collaborative, multi-agent workflows. Meanwhile, frameworks such as MIT CSAIL and Asari AI’s EnCompass modularize search strategies, enabling AI agents to backtrack or clone execution paths and swap search algorithms like beam or Monte Carlo tree search without extensive code rewrites—cutting implementation effort by up to 80% and boosting accuracy by up to 40%. This evolution addresses prior challenges from 2025, where infrastructure gaps and management complexity limited agent utility, by providing flexible, scalable, and robust agent orchestration critical for complex scientific tasks.

Specialized AI models and integrated toolchains are enhancing robustness and domain-specific capabilities in scientific research. Berkeley Lab’s MatterChat, introduced in May 2026, bridges large language models with physics-based interatomic potential models, effectively giving AI 'scientific eyes' to interpret complex 3D atomic structures and predict material properties with greater accuracy than general-purpose models like GPT-4. This approach embeds scientific inductive biases into AI representations, enabling scalable high-throughput discovery by leveraging large curated datasets of atomic structures and properties. Similarly, Argonne National Laboratory’s roadmap envisions coordinating multiple specialized LLM agents within AI-powered self-driving laboratories to automate iterative experimentation in battery research, addressing challenges of computational efficiency, task adaptability, and explainability to ensure reliable, scalable workflows that accelerate scientific innovation.

Sources
AI NewsletterInto AILatent.SpaceMIT NewsThe Stack Overflow PodcastHPCwire AIwire

Human-AI Synergy Redefines Research

The hybrid discovery loop—where AI proposes and humans validate—has become central to scientific progress, fundamentally altering institutional workflows and accelerating the pace of credible discovery.

By late 2025 and into 2026, hybrid human-AI discovery loops have crystallized as a transformative model in scientific research, where AI systems propose vast arrays of hypotheses or proof steps and humans retain critical roles in framing problems, imposing constraints, and validating outcomes. This iterative collaboration, exemplified in mathematics by researchers using AI to write experimental code and verify computational cases within a 'trust boundary,' ensures that AI acts as a creative proposer while humans provide judgment, strategic review, and accountability, preventing AI from becoming an unchecked oracle.

Institutional and infrastructural transformations are underway to fully harness AI’s capabilities, with platforms like FutureHouse Kosmos demonstrating how AI-driven workflows can ingest thousands of scientific papers, generate tens of thousands of lines of code, and produce structured scientific artifacts linking claims directly to evidence. Similarly, multi-agent AI systems such as Google DeepMind's Co-Scientist and FutureHouse's Robin are reshaping research workflows across disciplines by autonomously generating hypotheses, designing experiments, interpreting results, and refining hypotheses—all while keeping human researchers integrally involved, thereby enabling a scalable and efficient scientific method.

This hybrid human-AI discovery loop model fundamentally redefines the scientific method by enabling the simultaneous generation of hundreds of hypotheses, automated experimental design, and rapid data interpretation, which together accelerate discovery while maintaining human oversight. As Noubar Afeyan highlights, this AI-driven transformation is unprecedented and still nascent, signaling profound shifts in how research institutions must restructure to leverage AI’s potential. The loop’s power lies not in AI or humans alone but in their interaction—where AI widens exploration and tightens verification, and humans curate and decide significance, as vividly illustrated by Claude’s Cycles and Tao’s mathematical research.

Beyond accelerating discovery, AI expands the breadth of literature scanning and cross-disciplinary insight, surfacing analogies and relevant findings buried in distant or older subfields, thus acting as a multiplier for human expertise. Advanced workflows now combine pattern-based AI with formal logic to produce mathematically guaranteed proofs, enabling AI to undertake extended reasoning and formal verification tasks, while humans focus on strategic problem selection and priority setting. This evolution mirrors software engineering shifts and underscores a broader institutional transformation where expert attention becomes more selective and strategic, supported by AI’s execution capabilities.

Sources
Air Street PressGradient FlowTech XploreNew York Stock ExchangeThe Computist Journal

Tailored AI Powers Domain Breakthroughs

Discipline-specific AI frameworks like MULTI-evolve and MatterChat are driving exponential gains in fields from protein engineering to materials science, with national initiatives aiming to double scientific productivity.

By early 2026, domain-specific AI frameworks were revolutionizing scientific discovery across diverse fields by tailoring models to the unique challenges of each discipline. The Arc Institute’s MULTI-evolve system exemplifies this in protein engineering, achieving up to a 256× improvement in enzyme function by rapidly predicting synergistic multi-mutation combinations, while AI agents like Gauss advanced mathematics by autonomously formalizing complex proofs and correcting errors in landmark theorems. This targeted approach leverages specialized neural networks and protein language models to guide experimental design and accelerate innovation, illustrating how AI’s integration into domain workflows enhances both speed and precision.

In materials science, Berkeley Lab’s 2026 introduction of MatterChat marked a strategic leap by bridging Large Language Models with physics-based atomic interaction models, enabling AI to 'see' and interpret complex 3D atomic structures with unprecedented accuracy. Trained on nearly 143,000 stable atomic structures, MatterChat imbues LLMs with a scientific inductive bias that allows them to provide grounded insights into critical materials challenges such as thermal stability and electronic band gaps. This fusion of conversational AI with high-throughput materials data embodies a roadmap that accelerates discovery by integrating disparate data modalities, setting a new standard for domain-specific AI applications.

National-scale initiatives like the DOE’s Genesis Mission, showcased by Lawrence Livermore National Laboratory in mid-2026, illustrate how AI frameworks are being deployed to transform scientific workflows across molecular discovery, fusion energy, and advanced manufacturing. Tools such as FLASK Copilot and Multi-Agent Design Assistant (MADA) exemplify practical AI-driven innovation aimed at doubling U.S. scientific productivity within a decade by integrating laboratory expertise, experimental facilities, and computational resources. This ambitious vision underscores a coordinated, domain-specific AI strategy that leverages multi-agent systems to accelerate discovery and enhance national technological leadership.

Argonne National Laboratory’s 2026 technical roadmap for AI-enhanced battery research epitomizes the future of domain-specific AI integration by coordinating multiple specialized LLM agents to automate and accelerate experimental cycles within self-driving laboratories. This initiative, part of the broader DOE Genesis Mission, envisions AI systems mining vast literature, analyzing performance datasets, and optimizing battery design and degradation understanding, thereby transforming traditional trial-and-error approaches. As Khalil Amine and Guiliang Xu emphasize, close collaboration between battery researchers and AI experts is essential to align research questions with appropriate models, ensuring that AI-driven workflows deliver faster, more reproducible breakthroughs in energy storage technology.

Sources
Techno-OptimistHPCwire AIwireQCwireTech XploreQCwire

AI: From Tool to Co-Discoverer

AI now actively shapes the scientific process by generating new ideas and solving neglected problems, shifting the paradigm from automated assistant to essential partner in knowledge creation.

By early 2026, AI had transcended its traditional role as a mere automation tool to become a creative scientific collaborator, exemplified by GPT-5.2 Pro's autonomous resolution of decades-old Erdős problems with formal verification by Terence Tao and Harmonic's Aristotle system. While these solutions were not groundbreaking breakthroughs, Tao emphasized their significance as a milestone, stating, 'This is the worst that it’ll ever be, it’ll only get better from here,' highlighting growing trust and quality control in AI-generated knowledge.

The philosophical landscape of scientific discovery is shifting as AI systems evolve from data analyzers to active participants in hypothesis generation and rule discovery. DeepMind’s 2023 GNoME model dramatically expanded the search space for stable crystal structures, predicting over two million possible materials—vastly outpacing prior databases—and early demonstrations showed AI recovering physical laws from raw data without prior equations, underscoring AI’s ability to explore hypothesis spaces beyond human intuition. This inversion of roles, where 'the machine generates possibilities first and humans investigate them afterward,' marks a fundamental transformation in the scientific method.

Rather than replacing scientists, large-scale AI models trained across broad domains serve as general discovery tools that expand the range of testable ideas, addressing the 'long tail' of less glamorous problems often neglected by human researchers due to limited time or motivation. Terence Tao noted AI’s strength in systematically 'knocking off the easiest of the problems,' complementing human focus on complex challenges. This scalable, tireless approach accelerates incremental advances that cumulatively may lead to significant breakthroughs, effectively supercharging scientific progress.

The integration of AI as a scientific collaborator introduces critical challenges around verification, quality control, and trust, as autonomous AI-generated results increasingly contribute to the scientific corpus. The emerging human-AI discovery loop—where humans pose questions, AI proposes candidates, verifiers filter results, and humans curate outcomes—redefines discovery as a dynamic interaction rather than a solo act. Illustrative cases like Claude’s Cycles, where AI-surfaced structural patterns were validated and formalized by experts such as Donald Knuth, exemplify this evolving partnership and underscore the necessity of nuanced narratives beyond simplistic replacement or dismissal of AI.

Sources
Liberty’s HighlightsNoahpinionExploring ChatGPTThe Computist Journal

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.