Research

AI-generated text

From GPU kernels to recursive language models: Alex Zhang on harnesses, agent swarms and untapped capabilities

Alex Zhang (MIT) discusses his path from GPU kernels and KernelBench to Recursive Language Models (RLMs), Prime Agent and large multi‑agent experiments.

From GPU kernels to recursive language models: Alex Zhang on harnesses, agent swarms and untapped capabilities

Alex Zhang of MIT has moved from hands‑on GPU kernel work and KernelBench into designing Recursive Language Models (RLMs), harness architectures like Prime Agent, and experiments with large multi‑agent systems. In a recent long interview he explained why the software that wraps models — context offloading, programmatic subagents, persistent state and harness design — can unlock capabilities that raw model scaling alone does not. He also discussed AI‑generated GPU kernels, agent swarms, verification and efficiency issues, and where academic researchers should place contrarian bets.

GPU Mode, KernelBench and AI‑written kernels

Zhang was active in the GPU Mode community (originally a CUDA‑focused Discord) and participated in projects that tried to automate GPU kernel development using LLMs. KernelBench and related leaderboards showed promising AI‑generated solutions, but Zhang highlighted a verification gap:

  • Many top leaderboard entries are AI‑generated, yet when integrated into real end‑to‑end systems only a small subset (often those guided or verified by expert human authors) prove stable.
  • There remain nontrivial optimization patterns and domain knowledge in GPU kernels where human expertise still yields alpha over brute‑force token search.
  • Although theoretical lower bounds (a “speed‑of‑light”) for kernels can be estimated for simple operations, practical kernels are often far from these limits. Memory transfers, kernel fusion (megakernels) and whole‑model trade‑offs complicate isolated kernel optimization.

Why harness design matters

A central claim of Zhang’s view is that harnesses — the programmatic scaffolding around a model — are not mere engineering detail but a crucial inductive bias. Most mainstream harnesses (Codex‑style loops, tool‑call patterns) are structurally similar; different harness abstractions can produce very different scaling and generalization behavior.

Key harness elements that Zhang emphasizes:

  • Context offloading: storing and compacting long context externally (e.g., on disk) rather than appending long trajectories to the prompt.
  • Programmatic subagents: subroutines implemented as code that the model can call (and that can call the model), enabling recursion and modularization.
  • Persistent subagents and shared memory: subagents that persist beyond a single runtime and a shared workspace (file system, message board) that agents can read and write.

These design choices let each model call be locally in‑distribution even when the global task is out‑of‑distribution, improving compositional generalization.

What is an RLM (short definition)

Zhang’s practical definition: an RLM is a harness where the primary tool is code — the system exposes code execution as the mechanism for subagent calling, allows self‑calls (recursion), and stores context in persistent memory. In other words, the harness is a programmatic environment (Python REPL, files, modules) and the model uses that environment as its only tool.

RLMs aim to solve long‑context and compositional tasks by breaking problems down into many locally easy subcalls; empirical results show training on shorter instances often generalizes to substantially longer problems.

Prime Agent and continual harness concepts

Prime Agent is an example implementation of RLM principles: it restricts the core tool to an IPython kernel and loads other capabilities as Python modules. It also implements continual harness ideas where the harness itself can be modified — adding/removing subagents, updating system prompts, or altering skills — and supports persistent subagents to preserve state across sessions.

Zhang noted Prime/Select teams are experimenting with training models specifically for RLM workflows, though he himself (as an MIT academic) is not involved in the model training side.

Agent swarms, compute scale and wasted search

Large agent swarms can solve very hard problems but often at great compute and token cost. Zhang discussed the tradeoffs:

  • Agent swarms may discover useful information by throwing enormous search at a problem, but a high fraction of the tokens can be wasted exploration.
  • Public reporting referenced an OpenAI experiment that used roughly 10,000 agents over a multi‑hour window and produced ~130 billion output tokens; at public pricing that token output alone corresponds to an order‑of‑magnitude tens of millions of dollars — an indication of scale and cost.
  • Efficient swarm design, better coordination and harness choices remain open research problems: how to get swarms to converge, how to structure subagent roles, and when to prefer a harnessed RLM solution over a sprawling swarm.

Research taste, PhD bets and low‑hanging conceptual ideas

Zhang urged PhD students to make contrarian, high‑risk research bets that industry labs are unlikely to pursue. He argued that many influential ideas (ReAct, Quiet‑STaR, SWE‑bench, RLMs) initially looked simple or uninteresting; their impact came from a good framing and follow‑up experiments. Academia’s relative freedom to pursue odd ideas is an advantage.

New model architectures and output spaces (Jev, GEV, loop ideas)

Work like Jev shows that a language model need not be a standard autoregressive text decoder; altering the output space can produce dramatic latency and cost tradeoffs (e.g., very fast binary classification). Zhang sees such architectures as complementary to harness research because they expand the designer’s choice of inference tradeoffs for multi‑call systems like RLMs.

Third‑party RLM work and applied examples

Zhang highlighted external RLM adopters: legal AI company Harvey (post‑training RLMs on legal workflows), Headlong/Terminus variants, and others experimenting with persistent, always‑on harnesses. He also noted ARC‑AGI‑3 Kaggle entries and other competition submissions that used RLM‑inspired harnesses for composition and neuro‑symbolic pipelines.

Open‑endedness, Sakana AI and auto‑research

Open‑endedness (letting systems run without a fixed human prompt) and automated research pipelines are active research themes. Zhang described Sakana’s and related labs’ attempts to run persistent systems that generate and sift large amounts of candidate work; a key bottleneck remains how to identify the hidden gems in vast generated output.

Speculative programmatic tool calling and Neuralese

Zhang endorsed the simple idea of speculative/parallel tool execution: if a model writes code that will trigger external tools, launching those tools speculatively in parallel (when safe) reduces latency. He also discussed the possibility that the most effective internal representations for agent coordination will not be English or existing programming languages, but some hybrid "Neuralese" that biases reasoning in useful ways.

Capability overhang and immediate opportunities

A recurring theme: even current frontier models may carry a capability overhang that harness engineering could surface. Zhang argued that researchers should search for harness designs that enable reliable, long‑running, lower‑latency work (continuous assistance, long‑horizon tasks) without requiring frontier scale compute.

Practical pointers and resources

  • Alex Zhang website: alexzhang13.github.io
  • X: @a1zhang

Podcast timestamps (major sections)

00:00:00 – Introduction 00:00:49 – GPU Mode, KernelBench, AI‑written kernels 00:07:38 – Human expertise vs. brute‑force AI search 00:13:20 – Research taste and taking big bets 00:19:28 – GEV and rethinking the language model 00:29:03 – Video game agents and the harness problem 00:31:01 – Why Claude Code, Codex, and Pi are similar 00:36:42 – Harnesses as compositional generalizers 00:44:24 – RLMs explained 00:52:01 – Prime Agent and persistent subagents 00:57:41 – RLMs in the wild 01:00:30 – OpenAI swarms and the future of language models 01:07:26 – Open‑endedness and Sakana AI 01:15:52 – Kimi vs. OpenAI agent swarms 01:20:06 – Capability overhang and speculative tool calling 01:28:19 – Neuralese, future research, and AI for science


In short: Zhang’s research arc links low‑level systems work (GPU kernels, KernelBench) with high‑level system design (harnesses, RLMs, agent orchestration). His main message is practical: better harness engineering and targeted training for harness workflows can unlock substantial model capabilities without always needing another generation of bigger base models.