Jean‑Stanislas “JS” Denain, senior researcher and head of the Insights Team at Epoch AI, joined Nathan Lambert for a public podcast episode to dig into signals about rapid self‑improvement (RSI), robotics, and differences between US and Chinese frontier AI labs. The episode ran through a sequence of topics and chapter timestamps including RSI predictions, robotics’ role in acceleration, how far behind Chinese models are, whether distillation explains gaps, what Chinese job postings reveal, open vs. closed model safety, and how Epoch AI operates.
Main takeaway: substantial uncertainty
Both Denain and Lambert emphasized that there is a lot of epistemic uncertainty about near‑term trajectories. There are internal signals and published blog posts from organizations such as OpenAI and Anthropic that point to increased use of AI tools during model development and deployment, but neither guest treated those signals as conclusive proof of an imminent software‑only takeoff. Denain pointed to examples like an OpenAI chart showing a 2x per month rise in researchers’ Codex spending as suggestive of practical value, but not definitive evidence of an immediate, self‑sustaining acceleration.
What might internal teams see that outsiders do not?
The conversation noted that internal KPIs (for instance pre‑training compute multipliers, internal usage metrics, or team‑level indicators) can function as early warning signals. Denain said he does not know of a specific internal metric that would justify a much stronger alarm, but acknowledged that companies plausibly track indicators that outsiders cannot see. Those metrics combined with organizational intuitions can produce different forecasts on system performance trends.
RSI and existential risk
Denain stated he takes long‑term risks seriously and is uncertain but worried. He referenced other commentators’ estimates (for example Evan Hubinger’s cited figure of at least a ~10% x‑risk within some decade‑scale timeframe) as plausible. Denain framed many disagreements about risk as disagreements about how large and how fast AI capabilities will grow: if capabilities scale very far and fast, the economic and safety implications are much larger.
What does being a capabilities maximalist look like?
Denain described a capabilities‑heavy scenario as one in which AI automates much of AI research and robotics achieves large‑scale practical competence. That combination could create rapid industrial growth where AI systems manage factories and accelerate scientific progress well beyond historical rates. The key uncertainties are how intelligent systems become both in abstract (book) knowledge and in practical affordances for planning and project management, and what real‑world bottlenecks remain to fast R&D and hard power accrual.
Robotics: timelines and bottlenecks
Lambert and Denain were cautious about fast, broad robotic roll‑outs. Lambert suggested robotization might be a multi‑year to decade process for wide societal impact and compared expected robotics trends more to self‑driving cars (slow, bumpy progress) than to large language models (faster capability changes). His concerns focused on the difficulty of building reliable physical systems and on the breadth of tasks in real facilities versus highly repetitive jobs.
Denain pushed back slightly on timing pessimism by pointing out strong financial incentives: if robots can economically substitute for many blue‑collar tasks, the build‑out of factories and the supply elasticity could be large — data‑center construction is an example of very rapid industrial expansion when demand is high. Both agreed robotics is a critical variable to track but also a domain where they have less firm evidence.
How far behind are Chinese frontier labs?
Denain offered a rough public‑release gap estimate of about four to eight months between top Chinese labs and OpenAI/Anthropic on public ECI‑style comparisons, with some uncertainty depending on measurement methodology (e.g., whether you measure model release dates or training completion). Lambert noted that certain Chinese releases (Kimi K3, GLM 5.2) can appear closer (2–4 months) on some indices. They debated causes: compute and capital differences, distillation and router‑based data advantages, faster public releases in China, labor organization, and access to tested data or RL environments.
Denain reported internal checks suggesting contamination/overfitting effects in benchmark‑based lag analyses were not statistically significant in his tests, but acknowledged that ECI might still underestimate some capability differences where frontier labs optimized for broad serving and non‑benchmarked capabilities.
Is distillation a major factor?
Both speakers considered distillation (leveraging public API usage, router traffic, or generated data to train smaller or more targeted models) as a potentially important explanation for how some labs close gaps. Denain said he has updated toward distillation being a substantial factor — partly because more papers and accounts emphasize it, and partly because router‑driven prompt distributions (e.g., Claude routers) provide realistic usage data that can materially improve capabilities. Lambert remained skeptical that distillation alone explains all observed differences, especially when non‑US models also score strongly.
They also discussed training regimens: whether mid‑training supervised fine‑tuning (SFT) plus repeated cycles of RL and distillation, rather than one massive RL run, could be the effective industrial pattern. That looped approach could magnify the value of high‑quality usage data.
Organizational signals and Chinese job postings
Epoch AI has analyzed Chinese job postings to infer strategies and hiring trends. Those signals are best understood as anecdotal or suggestive rather than strict trendlines, but they can reveal things like increased hiring for B2B sales at some firms (e.g., Zhipu), local data‑center construction, and an uptick in internships and junior hires that may link to data collection. Denain emphasized that data markets are opaque and often mediated through private channels, making deeper reporting necessary to clarify compute/data flow and purchasing behaviors.
Open models, safeguards and proliferation
They discussed recent incidents where public model usage (or API interactions) allowed extraction or jailbreaks of sensitive behavior, noting cases such as the so‑called Stolen Thoughts line of work and a reported Claude‑based breach. Both argued that inference‑time safeguards and synchronous monitoring are imperfect; attackers can often choose the least protected provider. Denain thought that, at present, open models matter less because many poorly secured APIs exist anyway; Lambert observed that the evidence of misuse so far is concerning but not an avalanche.
They also distinguished between two proliferation routes: (1) actors hosting open models or self‑hosting weights without provider oversight; and (2) actors who fine‑tune closed models to remove safeguards. Currently, Denain judged route (1) to be more pressing because many inference providers aren’t exerting the same level of guardrails as frontier labs.
How Epoch selects research and what’s trackable
Denain described Epoch’s project selection as curiosity‑driven, aiming to fill a gap between industry and academia by producing well‑maintained public tracking on key trends (compute, data centers, job posting patterns, model compute estimates). He noted some verticals (compute, data centers) are relatively easy to observe; others (data procurement, internal post‑training recipes) are much more secretive.
Post‑training recipes at frontier labs
Both speakers agreed that actual post‑training practices are more complex than simple three‑stage approximations. Denain sketched a plausible structure: domain‑specific post‑training teams curate data, design RL environments and evals, and then bid to contribute their curated data and hyperparameter preferences into a large final RL run. There is heterogeneity across organizations — some use multi‑teacher expert‑style distillation, others more centralized large RL runs — and these design choices affect how quickly capabilities improve.
Closing remarks
The interviewers left the discussion with a cautiously vigilant stance: there are multiple meaningful signals (increased researcher tool use, distillation‑style workflows, job‑market and infrastructure indicators), but none of them alone proves an imminent, explosive RSI outcome. Many pressing unknowns remain — especially around robotics capability growth, the scale and structure of data markets, the efficacy of distillation, and model safety in hosted environments — and these deserve continued, careful measurement and threat modeling.
Source: Nathan Lambert interviewing Jean‑Stanislas “JS” Denain (Epoch AI) on the Interconnects podcast; topics and chapter timestamps referenced from the episode transcript.



