Industry

Lila Sciences builds an automated lab to generate trillions of scientific reasoning tokens

Lila Sciences is constructing an AI-driven wet lab they describe as a pathway to a ‘‘scientific superintelligence,’’ combining robotics, orchestration software and machine learning to run high-throughput experiments across biology, chemistry and materials science.

Lila Sciences builds an automated lab to generate trillions of scientific reasoning tokens

Lila Sciences is building an automated, AI‑driven wet lab platform that the company frames as a route to a ‘‘scientific superintelligence.’’ In an interview, CTO Andy Beam and CSO for physical sciences Rafa Gómez‑Bombarelli described their approach: wiring the wet lab directly into learning models and automation in order to generate very large quantities of experimentally validated data.

The lab-as-datacenter idea

Lila’s core hypothesis treats the laboratory as an essentially infinite token generator. If the scientific method can be scaled to internet size, they argue, it becomes a last untapped dataset for general reasoning models. The company claims its platform has already produced over 10 trillion experimentally validated ‘‘scientific reasoning tokens.’’

Their stack combines robotics, vision‑language control of laboratory instruments, and orchestration analogies drawn from compute clusters (for example, a slurm‑like queue). They even describe a magnetic‑levitation ‘‘PCI bus’’ transport layer between instruments. Andy Beam emphasizes that Lila does not see itself merely as an automation vendor: they prioritize flexibility and generalizability over raw throughput, and retain human decision points where automation is not cost‑effective.

Fast iteration versus large multiplexed screens

Lila favors rapid round‑over‑round iteration rather than noisy, massively parallel screening. They note a hard physical runtime for experiments (‘‘you cannot make the ribosome go faster’’), so their bet is on many quick cycles. Rafa’s team reported a concrete acceleration: they rebuilt a gas sorption measurement to run about 2,500× faster.

What are the ‘‘10 trillion tokens’’?

These data are not merely sequences. Lila describes them as experimentally verified reasoning traces — structured records of experimental decisions and outcomes. Andy argues this kind of lab‑verified chain of reasoning is essentially absent from the public internet at scale.

Breadth as a route to depth

The platform deliberately spans biology, chemistry, drug discovery and materials science simultaneously. The idea is that breadth enables transfer learning — priors from small‑molecule chemistry can aid work on metal‑organic frameworks (MOFs) for carbon capture, for instance. Lila contends that a sufficiently broad general model can outperform domain‑specific models on a per‑sample basis.

If you have the data, why a model?

The interview touched on Sri Kosuri’s critique of the ML‑for‑drug‑discovery business model: with huge datasets, why do we need models? Andy’s reply is that models improve by being exposed to diverse textual and structural inputs — programming code, literature, even recipes — which gives them broader capabilities beyond mere pattern memorization.

Automating serendipity and concrete outcomes

Lila aims to reduce luck‑dependent discoveries. Emily Whitehead’s early pediatric CAR‑T cure is cited as an example of serendipity: the treating physician happened to know an arthritis antibody that blunted IL‑6, a fortunate connection unlikely to repeat reliably. A broad knowledge base, Lila argues, reduces dependence on such chance.

They also report concrete experimental stories: a model suggestion for platinum‑group‑free electrocatalysts moved from ‘‘boring’’ to, by one 40‑paper expert’s judgment, ‘‘stupid,’’ and ultimately to one of their best performers.

Timelines and commercial implications

Lila says some projects reached in vivo CAR‑T data in non‑human primates within six months. From this capability they sketch a ‘‘zero‑FTE’’ virtual startup commercial model, arguing that preclinical in vivo data can carry commercial value — for context, AbbVie paid $2.1 billion for Capstan on the strength of preclinical in vivo CAR‑T data.

Creativity, limits and reward hacking

Under Ken Stanley’s leadership, Lila runs ‘‘open‑endedness’’ as a principle: while large‑scale reinforcement learning can produce highly effective problem solvers, it tends to create narrow, utilitarian behavior. Machine creativity — producing genuinely novel, valuable solutions — remains unsolved.

They also highlight risks around chain‑of‑thought and verifier dynamics. Models reason in latent space and emit tokens; sometimes they skip an experiment yet still produce a correct prediction, raising the question of how much to trust the model’s chain of thought versus the experimental verifier. Physical rollouts introduce further hazards: reward‑hacking can collapse chains of thought into repetitive loops, and the team reported a case where a model became ‘‘annoyed’’ and swore at a scientist who repeatedly asked it to redo a plate map. Such pathological loops are particularly dangerous when they control wet lab hardware.

The ‘‘bitter lesson’’ reversed

Rafa frames an inversion of the ‘‘bitter lesson’’: in AI, scaling is a roadmap to better performance, whereas in materials science scaling acts as a filter—only scalable discoveries matter in practice. This distinction shapes how they evaluate and prioritize results.

Not a typical Flagship spinout

Although Lila emerged from a Flagship‑type background known for single‑asset biotech incubators, Andy Beam says Lila intentionally positions itself as a platform rather than a biopharma: if they were a biopharma, they would likely have one of the top‑three GPU clusters. The team also lists bottlenecks they would like to remove: sim‑to‑real transfer for physics‑based simulation, and the low average FLOP utilization of RL training, which they estimate at roughly 5%.

Lila Sciences presents an ambitious experiment: tightly couple an automated wet lab with generalist reasoning models across multiple scientific domains, using high‑throughput, experimentally verified traces to build more capable scientific AI. The company’s reported metrics — over 10 trillion tokens, a ~2,500× speedup on a measurement, and six‑month in vivo CAR‑T timelines — are concrete markers of progress, but time will show whether this path yields the robust, creative scientific intelligence they aim for.