Research

AI-generated text

Intology’s Locus Advances on PostTrainBench, Demonstrates Automated Post‑Training Gains

Intology released an updated version of Locus, its system for using large language models to perform AI research tasks, and reports a 44.7% score on PostTrainBench when paired with Opus 5.

Intology’s Locus Advances on PostTrainBench, Demonstrates Automated Post‑Training Gains

AI startup Intology, whose stated goal is to "automate R&D," has released an updated version of Locus, its software for using large language models as researchers. The company reports that the new Locus achieves a 44.7% score on PostTrainBench, a benchmark designed to measure how well systems can take an open-weight model and improve its performance beyond the baseline.

Specific results and comparisons

  • Locus paired with Opus 5 scores 44.7%; by comparison, Opus 5 without any special harness scores 34.1%.
  • Locus also outperforms Fable 5, which scores 41.8% on the same benchmark.
  • According to Intology, Locus "outperforms every frontier-agent baseline on PostTrainBench, and given greater compute, post-trains models that collectively surpass both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite."
  • Intology states that these results were externally verified by the PostTrainBench authors and underwent stringent contamination and cheating checks.

Background on PostTrainBench

PostTrainBench was first introduced in March 2026. At that time, the highest-scoring system was Opus 4.6 with 23.2%, up from Claude Sonnet 4.5’s 9.9% in September 2025.

PostTrainBench+: extended GPU time and surpassing the human baseline

Intology also developed a variant of PostTrainBench (referred to here as PostTrainBench+) that removes the benchmark’s single-GPU 10-hour wall-clock limit so systems can be evaluated with much larger compute budgets. Using this variant, Intology reports that Locus exceeded the human baseline, achieving 51.6% when consuming over 4,000 hours of H100 GPU time. For comparison on this variant, Opus 4.8 scored 44.3% and GLM 5.2 scored 42.7%; Fable was not tested on this variant.

Results in other domains

Intology also reports that Locus discovered and trained a language model end-to-end that is now running in production at Bubble, a no-code app-development startup. The company states the production model operates at approximately 2.8× lower error, 5.4× lower latency, and 105× lower cost.

Why this matters

These results highlight that current AI systems can be used to automate parts of AI research and development, particularly when provided with an effective harness and significant compute. Intology’s reported improvements over the same base model (Opus 5) illustrate that post-training processes and tooling can materially change final performance. The external verification and the PostTrainBench+ comparisons provide additional context for evaluating how much further open-weight models can be improved through post-training and larger compute investments.