Research

AI-generated text

Latent Space releases AEO tracker mapping frontier LLM recommendations

Latent Space published the Frontier AEO tracker, a dataset and analysis that collects product recommendations from seven large language models across 161 categories using six prompt variants.

Latent Space releases AEO tracker mapping frontier LLM recommendations

Latent Space has launched the “Latent Space Frontier AEO tracker,” a system and analysis that runs six prompt variations across seven large language models to collect and rank product recommendations in 161 categories.

How it was built

The team spent a few billion tokens worth of prototyping, alignment and scaling work to assemble the pipeline. Answer extraction was performed by Astra, and responses were scored using a proprietary AEO score. That score weights first-choice recommendations, alternative suggestions and mentions, while applying negative weight to mild and strong anti-recommendations.

They also extracted the top-cited sources that influence model recommendations and produced an analysis of common failures. In the interest of transparency, the project makes every prompt–answer pair inspectable.

Key findings

  • Out of 161 categories, 28 have a single, universally dominant primary choice across all surveyed frontier models. Many other categories are close contests or show consistent stylistic biases that make them potential AEO battlegrounds.
  • Some familiar vendor names appear frequently in the results, which raises the question of contamination; the team says it checked for this and publishes prompt–answer pairs for inspection.

Cross-model shifts and lab comparisons

The researchers prepared special reports on Opus→Fable and Sol→Astra transitions and observed consequential flips in model choices between generations from the same lab. For several flips they provided neutral analyses of what competing models did better in each scenario.

Source usage, efficiency and confidence

One notable finding is that models differ in how many sources they consult: Sol has a median of 9 sources, Astra a median of 5, Opus a median of 11, and Fable a median of 15. Astra appears generally more “confident” or “efficient,” being far less likely to change its answer when the question is lightly paraphrased. The authors argue that as choice randomness declines, the value of AEO rises.

Source analysis also somewhat predicts the priorities of different labs, but the sample is limited: the dataset reflects what could be scraped from attempted toolcalls, not the full pretrain datasets.

Failures and practical impacts

The tracker includes an analysis of top failures. The team notes that AEO practices (illustrated by examples such as Ora and Vercel markdown content-negotiation) matter in practice: failures discourage models from reading your content.

Other notes and extras

  • The tracker covers a broad range of categories, from coding agents, AI podcasts, AI sandboxes and managed databases to ASR models and less typical categories like angel investors, corporate spend and payroll software.
  • Latent Space also built a lighthearted “family feud” style game so users can compare their priors to the aggregated model recommendations.

Methodology and coverage caveats

As noted in the methodology post, the team tried to include Gemini/Antigravity, GLM/Zcode and DeepSeek/DeepCode, but errors and rate limits prevented their inclusion in this first run. The authors invite representatives of those companies to contact them to raise limits.

Contact and business inquiries

Latent Space is open to suggestions and business enquiries. Contact: @latentspacepod or business@latent.space.