Glean, co-founded and led by former Google Distinguished Engineer Arvind Jain, has seen rapid growth: it reached $300 million in annual recurring revenue (ARR) this year, a threefold increase over 15 months. The company was last valued at $7.2 billion following a $150 million Series F round in June.
Why model routing matters for enterprises
With intense competition among frontier model providers and the growing capability of open-weight models such as Kimi K3 and Qwen3.8-Max, model routing has become central to AI deployments. Glean’s role is to decide which model to use for each task—or whether a large language model is needed at all.
Jain told Latent Space: “A big goal of Glean is to avoid using LLMs for tasks where we don’t need them.” He gave the example of simple arithmetic queries that are better served by a calculator than an LLM.
Delivering a ‘‘personal coworker’’
Glean aims to provide employees with what Jain describes as “one really powerful personal co-worker,” acting as a meta-harness for leading LLMs. The company introduced its third-generation Glean Assistant last September, and agents now play a major role in the system.
“You can think of Glean today as a superset of ChatGPT, Claude, Gemini, Grok,” Jain said. “All these different AI products that we’ve been using day to day, Glean combines the power of all of them into one experience.”
Bringing organizational knowledge into AI
Deploying AI technology is only part of the challenge for enterprises; the other half is integrating organizational knowledge. Jain said Glean’s business is to deeply understand a customer’s data, knowledge, and how work actually happens inside the company.
How model routing works in practice
Glean offers three levels of model selection:
- Employees can explicitly choose a model.
- Administrators can restrict models or set usage limits.
- Glean’s automatic mode dynamically selects a model per task.
Customers largely choose the automatic mode for economic reasons. Jain explained: “Why are people talking about model routing? Why are they excited about it? It’s mostly because of cost.”
Tony Gentilcore, Glean co-founder and engineering lead, recently claimed that Glean “is 4x more cost-effective” than Claude Code, averaging $0.45 per task versus $1.84 for Claude Cowork, attributing that to Glean’s harness and routing capabilities.
Jain noted that while individuals get value from $20–$200 monthly subscriptions, enterprise per-user costs can escalate quickly. The most advanced models are more capable but also costlier per token—sometimes double or quadruple previous rates—and users run longer tasks, leading to 10x–20x increases in per-user spending compared with the prior year.
The human feedback loop
A key advantage for Glean is visibility into how ordinary business users employ AI. The product can be deployed to every employee as a ‘‘coworker’’ and is used to build and run agents across departments.
Glean cites customers such as Zillow, which reports 80% adoption across 7,000 employees, and Booking.com, where Glean was “the first AI platform adopted company-wide.” Such penetration gives Glean broad insight into what people do with AI, which models they select first, and when they upgrade to other models.
That real-world usage feeds a human feedback loop: Glean runs the router’s chosen model for users but also, behind the scenes, attempts the same task with other models. AI-based judges then assess how accurate the router’s choice was. Although this auditing covers only a small fraction of traffic, Glean’s scale provides enough examples to continuously train and improve the router.
Waldo: gathering raw materials and filtering
Part of Glean’s architecture is Waldo, introduced in April as “Glean’s first agentic search model.” Waldo sits above large language models and functions as a filter: it decides how to break down a question, which tools to use, what to read next, and when there is enough evidence to hand off to a frontier model for a high-quality answer.
Glean claims Waldo reduces latency by 50% and tokens by 25%, reserving advanced models for work that needs them. Jain said Waldo helps assemble the ‘‘raw materials’’ needed for a task without burning LLM tokens, and noted that a cheaper model with better context can outperform a frontier model burdened with irrelevant data.
Rapid adoption of open-weight models
Jain confirmed that enterprise interest in open-weight models has risen sharply in recent months, driven primarily by cost concerns. He said usage of open-source LLMs was negligible last year and many companies were not seriously considering them, partly due to stigma around models developed outside the U.S.
“But in the last three months, because AI got so expensive, businesses have started to find it untenable to maintain these AI investments,” Jain said. Given that open source can be an order of magnitude cheaper for some tasks, many enterprises now view open-weight models as a key part of their AI strategy. Jain added that organizations are no longer willing to rely on just one or two providers and increasingly see open source as essential.
Evals and quality monitoring
Evals—assessing the quality of LLM outputs—are central to Glean’s approach. The company runs internal testing that compares real-world workloads across query classes and parallel attempts of the same task with less expensive and more expensive models. AI-based judges evaluate the router’s choices to maintain and improve quality.
From enterprise search to end-to-end AI platform
Founded in early 2019 as an enterprise search tool, Glean was one of the first enterprise-focused companies to work with transformers and language models. By 2026, AI has become integral to everyday workflows, expanding Glean’s role beyond search into a broadly used end-to-end AI platform.
Jain described Glean as an ‘‘end-to-end AI platform’’ that is used heavily by enterprise customers and — by virtue of its access to organizational data — is positioned to perform effective model routing.
Market implications
Model routing, agentic layers like Waldo, and growing enterprise adoption of open-weight models together are reshaping how large organizations deploy AI. Cost efficiency, continuous real-world feedback, and multi-model strategies indicate that AI model selection has become a strategic and economic decision, not just a technical one.



