Industry

How Expedia Group Turns AI Principles into Scalable, Trustworthy Systems

Expedia Group has codified machine learning and AI principles to ensure models deliver business value, scale across teams, and operate safely.

How Expedia Group Turns AI Principles into Scalable, Trustworthy Systems

Expedia Group draws a clear line between AI that simply works today and AI that remains reliable at scale. Many organizations optimize for quick results without asking whether those solutions will hold up across teams, use cases, and time.

The company’s experience shows the hardest part is not making a model work once, but creating systems that continue to work, scale beyond the teams that built them, and improve consistently. Modern AI systems do more than predict or optimize: they converse, reason, and increasingly act on users’ behalf. An autonomous system that makes decisions for a traveler creates very different expectations for reliability, governance, and accountability.

Expedia Group has applied machine learning and AI across the traveler journey for years — from personalization, ranking, and recommendations to fraud prevention, customer support, and more recently generative and agentic AI experiences. That depth of experience led the company to codify a set of ML and AI principles to guide how it builds, deploys, and evolves AI systems.

Purpose and principles

The objective is straightforward: ensure systems deliver real business value, scale, and operate safely. The principles define how the company measures, designs, governs, and operates its AI systems.

From principles to practice

Publishing principles is the easy part; the harder work is turning them into operational mechanisms: recommendations, requirements, tooling, and release processes teams actually use. Expedia Group has introduced “Agentic Release” tollgates — a series of recommended and in some cases required checks before launching agentic AI features. These tollgates translate principles such as clear ownership, risk-based governance, evaluation, safe rollout, and monitoring into concrete expectations for teams.

Some recommendations and requirements are already being automated and integrated into the software development lifecycle (SDLC). Over time, the goal is for these expectations to be embedded in how systems are designed, evaluated, approved, launched, and monitored from the start.

Outcomes: measure what matters

The first test for any model is whether it improves a business outcome and ultimately the traveler experience, not merely a technical metric.

  • Align models to business-impact metrics: every ML effort must tie directly to a key business outcome or traveler experience metric. Technical optimizations are useful intermediates, not end goals.
  • Optimize for return on cost: the value a model creates must justify the costs to develop, train, and monitor it, plus any added operational complexity. Favor solutions that deliver lasting impact relative to their operating cost.
  • Justify complexity with strong baselines: complexity should be earned, not assumed. Start with strong baselines — an existing general model, a simple heuristic, or an off-the-shelf solution — and pursue specialized models only when simpler options cannot meet the bar.
  • Require both offline and online evaluation: no model should go to broad deployment on offline validation alone or skip straight to A/B testing. Every model must be evaluated offline and online, and offline evaluations should increasingly predict online behavior.

Design: build to scale beyond the originating teams

Making a model work is one challenge; making its value extend beyond a single team or use case is harder.

  • Build on shared foundations; specialize only when justified: favor platform-wide foundations for core capabilities, data representations, and model building blocks. Specialization should extend these foundations rather than create isolated stacks, so improvements in the foundation benefit the entire organization.
  • Treat data as a first-class product: model quality is bounded by data quality. Maintain robust pipelines, clear lineage, reproducibility, and reusable features with documented ownership, clear schemas, and SLAs other teams can rely on.
  • Prioritize generality over local optimization: when two approaches perform similarly, prefer the one whose learnings, assets, and operating patterns are reusable across teams, brands, and use cases.
  • Minimize and sunset manual business rules: manual rules may be necessary for policy, safety, or compliance, but they should be explicit, regularly reviewed, and not become permanent maintenance debt.
  • Reproducibility and traceability by default: training data, features, configurations, evaluation results, deployment versions, and key decisions should all be documented and recoverable to enable debugging months later and smooth ownership handoffs.

Trust: ownership, governance, and responsible operation at scale

The question for deployment is not only “does it work?” but “can we stand behind it?” Trust is earned over time and must be sustained across a model’s full lifecycle.

  • Assign clear ownership and accountability: every model needs defined ownership across its lifecycle — a business owner, product owner, AI owner, and operational owner. These roles need explicit responsibilities: who is accountable for outcomes, who responds to model drift, and who handles incidents off-hours.
  • Adhere to standards and governance: AI/ML models should use approved platforms and comply with company standards, release gates, and governance processes. Operating outside those guardrails requires a defined remediation or deprecation path.
  • Govern proportionally to risk: the rigor of review, evaluation, and human oversight should scale with a model’s impact. Customer-facing models that affect pricing or availability for millions of travelers require a higher bar than internal tools used by small teams; high-impact or autonomous systems should include human-in-the-loop checkpoints from the start.
  • Design for fairness, privacy, and transparency: actively test for unintended bias, enforce data guardrails, and favor explainability when decisions meaningfully affect users.
  • Design for safe rollout, rollback, and control: deployments should be progressive and include rollback paths, fallback mechanisms, and circuit breakers prepared before launch.
  • Monitor continuously and adapt: once live, teams must monitor quality, drift, latency, cost, and business performance and retrain or recalibrate when data shifts. Teams should always be able to explain current model performance, not just launch-time results.

These principles do more than guide how Expedia Group builds AI; they define what the company is willing to release and how it stands behind those systems. In a world where AI systems increasingly make consequential decisions for real travelers and partners, consistent application of these standards is essential to build responsible AI that lasts.

Xavi Amatriain, Chief AI and Data Officer at Expedia Group, will present further details about Expedia’s architecture at VB Transform on July 14, 2026 at 11:10 am PT. The talk is titled: “Expedia's blueprint for building autonomous agents for high-stakes transactional systems.” Information about registration and a limited number of complimentary passes for senior technology leaders is available through the VB Transform program.