Writer, the enterprise AI agent platform used by firms including Accenture, Uber and Vanguard, announced Palmyra X6 today alongside a redesigned orchestration layer (the Writer Agent harness) and new governance features aimed at giving IT leaders control over token consumption and unexpected bills.
Headline claims: cost, speed and quality
Writer reports that their agent product now runs at an average 52% lower cost, delivers results 48% faster, and shows a 10% quality improvement when paired with Palmyra X6. Those figures are prominent, but the more consequential aspect is how Writer achieved these gains and what that implies for enterprise AI.
Base model and provenance: Palmyra X6 and GLM‑5.2
Palmyra X6 is not trained from scratch. According to Writer’s technical report, it is a post‑trained version of GLM‑5.2, the open‑weight mixture‑of‑experts model released by Beijing‑based Z.ai (formerly Zhipu AI). Writer discloses this fact and emphasizes that fine‑tuning and operations run entirely on U.S. infrastructure.
Matan‑Paul Shetrit, Writer’s director of product management, told VentureBeat that the model "is in no way, shape, or form connected to any of its original developers" and that everything runs on U.S. infrastructure. Dan Bikel, who leads Writer’s AI research, said the company took the floating‑point weights as a starting point and trained from there, producing what they call a Palmyra model.
Why it matters: agents and exploding token bills
Agentic AI consumes many more metered tokens than a typical chatbot because a single user request can trigger multiple rounds of planning, retrieval, tool calls, validation and retries — the user sees one answer, but the invoice reflects the whole loop.
Goldman Sachs forecasts token consumption could multiply 24× between 2026 and 2030 to reach 120 quadrillion tokens per month, driven by always‑on enterprise agents. Their analysis cautions that lower per‑token prices do not guarantee lower bills: if an agent draws 20× more tokens while unit prices fall 75%, total charges still rise fivefold.
Waseem AlShikh, Writer’s CTO and co‑founder, summarized the tension: enterprises want token consumption to grow because it signals adoption, but they also need costs to flatten.
Shetrit framed cost as the main barrier to enterprise AI adoption in many cases, arguing that the alternative to AI is often human labor. He also rejected the idea that reducing customers’ token consumption would harm Writer’s revenue, saying lower per‑task costs expand the addressable opportunity inside organizations and thereby grow the company’s business.
Palmyra X6 technical approach: 744B parameters, 626 synthetic trajectories
Palmyra X6 is a 744‑billion‑parameter mixture‑of‑experts model, with about 40 billion active parameters per token and the same architecture as GLM‑5.2. Writer’s contribution is a conservative post‑training recipe called anchored supervised fine‑tuning (ASFT). The company applied ASFT to a very small, curated corpus of 626 fully synthetic agentic trajectories, training for a single epoch at a low learning rate.
ASFT combines a token‑weighting scheme with a KL‑divergence "anchor" that penalizes the fine‑tuned model for drifting too far from a frozen copy of the base model — enabling new tool‑use behaviors without eroding base capabilities. Writer also replaced the standard Adam optimizer with Muon on the model’s core weight matrices.
The training data are entirely synthetic: teacher models generated every plan, tool call and final answer; these were then filtered by structural quality gates, a model‑based verifier and a two‑model LLM judging panel before entering training. Writer previously trained Palmyra X004 mostly on synthetic data for roughly $700,000 in 2024, and Palmyra X5 required about $1 million in GPU hours, according to earlier reporting.
Benchmarking and pricing
In Writer’s internal evaluations across nine capability categories (grounding and retrieval, tool use, content generation, sub‑agent delegation, brand voice, etc.), X6 averaged 0.87 out of 1.00. That slightly outperformed Anthropic’s Claude Opus 4.8 (0.86), Claude Sonnet 4.6 (0.85), OpenAI’s GPT‑5.5 (0.80), and Google’s Gemini 3.1 (0.77). More materially, Writer prices X6 at $2 per million input tokens and $8 per million output tokens versus Opus 4.8’s 15/75 pricing.
Writer acknowledges internal benchmarks invite skepticism; their technical report details both the public benchmarking protocol and internal evaluations. The company uses public benchmarks as sanity checks rather than slavish targets, they said.
The China question: security and trust when building on GLM‑5.2
Two years ago, using a China‑origin open‑weight model would have been unlikely for a U.S. enterprise vendor. Today GLM‑5.2 (released in June under an MIT license) is among the most capable openly available models. Independent analyses have scored it highly on agentic tasks while offering much lower cost than many U.S. flagship APIs.
Open‑weight availability brings baggage: an August report from AI safety nonprofit SaferAI found GLM‑5.2 did not refuse offensive cyber or biological tasks via Z.ai’s public API, and Z.ai published no safety framework or pre‑deployment risk assessment — risks that grow once anyone can download and modify the weights.
Writer’s response is that provenance and post‑training matter more than origin. The company says it obtained the weights from the U.S. Hugging Face mirror, ran all synthesis and storage in the U.S., and trained on U.S. hardware. It also ran a pre‑registered model risk evaluation covering political bias, censorship, factuality and refusal behavior: 19,674 responses were scored by blinded judges comparing X6 to GLM‑5.2 and four frontier control models.
On The Washington Post’s ModelSlant political‑bias evaluation, Writer reports X6 presented both sides of hot‑button questions 80% of the time — the highest rate in their test set — and answered politically sensitive prompts that DeepSeek V4 refused. On the FORTRESS adversarial safety benchmark, X6 with its deployment system message scored 8.6 points higher than raw GLM‑5.2 on adversarial safety, with negligible cost to benign helpfulness. The report does note behavior varied by language, a candid admission that 626 fine‑tuning trajectories do not erase all traces of the base model’s training.
The harness effect: orchestration may matter more than the model
Perhaps the strategically most important claim is about the harness rather than the model. Writer says its rebuilt Writer Agent harness reduces costs by 41% and speeds task completion by 44% across every tested model, including third‑party models from Anthropic and OpenAI, while maintaining quality. The company published these findings in a paper they call "The Harness Effect."
That leads to an obvious question: if the harness delivers most savings regardless of model, why build a model? Shetrit answered that building gives the company control over deprecation, data provenance and the answers enterprise customers demand. Bikel added that the model and harness were developed together and co‑evolved to work optimally with Writer Agent.
At the same time, Writer is hedging: Writer Agent now supports multi‑model configurations so admins can enable Anthropic, OpenAI and cloud‑provider models (Microsoft Azure, AWS Bedrock, Nvidia NIM) and even image‑generation models. The platform message is simple: use our model if it’s the cheapest and best fit, but the harness saves money either way.
New governance features to prevent surprise AI bills
The release’s third pillar addresses a quieter but widespread pain point: senior leaders often don’t know what agents are spending. New governance tools provide centralized visibility of agent usage across the business, per‑workflow analytics for shareable "Playbooks" and "Skills", and consumption controls with alerts and spending limits.
Shetrit framed these features as adoption enablers rather than damage control: giving CIOs and CISOs visibility and spend/security controls makes them comfortable expanding AI usage. Firms with clear cost data are more willing to automate use cases they would previously have avoided.
The features reflect a broader shift in enterprise AI budgeting: sophisticated buyers are modeling cost per successful task — including retries, tool calls and escalations — rather than simply multiplying expected calls by a rate card. Writer effectively productizes that financial discipline as a native platform capability.
Company focus and strategic signal
Writer was founded in 2020 by May Habib and Waseem AlShikh, and raised $200 million at a $1.9 billion valuation in late 2024. The company has focused on regulated, high‑stakes enterprise deployments rather than consumer scale.
Shetrit was explicit about focus: as an enterprise company they do not need general capabilities like writing a French sonnet; instead they optimize for customer problems and needs. He emphasized that Writer is pragmatic: if building from scratch is the right answer they will do so, but if better market alternatives exist they will use those.
That pragmatism may be the most important message of the release. A well‑capitalized American AI company with five years of model experience concluded that much of the frontier value now lies in post‑training open weights, engineering orchestration around them, and giving finance teams a dashboard. If Writer is correct, the moat of frontier labs narrows to workloads where quality justifies a sevenfold price premium — for everything else the winning model may be the one someone else already pretrained.
In an industry long focused on whose model is the smartest, Writer is betting the enterprise AI race will be won by the company that best knows how to use models and control their costs.



