
Writer, the enterprise AI company, has launched Palmyra X6, its new flagship model, alongside an upgraded “harness” built to do something the industry has been curiously reluctant to promise: spend fewer tokens. The model is a post-training variation on Z.ai’s open-source GLM-5.2, and it is available to Writer’s clients from today.
The pitch arrives at a nervy moment. As agentic AI quietly multiplies the steps, and therefore the tokens, behind every task, enterprise bills have started to sting. Writer is wagering that a cost backlash has arrived, and it is far from alone in scenting the opportunity, with Baseten raising $1.5bn on the theory that AI’s profits lie in cheap inference.
What is a harness?
A harness, in Writer’s telling, is the orchestration layer wrapped around a model, the machinery that decides how a multi-step agent actually executes each request. Optimise that layer, trimming the redundant calls and bloated context that agents tend to accumulate, and you cut token consumption without ever touching the model itself. The savings, Writer says, are substantial. The company estimates that the new model and harness together cut costs by as much as 50% for basic tasks, while a Writer research paper found that harness-efficiency changes alone trimmed costs by roughly 40% on average across testing.
What makes the harness interesting is that it is model-agnostic. It works with Writer’s own models, but also with external ones served through Microsoft Azure and Amazon Bedrock, so the efficiency gains are not locked to a single vendor’s roadmap. “The harness is the one component whose efficiency multiplies across every model an organization runs, present and future,” Writer’s researchers write, framing orchestration rather than raw model quality as the lever enterprises have been ignoring.
A pointed message to big labs
It is also a pointed one. Writer’s underlying claim is that the big labs have little incentive to help you spend less, because their revenue rises with every token you burn. Distrust of that arrangement is doing a lot of work in the company’s messaging. CEO May Habib put it bluntly. “The enterprise is absolutely sick of chasing the next benchmark,” she said. “They want flattening cost.” It is a neat inversion of the usual sales script, in which each new model is faster, cleverer and, invariably, hungrier.
A wider thrift-maxxing trend, in which buyers reach for cheaper Chinese models, has already started to unsettle the valuations underpinning eventual OpenAI and Anthropic IPOs, a sign that price sensitivity is no longer a fringe concern. Building the flagship on GLM-5.2 is telling in its own right. Rather than train a frontier model from scratch, Writer took a capable open-source base and specialised it through post-training, a route that keeps costs down and neatly matches the frugal story it is selling.
Appealing to European buyers
For European buyers, the framing should feel familiar. Scepticism about American hyperscalers’ incentives, and a preference for architectures that are not hostage to one provider, has been a running theme on this side of the Atlantic, and Writer is pitching straight into it. The company’s harness approach also resonates with a broader movement in enterprise technology toward vendor-neutral abstraction layers, where companies can switch models without rewriting their entire stack.
The timing is no accident. Agentic AI has moved from demo to production in many organisations, but the operational costs have become a board-level concern. Each agentic workflow can involve dozens of model calls, each with its own prompt and context window, and the cumulative token burn can far exceed the cost of a single answer. Enterprises are increasingly demanding observability into these token flows, and tools that can reduce waste without compromising output quality are gaining traction.
Writer’s bet on harness engineering is an attempt to redefine where AI value creation lies. Instead of competing solely on benchmark scores, the company is positioning itself as the efficiency player, the one that helps clients do more with less. This is a stark contrast to the prevailing narrative in the industry, where larger models and longer context windows are marketed as inherently better, even if they cost more to run.
The underlying economics are worth unpacking. Token prices have been falling across the industry, but the volume of tokens consumed has been rising even faster, driven by agentic workloads that loop through multiple reasoning steps. For many enterprises, the net effect is that their AI spend is still climbing, and finance teams are starting to push back. Writer’s research paper claims that its harness can reduce the number of context tokens sent to a model by pruning irrelevant information and storing intermediate results more efficiently. These optimisations are particularly valuable for complex tasks like document analysis, coding assistants, and customer support automation, where agents often re-read the same source material multiple times.
There are caveats worth keeping in mind. Efficiency figures produced by a vendor about its own product deserve the usual pinch of salt, and “up to” 50% is not the same as a guaranteed 50%, but the direction of travel is hard to dispute. Independent testing will be needed to verify Writer’s claims, and the company has not yet published detailed benchmark results for Palmyra X6 on standard efficiency metrics. Still, the broader industry is moving in the same direction. AI infrastructure companies like Baseten are raising huge sums to provide low-cost inference, and cloud providers are introducing new pricing models that reward efficient usage.
For Writer, the move also signals a strategic shift in its relationship with open-source models. By building Palmyra X6 on GLM-5.2, the company acknowledges that the frontier is no longer the exclusive domain of closed labs. Open-source models have become so capable that bespoke fine-tuning can yield competitive performance at a fraction of the cost. This is a playbook that several other enterprise AI vendors have adopted, and it is likely to become more common.
The deeper shift is what Writer calls harness engineering: the discipline of spending fewer tokens per task rather than chasing another point on a leaderboard. If it takes hold, and the incentives increasingly suggest it will, the era of the ever-hungrier model may finally be meeting its budget. Enterprises are beginning to demand that every AI purchase justify itself through measurable cost savings, and vendors that can demonstrate tangible ROI will have a clear advantage. Writer’s new offering is a bet that the future of AI is not just smarter, but leaner.
