Zhipu AI's open-source GLM-5.3-Flash matches Anthropic's Claude Opus 4.8 on the AA Intelligence Index at 1/40 the price.
Zhipu AI's open-source GLM-5.3-Flash matches Anthropic's Claude Opus 4.8 on the AA Intelligence Index at 1/40 the price.

Zhipu AI's open-source GLM-5.3-Flash matches Anthropic's Claude Opus 4.8 on the Artificial Analysis Intelligence Index at 1/40 the price, intensifying a price war that has pushed Chinese models past 60 percent of global token usage.
"For high-volume, repetitive tasks such as coding, data extraction and customer support, leading Chinese models have nearly closed the performance gap with proprietary Western systems while delivering the same outputs at a far lower cost," Henry Mascot, chief executive officer of Nigerian insurtech Curacel, said.
The 320-billion-parameter model, with 18 billion active parameters, is the first native multimodal release in the GLM-5 series and scores 57 on the AA Intelligence Index, matching Claude Opus 4.8. Zhipu priced GLM-5.3-Flash at one-tenth of its GLM-5.3 flagship, one-twentieth during a limited-time discount, and one-fortieth of Opus 4.8. The company confirmed the model behind the "Ox Alpha" name that swept to the top of online usage charts is a new iteration of its GLM series, with weights released Wednesday.
The pricing lands as Chinese models' share of global monthly token usage on aggregation platform OpenRouter jumped from under 10 percent early last year to more than 60 percent in July. Alibaba's Qwen 3.7 Max built an identical e-commerce website for $4.08 versus $48.99 for Anthropic's Fable 5, a twelvefold gap on comparable output.
The launch deepens a competitive squeeze on US AI leaders. OpenAI cut API prices for its Terra and Luna midrange models by 20 percent and 80 percent in late July, while Anthropic released Claude Opus 5 at the same price as Opus 4.8 — $5 per million input tokens and $25 per million output tokens — claiming it costs half as much as its flagship Fable 5 while scoring within 0.5 percent on coding benchmarks. Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens.
The open-weight architecture is the decisive factor for enterprises handling sensitive data. Curacel, Africa's largest insurtech, added Zhipu's GLM-5.3 to its insurance claims and fraud-detection system alongside US models, citing the ability to run bank transaction records and know-your-customer documents on in-house servers rather than transmitting them to overseas APIs. Ugandan researchers built "Sunflower," a farming advisory tool in more than 30 local languages, on Alibaba's Qwen3 foundation model after comparing US and Chinese alternatives.
The price gap is not uniform. AlphaSense research found OpenAI's GPT-5.6 Sol and Anthropic's Opus 4.8 delivered higher-quality responses at lower total cost than Moonshot AI's Kimi K3 and Zhipu's GLM-5.2 for certain tasks, because Chinese models required more tokens to complete answers. The median total cost per query for GPT-5.6 Sol was 13 percent lower than Kimi K3 while its quality score was 20 percent higher. Artificial Analysis, using a different method, estimated DeepSeek V4 Flash recorded the lowest per-task cost among high-performance models at $0.03.
The pricing pressure is reshaping procurement. US workflow startup Lindy switched its primary model from an Anthropic product to DeepSeek in June, cutting costs 90 percent, while Palcia assigns overnight tasks to a Chinese MiniMax model. For investors, the question is whether US AI leaders can defend premium pricing as open-weight rivals close the performance gap — Anthropic and OpenAI have already begun discounting midrange tiers, compressing the unit economics that underpin their valuations.
This article is for informational purposes only and does not constitute investment advice.