Did Kimi Just Start Another DeepSeek Moment?
The gap between Chinese and American AI just got a lot smaller.
On January 20, 2025, a relatively unknown Chinese AI lab called DeepSeek dropped a model that briefly erased nearly $600 billion in market cap from NVIDIA in a single session. If a Chinese lab could train a frontier model for a fraction of the cost, the entire justification for the hyperscaler’s insane AI spending binge was immediately in question
However, we moved on. Tariffs arrived and the news cycle shifted. The AI trade kept running regardless.
Today, Moonshot AI, startup from Beijing valued at only $31.5B, released Kimi K3, and the comparisons to that January are already starting.
What Kimi K3 Actually Is?
Kimi K3 is a 2.8 trillion parameter Mixture-of-Experts model, the largest open-weight model ever released from any country. It runs on a novel architecture Moonshot calls Kimi Delta Attention, which:
Reduces KV cache memory usage by 75%, and
Delivers 6.3x faster decoding at 1 million token context lengths compared to standard attention.
It also ships with native vision capabilities and always-on reasoning, alongside that same 1 million token context window.
On benchmarks, it landed 3rd on the Artificial Analysis Intelligence Index with a score of 57 right behind Claude Fable 5 at 60 and GPT-5.6 Sol at 59.
It beat every other model on Frontend Code Arena. It also placed 2nd on AA-Briefcase for long-horizon agentic knowledge work, and it led all models on AutomationBench and BrowseComp.
Full model weights drop July 27.
Nevertheless the capability most are missing isn’t a benchmark at all.
During its own development, an early version of Kimi K3 handled the majority of the team’s kernel optimization work. In a separate demonstration, K3 built MiniTriton, a tool that translates AI code into instructions a GPU can actually run, basically recreating a piece of software (Triton) that NVIDIA and others rely on for this job.
It did the same job from scratch and it matched Triton’s performance on certain benchmarks. K3 also designed a chip layout for a small version of its own architecture in 48 hours.
Anthropic and OpenAI don’t have a disclosed capacity like this, and it puts self-optimization question at the center of everything below.
Why Markets Are Scared
Whole AI trade, covering the semiconductor boom, the neocloud rerating, and the power and datacenter infrastructure buildout, rests on a single assumption:
Training better AI model is so profitable that it justifies exponentially increasing capex.
Anthropic trains Fable 5 → charges a premium → uses that revenue to fund Fable 6.
OpenAI raises $122 billion at a $850 billion valuation because the market believes the next model will be worth even more.
The hyperscalers commit 100s of billions per year in capex because the ROI on owning the infrastructure that serves this intelligence is enormous.
That logic holds as long as the premium holds, and the premium only survives as long as the capability gap does.
Moonshot AI is valued at $31.5 billion and Anthropic is reportedly valued north of $1 trillion.
The model Moonshot just released is three points behind Anthropic’s best on an independent benchmark. We’ll let you draw your own conclusions.
Also, as SemiAnalysis noted, K3 is so large it won’t even fit on a single NVIDIA DGX B200 at FP4.
You need a GB300 NVL72, B300, or MI355X each with 288GB of memory just to load the weights. That’s the kind of hardware requirement that keeps the semiconductor trade alive. The real risk sits a generation ahead, in what happens when the next version of this model trains itself to be more efficient and fits on something smaller.
The Bear Case
Strip away the noise of the market and the bear case rests on a single question:
What happens to the capital equation if intelligence becomes a commodity?
Start with open-weight models from China closing the capability gap. That pressure alone collapses token pricing for closed models, which means less revenue frontier labs use to justify their next training run. Shrink that revenue enough and the economic case for $7-25 billion training runs analysts are projecting for GPT-6 class models gets much harder to make.
Without those training runs happening at the scale the market is pricing in, the semiconductor names currently priced on 2027-2028 estimates face a serious multiple compression problem. Starting with the most leveraged, most momentum-driven names in the AI basket:
Neoclouds,
Power Plays, and
Data Center REITs
The even more unsettling version of this scenario is what K3’s self-optimization capability implies. If each model generation uses the previous one to compress its own development cost, and those savings don’t get fully recycled into more ambitious models, the compute required per frontier model generation could start scaling slower than the market’s current projections embed. Resulting in a quiet erosion of the assumption that got everything priced where it is.
The Bull Case And Why Jensen Keeps Talking About Open Source
Gavin Baker, CIO of Atreides Management, put the bull case cleanly:
Cheaper models increase the ROI on AI spend for end customers, which drives more token consumption, which pushes more margin down to the infrastructure layer. That shift benefits whoever has the lowest cost per token at scale, NVIDIA and the hyperscalers most of all.
This is why Jensen Huang has been one of the loudest advocates for open source despite running the company that sells the hardware to train closed models. A world with many competitive model providers (open and closed) spreads buying power across a wider base and eliminates the monopsony risk of two or three frontier labs having pricing power over NVIDIA as a supplier. Open source growing the ecosystem is good for the company selling shovels to every participant in it.
And there’s a harder version of the bull case specific to K3 itself. An open model at this scale that gets widely deployed actually increases demand for the highest-end inference infrastructure. Every organization that wants to run K3 seriously needs GB300-class hardware, the same tier SemiAnalysis flagged as necessary just to load the weights.
Kimi K3 Isn’t Cheap
Kimi K3 is not cheap to run. Not even close.
At $3 per million input tokens and $15 per million output tokens, K3 is priced comparably to GPT-5.6 Terra, OpenAI’s mid-tier offering. Artificial Analysis puts:
K3’s cost per task at $0.94,
Claude Opus 4.8 at $1.80
GPT-5.6 Sol at $1.04.
It undercuts the most expensive closed models, but it isn’t a DeepSeek-style price shock. K3 is also token inefficient, which matters more than the pricing alone. It uses more tokens per task than GPT-5.6 Sol and Grok 4.5, which are both significantly more token efficient. Enterprise buyers care about what they get per dollar spent on a task, not what they get per token, and on that basis the per-task cost ends up comparable.
The real DeepSeek moment, the one that would actually break the pricing thesis for Anthropic and OpenAI, would be an open-weight model at frontier quality that is also token efficient.
K3 isn’t that model yet.
There’s also a deeper structural point that Gavin Baker raised: the most dangerous scenario for closed frontier labs comes from vertical integration, rather than OS. Meta owns Llama, and SpaceX/xAI owns Grok. These companies own both the model and the infrastructure, so they can charge a premium elsewhere and never need to extract margin at the model layer.
Grok 4.5 at $2 per million tokens, benchmarking close to Fable 5, does more damage to Anthropic’s pricing power than K3 does.
The Four Scenarios
At this point there are four distinct ways this plays out and none of them have resolved yet.
Commoditization.
China keeps closing the capability gap, token pricing collapses, the ROI on US capex commitments stops making sense, and multiples compress starting with the most leveraged names in the AI basket.
The Baker scenario.
Cheaper intelligence expands total token consumption via the Jevons paradox, margin migrates from the model layer to the infrastructure layer, and the picks and shovels names win regardless of who builds the best model.
The quiet erosion scenario.
K3’s self-optimization capability means each model generation gets cheaper to develop. If the labs respond by spending less rather than training more ambitiously which would be historically anomalous but not impossible the capex cycle slows without a dramatic trigger event. No crash, just a slow fade in the assumptions that got everything priced where it is.
The compounding loop.
K3 helped build itself. If K4 helps build K5, and K5 helps build K6, the cost curve for frontier AI bends faster than any model the market is currently using to price these stocks. At some point the compute required per frontier model generation stops scaling at the rate that $700 billion per year in hyperscaler capex requires to make sense.
The Bottom Line
Kimi K3 threatening Anthropic today isn’t actually the scenario worth pricing. A world with more competitive models at lower margins is better for every layer underneath the model:
semiconductors,
hyperscalers,
power,
and data centers.
What breaks the pricing thesis for Anthropic and OpenAI is an open-weight model that is simultaneously at the frontier and token efficient. K3 is neither. K3 runs 50-70% more expensive per task than GPT-5.6 and Grok 4.5. Intelligence per dollar, not intelligence per parameter, is the only metric that matters for enterprise adoption. On that metric the closed labs still win.
What Kimi K3 tells you is that the model layer is becoming more competitive. That is unambiguously good for Jensen, good for the hyperscalers, and good for anyone building on top of AI. It’s only bad for the two or three labs that were hoping to extract 90% inference margins forever. And those labs were never going to be allowed to do that anyway.
We’ve been here before, and the trend hasn’t broken yet. Infrastructure keeps winning.

















Unlike DeepSeek, I don't think there is any evidence Kimi is not living up to its claims.
The other issue is a massive circle jerk OpenAI and Anthropic are involved in with the semis/hyperscalers.
Nice one! Question is what happens to the market if openai for example could not win the price war and could not meet the RPO’s, curious what Oracle and MSFT stock would do in this scenario and if investors would like them to keep on spending on capex