A year ago, Kimi's founder described this week
I watched a viral clip and hit a year-old interview that maps onto what shipped this month. The snow mountain, the two stones, and why one model was never the plan.
A friend dropped a video in my chat. Fifteen minutes, a guy in glasses talking into a mic, Chinese audio with English subtitles. The caption said something like "watch this instead of Netflix". I almost scrolled.
Then I watched the whole thing. It turned out to be a cut-down of a much longer interview, recorded a year ago right after Kimi K2 shipped. A year on, it plays like a forecast.
The full interview the short clip was cut from - Yang Zhilin with Zhang Xiaojun, about an hour, straight from her own channel.
TL;DR The clip is a re-edit of a 100-minute interview with Yang Zhilin, founder of Kimi, recorded a year ago right after K2. In it he draws a map: progress is an endless climb, "reasoning" and "agent" are separate axes, and no single model has to be best at everything. This month K3 shipped and the market quietly proved the map. The frontier is now routing between an open default and a frontier specialist. The single-model era is ending, and two vendors who profit from that are the ones announcing it.
Most AI arguments this month are a fight over which single model is king. I stopped finding that fight interesting a while ago.
1. The mountain has no top
Yang's metaphor is a snow mountain. You climb, you reach a ledge, and the view is just more mountain. Every problem you solve lifts you a few hundred meters, and from the new height you see problems you could not see before.
He got it from a book. "The Beginning of Infinity", David Deutsch, 2011. Yang says there are two sentences worth carving into stone. Problems are inevitable. Problems are solvable.
The standard pitch treats AGI as a finish line - cross it, collect the trophy, done. Yang treats it as a heading you steer by for years. In some narrow fields the models already beat 99% of us. In others they trip on things a child would not.
2. Reasoning and agent are not the same axis
A year ago Yang made a technical aside. He said "reasoning" and "agent" are two different things, and a model can be strong at one while ordinary at the other. His example was Claude. Great at acting as an agent, running long tool loops, getting real work done, even when its raw reasoning score was not the top of the chart.
So he refused the neat ladder. You have seen the L1 to L5 levels of AI, the chatbot then reasoner then agent then innovator then organization progression. Yang says those are milestones you can hit out of order. You can be a serious agent before you are a serious reasoner. The two lines rise on their own clocks.
Innovation, in his telling, is when a model helps build the next model. His words, roughly, we hope K2 can help develop K3. A year ago that was a wish in an interview.
Then this month K3 shipped.
3. The market shipped his map
While Kimi K3 was landing, a blog from Fireworks made the rounds. They pit K3 against Fable, the frontier model from Anthropic, across about a thousand agentic tasks. Fireworks sells inference. They have a reason to want this result, so read the numbers as their numbers.
Their headline: the two models basically tie on accuracy, and on long agentic loops K3 comes in up to fifty times cheaper. On software tasks they report 92.4 for K3 against 92.6 for Fable. K3 pulls ahead on terminal work, on crypto, on dev tooling. Fable pulls ahead on web, on data visualization, on breadth across languages.
Cost against accuracy across five task families. Chart: Fireworks AI.
Their actual conclusion: K3 plus Fable beats either model alone. Route each task to whichever one owns it. An ideal router, they say, sends most work to the cheap open model and saves the expensive one for the cases that earn it.
Same week, a "Show HN" post called Echo made the same claim from the other side: frontier-level results at a third of the cost by routing across open-weight models. The idea is not new, OpenRouter and NotDiamond got there first. Nobody shows you which model actually answered. Cache fragments when you bounce a conversation between models. And on a subsidized 200-dollar Claude plan the math changes anyway.
Strip the pitch and the skepticism both, and the shape underneath is the one Yang drew. The old assumption was one model best at everything. What is arriving instead is an organization of models, each doing what it is good at. He used that exact word a year ago. Organization.
4. What I actually take from this
I run a company that builds content tools with AI. Full disclosure, we have run Kimi K2 inside those products since around July last year, right when it shipped, picked for its creativity and its high EQ-Bench score. I want the "many models, routed" world to win. It is cheaper for people like me, and it means no single vendor holds the whole board.
When the sellers and the skeptics agree on the shape and only argue about the price, the shape is usually real. And the founder who said a year ago that his model would help build the next one just shipped it. You do not have to trust the benchmark numbers to notice that. He said it, then he did it, while the rest of us argued about the summit.
If you build on these models, the move this year is boring and it works: route well, and reach for the expensive one only where the task earns it. The mountain has no top. Get used to that.
Three things, if this landed:
Working with the new open models? Tell me what you are routing and how you decide, I am collecting real setups, not vendor decks.
Want the source? The full one-hour interview holds up. Tell me which of Yang's bets you think ages worst.
Building content with AI at any scale? That is what I do at Co.Actor. Write me if the "many models" world is breaking your stack.
Next issue: what routing between an open model and a frontier one actually costs once you leave the benchmark and run real work through it. Subscribe if you want the real numbers.






