Kimi K3 Tops Fable as Chinese Open-Source Models Sweep the Global Top Five

Source: 20VC with Harry Stebbings | Published: 2026-08-03T14:06:48Z

An AI-generated fake candidate fooled top engineers in an Arena technical interview, sailing through every round — only for the team to discover the person didn't exist when they moved to extend an offer.


Kimi K3 recently beat top American closed-source models — including Fable — on frontend coding tasks. The debate it sparked goes far beyond "yet another Chinese model."

Arena founder Anastasios called this a "narrative fracture" — not a shakeup in leaderboard rankings, but a challenge to the entire set of assumptions the American AI world holds about its own technical superiority.


Kimi K3 Didn't Break a Record — It Broke a Story

Before Kimi K3, there was a prevailing consensus in American AI circles: Chinese models kept pace primarily by distilling outputs from American models — essentially living off American scraps.

Kimi shattered that assumption. It surpassed some of the best American closed-source models on certain tasks — including Fable — and did so in a real, high-demand scenario: frontend web development. Anastasios's view is that this doesn't rule out distillation — distillation is just one substep in Chinese labs' training pipelines, not the whole story. "On top of distillation, they're doing something else that's pushing performance above American labs."

If you look at Arena's leaderboard today, the top five open-source models are all Chinese. Thinking Machines' Inkling — the top-ranked American open-source model — sits around tenth globally.

Why American Open Source Keeps Falling Behind

Anastasios argues this is a business model problem, not a technical one.

Open-source models don't generate revenue directly, so no one has figured out how to make enough money from them to sustain a company. Two paths are gradually becoming clear: revenue sharing — licensing open-source models to inference providers like Fireworks and Together and taking a cut past a certain scale — and using open-source as a top-of-funnel, where the real business is helping enterprises with fine-tuning, AI strategy, and implementation. That's the path Mistral and Thinking Machines are walking.

The second path taps into what he sees as a severely underestimated market: "One of the biggest markets of the next decade is AI transformation entering every company on earth — helping them restructure their data, integrate models into workflows, train employees on how to use them."

His prediction: the US will produce at least one company worth hundreds of billions — or even a trillion dollars — focused on "American-first open source."

Enterprises Want AI Sovereignty, Not the Best Model

A structural trend is taking shape on the enterprise side: companies don't want to hand their data to third parties that might one day compete with them.

Anastasios calls this impulse "AI sovereignty" — take an open-source model, fine-tune it on your own data, run it on your own infrastructure, keep the entire supply chain in-house. He sat in a meeting with a Fortune 50 company and was asked point-blank: "Does your stack use Qwen? Can you swap in an American model?"

This isn't an isolated case. His conclusion: a data moat combined with a self-improving model is one of the few viable ways for enterprises to maintain a competitive edge in the AI era — because software itself is becoming something that can be instantly replicated.

Local Deployment Doesn't Solve the Backdoor Problem

There's a common assumption that hosting a model on your own servers makes backdoor threats disappear. Anastasios says that's a misconception.

He laid out a concrete attack scenario: suppose you deploy a Chinese-trained model on your own infrastructure to run a chatbot with access to all your company's data. If the model was implanted with a trigger mechanism during training — a specific character sequence or keyword — an adversary only needs to craft a single conversation to make the model dump everything it knows. "This can absolutely be trained into a model, then let companies host it on their own infrastructure. It's an attack vector, and there are many possibilities."

His call: within three years, the US will very likely impose restrictions on the use of Chinese open-source models. "I'm not saying I support this, but if I had to bet, I'd bet on that side."

Fake Candidates Are Applying for Your Engineering Roles

Arena discovered something unsettling internally: they posted a job, someone applied, passed the technical interview, had one-on-ones with their top engineers — and then when it came time to make a formal offer, the person didn't exist.

"Not an AI-generated résumé. An AI-generated interviewer." Anastasios said his engineers sat across from this person and believed they were real. "Our engineers are world-class, and they thought the candidate was genuine."

The motive could be stealing the codebase, accessing company data, or something like that case from a year ago where someone collected four salaries simultaneously — except this time, even the person was AI-fabricated. Arena is considering moving the entire hiring process to in-person, requiring physical presence and a handshake. Figma has reportedly already done this.

"This is something the entire industry is going to go through. We're going to redesign the whole hiring process because of this, and it genuinely worries me."

75 Neolabs, Two-Thirds Going to Zero

Anastasios is blunt: there are at least 75 Neolabs right now, and two-thirds of them will end up either worthless or acqui-hired for parts.

The math backs it up. If a Neolab is valued at $10 billion and investors want a 10x return, the company needs to reach $4 billion in annual revenue within two to three years (working backward from a 25x revenue multiple). Many Neolabs are carrying multi-billion-dollar valuations with zero revenue. "The market has become P&L-driven. Just building a model and hosting a launch party — that's yesterday's game."

But he acknowledges a counter-logic: worst case, you get acquired via liquidation preference, the team itself is worth a billion, and investors who put in $200 million feel like they came out ahead. That reasoning is turning the next funding round into a trap for many Neolabs — "the next round is a pit."

Data Is a "Scaling Complement"

Anastasios offers a clean economic framework for why the data market is a hypergrowth market.

He calls data a "scaling complement" — like the relationship between cars and gasoline: sell more cars, demand for gas goes up. The larger the model, the more training data it needs; the more enterprises start training their own models, the more data demand multiplies.

When does data become a commodity no one cares about? When humans become irrelevant. In other words, not until AGI. He predicts the data market will be at least $100 billion by 2030, potentially reaching a trillion.

To the common criticism — "data providers are too concentrated on OpenAI, Anthropic, and Meta" — his response: TSMC has revenue concentration. Anduril has revenue concentration. Plenty of publicly traded companies worth hundreds of billions have revenue concentration, and it doesn't stop them from being great businesses. Silicon Valley investors, he thinks, have been spooked on this issue.

Frontier Labs Are Eating the Application Layer

Claude's design-focused products have already started cannibalizing Figma's market. The threat Harvey and Legora face isn't just from competitors — it's from the strategic direction of their largest supplier.

Anastasios says friends of his running multi-hundred-million-dollar businesses are all experiencing the same thing: their biggest customers come in and say, "OpenAI is moving into this space and we're going to work with them because they're more AI-forward."

He believes the risk is real, and the structural logic holds: as inference itself trends toward commoditization, model providers need to move up the application layer to stay closer to end users in the value chain and avoid being commoditized themselves. Harvey's CEO has acknowledged that its biggest competitive threat comes from the model labs.

On the apparent contradiction — enterprises both fearing and embracing Frontier Labs — his explanation comes down to deployment complexity. Claude building a design tool is something designers can pick up and use immediately. Harvey doing legal work means penetrating the relationship networks of fifty-year-old partners and convincing junior associates to adopt a tool they think will take their jobs. The GTM heavy lifting is real, and that's currently Harvey and Legora's most durable moat.

Evaluation Is the Real Bottleneck for AI Deployment

Arena has over 30 million monthly active users and annualized Q2 revenue above $100 million, growing fast. But Anastasios says their core business thesis isn't data — it's evaluation.

"Every enterprise needs evaluation. That's non-negotiable, and it's the biggest bottleneck to AI deployment." Cost reduction is easy to quantify — switch to Gemini Flash and token spend drops immediately. But how do you define "performance"? For a healthcare company and an e-commerce company, the answer looks completely different.

Arena is bringing the methodology it built on public benchmarks into the enterprise: by analyzing AI agent execution traces, extracting performance metrics from real workflows, and helping companies use their own data to determine which model works best for their context — rather than relying on externally purchased datasets.

This, he believes, is a need that will inevitably emerge as every enterprise begins training its own models and accumulating its own data: someone to tell them whether this model actually works for their specific business.

More articles on TLDRio