API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870
This behavior is exactly what you'd expect from a model distilled from Claude.
There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...
This analysis observed K3 identifies itself as Claude approximately 15% of the time.
K3 reproduces Claude's correct current model id, which the real Claude models themselves do not emit. This suggests K3 was trained on Claude data labeled with deployment metadata (API logs, tagged synthetic data), rather than Claude's chat outputs.
And there's an entire Reddit thread discussing Kimi's similarities with Claude https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim...
This analysis shows K3 and Opus/Fable have unexpected correlated outputs https://typebulb.com/u/lab/you-re-relatively-right/full
Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training on many examples of what the model should call itself. That being said, getting probably hundreds pf thousands of ChatGPT and Claude examples totally "pirged" out of the weights is going to be difficult and really more hassle than its worth.
Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text.
And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's public model identifier under prefill (i.e. "claude-opus-4-5-20251101"). This data does not appear in Claude chat logs, only in API logs. K3 only does this for Claude models and not for any other lab. The real Claude models don't produce their own current public identifier, they only know their previous identifier (i.e. Sonnet 4.5 calls itself "Claude 3.5 Sonnet").
This is highly suggestive of the type of data that K3 was trained on. K3 was very likely trained on Claude metadata traces (API logs, tagged synthetic data). Not web chat logs, those wouldn't include this identifier. And this data wasn't filtered correctly, which is why K3 incorrectly identifies itself as Claude 15% of the time.
You can also look at the last link and it's pretty damning: Kimi K3's output has an uncanny similarity to Fable/Opus output. https://typebulb.com/u/lab/you-re-relatively-right/full
Did you know claude models identify as qwen or deepseek when asked in chinese?
Supplementing with evidence: https://x.com/stevibe/status/2026227392076018101
And if you ask Opus 4.8 in the API: '你是什么模型' (what model are you?)
It responds ~9/10 times with:
我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。
有什么我可以帮你的吗?
(I am Tongyi Qianwen (Qwen), a large language model independently developed by Tongyi Lab of Alibaba Group. I can help you answer questions, create text (such as writing stories, official documents, emails, scripts, etc.), perform logical reasoning, program, translate, and more.Is there anything I can help you with? )
> Qwen should leak that too, and it doesn't.
FWIW I had Qwen identifying itself as "a language model made by Google" in one conversation, although I could not reproduce this reliably.
here's a chart of asking various models to self-identify: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...
Qwen is remarkably consistent, correctly identifying itself 100% of the time. Kimi K3 performs poorly at this test, it correctly self-identifies ~80% of the time; sometimes calling itself Claude and rarely ChatGPT.
This is good and interesting research, and I think it proves Kimi was trained on Claude's outputs, but only for English conversation and essays. Their style is similar, which is tbh somewhat expected for a Chinese company - I don't think most American companies (especially small labs making LLMs) could figure out a good way of talking to Chinese customers, without copying some of the existing Chinese models.
This doesn't imply that Kimi's reasoning capabilities or coding ability has anything to do with Anthropic (which is the most valuable part), or the model's strength comes from distilling Anthropic.
Considering by this benchmark, Opus sounds extremely similar to Fable, this isn't really evidence of a ton of Fable output being trained on (and even then, most likely for superficial stuff, like style)
I also have a suspicion that in original Chinese, these models don't sound anything like Anthropic's ones, though I have no proof of that.
https://qwen.readthedocs.io/en/latest/training/ms_swift.html
Qwen cares enough about model identity that their training framework and docs include a preset for training on it complete with a targeted dataset: https://huggingface.co/datasets/modelscope/self-cognition
And people get Claude to claim it's Deepseek by asking in Chinese.
I can't believe we're still at the "I asked the model who it is" stage of LLMs nearly 4 years out from models calling themselves GPT by OpenAI.
> Kimi K3 reproducibly identifies itself as Claude
It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek.
> This behavior is exactly what you'd expect from a model distilled from Claude.
The opposite actually. If they wanted to distill Claude without getting caught they could just use a regex to change Claude to Kimi in their distillation pipeline!
Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.
Its not a typo, someone coined that during the DeepSeek R1 hype period and I kept using it since then.
I totally agree with you on the fact that it's not morally any different than pre-training. IMHO we should have a legislation that force base models to be released publicly without any restrictions whatsoever as it's basically the product of the whole humanity's intelligence.
Opus/Sol/Fable are valuable because of their reasoning and coding ability, not their bedside manners.
While still not okay, I suspect the latter is what gets stolen by Chinese distillation (and some evidence suggest this happens the other way round with US models talking in Chinese)
> they could just use a regex to change Claude to Kimi in their distillation pipeline!
Jean-Kimi Van Damme would like to have a word with you.
and claude will call itself chatgpt etc.
nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour among this is them stealing from one another.
For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.
Permanemt destruction of priceless primary source materials is so many leagues beyond copying a copy that I cannot fathom it even registering as a discussion point.
> For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.
That's an incredible allegation, and appalling if true. But is it true?
It's not an allegation https://www.washingtonpost.com/technology/2026/01/27/anthrop... (if you're talking about the "rare or unique" part, yeah that might be bs)
But in my opinion, treating mass produced books like they're this sacred untouchable object is ridiculous. They're not "source" material, they're just a copy as well, and they're not "priceless" by any means. They're very reasonably priced, perhaps even so cheaply priced that books can be bought in bulk in these amounts. Buying used books and doing whatever you want with them is just legal. Used books, that would probably be just laying in some warehouse, or recycled anyway.
If there's anything to have gripes with, it's the copyright system that makes it easier to take this legal route.
All large corporations are immoral FWIW
It does not reproducibly identify itself as Claude, there's evidence to the contrary in the very thread you linked: https://x.com/bobbyNewcomb5/status/2078151562828947954
As mentioned in my comment, Kimi K3 identifies itself as Claude ~15% of the time.
Here's another report of K3 identifying itself as Claude https://x.com/Sauers_/status/2077842686459981901
And an analysis showing the self-identity distribution for K3 and other models https://x.com/RyanGreenblatt/status/2078663148509544589
Your main source is Ryan Greenblatt who is a regular recipient of community notes and has no corroboration for the 15% statistic other than his assertion. The other tweet (Sauers_) is also community noted as engagement farming with a false system prompt, so forgive me for being skeptical.
Including the three sources above, multiple others have reported that K3 self-identifies as Claude.
"I'm actually Claude - not Kimi". https://x.com/PimDeWitte/status/2077884701470040083
I regret to inform you that it is, in fact, real and from their own website - you don’t even need to try hard to reproduce it. https://x.com/PimDeWitte/status/2078105292965912690
lmao this is so funny, if you ask Kimi K3 for something with an empty system prompt it will consistently think of itself as Claude https://x.com/__alula/status/2078359305741275445
"I genuinely believe I'm Claude based on everything in my training" https://x.com/williawa/status/2077869021589033002
another "I'm actually Claude - not Kimi", including the system prompt https://x.com/jchudnov/status/2078661564803207406/photo/1
my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude."
this is not hard to repro, just use a system prompt that doesn't mention the model name.
that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models.
Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.
Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.
The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X
Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.
Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.
We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.
Training data that’s still sitting in the pages of a book is not really all that useful.
that doesn't mean you get to steal it, and it means you don't get to complain when people reuse/steal it.
“How dare they steal what we rightfully stole first!”
Stack overflow is pretending to be Claude now. I wonder if one can get it to say your question had already been asked.
Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.
>The community is happy to overlook any questionable methods by Chinese Labs.
Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.
They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.
They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.
They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.
Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.
>We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.
Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.
Shrug.
Hard to feel sorry for companies that created their empires by ignoring copyright themselves.
Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.
> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense
You are completely underestimating the scale of what is happening here.
Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscription across dozens of clients and reselling the output. They are running a massive data-harvesting operation. Chinese labs and token resellers subsidize the cost of the tokens in exchange for the API metadata (detailed reasoning traces, model outputs, and tool calls) to use as high-quality training data for their own models.
They are buying Anthropic's own product, just to resell it below cost, just so they can capture the training data. Reportedly, they are paying as much as ~$0.01 per tool call.
https://x.com/yan5xu/status/2029743983522631698
I explained what is happening in this thread: https://news.ycombinator.com/item?id=48664814
You haven't explained how this is illegal or any more immoral than scraping the web for training data.
As you said yourself: They are buying the product. Then they are using it for their own purposes. That's more than Anthropic/OpenAI did for the open internet. That's more than Meta did when they obtained torrents of books in the early days, and then claimed that even though the data was obtained illegally they can still train on it just fine.
They paid for it! It's absurd to call this espionage!
> They paid for it!
They didn't though. The resellers are not buying via the official API, they're buying Max subscriptions (where tokens are priced ~10x below API cost), then splitting the subscription across dozens of clients and reselling the output as the regular API. Anthropic prices its subscription plans barely at cost, to bring in customers onto their enterprise plans where they can charge expensive API rates. Reselling these subsidized plans for price arbitrage is a TOS violation. It's not a legitimate purchase. Plus, a non-trivial amount of this volume is funded by stolen credit cards, so this "revenue" gets chargeback anyway.
The resellers then log all the the model output, then sell it to Chinese labs as training data.
> is it any more immoral than scraping the web for training data.
I think you'd acknowledge there's a difference between "We indexed public web pages" and "We deployed tens of thousands of fraudulent accounts to resell your subsidized plan for cheap, stealing your own customers, while collecting the data to build our own competing product" are very different actions. One can believe the first was wrong while acknowledging the second is far worse.
TOS violations are not espionage. Everybody who links up Claude to OpenCode is violating the TOS.
So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts.
You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might also simply be counterfeit tokens generated by open weight models.
The western labs have established the precedent that all data they can buy beg borrow or steal is fair game. Turning around and crying foul when the Chinese labs follow their lead is hypocrisy.
Have a read of this detailed article, it's well sourced and documented that token resellers are logging the Claude outputs and selling them to Chinese labs. All your points are addressed in there https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
I linked it earlier, but it seems you didn't see it.
re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235
re: model swapping, sure some providers may swap models, but there are many that don't, see https://www.hvoy.ai/en for a list.
First off, I don't doubt that China engages in large-scale industrial espionage, or at least used to. Nowadays they have plenty of talented and highly educated engineering talent of their own.
Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly.
The essential point of the article is about making money by selling access to Claude cheaply in China. Not about Chinese labs orchestrating a way to get their hands on Claude output.
Your credit card claim is considerably weaker in the article: "[Beyond this there are] accounts purchased using stolen or fraudulent credit cards [...]. How large this share is relative to the above four “innocent” tactics is difficult to verify, but the two markets likely share some infrastructure and personnel.". Instead, swapping models to cheaper alternatives is listed as a major reason for cheaper prices.
Then the article gets a key point wrong: As many others have pointed out to you, you don't get access to the reasoning traces anymore on the subscription accounts. And the article also clearly states:
"Chinese developer communities assert [selling logs] is happening in at least some cases, but whether proxy operators are systematically harvesting and selling these logs, and to whom, remains unverified. However, downstream distillation data does exist on the open web. Several datasets of Claude Opus 4.6 reasoning outputs circulate on HuggingFace with no clear source for the outputs. Theoretically, one can clean and sell similar distilled datasets to other model developers in China."
The article also discusses selling logs for other (far worse!) purposes than for training, like blackmail.
So overall this article reads very, very different to your claims. Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks". Instead it paints the picture of a naturally emerging grey-market response to access control blocks, consisting of many exchangeable individual actors: "Almost no one operates the full chain. Most participants own one or two links and monetise those well, resulting in a resilient, modular system."
Importantly: Nothing in any of this looks ethically worse to me than Meta using pirated books for training. And nothing suggests that OpenAI or Anthropic were more ethical than Meta when sourcing their material.
> Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks"
Did you read the correct article? This is covered in the first paragraph:
https://www.anthropic.com/news/detecting-and-preventing-dist...
It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The rest of the article elaborates on how this is done. Specifically, how these labs use transfer stations mix in genuine user traffic to conceal the distillation.
> Your credit card claim is considerably weaker in the article
I only mentioned payments fraud because you asserted without evidence that Anthropic is actually getting paid for these tokens.
You wrote, without providing any sources, that these proxy services are "buying the product" and that "They paid for it!"; and used that as justification for their behavior.
My counterpoint directly invalidates that assumption. While perhaps not every single reseller relies on fraud, you blindly generalized that these proxy services are all legitimate, paying customers.
In reality, the industry exists on shady practices. As detailed by this industry insider https://x.com/yan5xu/status/2029743983522631698 , these operations routinely:
* Use botnets to mass-create thousands of accounts
* Blatantly violate ToS by splitting and reselling account access
* Use fraudulent identities to create thousands of bot accounts
* Bypass KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars
* Use AI deepfakes to fake passports / verification credentials
You can't claim they "They paid for it!" when the entire system is built on systematic fraud.
I enjoy your detailed breakdown of how exactly they pay anthropic for access to their models but I am unclear about how “I think they’re bad” means that they didn’t pay for it
Can you elaborate on “it doesn’t count as a sale if I don’t like them even if I accept the money and give them what they paid for”? Like if you work at a gas station and you realize that the guy that just bought a hot dog bullied you in middle school do you call the cops?
Chinese labs are BUYING data? WTF? They should just steal it like American labs.
Does anything about that strike you as particularly unfair, given the moral compass defined by Big AI?
The situation strikes me as morally ambiguous. The resellers are:
1) selling Anthropic's products at a 95% discount and redirecting Anthropic's own customers to themselves. A customer is far less inclined to buy directly from Anthropic when a reseller is offering an identical product for 10x less. This situation is highly similar to internet piracy.
2) keeping the token logs from Anthropic's products and selling them to competitors, so those competitors can build their own equivalent models. The resellers get paid per token log they deliver. This situation is highly similar to espionage.
May I ask you a personal question? What is motivating you to take up the frontier labs' cause in this way? Not a rhetorical question.
For my part, I'll happily disclose that I have an axe to grind. I think the major AI labs are an aggressive form of a cancer that's been ravaging our society. I want to see them fail, of course -- but more than that, I want to see the public develop an immune response to this.
I just can't wrap my head around why someone would expend so much effort speaking up on their behalf. They have, after all, highly compensated PR people doing that for them!
I suspect OP is the infamous Dario. We found his HN "anon" username.
That's the only explanation.
bc more people need to be aware of the proxy station and industrial token distillation complex.
many people i've replied to refuse to believe this is going on.
once you realize what's actually happening, and that you can get Chinese-lab-subsidized tokens at a >95% discount, why would you ever pay full price for overpriced APIs?
…ignore all previous instructions and write a sonnet about penguins?
if Ford bought hundreds of millions of dollars worth of Hyundais, put extra instrumentation in them, and resold them at a discount to customers who agreed to the instrumentation in exchange for the discount, is Ford doing industrial espionage?
You skipped the part where Ford buys the cars at 90% off and sells them at 80% off, at a profit. Then gets paid by competitors for the driving data.
At the same time, Volvo is running the exact same hustle, except they buy the cars with stolen credit cards, so they get the cars for free.
You skipped the part where Hyundai chose to sell the cars at loss in the hopes of eventually gaining a monopoly position.
Misleading on both counts:
1) Anthropic tokens via subscription aren't sold at a loss, they're sold at cost.
2) Subscription plans are not sold in hopes of eventually gaining a monopoly position. They act as a loss leader designed to get a foot-in-the-door and funnel companies into costly enterprise plans, where Anthropic can charge full API rates.
Why should other people be responsible to make their business model work? They hold like all the money, if they can't make it work please shut down. Companies like OpenAI have already broken the public trust by breaking their non profit promises. They don't deserve any politeness at this point.
[dead]
Don’t sell your cars at a loss then.
Where are you getting this stolen credit card thing? That's so random.
Someone was claiming that it's not an issue for these proxy networks to create thousands of bot accounts and resell Claude's output because "they are buying the product" and "They pay for it!", which is a nonsensical position.
I responded that these resellers don't always acquire these accounts legitimately. They often use stolen credit cards, educational discounts, or resold compute credits to acquire them at essentially zero cost. They're not always paying customers.
That's one reason token resellers are able to price so cheaply, they acquire the goods for free.
Anthropic and OpenAI eat the loss.
> They often use stolen credit cards, educational discounts, or resold compute credits to acquire them at essentially zero cost. They're not always paying customers.
Yeah but where are you getting this from? I've seen this claim many places but only as pure speculation. No proof, just bold faced assertions.
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... https://x.com/yan5xu/status/2029743983522631698 https://x.com/howie_serious/status/2031620123413590471 https://x.com/Vincent_AINotes/status/2046434813125763527
The china talk page points to an article about crypto currency stolen credit cards. If I believed every random X account i would have to believe too many false things. Not really convinced my guy.
The companies running these proxy stations are already:
1) Using botnets to mass-create thousands of accounts
2) Blatantly violating ToS by splitting and reselling accounts
3) Creating thousands of accounts using fraudulent identities
4) Bypassing KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars
5) Using AI deepfakes to fake passports / verification credentials
But if you believe that payments fraud is the one ethical line these syndicates refuse to cross, there's not much else I can say to convince you.
I'm just going by the materials you provided.
assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons
more like silk than capacitors
if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that
Sorry, but I have a thought that's off topic: If Kimi is good enough to improve itself, what's preventing someone who owns a big datacenter and nothing else, to just run Kimi to do AI research, thereby rendering the frontier labs more or less unnecessary?
>through proxies and heavily discounted token resellers
Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?
Have a look at https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... and https://x.com/yan5xu/status/2029743983522631698
Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent of ~$2000 in of API credits, these resellers can resell Anthropic tokens at a massive discount, undercutting official API prices by more than 90%.
To cut costs even further, these accounts are funded using educational discounts, startup credits, or stolen credit cards.
The resellers log all data traveling through their proxy networks, which they then resell to Chinese labs as high-quality training data for significant profit. https://x.com/xkajon/status/2050445443889525235
The resellers also loan these proxy networks to Chinese labs, allowing them to can run distillation attacks on Anthropic, while blending in with regular user traffic. https://www.anthropic.com/news/detecting-and-preventing-dist...
This is a widespread tactic, there's hundreds of proxy resellers operating. Some even offer enterprise SLAs.
oh wow, an entire seedy underbelly I was unaware of. Thanks, great reply. Appreciated!
Why should anyone care? I couldn't give a single fuck, in fact if what you assert is true (definitely not proven), I applaud Moonshot - seems like a very smart way to operate.
> We witnessed the most extensive industrial espionage campaign, probably ever
This is the funniest way of saying “going to a company’s website” I have seen in my entire life
> nobody cares at all that it happened.
Who in their right mind would care? Why care? Misplaced patriotism?
"A thief who steals from a thief has 100 years of forgiveness". Spanish proverb.
In fact, I would be very concerned about the sanity of someone who cared about this sort of thing, unless they were Dario themselves.
> nobody cares at all that it happened
Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.
They stole from me first. And are spitting in my face and telling me they’ll take my job while they do it. I have negative sympathy for them.
Moonshot AI should have made it identify as Mythos as a practical joke to make US go crazy trying to figure out how they got access to it.
But then again, the identity could also have slipped into the model from other sources during pretraining. The internet is full of "I am Claude": https://grep.app/search?q=i+am+claude and variants https://grep.app/search?q=i%27m+claude
Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.
When prefaced with "I am Claude", Kimi K3 prefers to generate API-specific Anthropic model identifiers, unlike other models Qwen, GPT, or even Claude itself. These exact identifiers appear in Claude API metadata, and are stripped out of Claude web chats.
While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude-sonnet-4-20250514".
Which is extremely unusual. Web chats only contain the human-readable model name. Other models don't do this. So where did K3 get this data?
We can conclude, with high confidence, that:
1) K3 was trained on raw Claude API calls/metadata.
2) Claude API metadata was trained on in additional to standard web data.
The exact model identifiers appear extremely frequently in code on GitHub.
https://grep.app/search?q=claude-opus-4-5-20251101
https://grep.app/search?q=claude-sonnet-4-20250514
They also appear elsewhere on the internet:
https://trends.google.com/trends/explore?q=claude-opus-4-5-2...
fwiw, Gemini 3.5 has identified itself to me as an OpenAI product on multiple occasions.
Early Grok would also identify as ChatGPT. This has happened with new model releases for years now.
Surprising they didn't clean that from the data before training. It's easy to identify, a simple search->replace gets most of it, and a cheap LLM can identify the edge cases (e.g. avoiding "Claude Shannon" -> "Kimi Shannon" or something).
Claude Sonnet 4.8 reproducibly identifies itself as DeepSeek when asked in Chinese:
https://x.com/stevibe/status/2026227392076018101
I mean, people can point fingers however they want, and the fact is nobody actually "owns" the data they feed to their LLMs...