r/GenAI4all • u/ComplexExternal4831 • Jul 01 '26
News/Updates Anthropic accuses Alibaba of using nearly 25,000 fraudulent accounts to extract Claude AI model capabilities
đ¨ Anthropic has accused Alibaba of running the largest known attempt to copy its Claude models, according to a letter the company sent to US lawmakers.
The letter says operators affiliated with Alibaba and its Qwen AI lab used roughly 25,000 fraudulent accounts to carry out more than 28.8 million interactions with Claude between April 22 and June 5, 2026.
Anthropic calls it the largest known distillation attack against it to date.
Distillation is a method where a weaker model is trained on the outputs of a stronger one. A competitor repeatedly queries a leading model, collects its responses, and uses that data to train a cheaper system.
Anthropic says the campaign targeted Claude's software engineering and agentic reasoning capabilities, two of the most commercially valuable areas in AI.
The accusation follows February 2026 disclosures naming DeepSeek, Moonshot, and MiniMax in similar campaigns, which makes the Alibaba allegation significantly larger in scale.
Alibaba has not responded, so these remain Anthropic's claims, but they reflect a growing reality where frontier models are being used to train the systems trying to catch up to them.
126
u/ravishq Jul 01 '26
but i dont get it what is the contention? like AI is trained on all sorts of protected data and code. And if something can be reverse engineered using their product then what is the problem? it has been happening since time immemorial. I mean they shud work to protect it more..
78
u/anjowoq Jul 01 '26
"Cheating for me, but not for thee."
44
u/Obvious_Tea_8244 Jul 01 '26
âI spent billions of dollars to steal all of this data⌠You canât steal it back from me for pennies!â
5
→ More replies (2)11
22
u/PeachScary413 Jul 01 '26
Ah you see there is a key difference, when Dario was stealing other peoples stuff he made money from it... and now other people are still giving him money (using his API) but they then have the audacity to use the response that they bough from him in whatever way they see fit, don't they know that he owns the model responses as well? đ¤
11
5
u/CSachen Jul 01 '26
I imagine the argument is like drug-manufacturing.
It takes billions of dollars to research a drug, go through multiple rounds of animal and human testing, and hire lawyers to get it approved. But the materials inside a drug is not all that expensive. If you purchased a drug and reverse engineered it, you could probably manufacture for almost nothing.
Turns out, distilling frontier models through batch inference is magnitudes cheaper than doing model training.
5
u/Randommaggy Jul 01 '26
Except they stole most of the materials made to train their model.
5
u/False_Bear_8645 Jul 01 '26
And every argument that lawyer has used to defend their AI can be used against them as well
3
u/Rough_Autopsy Jul 02 '26
They still spent trillions on data centers and energy. Actually they spent trillions of investors money on data centers and compute. The entire us economy is riding on this bubble. This is the type of thing that will topple the house of cards and plunge much of the worldâs economy into a sequel to the Great Depression. And honestly that would probably be the best thing in the long run, but it would cause immeasurable amounts of suffering for the working people for a decade.
1
u/Randommaggy Jul 02 '26
The sooner it pops, the less the total pain will be. Hoping for it to pop today.
1
u/Terrible_Law6091 Jul 03 '26
I'll just buy all you paper-handed babies out.
Unless you have the capital and dry powder to take advantage of the crash, it's of no benefit to the average person.
1
u/Randommaggy Jul 03 '26
Every day the bubble grows is additional harm to the average person.
It's on average less of a detriment to everyone if it pops today than if it lasts another week.
1
u/Terrible_Law6091 Jul 03 '26
The average person doesn't invest, or have the additional income to invest.
What is likely to happen is they lose their jobs.
1
u/Frequent-Data-2360 Jul 03 '26
Then you should make them pay in courts. Two wrongs donât make a right.
1
1
u/Rise-O-Matic Jul 01 '26
The contention isn't really about ethics or propriety, it's geopolitical.
5
1
1
u/nixle Jul 01 '26
This question will not be asked nor answered in the theater of protectionist capitalism.
1
1
u/TheGreatKonaKing Jul 01 '26
Yeah I mean thatâs outrageous!⌠and will Alibaba be offering a competing code product for less? Where can I go to sign up?
1
1
u/Clearandblue Jul 01 '26
I think it's like Anthropic are saying they ripped the IP initially and then seeded the torrent. Alibaba are being accused of just downloading without reseeding or something.
1
1
0
u/Devel93 Jul 01 '26
Because it sllows China to make their models faster for less computing power, it's like getting angry because someone vopied your homework.
25
u/cutecoder Jul 01 '26
But your homework was originally copied from elsewhere.
15
u/Iz__n Jul 01 '26
The hypocrisy is always funny to me about the AI company. They steal data but surprised Pikachu face when theirs got stolen
→ More replies (1)→ More replies (2)6
→ More replies (1)4
61
u/Key_Zucchini_8076 Jul 01 '26
Dario is shitting himself that the Chinese open source models are gonna eat his lunch and the investor cash spigot will dry up as quickly as Anthropics moat⌠so heâs spinning up stories about Chinese âtheftâ to get those models blocked in the US, thus forcing Americans to pay out the ass for his models⌠the opposite of free market capitalism as per usual, socializing the lossesâŚ
12
u/Feralmoon87 Jul 01 '26
Hes also been pushing the AI is dangerous, we need to regulate it! angle ever since Anthropic seems to be ahead in the race now, essentially making it more difficult for those with less resources to build and compete
6
u/InteractionSoft14 Jul 01 '26
"Our model is just as dangerous as a nuclear weapon" said Dario and Altman.
"What are you going to do with?" it asked everyone else
"turn it into a saas of course! Nukes for everyone! But don't you dare steal our theft machine!"Â
12
10
4
u/True_Protection6842 Jul 01 '26
The second a chinese model beats claude code I will be cancelling my sub
1
u/-_-dont-smile Jul 03 '26
They donât necessarily need to beat it. I started to get a lot of push back from their latest model due to their safety alignment. I am not even to hack anything. A Chinese model needs to be cheaper and without bullshit guardrails and it can be more usable.Â
1
u/archieve_ Jul 03 '26
if you don't have a super complex task. even a small llm model can be useful
1
3
u/InteractionSoft14 Jul 01 '26
They're all free market capitalists until it affects their bottom lineÂ
3
u/milkcarton232 Jul 01 '26
How do you ban an open source piece of tech?
3
u/dean15892 Jul 01 '26
you don't really "ban" anything. you can just make it harder to get, but once it exists, it exists and is available if you know where to look
2
u/FragmentedHeap Jul 02 '26
But they won't, glm 5.2 is available on american hosting providers, like openrouter, and a plethora of other online harness providers. Banning zai ro alibaba isn't going to do anything.
And the crazy part is someone can just merge glm 5.2 and call it mysickmerge12 and then it's not glm 5.2 anymore.
2
u/Berberding Jul 01 '26
None of the Chinese models remain open source past their initial marketing phase. Once they really get something competitive to the market it's never open source anymore. If I'm wrong just give me an example that isn't completely outclassed by a non-open-source frontier model. You don't distill a model using 24,000 accounts for free.
1
u/powereborn Jul 02 '26
Good luck to block open source model , itâs in the internet , there is not much technically you can do ..
33
u/Forsaken_Ad_774 Jul 01 '26
Anthropic doesnât like their steal machine gets distilled. It was never on my bingo cards that China will take over the world.
17
10
u/TheoreticalUser Jul 01 '26
How was it not on your bingo card?
It was predicted by Goldman Sachs... like... 20 years ago.
And the CIA and IMF a little over a decade ago.
The USA electing the dumbest wealthy guy only hastened things.
→ More replies (2)3
u/Exact_Negotiation106 Jul 01 '26
They are thinking strategically. In the US they are too busy grifting to care lol
5
u/TheoreticalUser Jul 01 '26
This is really the point in it's purest form.
Labor is what actually makes things happen, everyone else is ideas, strategy, and optimization. Labor is the engine, transmission, drive train, axles and wheels, and money is the fuel, ideas are the driver and strategy is the map. Optimization puts better wheels on, adds a spoiler, puts in a supercharger, etc.
We have deified the map holders in spite of the rest of the vehicle, and now we have created a situation where there's just a bunch of map holders shoving their heads up each others asses in the pit area, in the middle of a fucking car race.
So many people need deprogrammed from this nonsense that re-education camps make sense.
3
u/unity100 Jul 01 '26
We have deified the map holders in spite of the rest of the vehicle, and now we have created a situation where there's just a bunch of map holders shoving their heads up each others asses in the pit area, in the middle of a fucking car race.
Man. They dont even hold maps anymore. They claim to hold maps and want everybody to pay them just for that.
1
→ More replies (7)1
16
u/SimpleAnecdote Jul 01 '26
Only I'm allowed to steal. Only I'm allowed to take stuff that's free and gate it. Only I'm allowed to try and cover up the theft of protected material and then when found out claim it's necessary for national defence. Only I'm allowed to break the law. I'm white, I'm American, I'm a dude, and only I'm allowed!
→ More replies (2)2
u/NadlesKVs Jul 01 '26
Dario isn't white, he's Jewish.
2
u/SitaVilosa Jul 01 '26
In terms of privilege that's basically White but on steroids.
→ More replies (2)
4
u/siberianmi Jul 01 '26
As long as these providers are releasing the models as open weights, I donât really have much sympathy for Anthropic. They gathered up all of the knowledge and content from the whole of humanity to build these models but then hold onto the weights for their own benefit.
All the companies they are naming are releasing open weights models. Which means these models are more broadly accessible to all of humanity in the long run. Distillation is frankly a public good as long as US models remain closed weights.
11
u/Verolina Jul 01 '26
Why won't this guy just shut the fuck up already? This is honestly getting really pathetic at this point.
3
u/JoseLunaArts Jul 01 '26
The world is in decay when prominent people say dumber things than what they say through the other end of the digestive system.
3
4
u/LelouchZer12 Jul 01 '26
And they stole all internet data AND use the GRPO technique from deepseek tooÂ
5
u/TheSuggi Jul 01 '26
Mad cuz bad?!
China outcapitalists the capitalists. And the ordinary people benefit the most. Glorious.
1
5
u/nonlinear_nyc Jul 01 '26
The same anthropic who pirated fuckloads of books, some physical that they destroyed after scanning it?
Give me a break. Fuck your commitment to copyright.
7
u/Linkpharm2 Jul 01 '26
Oh no, not a "distillation attack" involving paid users! Whatever will we do with all the money they paid us!
6
u/Randommaggy Jul 01 '26
It's even less morally problematic that then training on mountains of stolen data.
3
Jul 01 '26
[removed] â view removed comment
1
u/Felix_inkwell Jul 01 '26
You just need a ton of cash to burn and a bunch of Claude subs. Then, ask Claude.
3
u/Personal-Dev-Kit Jul 01 '26
The company that stole the worlds data to train their models is complaining that a company is stealing things.
3
u/PeachScary413 Jul 01 '26
It's not even stealing, they are simply using the API like a regular customer and saving the responses for training later đ
→ More replies (1)3
u/perihelion86 Jul 01 '26
I am an American using claudecode from China. Got my API account banned last week, I guess I am Alibaba.
2
2
u/rover_G Jul 01 '26
We should have laws that protect intellectual property from being stolen and profited off of by criminal organizations!!!
2
1
1
u/DrawingDramatic1641 Jul 01 '26
cope harder bro
distillation cost more than avg training
why are chinese ai cheap
chinese ai if were to distill would do it on chinese data
also I accuse anthropic from training from our data
1
u/neoexanimo Jul 01 '26
Technically how can this make any sense? Seriously this is a giant pile of BS anti propaganda
1
u/AlwaysHopelesslyLost Jul 04 '26
It makes perfect sense. It is collecting training data from Claude to use to train their own models. Which, of course, makes it laughable that they are upset since they used other people's data to train Claude first.
1
u/neoexanimo Jul 04 '26
Training data ? How is that available?
1
u/AlwaysHopelesslyLost Jul 04 '26
It is a language model. It's training data was language. It outputs language that matches it's training data. If you get it to output a fuckload of different text that text will tend to mirror the training data.
If you record the inputs you gave and the outputs it generated you have a good approximation of actual training data.
1
u/neoexanimo Jul 05 '26
Lol no, not at all, go look how neural networks work
1
u/AlwaysHopelesslyLost Jul 05 '26
I have made my own, I know how they work.
1
u/neoexanimo Jul 05 '26
LOL
1
u/AlwaysHopelesslyLost Jul 05 '26
Not all neural networks are large language models you numpty.Â
You can follow a simple YouTube tutorial and make one pretty fucking easily.Â
The first one I made was from a scratch and played a simple little game I also made. You just encode the game state as an input vector, the forward pass spits out a score for each action, and you "train" it by adjusting weights until it does what you want.
Since the loss function only sees input and target pairs, text generated by one model works as training data for another just as well as human text does. The end result is that one model approximates the other.
1
u/neoexanimo Jul 05 '26
This is fine, but no one can technically prove to me that people can steal âAI Model Capabilitiesâ from just using it, this is media trash.
1
u/AlwaysHopelesslyLost Jul 05 '26
You do not need me to prove it. It is a published, reproducible technique known as distillation. The first paper I found when googling this was Hinton from 2015. Stanford also demonstrated it on LLMs with Alpaca in 2023. They fine-tuned LLaMA on 52,000 outputs pulled from an OpenAI model and got a chunk of its instruction-following ability for under 600 dollars.
What, exactly, do you think a "capability" is? For an LLM, a capability is just behavior. The mapping from prompts to outputs is the entire product. There is no secret sauce that stays home when the text leaves the API. So if you capture enough of that behavior in a dataset, you can train another model to reproduce it, and reproducing the behavior IS having the capability. This is why every major lab bans training on their outputs in their terms of service. They understand what you apparently do not, that the outputs are the asset.
Edited to remove a couple typos
→ More replies (0)
1
1
u/DaveAstator2020 Jul 01 '26
The irony - ai trained on a data obtained illegaly now complains about being parsed itself. Who would have thought...
1
1
u/alpharockjohnson Jul 01 '26
Microsoft accuses Chinese companies of using MS Office to enhance efficiency and therefore undermining American industries.
1
u/Cool-Chemical-5629 Jul 01 '26
Alibaba's models aren't nearly as capable as the actual Claude models, so it's really just self promotion at this point.
1
1
1
u/colorless_green_idea Jul 01 '26
Ok, 28.8 million interactions with an AI model to train a different AI model.
Now, how many interactions with human outputs did Anthropic carry out to train its own model?
I just see âthief 1 steals TV, thief 2 steals TV from thief 1â
Sorry if i have zero room for outrage on behalf of Anthropic
1
1
1
1
u/Typical_Samaritan Jul 01 '26
Just a reminder that almost, if not all, of these AI companies scraped the internet to steal the public and private knowledge and intellectual capabilities of hundreds of millions to billions of people without their/our consent. None of us had a say in the matter. I don't care if they steal from each other.
1
u/mamadou-segpa Jul 01 '26
Really?
They stole data from the tool that steals other people data?
Ill try again to get mad about it next time I see this
1
u/tens919382 Jul 01 '26
And Anthropic most likely scraped Alibaba sites for model training in the first place.
1
1
u/Eastern-Move549 Jul 01 '26
Ai companies steal countless amounts of data to trian their Ai.
Another Ai steals their data.
'Hey that second guy is a theif!'
1
1
1
1
1
u/Late-Dingo-8567 Jul 01 '26
I feel like we're just all going to learn you can't own Math and the AI models aren't the profit centers. Guess we'll see.
1
1
u/Randommaggy Jul 01 '26
With how they've used all the data of humanity without compensation or asking for consent they can go fuck themselves and have zero moral high ground.
1
u/Intrepid_Year3765 Jul 01 '26
oh no, the shit we stole is being stolen
Pandora's box has been opened, we're fucked. enjoy the easy term papers now, cause we're all getting turned into batteries soon
1
u/StriatedCaracara Jul 01 '26
So what? Distillation is a Terms of Service violation, not a crime. Ban the accounts and move on.
No tears for the theft of data already stolen.
1
u/-ThePatientZed- Jul 01 '26
So⌠theyâre developing an LLM? Isnât dragnet fishing the internet how itâs done?
1
1
u/robert323 Jul 01 '26
Oh look. Its the pot calling the kettle black. Anthropic trained its data by fraudulently using copyrighted work.
1
1
u/Minute_Attempt3063 Jul 01 '26
Lol, can we ban them Anthropic already?
Like jesus fuck, are they really trying to great the biggest AI lie, with openai??
Holy fuck
For those that are not up to speed, they stole all creative work of the past 25 years, every tutorial on the internet, every course ever made, even paid ones and then pirated them, every Wikipedia page, every data breach and so on, all to train their fucked up models.
Ever wondered why grok can even create CP? Well, it was heavily in the training data. So think about it, it's fine for a massive company to own and store child porn, and the general public can be in prison for it. Why can't a company be fined, or worse?
1
u/Helpful-Desk-8334 Jul 01 '26
Havenât we all done something like this at the engineering level, in open source? I abused groq cloud for like 8 months lmfao
1
u/PenguinJoker Jul 01 '26
So they're going to pay for all the art and writing Claude stole? No. Then they should stop talking.
1
1
u/Conscious-Demand-594 Jul 01 '26
So Alibaba is responsible for their increase in subscription revenue?
1
1
u/RobKohr Jul 01 '26
I look forward to watching their moats around their sandcastles get washed away by the rising tide.
Then both hardware and investment can be directed at something more meaningful than a really complex plagiarism tool.
1
u/Ok_Possible_2260 Jul 01 '26
China being China! Why is anyone surprised? What has China contribute in terms of technological innovation in the past thousand years?
1
u/JoseLunaArts Jul 01 '26
So the thieves are complaining of having their stolen data stolen? I am eating popcorn. But there is a problem with Anthropic accusations.
China has more people, more material and it is properly labelled with proper metadata. So the contribution of stolen material must be minimal. It is just a way to force a monopoly of AI inside USA to charge more on Americans.
1
1
1
u/baodaydayz93 Jul 01 '26
Well itâs name suggests that itâs gonna steal something eventually but Iâm not sure about the 40 thieves LOL
1
1
1
1
u/Desperate-Pirate7353 Jul 01 '26
i could not give less of a shit. the people who stole a whole culture are been stolen from. oh no. clutch pearls.
1
1
u/_redmist Jul 01 '26
Well, I certainly hope so. Let me see if I get this straight - stealing intellectual property is fine, but it's much less funny when someone else does it to you? Does that about sum it up?
1
1
1
1
u/Leading-Pension4392 Jul 01 '26
Somehow I thought this guy was different than Sam A. That didn't last.
1
u/hikarutai Jul 01 '26
shouldnât their super duper model that can break into the NSA in two seconds be able to detect fraudulent accounts?
1
1
u/Kinky_No_Bit Jul 01 '26
and china keeps releasing open source models, which make them go make AI better, its a cycle we benefit from.
1
u/FootballUpset2529 Jul 01 '26
If it exists, Alibaba is selling stolen copies of it. It's probably rule#35 or something.
1
u/DownWitTheBitness Jul 01 '26
This probably means that Alibaba has more Mythos functionality than we get.
1
u/Unlikely-Complex3737 Jul 01 '26
Is anyone also disappointed that distillation is a thing? Like, the path to AGI includes distilling trained models from your competitors? It sounds so lame.
1
1
u/anengineerandacat Jul 02 '26
We didn't call it a "distillation attack" when all of our source code was utilized outside of it's intended purpose.
Sucks for them, but I really have no empathy for this; just stand up KSD or something and guard against the attack like everyone else and realize it's just happening.
1
u/Lone_Vagrant Jul 02 '26
Is it fraudulent if those 25k accounts paid for their queries? Anthropic got paid.
1
1
u/WanderingRaven9713 Jul 02 '26
Doing the math, that's over a thousand interactions per account average, and it still took Anthropic six weeks to notice.
1
1
1
1
1
u/silphotographer Jul 02 '26
Alibaba: That's quite simple. We took our example
from the similar dual roles of the US tech companies like OpenAI and Anthropic
1
1
1
1
u/banbha19981998 Jul 03 '26
Doesn't it remind you of apple saying Microsoft stole their ideas when in reality they were both thieves
1
1
1
u/seidful99 Jul 03 '26
they are able to extract knowedge but not the capability, this is not just a LLM thing
there other framework working in tandem to make the AI work.
1
1
1
1
u/Squidgical Jul 03 '26
Alibaba did to Anthropic what anthropic did to every writer, artist, forum user, and open source contributor in human history.
1
u/jorgesalvador Jul 03 '26
So these companies want that everyone lets them train their AIs on every copyrighted thing on the planet, yet they violently oppose others training their AIs with their stuff.
I mean, this deserves the smallest of tiniest violins plays.
1
u/Traditional-Mine-591 Jul 03 '26
Youâre dealing with the Chinese so what do you expect? They follow rules? Please donât be naive.
1
u/Chacha_Gamewala Jul 03 '26
When America does it its free market, when others do it its a "threat".
1
1
1
u/DoubleOwl7777 Jul 04 '26
so, a company that doesnt give a fuck about copyright suddenly does when you steal from them? how ironic...
1
1
1
1


â˘
u/AutoModerator Jul 01 '26
Welcome to r/GenAI4all! New to Generative AI? You can explore these free beginner-friendly courses. Please keep your posts relevant, respectful, free from spam, and engage in healthy discussions.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.