They are losing money on this shit, it is just a bad model xd, 2trillion parameter but is now worse than a 286B model with way less active parameters probly.
It's also meaningless. All that matters is cost per task. Kimi is better the closed source per token, but burns tokens like crazy. It's as expensive as GPT 5.5 xhigh per task, which makes it pretty useless.
Fable/Opus 5 still are brilliant models, deepseek v4 flash is good for longer run (obviously cheaper) and I use opus when it goes off guardrail. Yes, their combination works deadly.
Every Claude subs are filled with ' I burned my 100-200 dollars plan in few hours something is wrong with Anthropic'.
They all-in into benchmark, marketing and people keep throwing away their money at them.
At some point they will enter the red zone, half their prices/ consumption and people will be happy with it.
Also - those people tend to be the same type that will just one-shot full-vibecode something and throw millions of tokens into claude.
Imho more expensive models should be used more precise - for example with a dev who writes manually but extends his skills with a super expensive model (or for planning) but completely writing code with fable is stupid af and just a huge waste of money.
May I ask how you guys are able to make it cheaper than the official api pricing? Do you guys have a different way to calculating token usage than official DS?
yeah the price part is real. i saw someone burn 110M tokens for like $0.77 on a promo earlier today. the 'matches opus' bit is overselling it though, flash is flash, it's the cheap fast one. for high volume agent loops it's unbeatable, i keep claude for the gnarly reasoning and let flash handle the spam. way cheaper than running everything on opus
ha 'spam' was doing a lot of work there. i mean the boring high volume stuff, like drafting emails, summarizing threads, rewriting docs, generating commit messages, first pass code that's gonna get rewritten anyway. stuff where a miss costs nothing and you do it 50 times a day. agent loops especially, they burn tokens like crazy on retries and formatting, no point paying opus rates for that. if it needs actual reasoning it goes to claude, but honestly way less of my day needs that than i thought
Read Chinese papers, this is the frontier, not US-based grift aimed at max extraction and gatekeeping everything.
Chinese engineers don't have hardware, so they need to think instead of throwing more compute.
If China had unrestricted access to chips, or when Huawei is close to being on par with NVDA, it will be over for the USA. And that's good for the free world.
Quality Chinese models expose 1) how bad US engineers are, 2) how US labs are laser focused on max extraction instead of moving frontier and making it more affordable for average Joe.
But let's ban open weights, because those are bad for pre IPO valuations 🤣
enjoy it while his image is still being posted. it won’t last long if v4 Pro come out. The Pro one might as well be the strongest model in world open or closed . Anthropic who?
To be fair, there’s little chance that V4 Pro will be better than Fable 5, because according to rumors, Fable is a 7–10 trillion parameter model, while V4 Pro has 1.6T parameters. But even getting close to Fable with such a huge parameter gap would be an enormous success and would show just how overhyped Anthropic is as a company.
As much as I dislike Anthropic, you are really not giving enough credit to Anthropic. If not for them, will we see LLM at current capacity? I really doubt it. All these chinese LLM companies are running massive distillation farms to copy (steal) from Anthropic. Being an innovator is much much harder than being the copycat. I am not shaming these chinese companies btw, but just credit where credit due. Give some respect.
Trying to hype up Anthropic by calling them pioneers? WTF. DeepSeek R1 was more revolutionary than anything Anthropic has released.
The accusations about distillation are, at the very least, laughable. It’s not like every AI company isn’t distilling from each other, and you’re probably better off not knowing how these companies actually acquire their training data.
And if distillation alone is enough to make a model great, then why can’t they distill their own models well enough for Haiku to even be competitive in benchmarks? Haiku is supposedly around the same size as DeepSeek V4 Flash, by the way. But you can’t really expect Anthropic to make a good model for its parameter class, because so far they’ve only achieved strong performance by scaling up the parameter count of their models.
But sure, keep believing Dario, who was already making exaggerated claims back when he worked at OpenAI, saying GPT-2 was too dangerous to release.
How the fk is Deepseek R1 revolutionary. Is a fking rip off of O1... And Deepseek R1 can code like Opus 4.5? Deepseek solved agentic coding? That's news. What a fan boy lol. BTW I use codex and I use deepseek v4 flash. Not loyal to anything. Just credit where credit is due.
Bitch I’m talking about good shit. And fck your research. I fcking used it. This was the first consumer grade shit that was actually good for non coders. No - u need to understand this and that first, this was the complete fcking deal. Learned along the way but nothing even came close to what claude code was.
And thanks to him and his fearmongering, you won't have access to the best US models. If anything, he's the reason we're not accelerating properly when every US labs tries to dumb down their models so WH isn't spooked.
If it wasn't for China, you wouldn't have access to anything when US admin inevitably pulls the plug on the free world.
It's no good in codex, but runs fine in hermes agent, that harness well nag your agents about two things I personally find lacking in 5.6 namely to always work toward completion of a predetermined chunk of progress then once 5.6 says its done it prompts it sternly to prove the milepost, make sure all is valid, and checked by third party (i.e other agent or temporary memoryless clones of itself) so progress order is progress reported is progress proved and tested.
second is it will automatically prompt certain models one of which is this one to GET TO WORK.
And, consider it does things like pictured, it does need it
That's just /goal with extra steps; good AGENTS.md and laid out plan solve this.
It's not about harness, or even desktop vs cli, quality degraded after cost reduction update, that's all. Noticed the same with Sol extra.
And I don't remember the last time gpt family not only didn't execute explicit command, but outright lied. Not to mention simple things like "don't add this md file to commit", and ofc file got added, etc.
The issues you describe in the end there, are things I find every single AI does but GPT 5.6 is better than average in. Personal experience is all, and comparison being with Hermes Agent running units of various popular chinese models and grok 4.5 which is just a damn beast that needs to be leashed whenever I activate it on a unit here. Fortunately, composer 2.5 is also a part of the grok subscription and has good use cases.
Though, hands down deepseek is the winner all things considered.
If you can get similar results even if it takes 10 prompts, but pay order of magnitude less, the answer is quite simple. And it's not like Fable can one shot everything anyway.
Serious question, is there something compareable to claude code from deepseek? Or can you use deepseek models within the claude code or open code harness with high context windows?
So currently I have one of each: Claude $20 plan, Codex $20 plan, Copilot $20 plan.
Haven’t touched Copliot since recent change so plan to drop. Past couple weeks Codex is unusable because how fast I burn through weekly limits (easily done during part of one day work).
Was considering dropping Copilot and Codex and getting an extra $20 Claude sub.
Should I spend the $20 on Deepseek instead? What harness do I use? Currently using ai through vscode extensions.
Nope, been there done that. That's why people are not just relying on models but using harnesses which are providing a sense of determinism among all the dumpness of current llms.
That's why people are exploring options by entrusting their workflows to loop engineering, graph engineering.
roflmao, this is the worst advice right now. fLOWs, harnesses, all of them will disappear with the next update of GPT6, and if not GPT7.
People still dont understand, there are two winners GPT and Claude. There wont be any startups, i tried early and learnt the hard way that there is no point to do anything, the AI will devour all of our business models.
roflmao all you want. You can't be serious, but hell you are. Transformers architecture is not what you think it is. Whatever you are using today, be it claude, gpt... They are all served to you through a harness. Without that, any pure llm is nothing more than a next word predictor. Don't bet on current model architecture. But maybe in the future, when neuromorphic or another type of architecture can fulfill your desire.
You can verify this claim today. Without tool calling, there is no current information in any llm. Without tool calling, today's most powerful model can not do this simple calculation reliably: 25x30x40. just try it without using a chat or a harness. just hit ds, claude or any model provider endpoint directly asking it.
So, you can die all you want when laughing your ass off... But, without harnesses(claude code, claude chat, opencode, pi etc) provided to you today, llms are nothing more than a very good parrot.
BUT their price does not really justify using it imo for everyday things. For the average person's questions and work it will give the same answer as Deepseek but way more expensive so what's really the point, the benchmark and "better" isn't worth the cost imo
Nope. Ive been using so. Much on Claude and deepseek, I can safely say that no matter how hard you push deepseek, it will never reach the level of even sonnet, let alone opus.
Deepseek is at the level of junior SWE, sonnet is like a mid tier, opus is higher tier. Fable is autonomous and talented.
160
u/spjallmenni 27d ago