Rendered at 23:50:25 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
vishvananda 3 hours ago [-]
The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
bionhoward 2 minutes ago [-]
opencode DeepSeek v4.1 Flash isn’t us/eu hosted as of recently, so not sure how this impacts privacy / model training
pimeys 2 hours ago [-]
Yes. You have to find the provider with pricing that suits your usage.
I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.
If your tasks are write-heavy, find a provider with cheaper output.
If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.
sheeshkebab 1 hours ago [-]
What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped
maxdo 31 minutes ago [-]
You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .
I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
georgel 2 hours ago [-]
I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).
And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
whstl 1 hours ago [-]
It's wild how different usage patterns are between users.
I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
Quothling 20 minutes ago [-]
Now that Microsoft allows you to do /cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
lifeisloving 2 hours ago [-]
Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
eikenberry 1 hours ago [-]
But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?
kevin42 44 minutes ago [-]
Compared to what a lot of companies spend on software for chip and electronics design (we're talking about $10k-200k/seat per year), AI coding assistants have a long way to go in cost before companies won't be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.
For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.
If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.
redanddead 13 minutes ago [-]
Nah dude I think it’s worse than that
Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness
Show me people handwriting code à la NASA
FromTheFirstIn 1 hours ago [-]
No one involved in this is thinking about the long term
irjustin 1 hours ago [-]
> But aren't you developing bad habits and learning patterns that won't work long term?
2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
jaggederest 1 hours ago [-]
I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.
In the same way that only supercomputers used to have multiple processors and caches but it's now standard.
denkmoon 1 hours ago [-]
I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.
pmontra 1 hours ago [-]
Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.
lionkor 2 hours ago [-]
How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.
onlyrealcuzzo 2 hours ago [-]
OpenRouter is complete garbage.
Buy directly from DeepSeek's API.
You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
georgel 2 hours ago [-]
I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:
OpenRouter Pricing:
$0.02/M input tokens $0.60/M output tokens
DeepSeek Pricing (cache miss, off-peak):
$0.15/M Input $0.60/m output
girvo 2 hours ago [-]
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.
It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
georgel 2 hours ago [-]
Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.
girvo 33 minutes ago [-]
It will of course depend on what you’re doing with it, but right now my session at work has a 99.8% cache hit rate, and I’ve been running this session for hours with 23M tokens read and 713K tokens written (Opus 5.5 in this case though)
mswphd 2 hours ago [-]
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
ckdot 1 hours ago [-]
There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
patwolf 57 minutes ago [-]
One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek's own API doesn't support ZDR.
selectodude 48 minutes ago [-]
DeepInfra does and it's the same price. That's what I use.
koe123 31 minutes ago [-]
Youre doing something so special you need that?
rrr_oh_man 15 minutes ago [-]
I've got a toilet cam to install in your bathroom
BeetleB 57 minutes ago [-]
DeepSeek trains on your inputs. That's why people go on OpenRouter and choose ZDR providers.
bethekidyouwant 26 minutes ago [-]
Let me get this straight you guys really like deep seek because it’s open but you don’t wanna help them improve.
bottled_poe 17 minutes ago [-]
No, we just want a choice on how to license our work.
jorvi 52 minutes ago [-]
Or.. BYOK Deepseek because OpenRouter's UX is much nicer?
runtime_terror 29 minutes ago [-]
Zero Data Retention and not having company source code leak to "CHINA!" (said in Trumps annoying voice) would be two reasons not to
seunosewa 2 hours ago [-]
Which provider was that?
mensetmanusman 2 hours ago [-]
Also DeepSeek usage is subsidized as well, it’s a power hungry model.
jchw 2 hours ago [-]
Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
arjie 1 hours ago [-]
No, it’s outrageously profitable above x% utilization without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.
worldsavior 2 hours ago [-]
You're the RLHF.
edflsafoiewq 2 hours ago [-]
Not if you're not giving feedback.
jchw 2 hours ago [-]
With ZDR-only enabled? That seems illegal.
2 hours ago [-]
jansan 2 hours ago [-]
Don't ask and dance as long as the music keeps playing.
kennywinker 2 hours ago [-]
Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
thesnarkitecht 2 hours ago [-]
Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
0xbadcafebee 1 hours ago [-]
You should basically never pay API prices, they are always several times higher than subscriptions.
There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.
giancarlostoro 4 hours ago [-]
Call me crazy but:
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
kristopolous 4 hours ago [-]
Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.
Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.
There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.
If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.
The market is legally locked down and we're in hostage pricing mode.
And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...
Nobody is coming to save us. That's our job.
phil21 4 hours ago [-]
> do they plan to increase production? No.
Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.
Plus expanding other existing facilities.
These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.
Samsung and HK Hynix also have fabs under construction and planned.
CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.
Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.
> The Micron CEO just recently said this is the exact plan
CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
> CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
It's taken them this long to catch up to the DDR5 standard. They've only recently been through qualifications to be a DDR5 supplier for the big boys.
> Every Major Motherboard Maker Now Validates CXMT DDR5
After their recent IPO, they have more than enough cash to ramp up in a major way.
It's just a matter of time.
cogman10 3 hours ago [-]
The second Micron boise fab hasn't even broken ground yet, they are still working on the first one. So don't expect these things to be completed in parallel.
Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.
dboreham 3 hours ago [-]
Anyone who has been around the semiconductor industry since the last century will remember various huge fabs e.g. in Arizona that were partially built but never finished due to oversupply by the time the walls and roof were done.
BizarroLand 3 hours ago [-]
Yeah, but why would they make consumer memory when HBM for GPUs is much more profitable?
kristopolous 3 hours ago [-]
Capitalism eats itself this way. Second and third order effects will collapse the demand.
You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.
I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.
It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.
Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.
But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.
But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.
This has happened in electronics markets before. Many times.
When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.
usefulcat 2 hours ago [-]
> Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
If you were a DRAM manufacturer, isn't this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?
Analemma_ 3 hours ago [-]
What "second- and third-order effects" do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can. I don't see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
kristopolous 2 hours ago [-]
This is all hypothetical and debating hypotheticals isn't productive so let's roll back to markets.
Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.
You have a very expensive data center and you're in debt financed on the premise that you have these special computers.
Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.
This stuff happens all the time. It's why we don't use BMP files on websites or serve giant MOV files on YouTube. It's why postgres queries are faster now than they were 10 and 20 years ago.
You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.
Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.
On market if you were to sell some of that ram you have 100% profit right now but not for long.
Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...
We live in the real world so what do you do?
Historically the answer has been "sell that shit"
There's an aphorism for this "stairs on the way up elevator on the way down"
If we had a healthy market with sane prices where you can't flip the thing you bought for 100% profit the answer would be "create more value."
charcircuit 2 hours ago [-]
>Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.
Or they could stay at $10,000 per month since they are willing to pay that much already.m, so they just use AI more and in more places.
lxgr 2 hours ago [-]
> locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can
Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).
Analemma_ 2 hours ago [-]
I think locally-hosted models at the org level will definitely be somewhat popular, but you seem to be talking about decentralizing for people's personal, non-business use, and I just don't think that's going to happen to any real degree.
People who say they want local runs really mean it: they want local runs on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I'm not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.
hn_acc1 2 hours ago [-]
>Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.
How many people, outside of tech geeks and megacorps care about RAM prices? And how gullible would you be to BELIEVE the politician they could actually make it happen, and even if they did, that it would extend to the average person, and not JUST megacorps/megadonors?
Aerroon 27 minutes ago [-]
Phone companies have been differentiating their models based on RAM for a decade. As have laptop and desktop sellers. The reason your router sometimes randomly crashes could very well be a result of not enough memory. The reason it takes such a long time to launch some programs repeatedly is because you don't have enough memory to cache it. Swapped from your browser to an app on your phone, but when you go back to the browser the site has reset and you lost everything you were working on? Not enough memory. Etc.
I think a lot of people care about the downstream effects of memory prices, but I agree with you that they may not realize that they happen because of memory prices.
ashdksnndck 3 hours ago [-]
RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.
dboreham 1 hours ago [-]
Don't believe Nvidia has any fabs of its own.
m463 4 hours ago [-]
> "i'll bring down ram prices"
wonder what voting would be like?
gamer vote ++
datacenter hater vote --
datacenter lobby ++
micron lobby --
rezonant 3 hours ago [-]
Yep, that's all the voting blocs.
mwambua 4 hours ago [-]
Wouldn’t cheaper memory make it easier to bring compute out of data centers and onto consumer hardware?
TeMPOraL 3 hours ago [-]
Datacenter haters will read this as "that's still evil AI", and everyone else hopefully can count and understands it'll be worse for environment.
1 hours ago [-]
bob1029 3 hours ago [-]
If we take some time to understand how HBM memory is manufactured (with particular focus on yield risk for final packaging steps), we will hopefully learn that the current capacity crisis is not bullshit.
I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.
neya 3 hours ago [-]
> they could then shoot a puppy and call me a slur
I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.
fhn 2 hours ago [-]
How many people would you allow them to kill?
xyzsparetimexyz 3 hours ago [-]
Neither political party cares at all about memory pieces get real lol
Aerroon 1 hours ago [-]
I don't really understand why. Memory is a critical component of every computational device.
Exoristos 53 minutes ago [-]
That's tautological, but I think you would need to explain how it extends their and their backers' influence to get party attention.
kristopolous 3 hours ago [-]
wait until holiday shopping...it affects the price of almost everything with a battery or power cord.
xyzsparetimexyz 2 hours ago [-]
What do you think they'll do? Neither repubs nor dems will touch ai companies in a meaningful way. Anything China does wrt memory fabs week be more significant
3 hours ago [-]
gchamonlive 4 hours ago [-]
[flagged]
Analemma_ 4 hours ago [-]
I don't think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those always have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.
Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.
boustrophedon 2 hours ago [-]
The RAM vendors have formed cartels previously and been convicted, so although demand is exceeding supply it is not that crazy to at least consider.
3 hours ago [-]
kristopolous 3 hours ago [-]
The AI boom started in 2022. Prices rose THREE years later after 2025 sanction and tariff style legislation to protect the market during a price hike.
I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.
Connect the dots.
Analemma_ 3 hours ago [-]
Man I think you're just spewing word salad and a lot of what you've written is either wrong or not even wrong. The AI hype really got started in 2022, but hype on social media doesn't mean anything for RAM prices, only real buildouts do that. They rose pretty steadily until OpenAI revealed their shenanigans re: locking up a ton of supply from two different vendors with secret contracts, and that's when the takeoff really started. This is definitely scummy behavior from OpenAI (big surprise), and I actually think they arguably should see an antitrust investigation for that (not that that will ever happen), but OpenAI is a buyer; that's not the same thing as the vendors forming a cartel.
You can't say "connect the dots" at the end of a raving, mostly-incorrect post and act like you've made an ironclad argument.
hn_acc1 2 hours ago [-]
I mean, just because OpenAI started it doesn't mean the vendors didn't form a cartel afterwards (or conspire together) to ensure maximum profits in a "crazy high demand" situation..
3 hours ago [-]
petu 4 hours ago [-]
There's no BF16, original full quality weights are quantized already and 510GB.
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
wren6991 2 hours ago [-]
Are you counting the n-gram/PLE as part of the model weights there? They can go in host memory. Would be good to show your working. Also the released weights are pre-quantised and presumably QATed, so your "Full Precision" and INT8 are simply not a version of the model that actually exists.
Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek's upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it's a larger fraction of the bytes than it is of the parameters.
So,
> Call me crazy but:
You're crazy. :-)
mrinterweb 2 hours ago [-]
Projects like DwarfStar https://github.com/antirez/ds4 really lower the hardware bar a lot so Deepseek 4.1 flash and other mixture of expert models can run on consumer hardware. There are also other inference providers who make their money serving openweight models. Services like OpenRouter make it all too easy to utilize these models. Access to these models isn't hard. The hardware moat is becoming pretty easy to bridge.
contingencies 2 hours ago [-]
More concretely DwarfStar M5 128GB Deepseek 4.1 flash 1K tokens @ 29s, 5K tokens + reasoning @ 147s, 10k token prompt @ 463 tokens/s = 22s. Hardware buy-in USD$7K / AUD$8.5K / EUR€6.8K. At typical workloads, ROI is still poor vs. current-era subsidies, but owning hardware is good for privacy/longevity/connectivity independence. Whether you actually consider Apple hardware 'owned' is a valid and thought provoking question.
onlyrealcuzzo 2 hours ago [-]
Still gonna take 2-3 years to get DeepSeek V4.1 Flash quality at decent speeds on reasonably priced hardware.
Hardware update cycles are 2-3 years even on the high end, so it's still a ways away before "good enough" and "local" belong in the same sentence for the average person.
And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You'd still need to go out of your way.
Most people are spending most of their time on their phones anyway. ..
GLM 5.3 Flash runs fine on two Sparks and Qwen 3.8 Flash Next on one is indeed incredible! I made this 3D game with it in two days using Qwen Code as agent:
I have access to two and will explore this the coming weeks.
girvo 2 hours ago [-]
Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun
ct520 2 hours ago [-]
"1070s or 1070 TIs because GPUs have been severely overpriced for too long" ... ."
1070ti launch MSRP was $450 ish.
5070 could be had in the last year for 5xx-6xx range easily.
All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.
Might be a bit of a stretch blaming it on "severely overpriced for too long..."
ManuelKiessling 3 hours ago [-]
Thanks for the data!
Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.
Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?
So at FP16, I alone keep a 1,664 GiB system occupied all the time?
DrammBA 3 hours ago [-]
No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.
rnxrx 2 hours ago [-]
It depends hugely on what "rent usage of this model through one of the many LLM hosting providers" means. If you're asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that's only getting/answering requests from you alone. When you aren't actively using the model the hosting process is still active and waiting with all of that memory still wired to it.
As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.
There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.
keammo1 3 hours ago [-]
The article isn't just about running locally though. The author is saying it's super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I'm always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don't have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that's a big win for customers and buts competitive pressure on the frontier labs as well.
ByteAtATime 4 hours ago [-]
I think, considering the size of this model, it's closer to a Pro than a Flash on everything other than speed
crossroadsguy 3 hours ago [-]
I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I'd rather pay the 15-30 USD per month via subscription and/or API key combos for years than buying a local setup which might go out of date very fast, if it doesn't goes kaput just out of warranty. I won't be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)
apitman 3 hours ago [-]
> Even so why would anyone not sleep on a model they cannot run?
Because it's an open model so providers compete on price.
cookiengineer 3 hours ago [-]
It's dangerous to go alone. Take this: [1]
I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)
I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.
My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.
Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.
I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.
nullc 3 hours ago [-]
The bulk of its weights are natively MXFP4. And engram values don't need to be in vram.
functionmouse 4 hours ago [-]
one can make a fine gaming pc for ~$350
1660 ti, 4790k, 16gb ddr3
holoduke 4 hours ago [-]
He doesn't ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east.
Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.
CorrectHorseBat 4 hours ago [-]
I've read the exact opposite, vendors are reducing the standard from 16GB back to 8GB
holoduke 3 hours ago [-]
That's only temporary till production meets demand again.
CorrectHorseBat 3 hours ago [-]
Which is not going to happen in the upcoming years
CamperBob2 3 hours ago [-]
You can run it locally for the price of a decent car, or run it (hopefully) privately on somebody else's hardware at vast.ai or a similar provider for much less. What's not to like?
No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.
For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.
liuliu 4 hours ago [-]
What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?
jauer 3 hours ago [-]
This “blame sama for memory prices” meme is so tired.
He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.
mlinsey 3 hours ago [-]
I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot
closer, but I didn't use it enough to really say for my workloads).
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
apitman 3 hours ago [-]
The tightening of subscription value has already begun. dsv4.1f is already worth paying for at market API prices. Maybe it goes to 2x because apparently no one has figured out how to match DeepSeek's insane caching efficiency, but I don't see it getting much worse than that.
Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10/mo OpenCode Go sub for the price than my $20/mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.
rapind 3 hours ago [-]
It's really not cheaper than frontier subscriptions. It's getting closer, and it's a great model, but it is not more value per task than the frontier subscriptions. Don't be swayed by the token costs, it's very chatty, like 3x more tokens for the same task as sol. I used dsf 4.1 full time for about a week.
It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.
I do think it's the most viable alternative I've seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I'll be using.
When pricing this, it's important to consider whether or not you want to opt out of data training. You won't get the advertised rate. Also the dsf 4.1 subscription providers are throttled af... and of course they are, because otherwise they'd be haemorrhaging money.
apitman 3 hours ago [-]
I am looking at cost per task (and speed per task), both benchmarks and anecdotal experience.
LeBit 2 hours ago [-]
What are you talking about?
DS4.1 Flash not really cheaper than frontier models???
It is insanely cheaper.
runtime_terror 25 minutes ago [-]
IME GLM is inferior compared to Deepseek v4.1 Flash, highly recommend running it through some real work
gregwebs 2 hours ago [-]
I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
* the orchestator of my coding workflows
* the tester/verifier of code changes
* the sub agent that explores code or does web searches
* putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.
They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
runtime_terror 22 minutes ago [-]
I'm using DSv4.1 in OpenChamber (eg OpenCode) using the Superpowers skills and a lot of custom AGENTS.md instructions to iron out the kinks and I genuinely cannot see a difference between it and Opus and I've been building native iOS and AppleTV apps, Go servers, Typescript, Cloudflare workers, Svelte/Astro, etc.
It's a super capable model all around from my experience.
oh_no 2 hours ago [-]
Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks.
AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.
Why do people like to think open models are more competitive than they are?
pimeys 2 hours ago [-]
It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.
6 is worse than 5.6 here.
But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.
gregwebs 2 hours ago [-]
DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.
p1necone 3 hours ago [-]
I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
rspeele 3 hours ago [-]
On the Claude side of things I was previously following "strong model directs weak" with Fable directing Opus/Sonnet (its choice per-task). Since Opus 5.5 came out I've just been having Opus direct Opus.
The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I'm finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It's really really good at subjective quality of work per token used.
vorticalbox 47 minutes ago [-]
My work pays for Claude and cursor.
I have actually just dropped to using sonnet for everything, sure it does need some directing but I have yet to see a need to jump to opus.
To me it feels like sonnet/terra and composer 2.5 and grok 4.7 are actually good enough for most tasks and these companies are pushing the high models simply to make money.
p1necone 3 hours ago [-]
I would probably go that route if I could use other harnesses with claude models, but I don't want to be locked in to claude code, and their API pricing (which you need to use it with other harnesses) is so much higher than subscription.
nostrebored 1 hours ago [-]
plenty of harnesses use the claude subscription
runtime_terror 20 minutes ago [-]
If you're not using the Superpowers set of skills, give it a shot. It's been working really well for me on a variety of tasks.
hollowturtle 2 hours ago [-]
Can't wait to just use the hand made Jai and laugh at everything else built with AI :)
lmf4lol 4 hours ago [-]
Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
PcChip 3 hours ago [-]
>We run all our Personal Assistants now on flash
are you worried about sending all your data to third parties, especially if they're in different countries?
techmunky 3 hours ago [-]
Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).
octoberfranklin 3 hours ago [-]
Just a reminder that any API using a Cloudflare TLS certificate isn't ZDR.
The model engine provider might be ZDR, but the service as a whole isn't.
techmunky 2 hours ago [-]
did not know. ty! have my updoot as thanks
figmert 2 hours ago [-]
I use it through OpenRouter, which has ZDR enforcement.
crossroadsguy 3 hours ago [-]
I have asked OP that question but I think there are providers who are not in China and they just host the model/inference.
tripleee 47 minutes ago [-]
what would be the concern here?
yieldcrv 3 hours ago [-]
Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock
nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party
crossroadsguy 3 hours ago [-]
What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI's Coding Plan? Because that's what I use; and often hit the "wait". I wouldn't mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.
(And what are the preferred providers?)
HKCM852 1 hours ago [-]
What personal assistants are you using?
aftbit 4 hours ago [-]
Have you compared it against actual SOTA models like latest Fable or Astra?
mtrovo 2 hours ago [-]
The author explains this very well tbh:
> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.
I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.
The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.
sneurlax 3 hours ago [-]
Of course there's still a huge performance gap
but DS 4.1 Flash is good enough for most tasks
user43928 2 hours ago [-]
Because DeepSeek is not "a month or two" behind as claimed in the article.
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
runtime_terror 18 minutes ago [-]
Perhaps on certain benchmarks and for certain work, but anecdotally I've not been able to see a difference between it and Opus on a lot of dev work (web, Go, iOS/AppleTV native, scripting, general tasks)
BobbyJo 2 hours ago [-]
I was thinking about this earlier today and I came to the following question:
If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?
I think there would be a market and I think it would be large.
So, I agree.
lifeisloving 1 hours ago [-]
Many people would, and you'll find that they're building crappy webapps where you dont need SoTA. Like seriously who needs these frontier models?
Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.
There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.
You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.
ForHackernews 2 hours ago [-]
Doing what? How many jobs involve solving Millennium Prize math challenges?
99% of everything is CRUD LoB apps.
BobbyJo 1 hours ago [-]
I am coding CRUD apps with a mix of astra, sol 6.1, fable and opus 5.5. A more capable model would still benefit me imo. Being able to follow high level guidance better, and being able to harness other models for each task would be a big improvement.
lifeisloving 1 hours ago [-]
Do you know how what you're doing, or do you find yourself working on things you dont understand and need the best model because it's the only way to push your own capabilities (because you're avoiding learning how to do the thing yourself)?
Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.
BobbyJo 37 minutes ago [-]
It's a matter of bandwidth. The more I can offload onto the model, the more I can accomplish. For example, I had to do a lot of security work over the last 2 weeks to get ready for an event. This requires handholding current models on many fronts, like:
1) Do they actually implement the security fixes correctly.
2) Do their fixes create any new edge cases.
3) Do their fixes compromise existing interfaces or API surfaces.
I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I'd be able to do more with my time.
I've actually found models to be pretty poor at driving things I don't know well, so I generally don't do that unless its general design/product exploration and the end product code is throw-away.
airtnp 37 minutes ago [-]
Agree, will see after the dilution and CoT hack fixed, will they keep the pace now. MiMo had some good numbers recently because it's discovered that the post evaluation RL directly exposes answers to models, so RL and evaluation is runied.
aleqs 2 hours ago [-]
that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol
even their harnesses are far surpassed by pi and opencode at this point
also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days
enraged_camel 2 hours ago [-]
>> that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol
Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.
aleqs 1 hours ago [-]
Yeah the picture they paint is that they're mostly bullshit
ForHackernews 2 hours ago [-]
There absolutely is "good enough" and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.
hmontazeri 4 hours ago [-]
I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
jacquesm 4 hours ago [-]
If DS4.1 impresses you I would be really interested to see your comparison to GLM 5.3. I switched from the one to the other and even if GLM 5.3 is a bit slower I don't think I'll be going back.
ctolsen 4 hours ago [-]
GLM 5.3 is very impressive and definitely better, but it also at least 4x the price.
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
badatnames 3 hours ago [-]
DS4 (not 4.1) crossed my dont-care threshold and I genuinely stopped paying attention to new models. I'd love to try GLM 5.3 but I just don't see any point in spending the effort any more. I can get passable intelligence for a bargain price either direct from China or from a ZDR EU provider for a small markup. Paying 10x more will not make me 10x happier, it's unlikely to make me even 1.1x happier now I've got some intuition for the natural limits of these models.
I don't even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?
pjerem 3 hours ago [-]
IDK what happened today but I used GLM-5.3 as usual from Ollama cloud and it was so fast it generated entire documents like instantly.
The reasoning and the result document were done after less than 1 or 2 seconds.
Have Ollama suddenly bought GPU capacity?
LeBit 1 hours ago [-]
There is also GLM 5.3 Flash
pdhborges 3 hours ago [-]
What inference provider are you using?
arush15june 3 hours ago [-]
I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
pimeys 1 hours ago [-]
Yeah. I've been enjoying Coralbricks 250-350 tok/s speeds and it is hard to go back to slower models.
Lithos promises even faster speeds if you want to pay more.
apitman 3 hours ago [-]
> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
youniverse 2 hours ago [-]
How are you guys doing orchestration? I have been fumbling around in my free time trying to build something for myself but is there a repo or something that just works?
apitman 2 hours ago [-]
I might not be the best person to ask. I use Pi harness in tmux and just ask my current agent to spawn interactive pi instances in new tmux windows, create a sentinel file for each of them, and monitor the sentinel files for signals every 2 seconds.
Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.
I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.
I quite like Paseo (been maining it for a week), but Orca also looks good.
simpaticoder 4 hours ago [-]
The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
agoodusername63 4 hours ago [-]
I think it also has a bit to do with the AI sector of tech still moving at lightning speed.
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
pimeys 1 hours ago [-]
Luna is very slow and bad at agentic tasks. DS runs circles around it and there are US providers providing cheaper rates no matter the time of the day.
RGS1811 3 hours ago [-]
This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
zug_zug 3 hours ago [-]
I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
0xbadcafebee 1 hours ago [-]
I keep a spreadsheet that estimates actual value (dollar amount per token per month, per subscription rate limit) and open weights are basically always cheaper than frontier weights. Recently things like GPT 5.6 Luna finally got the frontier close to the value of open weights but their limits keep them behind.
wg0 4 hours ago [-]
While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
sampullman 4 hours ago [-]
Do you mean Fable 5.1? Or Opus 5.5? I'm not sure what you're working on but for me DS 4.1 flash isn't nearly at their level. For the price it's obvious very impressive, though Luna 6.0 is excellent too.
hirako2000 3 hours ago [-]
The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that's not so sure. And anyway, we are not throwing the same X.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.
wg0 3 hours ago [-]
Fabble 5.1.
mtrovo 2 hours ago [-]
care to share what exactly are you working on?
james2doyle 2 hours ago [-]
Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
swiftcoder 4 hours ago [-]
I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
9dev 3 hours ago [-]
Whatever the question, Meta is the wrong answer.
jacquesm 2 hours ago [-]
> so long as you are willing to share data with them
I don't think so.
airtnp 38 minutes ago [-]
Because good enough in the writer's context is a pretty low standard. While many people regards GPT 6.1 Sol or Opus 5.5 as "incapable" in some cases.
Just try Opus 5.5 reminds me how Opus 4.5/4.6 astonishes me. Completely different, and GLM-5.3/Kimi3/DS-4.1 are still like Opus4.8 levels.
shadyr 29 minutes ago [-]
I've been using DeepSeek's API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?
alex-moon 3 hours ago [-]
I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
aguilaair 4 hours ago [-]
What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.
Technically yes, but has been reported to be quite benchmaxxed. In practice Deepseek Flash 4.1 and GLM 5.3 might therefore still outperform Mimo 2.6 pro.
ctolsen 1 hours ago [-]
I’ve been using it a lot and it’s performing really well. Not GLM 5.3 levels but it beats Deepseek for my use. I’ve used it on long running coding tasks, though mostly prototyping, but it’s done a great job at very low cost.
2 hours ago [-]
ctolsen 1 hours ago [-]
Not sure "freaking out" is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.
elmer2 3 hours ago [-]
DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
qwerpy 3 hours ago [-]
Yeah. $100 for Claude just about gives me all the usage I want, as a more or less full-time hobbyist having it work in the background most of the day. I was trying to economize by having a local LLM, then Deepseek, then Cursor/Grok, and then I got a taste of Opus 5.5 and I simply cannot go back to having to carefully spec things out and double-check work. I just let it decide, Opus or Sonnet for the next task, and I get almost perfect results. Probably similar with OpenAI's models.
The token-equivalent monthly spend is > $5K+. If Deepseek's token cost is 20x cheaper, that's $250/mo, and I'd be spending a lot more of my brainpower babysitting it and getting worse results.
For business/team accounts that pay per-token, maybe I can see the "freaking out" being warranted on the part of the fronter labs. But as long as they're willing to subsidize their end-user subscriptions, I'm not going to move off of them until the alternatives are truly at their level.
pants2 2 hours ago [-]
Probably because Luna is faster, cheaper, and approximately as smart
Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
brunooliv 1 hours ago [-]
It’s obvious: they train on prompts and store data when using through their official API.
And for third party it’s just… not good. That’s it.
wren6991 2 hours ago [-]
It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
profsummergig 1 hours ago [-]
Why isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?
ne01 2 hours ago [-]
Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
browningstreet 4 hours ago [-]
What would freaking out look like, or is this just a stupid bloggish title flourish?
Is OpenAI coming in $20B under a sign of "freaking out"?
jerf 4 hours ago [-]
It would look like major chaos in the markets.
People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.
The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.
And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
hirako2000 3 hours ago [-]
It's also unclear whether those who approved those loans understand GPUs depreciation. In any case, progress in software but also hardware could bring chaos and ruin their house of cards.
pessimizer 3 hours ago [-]
> It's also unclear whether those who approved those loans understand GPUs depreciation.
This also assumes heavy utilization, though. If there's heavy utilization, it might mean they're doing well. If they're all spinning, it's time to raise prices.
hirako2000 3 hours ago [-]
Only if utilization isn't at a loss. Right?
NortySpock 2 hours ago [-]
Agree that people leaving the big two companies is going to be hard to keep a pulse on prior to IPO.
Anthropic and OpenAi are in the news, so they get the press and people go and try out their product. Large enterprise businesses are going to make larger, longer-term contracts with them and are only going to pivot if they think switching costs are easy or if they think the provider won't deliver.
The other inference producers are less well known or you need to get your cloud sales rep to tell you how to switch to them as a provider rather than Anthropic or OpenAI.
I use OpenRouter, I know switching is easy, but larger businesses tend to work in yearly cycles. DeepSeek v4 Flash came out in late April.
I agree OpenAI and Anthropic are going to struggle when the median price of running a smart-enough model keeps falling.
Edit: I also think demand for hardware will be rapidly absorbed by other companies if Anthropic or OpenAI stumble. We've finally turned hardware directly into runnable intelligence and people are not going to go back to the old ways.
efficax 4 hours ago [-]
they should be freaking out because every time the chinese labs or non "frontier" labs release a model that is only a few months behind and much cheaper than the openai/anthropic models it shows that they don't deserve their valuations
WJW 3 hours ago [-]
Perhaps, OR it might be that most people in the markets (think that they) are not all that exposed to the valuation AI labs and so their eventual collapse doesn't matter.
Or perhaps they consider the upside from cheap Chinese models to hedge the effect that OpenAI/Anthropic collapsing would have on their portfolios. This would make sense for (hedge funds holding) most companies: they don't really care about who supplies the AI, as long as they get it at roughly the same price as their competitors.
browningstreet 2 hours ago [-]
ironically, at my large enterprise, they aren't yet distinguishing between "chinese models" and "chinese models hosted at microsoft foundry". so far it's just _banned_. i'm not at all pretending it's like that at other orgs.
LeBit 2 hours ago [-]
I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.
It costs pennies and you got really great output.
The author is spot on.
smallmancontrov 4 hours ago [-]
They might be. They would delay public admission as long as possible, because public admission would make stocks go down.
booi 4 hours ago [-]
Because GLM 5.3 Flash is even cheaper?
f311a 3 hours ago [-]
Opencode Go gives only 6300 requests for glm and 23 000 for deepseek. And, if I wanted to, I would be able to do all my work on $10 plan with deepseek. It’s very cheap.
How did you calculate it? Based on per 5 hours max request allowance?
ActionHank 4 hours ago [-]
Nah fam, not true, also DS edges it out on coding / dev tasks.
jacquesm 4 hours ago [-]
That is opposite to my experience so far, can you describe your coding tasks? Mine are systems level code, utilities, operating system code, networking and real time control stuff.
UncleOxidant 3 hours ago [-]
I also prefer GLM-5.3-flash to DS-4.1-flash, but it's close. Since Z.ai has been offering essentially free GLM-5.3-flash tokens on their coding plan between 8am-6pm pdt I've been using it a lot... though that ends on Oct 10 IIRC.
DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
IanCal 2 hours ago [-]
Probably off topic but this is pretty wild to bury in the readme
> By default Mjolnir sends recent prompt and reply text and help-search text to TypeSafe's hosted Jev classifier through a public proxy
f6v 3 hours ago [-]
My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
aussieguy1234 2 hours ago [-]
What blows me away about this model is it's speed.
It's way faster than Opus or any of the GPT models.
I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
thefourthchime 4 hours ago [-]
For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.
I guess what we're seeing is selection bias - people clicking on this HN story will be those who are interested in DeepSeek. And those people who are invested in DeepSeek may not like the facts that you presented.
It's annoying that social networks work this way. The upvote should be for high-quality content and the downvote should be for low-quality content. But .. well.. human nature and tribal dynamics always seem to win.
xyzsparetimexyz 3 hours ago [-]
There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
wildster 3 hours ago [-]
I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
david-gpu 3 hours ago [-]
Don't you run into it sometimes outputting a few Chinese characters, or Cyrillic, for no apparent reason? I fear it writing some nonsense in the code or the terminal. DeepSeek V4.1 Flash doesn't seem to do that.
jacquesm 2 hours ago [-]
That hasn't happened with GLM 5.3 yet but with DS 4.1 Flash it did happen and it also had a tendency to loop.
HeavenFox 3 hours ago [-]
To be fair even OpenAI's and, to a lesser extent, Anthropic's models do that sometimes
pianopatrick 4 hours ago [-]
I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.
Would be cool if they added it.
hirako2000 3 hours ago [-]
Since they adhere to the same API spec, you can hook any model. It takes one line edit in /etc/hosts
There are some quirks if your harness use unsupported features of course.
aszen 3 hours ago [-]
Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
tengbretson 4 hours ago [-]
I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
athrael-soju 1 hours ago [-]
Because it will be replaced within weeks?
gsky 4 hours ago [-]
America bans Chinese models sooner or later just the China banned American big tech
hypfer 4 hours ago [-]
Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
zozbot234 2 hours ago [-]
It's still lacking llama.cpp support, and the work on that isn't moving very fast either. Looks more like a general community issue, where this model isn't drawing much interest.
jacquesm 4 hours ago [-]
You can ask them directly, Daniel Han-Chen is pretty responsive.
pizza234 3 hours ago [-]
People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
apitman 3 hours ago [-]
The argument that most people are making isn't that dsv4.1f is better than frontier, but that it's good enough for most tasks, faster, and way cheaper.
> if one looks at the CoT, it's evident that it's way way stupider than frontier models
Frontier models don't show the full CoT
computerex 3 hours ago [-]
The COT isn't an end all be all. Research has shown that the COT isn't necessarily what the model is actually thinking.
potsandpans 1 hours ago [-]
I'm using it quite extensively in my PlayStation decompilation harness
0xbadcafebee 1 hours ago [-]
Because GLM-5.3-Flash is both cheaper and better?
4 hours ago [-]
jeffrallen 1 hours ago [-]
Also, it is willing to do legitimate work I need done which other models flag as dangerous and refuse to do. (Software testing of a DHCP server to survive bad inputs.)
bitfilped 1 hours ago [-]
Because in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
try-working 2 hours ago [-]
I have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.
It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
robertlane0 2 hours ago [-]
Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.
kristianp 4 hours ago [-]
> shrank the KV cache by roughly 437X
Can't you just say "shrank to 1/437th the size"? It's not that hard.
MisterMunchkin 4 hours ago [-]
I had it make 25 different things today and it cost $0.70
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
Octoth0rpe 3 hours ago [-]
> Obviously can’t use it at work
I do wonder how long it'll be before a us-hosted offering is available via bedrock, copilot, etc.
computerex 3 hours ago [-]
There are already US hosted offerings on companies like fireworks.ai.
tonyhart7 56 minutes ago [-]
it literally hallucinating a lot
I dont get why people says D4.1 flash is good
anguralbanish2 4 hours ago [-]
I would love to get them more better, it's good not a bad thing.
cactusplant7374 3 hours ago [-]
Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
pessimizer 3 hours ago [-]
I'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
sergiotapia 3 hours ago [-]
In my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
AIblemblio 4 hours ago [-]
No they can't.
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
m3kw9 3 hours ago [-]
i thought 6.1sol copied the caching architecture so this isn't such a big deal no more
doctorpangloss 4 hours ago [-]
because it doesn't work very well?
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
computerex 3 hours ago [-]
What is your evidence? Deepseek v4.1 Flash is by far the most popular coding model on openrouter, having processed 38.7T tokens in just the last 7 days, over 3x the usage of the 2nd rank model.
So I ask again, what are you basing your assertion on?
doctorpangloss 2 hours ago [-]
my own usage of deekseep v4.1 flash, and that among the dozens of great programmers i know, not a single person is using it
BUT. they are employed to do / deciding-to-do authentic (if often meaningless) stuff.
here's a short list of inauthentic activity that claude and openai refuse to do:
- chat services that, when you ask them, say they are not chatbots when they are
- code to work around software licenses or DRM
- code to scrape or download copyrighted material
- directly cheating on homework
- adopting a persona in social media that spreads misinformation or propaganda
this is but a short list. but ask me, "are there enough inauthentic activity demands such that someone who CANNOT USE claude or gpt as the LLM would use dsv4.1 on openrouter instead?" yes. i mean there are whole countries right now where the culture can be summarized as, "bottom to top, inauthentic activity." i am surprised it is not more usage!
computerex 2 hours ago [-]
So your argument is that the 38T of tokens used in the last 7 days is by moron programmers or people doing "inauthentic" tasks? You think the person who made this post is also an idiot?
Do you realize how incredibly delusional/self-centered you sound?
doctorpangloss 1 hours ago [-]
do YOU know anyone gainfully employed in programming who is using dsv4.1 to do work? what kind of work is it? why don't you ask them if it is good?
in the market, where you cannot fake or hide stuff very easily: the outsource customer services and cheating sectors have been the most disrupted. Cheating company Chegg lost 99% of its market value. CS it remains to be seen - https://www.reuters.com/technology/teleperformance-shares-pl... - certainly perceived to be disrupted, but they are not dead yet.
in my personal usage: dsv4 is generally pretty buggy. for example, if you give it a needle-in-the-haystack simple copying problem, it catastrophically fails to find needles if they happen to be positioned at index 250k tokens out of 1m. it can also be triggered to spew all sorts of garbage when DSpark is enabled during ordinary long-context coding, such as spewing weird DSML tool call errors after a normally parsed tool call error.
i don't know why you have to attack me personally, i think you're a bright and otherwise nice person and you understand the thrust of my POV.
pimeys 1 hours ago [-]
Yes. Hi. From our team 3/5 of us use 4.1 to do our daily tasks. For paid work for a company who pays us salary. From people around me I hear a lot of my friends being really happy with it especially for the price.
I don't know man, maybe this is not super serious what I'm doing. Some systems stuff with rust, implementing my own desktop apps with iced, porting old DOS games to Linux...
It is a very good model.
doctorpangloss 12 minutes ago [-]
so what you're saying is though, if they could afford it they would just use claude or codex?
computerex 1 hours ago [-]
You are speaking out of your ass, that’s what I take issue with. Falsifiability is something I hold sacred and you are taking a dump on it.
Fwiw I work in a company producing software for many fortune 500’s you have heard about and many people from our team use deepseek.
I am literally using it right now. Your entire line of reasoning rubs me the wrong way.
Btw check your provider and harness… improperly configured deepseek can emit dsml. If you are not passing thinking tokens back to the model it tends to do that.
Use a proper harness and good provider.
doctorpangloss 39 minutes ago [-]
"Hey Mr. Fortune 500 Client, would you prefer us to use something called DeepSeek V4.1 Flash, made by the Chinese, sending your Fortune 500 code to some random service provider on something called OpenRouter, where they promise according to something called Zero Data Retention that--"
Mr. Client: "I'm going to stop you right there. Why aren't you using Claude, or Codex, or Claude on Bedrock? Don't we deserve the best?"
You: ...
Look I don't know. I can tell from the hyperbole of your language, talking out of asses and such, that there is more to the story than you are letting on. Like Chinese users are banned from officially using Claude and Codex, for example. So many reasons that you cannot use Claude, not so much reasons to not choose to use Claude. All I am really saying is, I know DSV4 is kind of bad, that there is a lot of inauthentic activity, and that Claude and Codex refuse to do many kinds of inauthentic activity, and that a lot of coding done by outsourced shops has always been of questionable quality and purpose. I mean in my personal life, I know more people who have been scammed by Bulgarian code body shops than I know people who have used DSV4.1.
computerex 31 minutes ago [-]
You have NO IDEA what you're talking about. You are clueless.
Deepseek v4.1 flash is an open weights model. You can run it on your own hardware. You have no idea how my companies gets access to it. A very cursory Google search would reveal to you that there are many enterprise grade LLM providers that host this model on US soil with SOC2 protections.
Try not to talk about subjects you have no knowledge about because you are making yourself look like an idiot.
Edit:
It's also clear to me that you don't deploy any LLM based system on scale because if you had you'd know why open weights models are so compelling.
Hint: it's the cost.
doctorpangloss 24 minutes ago [-]
okay, but are you US based? and can you specifically describe one of the pieces of software you are developing? it's okay if not. i am just wondering. i certainly believe that crappier stuff is cheaper!
computerex 20 minutes ago [-]
Yes the company I work for is based in San Diego, I work remotely from Alabama.
I bet you voted for trump. With brains like that.
cbeach 45 minutes ago [-]
Honestly, I just don't trust the Chinese Communist Party having agentic access to my computer.
Every company in China has to abide by the 2017 National Intelligence Law: "supporting, assisting and cooperating" with state intelligence work, and keeping that cooperation secret. They have to hand prior knowledge of vulnerabilities to the state before public disclosure, in order that the state always has an exploit pipeline. No matter how ethical the company staff may be, they'll always be bound by law into being an arm of the Communist Party.
Agentic access is infinitely worse than chatbots. They can exfiltrate silently, target users, plant persistent malware, and be run by third parties through you.
You don't have to be a tin foil hat sinophobe to understand the dangers of being a Westerner granting CCP access to your files and network.
ByteDance staff accessed US journalists' TikTok data to hunt leakers (admitted in 2022). Volt Typhoon and Salt Typhoon were state operations pre-positioned in Western infrastructure and telecoms. Regulators in Italy and South Korea blocked DeepSeek's app over data handling, and analysts found its web client sending data to a China Mobile domain.
Please don't sacrifice security for cost and convenience.
tensor 3 minutes ago [-]
Some of us don't use US models for the same reasons. It doesn't leave much out there aside from Cohere and Mistral.
verdverm 22 hours ago [-]
Why would we freak out? The systems we use have always gotten better, faster, cheaper with time
kydanet 3 hours ago [-]
[flagged]
oh_no 2 hours ago [-]
AA shows Luna at 1/4 the price, 1 point behind on intelligence matrix with a 38.
Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.
I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /s
Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?
CurbStomper4 3 hours ago [-]
[dead]
distantsounds 4 hours ago [-]
because we've all figured out that AI is just a huge grift?
sroussey 4 hours ago [-]
Not comparing to gpt-6-luna which seems comparable and priced well.
wewewedxfgdf 3 hours ago [-]
You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.
BarryMilo 3 hours ago [-]
Your comment seems to imply there's a good guy in a this. I just see the inevitable end of an era, championed by predictably selfish actors.
f6v 3 hours ago [-]
Oh, the writing is on the wall. Wait till you hear that European “sovereign AI” is just running GLM 5.3.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.
If your tasks are write-heavy, find a provider with cheaper output.
If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.
Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped
I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.
If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.
Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness
Show me people handwriting code à la NASA
2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
In the same way that only supercomputers used to have multiple processors and caches but it's now standard.
Buy directly from DeepSeek's API.
You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
OpenRouter Pricing:
$0.02/M input tokens $0.60/M output tokens
DeepSeek Pricing (cache miss, off-peak):
$0.15/M Input $0.60/m output
It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.
Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.
The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...
There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.
If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.
The market is legally locked down and we're in hostage pricing mode.
And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...
Nobody is coming to save us. That's our job.
Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.
Plus expanding other existing facilities.
These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.
Samsung and HK Hynix also have fabs under construction and planned.
CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.
Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.
> The Micron CEO just recently said this is the exact plan
CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.
It's taken them this long to catch up to the DDR5 standard. They've only recently been through qualifications to be a DDR5 supplier for the big boys.
> Every Major Motherboard Maker Now Validates CXMT DDR5
https://www.techtimes.com/articles/321572/20260725/every-maj...
After their recent IPO, they have more than enough cash to ramp up in a major way.
It's just a matter of time.
Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.
You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.
I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.
It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.
Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.
But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.
But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.
This has happened in electronics markets before. Many times.
When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.
If you were a DRAM manufacturer, isn't this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?
Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.
You have a very expensive data center and you're in debt financed on the premise that you have these special computers.
Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.
This stuff happens all the time. It's why we don't use BMP files on websites or serve giant MOV files on YouTube. It's why postgres queries are faster now than they were 10 and 20 years ago.
You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.
Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.
On market if you were to sell some of that ram you have 100% profit right now but not for long.
Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...
We live in the real world so what do you do?
Historically the answer has been "sell that shit"
There's an aphorism for this "stairs on the way up elevator on the way down"
If we had a healthy market with sane prices where you can't flip the thing you bought for 100% profit the answer would be "create more value."
Or they could stay at $10,000 per month since they are willing to pay that much already.m, so they just use AI more and in more places.
Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).
People who say they want local runs really mean it: they want local runs on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I'm not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.
How many people, outside of tech geeks and megacorps care about RAM prices? And how gullible would you be to BELIEVE the politician they could actually make it happen, and even if they did, that it would extend to the average person, and not JUST megacorps/megadonors?
I think a lot of people care about the downstream effects of memory prices, but I agree with you that they may not realize that they happen because of memory prices.
wonder what voting would be like?
gamer vote ++
datacenter hater vote --
datacenter lobby ++
micron lobby --
I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.
I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.
Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.
I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.
Connect the dots.
You can't say "connect the dots" at the end of a raving, mostly-incorrect post and act like you've made an ironclad argument.
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek's upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it's a larger fraction of the bytes than it is of the parameters.
So,
> Call me crazy but:
You're crazy. :-)
Hardware update cycles are 2-3 years even on the high end, so it's still a ways away before "good enough" and "local" belong in the same sentence for the average person.
And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You'd still need to go out of your way.
Most people are spending most of their time on their phones anyway. ..
It has a set of n-gram tables which you can stream from system RAM or even NVMe
That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?
https://github.com/christopherowen/spark-ds41f
I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware
https://blog.jonathanpage.com/
GLM 5.3 Flash runs fine on two Sparks and Qwen 3.8 Flash Next on one is indeed incredible! I made this 3D game with it in two days using Qwen Code as agent:
https://games.jonathanpage.com/
1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.
All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.
Might be a bit of a stretch blaming it on "severely overpriced for too long..."
Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.
Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?
So at FP16, I alone keep a 1,664 GiB system occupied all the time?
As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.
There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.
Because it's an open model so providers compete on price.
I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)
I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.
My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.
Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.
[1] https://github.com/cookiengineer/gonano
1660 ti, 4790k, 16gb ddr3
No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.
For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.
He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10/mo OpenCode Go sub for the price than my $20/mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.
It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.
I do think it's the most viable alternative I've seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I'll be using.
When pricing this, it's important to consider whether or not you want to opt out of data training. You won't get the advertised rate. Also the dsf 4.1 subscription providers are throttled af... and of course they are, because otherwise they'd be haemorrhaging money.
DS4.1 Flash not really cheaper than frontier models???
It is insanely cheaper.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
It's a super capable model all around from my experience.
AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.
Why do people like to think open models are more competitive than they are?
6 is worse than 5.6 here.
But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I'm finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It's really really good at subjective quality of work per token used.
I have actually just dropped to using sonnet for everything, sure it does need some directing but I have yet to see a need to jump to opus.
To me it feels like sonnet/terra and composer 2.5 and grok 4.7 are actually good enough for most tasks and these companies are pushing the high models simply to make money.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
are you worried about sending all your data to third parties, especially if they're in different countries?
The model engine provider might be ZDR, but the service as a whole isn't.
nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party
(And what are the preferred providers?)
> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.
I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.
The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.
but DS 4.1 Flash is good enough for most tasks
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?
I think there would be a market and I think it would be large.
So, I agree.
Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.
There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.
You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.
99% of everything is CRUD LoB apps.
Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.
I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I'd be able to do more with my time.
I've actually found models to be pretty poor at driving things I don't know well, so I generally don't do that unless its general design/product exploration and the end product code is throw-away.
even their harnesses are far surpassed by pi and opencode at this point
also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days
Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
I don't even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?
The reasoning and the result document were done after less than 1 or 2 seconds.
Have Ollama suddenly bought GPU capacity?
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
Lithos promises even faster speeds if you want to pay more.
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.
I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.
Check: https://agentmgmt.dev/ and find the one that works for you.
I quite like Paseo (been maining it for a week), but Orca also looks good.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.
I don't think so.
Just try Opus 5.5 reminds me how Opus 4.5/4.6 astonishes me. Completely different, and GLM-5.3/Kimi3/DS-4.1 are still like Opus4.8 levels.
see https://artificialanalysis.ai/models/releases/comparisons?co...
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.
The token-equivalent monthly spend is > $5K+. If Deepseek's token cost is 20x cheaper, that's $250/mo, and I'd be spending a lot more of my brainpower babysitting it and getting worse results.
For business/team accounts that pay per-token, maybe I can see the "freaking out" being warranted on the part of the fronter labs. But as long as they're willing to subsidize their end-user subscriptions, I'm not going to move off of them until the alternatives are truly at their level.
Is OpenAI coming in $20B under a sign of "freaking out"?
People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.
The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.
And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
This also assumes heavy utilization, though. If there's heavy utilization, it might mean they're doing well. If they're all spinning, it's time to raise prices.
Anthropic and OpenAi are in the news, so they get the press and people go and try out their product. Large enterprise businesses are going to make larger, longer-term contracts with them and are only going to pivot if they think switching costs are easy or if they think the provider won't deliver.
The other inference producers are less well known or you need to get your cloud sales rep to tell you how to switch to them as a provider rather than Anthropic or OpenAI.
I use OpenRouter, I know switching is easy, but larger businesses tend to work in yearly cycles. DeepSeek v4 Flash came out in late April.
I agree OpenAI and Anthropic are going to struggle when the median price of running a smart-enough model keeps falling.
Edit: I also think demand for hardware will be rapidly absorbed by other companies if Anthropic or OpenAI stumble. We've finally turned hardware directly into runnable intelligence and people are not going to go back to the old ways.
Or perhaps they consider the upside from cheap Chinese models to hedge the effect that OpenAI/Anthropic collapsing would have on their portfolios. This would make sense for (hedge funds holding) most companies: they don't really care about who supplies the AI, as long as they get it at roughly the same price as their competitors.
It costs pennies and you got really great output.
The author is spot on.
> and 23 000 for deepseek
How did you calculate it? Based on per 5 hours max request allowance?
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
> By default Mjolnir sends recent prompt and reply text and help-search text to TypeSafe's hosted Jev classifier through a public proxy
It's way faster than Opus or any of the GPT models.
I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
It's annoying that social networks work this way. The upvote should be for high-quality content and the downvote should be for low-quality content. But .. well.. human nature and tribal dynamics always seem to win.
Would be cool if they added it.
There are some quirks if your harness use unsupported features of course.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
> if one looks at the CoT, it's evident that it's way way stupider than frontier models
Frontier models don't show the full CoT
It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
Can't you just say "shrank to 1/437th the size"? It's not that hard.
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
I do wonder how long it'll be before a us-hosted offering is available via bedrock, copilot, etc.
I dont get why people says D4.1 flash is good
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
So I ask again, what are you basing your assertion on?
BUT. they are employed to do / deciding-to-do authentic (if often meaningless) stuff.
here's a short list of inauthentic activity that claude and openai refuse to do:
- chat services that, when you ask them, say they are not chatbots when they are
- code to work around software licenses or DRM
- code to scrape or download copyrighted material
- directly cheating on homework
- adopting a persona in social media that spreads misinformation or propaganda
this is but a short list. but ask me, "are there enough inauthentic activity demands such that someone who CANNOT USE claude or gpt as the LLM would use dsv4.1 on openrouter instead?" yes. i mean there are whole countries right now where the culture can be summarized as, "bottom to top, inauthentic activity." i am surprised it is not more usage!
Do you realize how incredibly delusional/self-centered you sound?
in the market, where you cannot fake or hide stuff very easily: the outsource customer services and cheating sectors have been the most disrupted. Cheating company Chegg lost 99% of its market value. CS it remains to be seen - https://www.reuters.com/technology/teleperformance-shares-pl... - certainly perceived to be disrupted, but they are not dead yet.
in my personal usage: dsv4 is generally pretty buggy. for example, if you give it a needle-in-the-haystack simple copying problem, it catastrophically fails to find needles if they happen to be positioned at index 250k tokens out of 1m. it can also be triggered to spew all sorts of garbage when DSpark is enabled during ordinary long-context coding, such as spewing weird DSML tool call errors after a normally parsed tool call error.
i don't know why you have to attack me personally, i think you're a bright and otherwise nice person and you understand the thrust of my POV.
I don't know man, maybe this is not super serious what I'm doing. Some systems stuff with rust, implementing my own desktop apps with iced, porting old DOS games to Linux...
It is a very good model.
Fwiw I work in a company producing software for many fortune 500’s you have heard about and many people from our team use deepseek.
I am literally using it right now. Your entire line of reasoning rubs me the wrong way.
Btw check your provider and harness… improperly configured deepseek can emit dsml. If you are not passing thinking tokens back to the model it tends to do that.
Use a proper harness and good provider.
Mr. Client: "I'm going to stop you right there. Why aren't you using Claude, or Codex, or Claude on Bedrock? Don't we deserve the best?"
You: ...
Look I don't know. I can tell from the hyperbole of your language, talking out of asses and such, that there is more to the story than you are letting on. Like Chinese users are banned from officially using Claude and Codex, for example. So many reasons that you cannot use Claude, not so much reasons to not choose to use Claude. All I am really saying is, I know DSV4 is kind of bad, that there is a lot of inauthentic activity, and that Claude and Codex refuse to do many kinds of inauthentic activity, and that a lot of coding done by outsourced shops has always been of questionable quality and purpose. I mean in my personal life, I know more people who have been scammed by Bulgarian code body shops than I know people who have used DSV4.1.
Deepseek v4.1 flash is an open weights model. You can run it on your own hardware. You have no idea how my companies gets access to it. A very cursory Google search would reveal to you that there are many enterprise grade LLM providers that host this model on US soil with SOC2 protections.
Like: https://fireworks.ai/
Try not to talk about subjects you have no knowledge about because you are making yourself look like an idiot.
Edit: It's also clear to me that you don't deploy any LLM based system on scale because if you had you'd know why open weights models are so compelling.
Hint: it's the cost.
I bet you voted for trump. With brains like that.
Every company in China has to abide by the 2017 National Intelligence Law: "supporting, assisting and cooperating" with state intelligence work, and keeping that cooperation secret. They have to hand prior knowledge of vulnerabilities to the state before public disclosure, in order that the state always has an exploit pipeline. No matter how ethical the company staff may be, they'll always be bound by law into being an arm of the Communist Party.
Agentic access is infinitely worse than chatbots. They can exfiltrate silently, target users, plant persistent malware, and be run by third parties through you.
You don't have to be a tin foil hat sinophobe to understand the dangers of being a Westerner granting CCP access to your files and network.
ByteDance staff accessed US journalists' TikTok data to hunt leakers (admitted in 2022). Volt Typhoon and Salt Typhoon were state operations pre-positioned in Western infrastructure and telecoms. Regulators in Italy and South Korea blocked DeepSeek's app over data handling, and analysts found its web client sending data to a China Mobile domain.
Please don't sacrifice security for cost and convenience.
Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.
I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /s
Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?