Rendered at 12:58:56 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
nchmy 17 hours ago [-]
The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
geniium 16 hours ago [-]
I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
jorvi 34 minutes ago [-]
As long as these models constantly keep switching up things like the temperatures at which a steak will be medium rare or at which temperature to season cast iron, you will never be able to trust them for cooking. My mother ruined a nice waterfowl for Christmas by listening to Gemini.
And this is inherent to how LLMs work.
ben_w 22 minutes ago [-]
While yes, their current reliability is too spiky to be relied upon for a lot of things:
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
adrinavarro 15 hours ago [-]
I share this feeling too. The latest models, even if not necessarily frontier, say Opus 5, Sol high and the likes, I could keep using these models forever even if they did not significantly improve beyond this point. I also believe we'll come up with new ways of using these very same models beyond the mainstream chat and agent interfaces, as the bottleneck is imho in harnesses/environments and not so much model intelligence anymore.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
drunkboxer 3 hours ago [-]
Do you have to give any special instructions to do this? I always want to do something like this, but any time I try I get so sick of listening to what it has to say, just long winded explanations of stuff that tends to go off the rails. Imo it's hard enough to read ai output when I can go back and forward between sentences to make sense of what's being said let alone listen to a continuous train of slop.
simonw 46 minutes ago [-]
The latest ChatGPT voice mode is really good at being interrupted - I'll often say "no, no, no, that's too much information" while it's talking to stop and redirect it.
CJefferson 5 hours ago [-]
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.
olmo23 4 hours ago [-]
I also kind of miss how "unhinged" the earlier models were.
ShinyLeftPad 4 hours ago [-]
What about censorship?
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
rrr_oh_man 3 hours ago [-]
Everyone censors for their core jurisdiction/audience. The Enlightened West just calls this guardrails
euroderf 2 hours ago [-]
You make it sound like the DNC.
59nadir 27 minutes ago [-]
US models censor and restrict more things than Chinese models by quite a margin.
lelanthran 1 hours ago [-]
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
afavour 38 minutes ago [-]
That sounds ideal for cooking but I do wonder about programming. Models frozen in ember won’t ever learn new APIs as they become available and development will end up in some weird kind of stasis.
PcChip 16 minutes ago [-]
embers are probably too hot to freeze anything
tmp10423288442 11 hours ago [-]
ChatGPT literally released a major update of their realtime voice model a month or two ago, going from gpt-4o-level (generously) to gpt-5.5 level performance. So at least 2026-level performance was necessary to provide a really good experience.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
trueno 2 hours ago [-]
this is how i felt about opus 4.6 i still use it it's just faster and does enough to be super helpful. i've used these later anthropic ones a few times but the word salad and slowness feels like it just opens the door to building shit that just stacks and adds on itself.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
r_lee 16 hours ago [-]
imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
josephg 15 hours ago [-]
Yeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.
somenameforme 10 hours ago [-]
IMO many are still missing a big part of the picture. We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winner. I think the historically precedented answer is none of them.
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
petra 5 hours ago [-]
I agree that the winners would be doing stuff in the physical world.
And there I think the winner would be China.
eru 5 hours ago [-]
We will all be winners.
petra 3 minutes ago [-]
Like all of us we're winners because of the internet?
josephg 4 hours ago [-]
Maybe - we'll have to wait and see.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
eru 4 hours ago [-]
Sorry, when I wrote 'all', I meant people all over the world (not just in China).
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
x______________ 6 hours ago [-]
> We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winn..
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
altmanaltman 2 hours ago [-]
Nvidia has been winning for decades, they got it right in gaming, they got it right in crypto, and they got it right in AI. People (outside of tech mostly) think they just got lucky but if that's the case, they have all the luck in the world.
ben_w 30 minutes ago [-]
Fair competition under capitalism necessarily drives down profit margins; high profits are either temporary, or due to a lack of competition (e.g. someone has a patent or other IP, or regulatory capture). For example, while a lot of the economy depends on electricity: where competition exists, the profit margin for making electricity is not high; where monopolies or government mandates exist, it can be otherwise. This means that assuming anyone wins (i.e. no doom scenario), the winners are probably going to be those who can make best use of the models. Even chip makers will probably not get a long-term boost out of this; there's plenty of room for more efficient compute, and competitive advantages from e.g. ASML last as long as it takes to reinvent their tech, it's not a law of nature.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
eru 5 hours ago [-]
Or perhaps customers / users?
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
stbede 4 hours ago [-]
I won financially. Encyclopedia sets were expensive.
eru 4 hours ago [-]
Your savings are real, but they don't show up in GDP or a profit-and-loss statement of any company.
KSteffensen 3 hours ago [-]
The cost of buying the encyclopedia becomes disposable income to be used on other consumer goods. In that sense in shows up in lots of other companies profit-and-loss statements.
eru 2 hours ago [-]
Maybe, but that's very diffuse and hard to attribute to Wikipedia.
And it would show up in real GDP, not necessarily in nominal GDP.
a2ff6eeb0 15 hours ago [-]
It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically. With the trillions of dollars that's going in through both investment and users, it's going to happen. I don't believe the human brain has fundamental magic that will make this impossible.
adrianN 11 hours ago [-]
True AGI would upend society in such a way that I'm not sure that being a shareholder of anything would be meaningful. Perhaps being a pitchfork manufacturer is the winning play in this scenario.
rustcleaner 4 hours ago [-]
BRB, longing Remmington and Winchester.
thelastgallon 11 hours ago [-]
The true followers (shareholders) of the AI messiah will be saved, everyone else is doomed.
altmanaltman 2 hours ago [-]
late stage christianity
georgemcbay 12 hours ago [-]
> It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically.
What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
icepush 7 hours ago [-]
The first AGI that decides it doesn't want any more AGIs is the last one that gets created.
r_lee 6 minutes ago [-]
that would require both AGI and physical bodies for the model, and no kills witches or anything that could stop it, e.g. the military
everything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
a2ff6eeb0 13 hours ago [-]
For the downvoters: What magic do you think the human brain has that makes it impossible to emulate acceptably?
pianopatrick 11 hours ago [-]
It's not about the feasibility of the technology.
If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
4 hours ago [-]
a2ff6eeb0 11 hours ago [-]
The AI is likely to control the ability to apply force (see all of the autonomous drone companies). There's a great deal of alignment work being done to ensure that the AI will continue to listen to the shareholders of these companies.
If that fails, who knows what things will look like.
pianopatrick 10 hours ago [-]
Are you sure that alignment work is aligning with the share holders and not the operators? Or not the creators? Or not the government? Which of these groups should the AI listen to when these groups disagree?
If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
a2ff6eeb0 9 hours ago [-]
To be honest: I don't know for certain, but I'd assume that the people who pay the bills get the strongest alignment. They may not be tech people, but I (so far) haven't got a reason to think that the AI engineers are going behind the backs of their corporate leadership and subverting what they're being asked to do; do you?
(I think it would be a good thing for humanity if they did)
pianopatrick 8 hours ago [-]
I think right now both the engineers developing AI and the share holders are more focused on beating coding benchmarks and gaining revenue than anything to do with alignment.
ThrowawayR2 10 hours ago [-]
The drones don't manufacture themselves, maintain themselves, reload their own ammunition, mine and refine the materials that are used to make them and their ammunition, or operate the power plants needed for all of the above. "AI" isn't going to control diddly squat.
a2ff6eeb0 9 hours ago [-]
There's a huge amount of research into embodied AI (and, also, people seem to be a lot more ok with manufacturing bullets than pulling triggers).
watwut 6 hours ago [-]
You did not read much about history, did you? Or law. If you embed AI into a gun ... you just created a bomb. It is still you who killed whoever it kills.
ksenzee 13 hours ago [-]
LLMs are not emulating the human brain. Somebody may well be able to do that someday, but right now nobody is even trying to.
josephg 13 hours ago [-]
Why would you need brain emulation to get superhuman intelligence?
ksenzee 12 hours ago [-]
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far? Or are you making the generic assertion that AGI is theoretically possible via means other than emulating the human brain? Because the latter is a strawman (nobody has asserted anything to the contrary), and I have seen no evidence at all to support the former.
josephg 12 hours ago [-]
I think we can compare the human brain and LLMs on a bunch of capabilities today, and see how we compare. By my reckoning:
- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
shawnb576 2 hours ago [-]
Because LLMs don't understand anything. That's the tech. They can only predict what they have been trained with and fail daily at the most basic tasks. Granted they can do amazing things, no question there. But they are not "smart".
For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
penteract 2 hours ago [-]
To me it sounds like you're repeating what gp said about the lack of online learning. Do you think that's insurmountble?
Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
OJFord 5 hours ago [-]
What is meant by 'superhuman intelligence'? Certainly it seems to be the case that LLMs are capable of a sort of 'polyhuman' intelligence, in that the same LLM that advances mathematics with a novel proof can add unit testing for a new software feature, design a recipe, and create an SVG of a pelican riding a bicycle. As generalists I'd say they're already 'superhuman'.
CamperBob2 11 hours ago [-]
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far?
Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
scotty79 5 hours ago [-]
> training LLMs on everything humanity knows so far?
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
lelanthran 1 hours ago [-]
> What magic do you think the human brain has that makes it impossible to emulate acceptably?
If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
glimshe 14 hours ago [-]
All it needs is Internet access to remain useful with few shortcomings.
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
matteoraso 16 hours ago [-]
>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
ColdStream 11 hours ago [-]
I have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.
RALaBarge 1 hours ago [-]
I ask DS4F to make a plan, then check it with grok/fable, build the code, check it with grok/fable, ship
KunYuan 5 hours ago [-]
A computer that costs $10,000 is impressive. A computer that costs $100 and reaches billions of people changes the world.
Maybe AI will follow the same path.
dsrtslnd23 6 hours ago [-]
I think it really depends - for a lot of things outside of coding and general knowledge tasks even the best models (fable 5 etc.) are not good enough yet: e.g. CAD, PCB design (though getting there on PCB design), ...
jimmydoe 9 hours ago [-]
Current AI is smart enough to help us, but the creators of AK want it to be smart enough to replace us.
nbardy 6 hours ago [-]
The real revolution is both. The cost and capability of frontier intelligence will go up AND the cost of "good enough" intelligence will go down.
lilbigdoot 17 hours ago [-]
If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand coding
nchmy 16 hours ago [-]
I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..
intrasight 11 hours ago [-]
> content if they never got smarter, and just kept getting even cheaper/faster.
I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
ksh09 16 hours ago [-]
I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.
ericd 11 hours ago [-]
It costs about as much as a cheap car to do this well, but it's attainable now, and qwen 3.8 seems to make it possible on a 5090.
nchmy 12 hours ago [-]
this is the holy grail
poincareball 16 hours ago [-]
Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.
Tuna-Fish 14 hours ago [-]
Please explain why you think cheaper/faster is not coming?
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
sipjca 14 hours ago [-]
what do you mean cheaper/faster is not really coming? the cost of the same level of intelligence steadily decreases year over year. computer hardware also advances at the same time enabling cheaper and faster serving (or move to local)
bad_haircut72 16 hours ago [-]
not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing
ACCount37 15 hours ago [-]
"Specialized models" are a bit of a doozy.
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
CamperBob2 15 hours ago [-]
And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.
ACCount37 13 hours ago [-]
Which are the kinds of tasks computers have been historically quite good at.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
CamperBob2 13 hours ago [-]
Computers have historically been good at answering word problems fed to them verbatim?
ACCount37 5 hours ago [-]
No, but they were good at answering formalized versions of the same word problems.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
scotty79 5 hours ago [-]
Computers were never good at math. They were good at pre-coded algebra.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
ACCount37 4 hours ago [-]
Moravec's paradox begs to differ. Things that are hard to humans are easy. Things that are easy to humans are hard.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Blazing fast...but terrible. Put Sol on silicon but will still need access to the internet...so it will be somewhat slow anyway
Tuna-Fish 2 hours ago [-]
It's terrible because it's Llama 3.1 8B. It's such a crappy model because HC1 was a relatively low budget proof of concept.
The team that built is working on a better implementation.
ACCount37 15 hours ago [-]
What "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.
redox99 15 hours ago [-]
Eh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities.
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
dmurray 2 hours ago [-]
Why are we not just in a free lunch moment but with harnesses, rather than models?
Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.
But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.
Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
Kinrany 45 minutes ago [-]
Hard to inagine the final outcome being anything other than the smartest model + cheap subagents. The only problem with that is user requests being pasted directly into the model's context, but that's got to be temporary.
peteforde 14 hours ago [-]
A few months ago folks were understandably annoyed when Microsoft dropped their heavily subsidized per-request pricing model because it was figuratively burning cash.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
robertjpayne 13 hours ago [-]
Going to be great to see the cash burn on SpaceX's next earnings report. Will the cult keep the stock price pumped?
Gareth321 3 hours ago [-]
Grok 4.6 XHigh uses 2.52x fewer tokens per task than Fable Max. It's much more efficient. It also has low market penetration, which is why SpaceX is selling so much of their compute to Anthropic et al. From a business perspective, they're capitalising on the market very well. If Grok becomes more popular we should expect to pay more.
rmast 16 hours ago [-]
Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
dd8601fn 8 hours ago [-]
There are whole classes of things I can’t thought exercise or really learn about because the “safeguards” keep tripping me down to haiku.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
rustcleaner 4 hours ago [-]
>They need to fix that.
There can only be one fix: send Amodei packing and release unguardrailed models.
lossolo 15 hours ago [-]
> At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
jdnier 11 hours ago [-]
I asked Fable to transcribe three short lines of Korean-language text in a small image. It suspected the image might contain song lyrics and refused. Haiku transcribed it with no issue.
danlugo92 12 hours ago [-]
No prompting is also prompting, young padawan
nicoburns 16 hours ago [-]
That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
11 hours ago [-]
pigpop 16 hours ago [-]
Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
r_lee 16 hours ago [-]
Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
TiredOfLife 15 hours ago [-]
The Cerebras version of 5.6 is available only to select customers
pigpop 15 hours ago [-]
You're right, I should have clarified that they are still slowly integrating it and it isn't the thing running all models. I meant moreso that since they are planning on moving more usage over to Cerebras wafers, they're able to relieve some pressure on their predicted expenses while also moving some current workload (ultrafast and codex spark) onto them freeing up Nvidia GPUs.
nottorp 6 hours ago [-]
Is the real LLM revolution the fact that every piece of news and opinion is now phrased as if it's the end of the world though?
mholm 17 hours ago [-]
As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
tyre 17 hours ago [-]
Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
jml78 16 hours ago [-]
I operate mostly in the devops arena. Lots of things opus is fine for. But there is just things where I can hand hold Opus through changes, or I can ask Fable to do it and it gets it right on the first try. People will say let fable plan and validate with opus doing the work. I found that burns fable tokens even faster because opus makes so many mistakes, fable has to review things 4-5 times before opus gets it right. A single fable implementation at medium or low effort would have one shot it.
ACCount37 15 hours ago [-]
Yep. Every time you get more intelligence, that buys you more autonomy, more reliability, more task complexity. Tasks done with less mistakes, less handholding, less interventions.
This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
a2ff6eeb0 16 hours ago [-]
Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.
vineyardmike 16 hours ago [-]
> Or even outside of software!
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
a2ff6eeb0 16 hours ago [-]
I suspect the focus will probably shift once software engineering is no longer the biggest cost center for most AI company's clients, and we'll start working on getting rid of the next cost center.
Foobar8568 5 hours ago [-]
Well for the last two years, chatgpt ( and I guess claude) were already better than most of my coworkers ( read IT in F500 style organizations ).
dgellow 16 hours ago [-]
It’s worth considering for companies paying API prices, and not relying on a subscription quota
DanielHall 4 hours ago [-]
What a clickbait title. I thought Fable was no longer included in the Max plan.
alasdair_ 7 hours ago [-]
I’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.
g42gregory 10 hours ago [-]
I have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.
gunalx 2 hours ago [-]
glm 5.3 gets awfully slow during peak hours. But you might not hit them to frequently.
blfr 16 hours ago [-]
What are all these rote coding tasks people do that they can farm it out to lesser models?
ihateolives 6 hours ago [-]
Add new route to API that displays additional information we need, work out query for it, update controllers/models/whatnot.
No need for top model for that.
denverllc 16 hours ago [-]
Write a detailed plan using a more expensive model and implement it using the cheaper one.
blfr 16 hours ago [-]
How much are you saving once the more expensive model already has all the context loaded and ready to go?
csullivannet 16 hours ago [-]
API calls get more expensive, not less, as you've loaded more context. This is exactly when you want to switch to cheaper models.
camdenreslink 13 hours ago [-]
There is caching to consider. Switching models throws away the cached tokens.
mattmanser 16 hours ago [-]
Are you genuinely asking?
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
nicoburns 15 hours ago [-]
One task I've found this useful for is writing example code. Release admin (updating version numbers, etc) as well.
janalsncm 11 hours ago [-]
This is essentially the anti-Bitter Lesson lesson which I feel has become a bit of a thought terminating cliche lately.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
zkmon 16 hours ago [-]
I guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
Zylokloto 16 hours ago [-]
He started with thinking were to send what.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
aabhay 16 hours ago [-]
This concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
wild_egg 13 hours ago [-]
I would love to pay for Fable at full API pricing but unfortunately it is blocked from working on any of my projects. Looking forward to the end of the year when the truly comparable open models will drop.
enraged_camel 17 hours ago [-]
>> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
tonyarkles 17 hours ago [-]
Something I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
unshavedyak 16 hours ago [-]
Those "present state" comments are the bane of my existence. It was present in 4.7/etc but i put in a ton of guards against that into my global memory and it worked quite well. Fable and Opus 5 regressed badly in this space though and i can't keep it from making those types of comments again.
Really frustrating.
senderista 16 hours ago [-]
I have Sol prune/revise those comments.
Jare 16 hours ago [-]
> my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by [others]
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
tyre 16 hours ago [-]
As a counterpoint (data point of one code base), I had Fable lead development of a complex system recently (an end-to-end insurance claims billing system) as a test project. It blew me away. Opus could not have done the same, given the feedback Fable had to give when Opus would implement individual features.
Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.
But the code it produced—and the implementations it guided Opus towards—were excellent.
robomc 16 hours ago [-]
> It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly
My brother, that's my job.
ericol 14 hours ago [-]
From my point of view the issue is that there are too many things wrong with Fable, making it seriously not worth the money.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1]
I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2]
"The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
zarmin 7 hours ago [-]
I agree completely. It's "I didn't have time to write you a short letter so I wrote you a long one"
bellowsgulch 17 hours ago [-]
Are people still using deepseek-v4-flash everywhere? I found after the price increases, mimo-v2.5 seems far more attractive.
farlight 17 hours ago [-]
It's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.
bellowsgulch 16 hours ago [-]
Awesome. Thanks for the heads up.
moltar 17 hours ago [-]
I just use Fable for reviews of specs and code then hand off to Opus to work on. Works well.
dude250711 16 hours ago [-]
Does it not silently degrade to Opus if it does not like some word?
lantry 2 hours ago [-]
Users have the option of silent/automatic degradation or a complete halt. I have it set to stop rather than degrade because I want to know when I've hit the safeguard.
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
dncornholio 5 hours ago [-]
Fable is only marginally better.
freepiai 16 hours ago [-]
I've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
m3kw9 16 hours ago [-]
looks like you haven't tried openai or Sol, or even luna (max)
uejfiweun 11 hours ago [-]
Seeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.
dbbk 3 hours ago [-]
I agree. I've been perfectly happy since Opus 4.6. I remember thinking at the time if it never improved I would have been fine there.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
resters 17 hours ago [-]
over time greater intelligence will be expressed in smaller and cheaper models. we are still somewhat near the beginning of this bc we are finally starting to understand what makes a model truly intelligent/capable.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
sudeepsd__ 5 hours ago [-]
[dead]
hypfer 17 hours ago [-]
[flagged]
dbreunig 17 hours ago [-]
[flagged]
hypfer 17 hours ago [-]
[flagged]
tyre 16 hours ago [-]
I don't think everything has to be Thought Leadership. OP compared the same generation of models to show that the latest open model—at the time of the latest closed model—was Good Enough.
I agree that the opener to their reply wasn't productive, but neither is "Weak."
dbreunig 17 hours ago [-]
I think it’s a fine response when you say, “Doesn't feel well informed enough to give advice,” because I said 5.2
hypfer 17 hours ago [-]
Idk man, but an engineer would've taken that and said something like: "Damn, yeah, good point, I shall add a sentence mentioning 5.3"
Because an engineer feels secure in their knowledge so that such an oversight doesn't make them suddenly defend their identity - it's just an oversight after all. Happens.
simonw 17 hours ago [-]
5.3 isn't available as open weights yet, and only became available via API three days ago. Prior to that the only way to access it was via a Z.ai subscription.
hypfer 17 hours ago [-]
[flagged]
dgellow 16 hours ago [-]
FWIW you’re not looking good in this engagement, feels very childish, looking for a gotcha that doesn’t mean much
hypfer 16 hours ago [-]
I think that depends on the audience. Thank you for caring though :)
kelnos 16 hours ago [-]
Audience member here: I agree with GP; you posts come off as petty and childish.
It seems natural to me to make comparisons only to open weight models where the weights have actually been released.
hypfer 16 hours ago [-]
[flagged]
gpjanik 16 hours ago [-]
"When Moore’s Law slowed in the mid-2000s" it did not, in fact, slow down in the mid 2000s, or at all.
You've selectively quoted the article. The full quote (emphasis added):
"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."
Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
selcuka 12 hours ago [-]
> Your link is talking about transistor count. The article is talking about single-threaded performance.
But Moore's Law has always been about transistor count, not performance.
12 hours ago [-]
gpjanik 5 hours ago [-]
And more cores means what exactly in terms of transistors count?
sscaryterry 16 hours ago [-]
It did in terms of the traditional more MHz (GHz) is better, but as you've correctly pointed out, not when it comes to actual compute.
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
And this is inherent to how LLMs work.
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
And there I think the winner would be China.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
And it would show up in real GDP, not necessarily in nominal GDP.
What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
everything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
If that fails, who knows what things will look like.
If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
(I think it would be a good thing for humanity if they did)
- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Model on a custom silicon: https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet
The team that built is working on a better implementation.
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.
But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.
Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
There can only be one fix: send Amodei packing and release unguardrailed models.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Really frustrating.
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.
But the code it produced—and the implementations it guided Opus towards—were excellent.
My brother, that's my job.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
I agree that the opener to their reply wasn't productive, but neither is "Weak."
Because an engineer feels secure in their knowledge so that such an oversight doesn't make them suddenly defend their identity - it's just an oversight after all. Happens.
It seems natural to me to make comparisons only to open weight models where the weights have actually been released.
https://ourworldindata.org/data-insights/moores-law-has-accu...
"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."
Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
But Moore's Law has always been about transistor count, not performance.