Rendered at 17:16:35 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
HighGoldstein 2 hours ago [-]
I find it somewhat funny that the author starts by talking about how Moore's law enabled inefficient software and that we then had to make it more efficient, when almost all software today is horrendously inefficient compared to even 10 years ago, let alone 20-30. We've somehow even achieved a state where it doesn't matter how fast your CPU and memory are, the software will just perform horribly on any machine.
fidotron 2 hours ago [-]
One of the funniest parts of the LLM wave is discovering that cron was so annoying to use that we will burn the planet to put an interface on it that people can actually work with.
The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.
dboreham 2 hours ago [-]
Assuming you are referring to using an LLM to generate crontab entries, is that a bad move? Seems like it removes the need for layers of UI that most of the time is never used. Same goes for regexes. Actually same for SQL. No need for layers to translate between what the user can specify and what's executed. Just type what you want to query for and the LLM generates the SQL.
phplovesong 1 hours ago [-]
Its slow because of human slop from the early to mi 2010s, and now slow because of AI was trained on the slop that existed.
Bottom line is the slop used to be manageable, but now there is 100x more code pushed, so that train has departed.
In the end its more bad code for features no one will use.
nchmy 21 hours ago [-]
The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
geniium 20 hours ago [-]
I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
jorvi 5 hours ago [-]
As long as these models constantly keep switching up things like the temperatures at which a steak will be medium rare or at which temperature to season cast iron, you will never be able to trust them for cooking. My mother ruined a nice waterfowl for Christmas by listening to Gemini.
And this is inherent to how LLMs work.
ben_w 5 hours ago [-]
While yes, their current reliability is too spiky to be relied upon for a lot of things:
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
lukan 4 hours ago [-]
Not if you let your LLM grounds its truth in established facts. Otherwise they would be useless for programming for example.
saghm 3 hours ago [-]
I don't understand what this means. I use LLMs daily for my work in programming things, and they regularly will assert things that are not accurate.
lukan 2 hours ago [-]
When you let the the agent do a test, or tell it to read that doc first, you will ground it in reality. Doesn't mean they are 100% reliable. But without and on their own without access to grounding information, they halluzinate wildly.
stevage 2 hours ago [-]
A lot like humans, really. People regularly cite something they read, or quote a stat that turns out to be just completely inaccurate. But if you look up the thing, then you have facts again.
elorant 3 hours ago [-]
So why isn't this the default mode then?
lukan 3 hours ago [-]
More expensive.
geraldwhen 3 hours ago [-]
Well I wouldn’t trust Gemini with most things.
ravishi 3 hours ago [-]
Really? I've essentially learned how to cook from Gemini. Not a great cook yet, but I can do the basics now.
CJefferson 9 hours ago [-]
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.
olmo23 8 hours ago [-]
I also kind of miss how "unhinged" the earlier models were.
ShinyLeftPad 8 hours ago [-]
What about censorship?
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
CJefferson 6 minutes ago [-]
I dislike all censorship, but US models are much more censored, I often find myself using Chinese models to get answers I want.
Now, of course I’d prefer no censoring, but I live in the world we live in.
I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.
rrr_oh_man 7 hours ago [-]
Everyone censors for their core jurisdiction/audience. The Enlightened West just calls this guardrails
euroderf 6 hours ago [-]
You make it sound like the DNC.
59nadir 5 hours ago [-]
US models censor and restrict more things than Chinese models by quite a margin.
ShinyLeftPad 3 hours ago [-]
So you choose bad over worse and pretend it's good
lelanthran 5 hours ago [-]
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
ShinyLeftPad 3 hours ago [-]
The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.
adrinavarro 20 hours ago [-]
I share this feeling too. The latest models, even if not necessarily frontier, say Opus 5, Sol high and the likes, I could keep using these models forever even if they did not significantly improve beyond this point. I also believe we'll come up with new ways of using these very same models beyond the mainstream chat and agent interfaces, as the bottleneck is imho in harnesses/environments and not so much model intelligence anymore.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
drunkboxer 7 hours ago [-]
Do you have to give any special instructions to do this? I always want to do something like this, but any time I try I get so sick of listening to what it has to say, just long winded explanations of stuff that tends to go off the rails. Imo it's hard enough to read ai output when I can go back and forward between sentences to make sense of what's being said let alone listen to a continuous train of slop.
simonw 5 hours ago [-]
The latest ChatGPT voice mode is really good at being interrupted - I'll often say "no, no, no, that's too much information" while it's talking to stop and redirect it.
tmp10423288442 16 hours ago [-]
ChatGPT literally released a major update of their realtime voice model a month or two ago, going from gpt-4o-level (generously) to gpt-5.5 level performance. So at least 2026-level performance was necessary to provide a really good experience.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
r_lee 20 hours ago [-]
imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
josephg 20 hours ago [-]
Yeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.
somenameforme 14 hours ago [-]
IMO many are still missing a big part of the picture. We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winner. I think the historically precedented answer is none of them.
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
23 minutes ago [-]
petra 9 hours ago [-]
I agree that the winners would be doing stuff in the physical world.
And there I think the winner would be China.
eru 9 hours ago [-]
We will all be winners.
josephg 8 hours ago [-]
Maybe - we'll have to wait and see.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
eru 8 hours ago [-]
Sorry, when I wrote 'all', I meant people all over the world (not just in China).
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
petra 4 hours ago [-]
Like all of us we're winners because of the internet?
x______________ 10 hours ago [-]
> We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winn..
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
altmanaltman 7 hours ago [-]
Nvidia has been winning for decades, they got it right in gaming, they got it right in crypto, and they got it right in AI. People (outside of tech mostly) think they just got lucky but if that's the case, they have all the luck in the world.
ben_w 5 hours ago [-]
Fair competition under capitalism necessarily drives down profit margins; high profits are either temporary, or due to a lack of competition (e.g. someone has a patent or other IP, or regulatory capture). For example, while a lot of the economy depends on electricity: where competition exists, the profit margin for making electricity is not high; where monopolies or government mandates exist, it can be otherwise. This means that assuming anyone wins (i.e. no doom scenario), the winners are probably going to be those who can make best use of the models. Even chip makers will probably not get a long-term boost out of this; there's plenty of room for more efficient compute, and competitive advantages from e.g. ASML last as long as it takes to reinvent their tech, it's not a law of nature.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
eru 9 hours ago [-]
Or perhaps customers / users?
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
stbede 8 hours ago [-]
I won financially. Encyclopedia sets were expensive.
eru 8 hours ago [-]
Your savings are real, but they don't show up in GDP or a profit-and-loss statement of any company.
KSteffensen 7 hours ago [-]
The cost of buying the encyclopedia becomes disposable income to be used on other consumer goods. In that sense in shows up in lots of other companies profit-and-loss statements.
eru 6 hours ago [-]
Maybe, but that's very diffuse and hard to attribute to Wikipedia.
And it would show up in real GDP, not necessarily in nominal GDP.
a2ff6eeb0 20 hours ago [-]
It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically. With the trillions of dollars that's going in through both investment and users, it's going to happen. I don't believe the human brain has fundamental magic that will make this impossible.
adrianN 15 hours ago [-]
True AGI would upend society in such a way that I'm not sure that being a shareholder of anything would be meaningful. Perhaps being a pitchfork manufacturer is the winning play in this scenario.
rustcleaner 9 hours ago [-]
BRB, longing Remmington and Winchester.
thelastgallon 16 hours ago [-]
The true followers (shareholders) of the AI messiah will be saved, everyone else is doomed.
altmanaltman 7 hours ago [-]
late stage christianity
georgemcbay 16 hours ago [-]
> It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically.
What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
icepush 11 hours ago [-]
The first AGI that decides it doesn't want any more AGIs is the last one that gets created.
r_lee 4 hours ago [-]
that would require both AGI and physical bodies for the model, and no kills witches or anything that could stop it, e.g. the military
everything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
a2ff6eeb0 17 hours ago [-]
For the downvoters: What magic do you think the human brain has that makes it impossible to emulate acceptably?
pianopatrick 16 hours ago [-]
It's not about the feasibility of the technology.
If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
8 hours ago [-]
a2ff6eeb0 15 hours ago [-]
The AI is likely to control the ability to apply force (see all of the autonomous drone companies). There's a great deal of alignment work being done to ensure that the AI will continue to listen to the shareholders of these companies.
If that fails, who knows what things will look like.
pianopatrick 14 hours ago [-]
Are you sure that alignment work is aligning with the share holders and not the operators? Or not the creators? Or not the government? Which of these groups should the AI listen to when these groups disagree?
If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
a2ff6eeb0 14 hours ago [-]
To be honest: I don't know for certain, but I'd assume that the people who pay the bills get the strongest alignment. They may not be tech people, but I (so far) haven't got a reason to think that the AI engineers are going behind the backs of their corporate leadership and subverting what they're being asked to do; do you?
(I think it would be a good thing for humanity if they did)
pianopatrick 12 hours ago [-]
I think right now both the engineers developing AI and the share holders are more focused on beating coding benchmarks and gaining revenue than anything to do with alignment.
ThrowawayR2 14 hours ago [-]
The drones don't manufacture themselves, maintain themselves, reload their own ammunition, mine and refine the materials that are used to make them and their ammunition, or operate the power plants needed for all of the above. "AI" isn't going to control diddly squat.
a2ff6eeb0 13 hours ago [-]
There's a huge amount of research into embodied AI (and, also, people seem to be a lot more ok with manufacturing bullets than pulling triggers).
watwut 10 hours ago [-]
You did not read much about history, did you? Or law. If you embed AI into a gun ... you just created a bomb. It is still you who killed whoever it kills.
ksenzee 17 hours ago [-]
LLMs are not emulating the human brain. Somebody may well be able to do that someday, but right now nobody is even trying to.
josephg 17 hours ago [-]
Why would you need brain emulation to get superhuman intelligence?
ksenzee 17 hours ago [-]
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far? Or are you making the generic assertion that AGI is theoretically possible via means other than emulating the human brain? Because the latter is a strawman (nobody has asserted anything to the contrary), and I have seen no evidence at all to support the former.
josephg 16 hours ago [-]
I think we can compare the human brain and LLMs on a bunch of capabilities today, and see how we compare. By my reckoning:
- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
shawnb576 7 hours ago [-]
Because LLMs don't understand anything. That's the tech. They can only predict what they have been trained with and fail daily at the most basic tasks. Granted they can do amazing things, no question there. But they are not "smart".
For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
penteract 6 hours ago [-]
To me it sounds like you're repeating what gp said about the lack of online learning. Do you think that's insurmountble?
Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
dboreham 1 hours ago [-]
I see this thread as progress because now there's three people saying this (seemed like it was just me for a year or two).
josephg 4 hours ago [-]
> Because LLMs don't understand anything.
How do you square that then? They can do amazing things, but they're also not smart? Do you think its possible to solve Erdos problems without any "smarts"? Can you do it without even understanding mathematics?
I find it very hard to hold the idea that LLMs don't understand anything. They can explain concepts, translate them, simplify them and implement them in code. From the outside, LLMs seem to understands most concepts better than most humans do. Do you understand anything? Couldn't I make the same argument? How would you prove that you understand what a for loop is, or that you know what calculus is? I assume you'd demonstrate your knowledge by using a for loop in a program, or explain calculus back to me. But LLMs can do that too.
> For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them.
Funny example, because lots of human struggle with this too. The number of meetings I've had with people in the US! "Lets meet on thursday morning australia time!". Only, they actually meant thursday night US time, which is friday morning australia time. "Oooh that's so weird! Its the next day for you!". ...... Yes, I know.
I think LLMs are just a different kind of intelligence than humans. They're better at some things than us, and worse than others. They can find latent security vulnerabilities in the linux kernel, but struggle to count the Rs in strawberry. They're not as smart as humans in many ways. But we're not as smart as LLMs in plenty of ways too. I didn't find those linux bugs.
OJFord 9 hours ago [-]
What is meant by 'superhuman intelligence'? Certainly it seems to be the case that LLMs are capable of a sort of 'polyhuman' intelligence, in that the same LLM that advances mathematics with a novel proof can add unit testing for a new software feature, design a recipe, and create an SVG of a pelican riding a bicycle. As generalists I'd say they're already 'superhuman'.
scotty79 9 hours ago [-]
> training LLMs on everything humanity knows so far?
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
ksenzee 23 minutes ago [-]
To clarify: We are not training LLMs on any information that humanity does not already have access to.
CamperBob2 15 hours ago [-]
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far?
Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
ksenzee 18 minutes ago [-]
Superhuman intelligence? Really? These leading-edge advances indicate intelligence beyond the level of humanity?
lelanthran 5 hours ago [-]
> What magic do you think the human brain has that makes it impossible to emulate acceptably?
If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
trueno 6 hours ago [-]
this is how i felt about opus 4.6 i still use it it's just faster and does enough to be super helpful. i've used these later anthropic ones a few times but the word salad and slowness feels like it just opens the door to building shit that just stacks and adds on itself.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
r_lee 26 minutes ago [-]
the jump from 4.7-4.8 to 5 is so bad in terms of the word salad
i just get fatigued from it, am I holding it wrong or something?
sometimes it's fine but the constant RLHFisms like the constant "worth flagging" and stuff is getting really old
glimshe 18 hours ago [-]
All it needs is Internet access to remain useful with few shortcomings.
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
afavour 5 hours ago [-]
That sounds ideal for cooking but I do wonder about programming. Models frozen in amber won’t ever learn new APIs as they become available and development will end up in some weird kind of stasis.
PcChip 5 hours ago [-]
embers are probably too hot to freeze anything
someothherguyy 2 hours ago [-]
eventually the novelty wears off and depression kicks in
bunderbunder 3 hours ago [-]
But they don't really have that option. They're trapped in a Red Queen's race.
The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
cosmojg 2 hours ago [-]
Newer models[1] are being trained in ways that prioritize coding and agentic performance over raw knowledge[2] such that they increasingly rely on external tools for accessing hard data and information.
Models don't need to keep retraining just to stay current. Harnesses give them access to the internet, internal systems use RAGs, and so on.
I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold.
Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :)))
Ferrari/Lambo numbers are microscopic
zdragnar 2 hours ago [-]
Even with a harness, models don't reach out for new information they don't know about. For some tech, I have to have a local model draft a plan, then I have to adjust the plan to update it with the new API and references for where to find it. Even if I include that updated information in the prompt for the plan, the model says "what the user says is wrong, they probably meant this instead" and goes off in its own direction with old APIs anyway.
ninahaberl 1 hours ago [-]
You are basically saying that some models (your local one, which is it?) in some setups (the API you mentioned) can fail to use fresh information if that conflicts with strong training priors.
I agree:)
BUT
That's a bad model.
My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
bunderbunder 15 minutes ago [-]
For starters, it's not just APIs. Like I pointed out with the C example, programming language syntax and semantics also evolve over time.
But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
zdragnar 43 minutes ago [-]
I most recently experienced this with Qwen 3.8 27b, though I've seen it on several other versions of their local models. It's also heavily biased towards digging into library source code rather than looking at API documentation.
To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
cj 2 hours ago [-]
> they don't really have that option
I imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
anthonypasq 2 hours ago [-]
yeah its called web search
matteoraso 20 hours ago [-]
>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
someothherguyy 2 hours ago [-]
i wish i was experiencing these things that everyone else is.
my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still.
even the interesting results in academic work seem to be more of a function of effort (proofs by exhaustion, fitting puzzle pieces in a search space, etc) than anything else. not to say people aren't using large language models to do impressive things, but the agents themselves do not seem very intelligent to me.
it feels like some engineering teams are aware of this fact and are driving agents using strict rule checks (like hooks on steroids), so they can drive some shape of output that aligns with what they require.
insanitybit 3 hours ago [-]
I'm really not convinced that these models are even that much more intelligent, as opposed to simply being more token aggressive. I do not find Fable that much smarter than Opus 4.6, and no Opus model seems to have improved things much at all.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
ColdStream 15 hours ago [-]
I have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.
RALaBarge 5 hours ago [-]
I ask DS4F to make a plan, then check it with grok/fable, build the code, check it with grok/fable, ship
dsrtslnd23 10 hours ago [-]
I think it really depends - for a lot of things outside of coding and general knowledge tasks even the best models (fable 5 etc.) are not good enough yet: e.g. CAD, PCB design (though getting there on PCB design), ...
KunYuan 9 hours ago [-]
A computer that costs $10,000 is impressive. A computer that costs $100 and reaches billions of people changes the world.
Maybe AI will follow the same path.
jimmydoe 13 hours ago [-]
Current AI is smart enough to help us, but the creators of AK want it to be smart enough to replace us.
ksh09 20 hours ago [-]
I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.
ericd 16 hours ago [-]
It costs about as much as a cheap car to do this well, but it's attainable now, and qwen 3.8 seems to make it possible on a 5090.
nchmy 16 hours ago [-]
this is the holy grail
lilbigdoot 21 hours ago [-]
If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand coding
nchmy 20 hours ago [-]
I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..
ipsod 3 hours ago [-]
I more have an agent that I drive to create features, and then a fleet of agents that turn those ad-hoc implementations into refined, integrated code.
nbardy 10 hours ago [-]
The real revolution is both. The cost and capability of frontier intelligence will go up AND the cost of "good enough" intelligence will go down.
intrasight 15 hours ago [-]
> content if they never got smarter, and just kept getting even cheaper/faster.
I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
poincareball 20 hours ago [-]
Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.
Tuna-Fish 18 hours ago [-]
Please explain why you think cheaper/faster is not coming?
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
naasking 3 hours ago [-]
> Evidence actually supports that capabilities are leveling off
What evidence?
sipjca 19 hours ago [-]
what do you mean cheaper/faster is not really coming? the cost of the same level of intelligence steadily decreases year over year. computer hardware also advances at the same time enabling cheaper and faster serving (or move to local)
bad_haircut72 20 hours ago [-]
not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing
ACCount37 19 hours ago [-]
"Specialized models" are a bit of a doozy.
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
CamperBob2 19 hours ago [-]
And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.
ACCount37 18 hours ago [-]
Which are the kinds of tasks computers have been historically quite good at.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
CamperBob2 17 hours ago [-]
Computers have historically been good at answering word problems fed to them verbatim?
ACCount37 9 hours ago [-]
No, but they were good at answering formalized versions of the same word problems.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
CamperBob2 36 minutes ago [-]
"Formalizing the problem" is pretty much the whole shooting match. It takes intelligence to do that. The rest is mere calculation.
scotty79 9 hours ago [-]
Computers were never good at math. They were good at pre-coded algebra.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
ACCount37 8 hours ago [-]
Moravec's paradox begs to differ. Things that are hard to humans are easy. Things that are easy to humans are hard.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
scotty79 3 hours ago [-]
> but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke"
Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation)
They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Blazing fast...but terrible. Put Sol on silicon but will still need access to the internet...so it will be somewhat slow anyway
Tuna-Fish 6 hours ago [-]
It's terrible because it's Llama 3.1 8B. It's such a crappy model because HC1 was a relatively low budget proof of concept.
The team that built is working on a better implementation.
ACCount37 20 hours ago [-]
What "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.
redox99 19 hours ago [-]
Eh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities.
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
dmurray 6 hours ago [-]
Why are we not just in a free lunch moment but with harnesses, rather than models?
Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.
But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.
Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
Kinrany 5 hours ago [-]
Hard to inagine the final outcome being anything other than the smartest model + cheap subagents. The only problem with that is user requests being pasted directly into the model's context, but that's got to be temporary.
peteforde 18 hours ago [-]
A few months ago folks were understandably annoyed when Microsoft dropped their heavily subsidized per-request pricing model because it was figuratively burning cash.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
chinathrow 4 hours ago [-]
Giving Elon cash seems still wrong to me.
Varelion 2 hours ago [-]
Yea. I can't stress this enough -- my life would need to be unquestionably on the line for me to give elon anything other than grief.
timmg 2 hours ago [-]
Worse than giving to Sama? (I mean that honestly.)
ipsod 2 hours ago [-]
How well does Cursor work with heavily agentic workflows?
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
robertjpayne 18 hours ago [-]
Going to be great to see the cash burn on SpaceX's next earnings report. Will the cult keep the stock price pumped?
Gareth321 7 hours ago [-]
Grok 4.6 XHigh uses 2.52x fewer tokens per task than Fable Max. It's much more efficient. It also has low market penetration, which is why SpaceX is selling so much of their compute to Anthropic et al. From a business perspective, they're capitalising on the market very well. If Grok becomes more popular we should expect to pay more.
rmast 20 hours ago [-]
Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
dd8601fn 12 hours ago [-]
There are whole classes of things I can’t thought exercise or really learn about because the “safeguards” keep tripping me down to haiku.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
rustcleaner 8 hours ago [-]
>They need to fix that.
There can only be one fix: send Amodei packing and release unguardrailed models.
kay_o 2 hours ago [-]
I don't even need to tiptoe ! Not being there and not prompting anything is enough to trigger safeguards.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ... if it does not go back and notice it has written bad code.
Since they hide thinking and reasoning from the user (who is also paying for those tokens) it is a black box what is triggering it, has the LLM this time thought of "Oh, this has XSS" and used a bad dangerous word such as XSS, while the previous conversation did not ?
lossolo 20 hours ago [-]
> At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
jdnier 15 hours ago [-]
I asked Fable to transcribe three short lines of Korean-language text in a small image. It suspected the image might contain song lyrics and refused. Haiku transcribed it with no issue.
danlugo92 17 hours ago [-]
No prompting is also prompting, young padawan
nicoburns 20 hours ago [-]
That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
15 hours ago [-]
pigpop 20 hours ago [-]
Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
bitmasher9 4 hours ago [-]
Anthropic -> OpenAI switcher here.
I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.
I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
r_lee 20 hours ago [-]
Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
TiredOfLife 19 hours ago [-]
The Cerebras version of 5.6 is available only to select customers
pigpop 19 hours ago [-]
You're right, I should have clarified that they are still slowly integrating it and it isn't the thing running all models. I meant moreso that since they are planning on moving more usage over to Cerebras wafers, they're able to relieve some pressure on their predicted expenses while also moving some current workload (ultrafast and codex spark) onto them freeing up Nvidia GPUs.
Jackson98Tom 3 hours ago [-]
I think that the publishing of K3 and Sol proved another thing, following the Moore's Law comparison: that the real deal was not only upgrading the number of neurons (as it was for the transistors) building larger models, but was (again, as for the transistors) making them also more efficient, more portable. That said, we're not in a slowing of the curve of developement, we're fastening, expecially with the chinese models. The free lunch is getting bigger, not ending, still.
nottorp 10 hours ago [-]
Is the real LLM revolution the fact that every piece of news and opinion is now phrased as if it's the end of the world though?
Schlagbohrer 4 hours ago [-]
Apocalyptic doomsaying as marketing strategy. Not great for the Zeitgeist honestly
mholm 21 hours ago [-]
As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
tyre 21 hours ago [-]
Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
jml78 21 hours ago [-]
I operate mostly in the devops arena. Lots of things opus is fine for. But there is just things where I can hand hold Opus through changes, or I can ask Fable to do it and it gets it right on the first try. People will say let fable plan and validate with opus doing the work. I found that burns fable tokens even faster because opus makes so many mistakes, fable has to review things 4-5 times before opus gets it right. A single fable implementation at medium or low effort would have one shot it.
ACCount37 19 hours ago [-]
Yep. Every time you get more intelligence, that buys you more autonomy, more reliability, more task complexity. Tasks done with less mistakes, less handholding, less interventions.
This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
a2ff6eeb0 21 hours ago [-]
Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.
vineyardmike 20 hours ago [-]
> Or even outside of software!
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
a2ff6eeb0 20 hours ago [-]
I suspect the focus will probably shift once software engineering is no longer the biggest cost center for most AI company's clients, and we'll start working on getting rid of the next cost center.
Foobar8568 9 hours ago [-]
Well for the last two years, chatgpt ( and I guess claude) were already better than most of my coworkers ( read IT in F500 style organizations ).
dgellow 21 hours ago [-]
It’s worth considering for companies paying API prices, and not relying on a subscription quota
g42gregory 15 hours ago [-]
I have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.
gunalx 7 hours ago [-]
glm 5.3 gets awfully slow during peak hours. But you might not hit them to frequently.
janalsncm 16 hours ago [-]
This is essentially the anti-Bitter Lesson lesson which I feel has become a bit of a thought terminating cliche lately.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
blfr 21 hours ago [-]
What are all these rote coding tasks people do that they can farm it out to lesser models?
ihateolives 11 hours ago [-]
Add new route to API that displays additional information we need, work out query for it, update controllers/models/whatnot.
No need for top model for that.
denverllc 21 hours ago [-]
Write a detailed plan using a more expensive model and implement it using the cheaper one.
blfr 20 hours ago [-]
How much are you saving once the more expensive model already has all the context loaded and ready to go?
csullivannet 20 hours ago [-]
API calls get more expensive, not less, as you've loaded more context. This is exactly when you want to switch to cheaper models.
camdenreslink 17 hours ago [-]
There is caching to consider. Switching models throws away the cached tokens.
mattmanser 20 hours ago [-]
Are you genuinely asking?
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
nicoburns 20 hours ago [-]
One task I've found this useful for is writing example code. Release admin (updating version numbers, etc) as well.
alasdair_ 12 hours ago [-]
I’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.
zkmon 20 hours ago [-]
I guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
DanielHall 8 hours ago [-]
What a clickbait title. I thought Fable was no longer included in the Max plan.
mwigdahl 4 hours ago [-]
It's still there. You can use up to half your Max capacity on Fable, then you have to downshift.
Zylokloto 20 hours ago [-]
He started with thinking were to send what.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
aabhay 20 hours ago [-]
This concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
wild_egg 17 hours ago [-]
I would love to pay for Fable at full API pricing but unfortunately it is blocked from working on any of my projects. Looking forward to the end of the year when the truly comparable open models will drop.
enraged_camel 21 hours ago [-]
>> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
tonyarkles 21 hours ago [-]
Something I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
unshavedyak 21 hours ago [-]
Those "present state" comments are the bane of my existence. It was present in 4.7/etc but i put in a ton of guards against that into my global memory and it worked quite well. Fable and Opus 5 regressed badly in this space though and i can't keep it from making those types of comments again.
Really frustrating.
senderista 20 hours ago [-]
I have Sol prune/revise those comments.
Jare 21 hours ago [-]
> my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by [others]
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
tyre 21 hours ago [-]
As a counterpoint (data point of one code base), I had Fable lead development of a complex system recently (an end-to-end insurance claims billing system) as a test project. It blew me away. Opus could not have done the same, given the feedback Fable had to give when Opus would implement individual features.
Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.
But the code it produced—and the implementations it guided Opus towards—were excellent.
robomc 20 hours ago [-]
> It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly
My brother, that's my job.
owen-hill 4 hours ago [-]
[flagged]
ericol 18 hours ago [-]
From my point of view the issue is that there are too many things wrong with Fable, making it seriously not worth the money.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1]
I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2]
"The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
zarmin 11 hours ago [-]
I agree completely. It's "I didn't have time to write you a short letter so I wrote you a long one"
freepiai 21 hours ago [-]
I've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
m3kw9 20 hours ago [-]
looks like you haven't tried openai or Sol, or even luna (max)
freepiai 4 hours ago [-]
Oh I have, they are good (sol terribly overbuilds though). Luna is good as well. That said Deepseek v4 flash is generally faster, and IMHO a bit smarter than luna, and it's definitely cheaper. (If you use my harness freepi.ai it's free!). So that tips the scales for me.
bellowsgulch 21 hours ago [-]
Are people still using deepseek-v4-flash everywhere? I found after the price increases, mimo-v2.5 seems far more attractive.
farlight 21 hours ago [-]
It's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.
bellowsgulch 21 hours ago [-]
Awesome. Thanks for the heads up.
moltar 21 hours ago [-]
I just use Fable for reviews of specs and code then hand off to Opus to work on. Works well.
dude250711 21 hours ago [-]
Does it not silently degrade to Opus if it does not like some word?
lantry 6 hours ago [-]
Users have the option of silent/automatic degradation or a complete halt. I have it set to stop rather than degrade because I want to know when I've hit the safeguard.
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
uejfiweun 15 hours ago [-]
Seeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.
dbbk 7 hours ago [-]
I agree. I've been perfectly happy since Opus 4.6. I remember thinking at the time if it never improved I would have been fine there.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
resters 21 hours ago [-]
over time greater intelligence will be expressed in smaller and cheaper models. we are still somewhat near the beginning of this bc we are finally starting to understand what makes a model truly intelligent/capable.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
dncornholio 9 hours ago [-]
Fable is only marginally better.
sudeepsd__ 9 hours ago [-]
[dead]
hypfer 21 hours ago [-]
[flagged]
dbreunig 21 hours ago [-]
[flagged]
hypfer 21 hours ago [-]
[flagged]
tyre 21 hours ago [-]
I don't think everything has to be Thought Leadership. OP compared the same generation of models to show that the latest open model—at the time of the latest closed model—was Good Enough.
I agree that the opener to their reply wasn't productive, but neither is "Weak."
dbreunig 21 hours ago [-]
I think it’s a fine response when you say, “Doesn't feel well informed enough to give advice,” because I said 5.2
hypfer 21 hours ago [-]
Idk man, but an engineer would've taken that and said something like: "Damn, yeah, good point, I shall add a sentence mentioning 5.3"
Because an engineer feels secure in their knowledge so that such an oversight doesn't make them suddenly defend their identity - it's just an oversight after all. Happens.
simonw 21 hours ago [-]
5.3 isn't available as open weights yet, and only became available via API three days ago. Prior to that the only way to access it was via a Z.ai subscription.
hypfer 21 hours ago [-]
[flagged]
dgellow 21 hours ago [-]
FWIW you’re not looking good in this engagement, feels very childish, looking for a gotcha that doesn’t mean much
hypfer 21 hours ago [-]
I think that depends on the audience. Thank you for caring though :)
kelnos 20 hours ago [-]
Audience member here: I agree with GP; you posts come off as petty and childish.
It seems natural to me to make comparisons only to open weight models where the weights have actually been released.
hypfer 20 hours ago [-]
[flagged]
gpjanik 20 hours ago [-]
"When Moore’s Law slowed in the mid-2000s" it did not, in fact, slow down in the mid 2000s, or at all.
You've selectively quoted the article. The full quote (emphasis added):
"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."
Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
selcuka 17 hours ago [-]
> Your link is talking about transistor count. The article is talking about single-threaded performance.
But Moore's Law has always been about transistor count, not performance.
16 hours ago [-]
gpjanik 9 hours ago [-]
And more cores means what exactly in terms of transistors count?
sscaryterry 20 hours ago [-]
It did in terms of the traditional more MHz (GHz) is better, but as you've correctly pointed out, not when it comes to actual compute.
The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.
Bottom line is the slop used to be manageable, but now there is 100x more code pushed, so that train has departed.
In the end its more bad code for features no one will use.
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
And this is inherent to how LLMs work.
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Now, of course I’d prefer no censoring, but I live in the world we live in.
I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
And there I think the winner would be China.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
And it would show up in real GDP, not necessarily in nominal GDP.
What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
everything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
If that fails, who knows what things will look like.
If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
(I think it would be a good thing for humanity if they did)
- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
How do you square that then? They can do amazing things, but they're also not smart? Do you think its possible to solve Erdos problems without any "smarts"? Can you do it without even understanding mathematics?
I find it very hard to hold the idea that LLMs don't understand anything. They can explain concepts, translate them, simplify them and implement them in code. From the outside, LLMs seem to understands most concepts better than most humans do. Do you understand anything? Couldn't I make the same argument? How would you prove that you understand what a for loop is, or that you know what calculus is? I assume you'd demonstrate your knowledge by using a for loop in a program, or explain calculus back to me. But LLMs can do that too.
> For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them.
Funny example, because lots of human struggle with this too. The number of meetings I've had with people in the US! "Lets meet on thursday morning australia time!". Only, they actually meant thursday night US time, which is friday morning australia time. "Oooh that's so weird! Its the next day for you!". ...... Yes, I know.
I think LLMs are just a different kind of intelligence than humans. They're better at some things than us, and worse than others. They can find latent security vulnerabilities in the linux kernel, but struggle to count the Rs in strawberry. They're not as smart as humans in many ways. But we're not as smart as LLMs in plenty of ways too. I didn't find those linux bugs.
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
i just get fatigued from it, am I holding it wrong or something?
sometimes it's fine but the constant RLHFisms like the constant "worth flagging" and stuff is getting really old
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
[1] https://artificialanalysis.ai/evaluations/omniscience?models...
[2] https://old.reddit.com/r/LocalLLaMA/comments/1vt7l3e/qwen382...
I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold.
Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :)))
Ferrari/Lambo numbers are microscopic
BUT
That's a bad model. My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
I imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still.
even the interesting results in academic work seem to be more of a function of effort (proofs by exhaustion, fitting puzzle pieces in a search space, etc) than anything else. not to say people aren't using large language models to do impressive things, but the agents themselves do not seem very intelligent to me.
it feels like some engineering teams are aware of this fact and are driving agents using strict rule checks (like hooks on steroids), so they can drive some shape of output that aligns with what they require.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
What evidence?
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation) They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
I don't agree it has. And I stand by it.
Model on a custom silicon: https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet
The team that built is working on a better implementation.
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.
But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.
Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
There can only be one fix: send Amodei packing and release unguardrailed models.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ... if it does not go back and notice it has written bad code.
Since they hide thinking and reasoning from the user (who is also paying for those tokens) it is a black box what is triggering it, has the LLM this time thought of "Oh, this has XSS" and used a bad dangerous word such as XSS, while the previous conversation did not ?
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.
I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Really frustrating.
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.
But the code it produced—and the implementations it guided Opus towards—were excellent.
My brother, that's my job.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
I agree that the opener to their reply wasn't productive, but neither is "Weak."
Because an engineer feels secure in their knowledge so that such an oversight doesn't make them suddenly defend their identity - it's just an oversight after all. Happens.
It seems natural to me to make comparisons only to open weight models where the weights have actually been released.
https://ourworldindata.org/data-insights/moores-law-has-accu...
"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."
Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
But Moore's Law has always been about transistor count, not performance.