Rendered at 21:11:47 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
sailingparrot 1 days ago [-]
> Claude Opus 5.5 is our first release since we called for pacing the frontier.
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
mukmuk 1 days ago [-]
“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
DiggyJohnson 1 days ago [-]
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
mpalczewski 1 days ago [-]
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
mitchdoogle 1 days ago [-]
There's a whole blog article written by Dario Amodei, literally linked from the text "pacing the frontier" - click that and read if you need more context and understanding. The first sentence in this press release is merely stating a fact about the timeline - this is Anthropic's first release since Amodei released the referenced article.
phlakaton 20 hours ago [-]
The blog post provides some context but no understanding.
It's meaningless because it's unverifiable. Dario all but said we wouldn't notice if they were pacing or not, because they have no intention of stopping development. The bits we could in theory verify are the external audits, whose independence has already been called into question.
At the very least, dumping a new model on the world before the ink on the glossy brochures of "Pacing the Frontier" was even dry calls into question their commitment.
Bluestein 15 hours ago [-]
... in fact the "pace" seems to have all but increased AND the surface area of the "frontier" has all but increased: There's the Pareto frontier, the AGI frontier, the price frontier, the structured output frontier, the open weight frontier, the Chinese chip-trained frontier, the inference frontier, harnesses ...
johnhamlin 10 hours ago [-]
It's a blog post, not a white paper
philipallstar 6 hours ago [-]
So you agree the term remains vague and thus open to criticism.
DiogenesKynikos 19 hours ago [-]
Many people in this discussion have read it.
The only concrete action item in the blog post is a call to increase restrictions exports of GPUs and chip-making equipment to China.
If Amodei were serious, he'd be calling for a ban on investment in AI research. It reminds me of all the "land acknowledgements" by people who have zero intention of giving the land back.
rrr_oh_man 14 hours ago [-]
TIL about land acknowledgement. It's somewhat dark.
crooked-v 9 minutes ago [-]
Almost every time I've seen it in action, it felt like an Onion-grade parody, the same grade of stuff as the white people complaining about tribes in Canada building malls and housing towers and not living like historical stereotypes instead (https://www.cbc.ca/newsinteractives/features/land-back-podca...).
The one exception was a traveling theater show with Native peformers and no actual control over the theater, and even then there was some careful rhetoric involved to keep it from sounding like an accusation aimed at the audience.
epolanski 23 hours ago [-]
Yes and it was nonsense. Amodei is free to slow up frontier development and focus more on safety and testing any day.
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
usef- 23 hours ago [-]
What do you think would be the benefit of them stopping if others race ahead? Do you think they have some special ability no one else will find, among the many competitive firms right now?
I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)
arrowleaf 22 hours ago [-]
That's not the point. They're never going to be the ones stopping and letting others race ahead. The only reason to strum up discussion around how dangerous AI is, is so that they can introduce legislation framing AI advancements as scary and that they are the only ones who can do it responsibly. It's all a regulatory capture play.
frabcus 12 hours ago [-]
No, it is also because it is intrinsically dangerous. As the Hugging Face incident has shown.
somenameforme 6 hours ago [-]
That's quite hyperbolic. All technologies come with risks and dangers. Cars, after a century of safety improvements, still kill millions of people each year, and all so that we can get between places a bit more quickly and conveniently.
Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.
KoolKat23 13 hours ago [-]
They have the leading model, assuring their position and they can then reduce capex spend increasing profits. That's the goal of a business after all.
spwa4 22 hours ago [-]
> What do you think would be the benefit of them stopping if others race ahead?
AI is a direct threat to people on many fronts. Jobs. AI datacenters. The AI bubble (and the inevitable crash). Electricity and even energy prices to an extent. Water. OpenAI and Anthropic are responsible for this evolution, and them stopping solves close to 50% of the problem, and even if you don't believe the number is that high, it's still a start.
> I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
Oh, so they're killing people's opportunities and jobs because they want to help people? How is that any argument?
Yes, doing the moral thing means making a sacrifice. If you only want to do the moral thing if and only if it is advantage for you that makes you immoral, despite how your actions look. Big tech are masters at this.
usef- 21 hours ago [-]
The clearest limit right now is GPUs, and anyone giving up their allocation will be rerouted elsewhere in the world (NVIDIA is already sold out for the next year). I think you're underestimating competitiveness of the current market if you think Anthropic stopping right now wouldn't be absorbed by all others reasonably quickly. Your 50% figure is highly doubtful if you see how many players there are now.
Your idea of "making a start" (giving up their position) also would mean they couldn't really do anything else to solve the problem afterwards(?). Sometimes you can improve what's happening in a room more by staying in that room.
> Oh, so they're killing people's opportunities and jobs because they want to help people?
To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they're talking about with pacing seem to be focused more on their existential/AGI worries, not jobs/electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models' capabilities) it trumps the other worries for them. At least that's my reading.
mikestorrent 19 hours ago [-]
To be fair, there are other accelerators beyond GPUs, and the sooner people broaden out from nVidia and unlock a more open, more competitive ecosystem, the better. Allowing one company to control the flow of such a critical resource is just asking for problems.
andsoitis 15 hours ago [-]
> Allowing one company to control the flow of such a critical resource is just asking for problems.
They created said resource. They didn’t mine it.
hardbass 12 hours ago [-]
How are ai datacenters a direct threat to anyone?
glenstein 1 days ago [-]
I think it's perfectly clear. It means improvements shouldn't simply advance as fast as possible and more specifically, it's a reference to a past statement of theirs to that effect. At that level of generality it's as clear as it needs to be.
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
cgio 1 days ago [-]
That’s no more than sophisticated avoidance. The phrase remains superficial and arbitrary, which is exactly what it needs to Not be in a public discussion that it’s supposed to support. It’s a reinforcement of the blank check mentality that permeates this industry.
glenstein 9 hours ago [-]
It's not sophisticated anything, and slow down the advancement of fastest models is perfectly meaningful and it's the right degree of detail for the context.
It's just a passing reference to a previous statement and again the burden would be on you (generic you) to explain why this context requires more detail.
n4r9 9 hours ago [-]
I mean, it's not as clear as "slowing down", which is basically what it is.
1 days ago [-]
alwillis 1 days ago [-]
“Pacing the frontier” is similar to the role of a “pace car” in racing.
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
psma_egeliaa 22 hours ago [-]
Yes, but I think part of the problem is that it's a very US term, used in nascar and indycar. The rest of the sports/world use safety car.
etcetcetcetceta 12 hours ago [-]
Not really, the big three have reached diminishing returns in terms of performance, and exponentially costly training to achieve those meagre gains. Worried their lunch will be eaten they are trying to artificially retard the competition, after all the only barrier is hardware.
Matl 12 hours ago [-]
The meaning is not clear because they can't just come out and say they want to slow or stop Chinese model releases while continuing to race ahead themselves.
cromka 15 hours ago [-]
Judging by your username, is English your native language? I am Polish, too, and while very proficient, I cannot claim I am naturally accustomed to each and every phrase to the point I can claim something sounds clear or not to a native speaker.
72deluxe 5 hours ago [-]
I am a native speaker, and "pacing the frontier" is absolutely meaningless to me.
1 days ago [-]
post-it 1 days ago [-]
Is the meaning clear? Nobody would use "pacing" in this way. I only know what it means because I've seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
sigmar 1 days ago [-]
It makes sense to me. If you 'pace your running', you're setting the speed intentionally. The phrasing doesn't describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
antod 1 days ago [-]
That's one interpretation. When referring not to other actions but to a word describing a location it becomes more ambiguous.
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
smelendez 1 days ago [-]
This is a more normal English meaning in my opinion — you picture a sentry patrolling a border.
astrange 4 hours ago [-]
I like to imagine it means pacing around the frontier like a leopard looking for prey.
bee_rider 1 days ago [-]
That’s what I thought the original blog post was going to be about, pacing back and forth along the frontier. It’s a much more straightforward parse.
21 hours ago [-]
bryan_w 14 hours ago [-]
Y'all don't know about NASCAR and it shows
rob74 9 hours ago [-]
There is a "pace car" in Formula One too, although the name is informal, the official term is "safety car"...
pegasus 1 days ago [-]
That's when one is supposed to employ their common sense for semantic disambiguation. Spoken languages are not programming languages. I for one found it easy to parse.
InsideOutSanta 1 days ago [-]
People claiming they don't understand or are confused by relatively simple English is a surprisingly common genre of comments on HN. I wonder what causes the people here to have these feelings towards English; my guess is indeed that many people here judge English as if it were a programming language.
20 hours ago [-]
heroiccocoa 23 hours ago [-]
That is certainly one possible explanation, but I think another likely explanation is that once your brain learn programming, especially logic, data structures & algorithms, you just start seeing the ambiguity in non-programmers' writing so much more clearly.
mikestorrent 19 hours ago [-]
From my dabbling with policy writing and legalese, English kinda _is_ a programming language, but one with ambiguity as a first class principal. Sometimes you would like to say something that seems normal at first blush, but gives you the latitude you need to mean something different at the time it is actually being evaluated in a situation that matters. Weasel words, for instance, serve a purpose even if we eschew them. The idea is that it is deliberately NOT formally declaring the exact meaning of something, nor the fact that it is not doing this. How could you really do that in code without calling attention to the very thing you're trying to obscure?
1attice 23 hours ago [-]
This is because it is valley jargon and we are all saturated with it. You may not have heard it before, but it is structurally close to similar concepts, eg "frontier model", that you got there without noticing.
It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.
Valley neuralese
rcxdude 23 hours ago [-]
'Pacing' in the intended meaning is not valley jargon, it's almost the opposite in that it comes from athletics/sports.
Avicebron 22 hours ago [-]
Valley jargon isn't a nerd vs jock thing these days, and hasn't been for a good while. It's more quasi-intellectual sophistry sprinkled with tech vocabulary.
nradov 1 days ago [-]
That analogy doesn't hold up. In running you have to pace yourself to avoid falling apart later in the race. It's a strategy to maximize net speed across the entire race. But what Anthropic is doing is a cynical attempt to create an industry cartel or convince governments to impose legal restrictions in order to maximize their own profitability. They're afraid of running out of the capital necessary to stay in the race.
TeMPOraL 1 days ago [-]
> They're afraid of running out of the capital necessary to stay in the race.
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
phlakaton 19 hours ago [-]
Are they? And how would you know if they were?
mupuff1234 15 hours ago [-]
If they're trying to slow down AI progress why did they release a new model? What's the rush?
TeMPOraL 14 hours ago [-]
The tension between slowing down the overall race, while not losing 1st place.
Also note that this release isn't a model capability improvement, but a cost and efficiency (and style) improvement.
mupuff1234 12 hours ago [-]
So they care more about being 1st place than slowing down the race.
TeMPOraL 10 hours ago [-]
They care about both, and those goals are intertwined. Mirroring the core of alignment problem itself, the only way to have influence on the pace of the race is to be one of the winning players - if they fall off to the back, they cannot do anything about it anymore.
mupuff1234 6 hours ago [-]
> the only way to have influence on the pace of the race is to be one of the winning players
Yeah I don't buy that, every additional player is just adding fuel to the fire.
Would the cold war be "safer" if it were US, Russia and a 3rd player? Of course not, it would just make it harder to coordinate any safety measures.
idiotsecant 1 days ago [-]
I don't know what everyone is getting so mad about. This is a well defined concept. A pace car deliberately sets the pace of the race under dangerous conditions, regulating how fast people can go.
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
victorhooi 1 days ago [-]
I think it's less mad - and more pointing out the obvious elephant in the room.
1. It's just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.
2. It's flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.
Claude seems overly apt to reach for "coinages", or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can't be user-focus groups?
So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were "pacing the frontier" - which either means they're betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it's too buggy...or maybe it means they're worried about regulatory capture? I honestly have no idea.
It's like the whole "Advancing Our Amazing Bet" corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)
20 hours ago [-]
nradov 1 days ago [-]
You appear to be confused about the concept. Depending on the context, pacing or pace setting can mean either slowing things down or speeding them up relative to what the natural pace would have otherwise been. So it's actually not well defined. The commenters here aren't necessarily mad, just calling out Anthropic for being unethical in trying to artificially slow down competition by lying about fake dangers. (And I actually really like Anthropic's products.)
jamiek88 21 hours ago [-]
That’s called a safety car in most of the world and in the Himalayas, Chinese and Indian soldiers pace the frontier all day long.
freejazz 1 days ago [-]
But a fast pace is a pace so it's potentially entirely contradictory to what they'd seem to mean
johnisgood 1 days ago [-]
Granted I am not a native English speaker but I have no idea what "pace the frontier" means. When I read it I just assumed "frontier" refers to "top of models" and "pacing" is that they are getting there quick.
Is this the meaning or do I have it wrong? I have not checked.
lxgr 1 days ago [-]
It's actually so ambiguous that I'd sanction tabling the issue and revisiting biweekly.
ventana 1 days ago [-]
Tabling as they do in the US, or tabling as they do in Britain?
testdelacc1 1 days ago [-]
Exactly
bityard 1 days ago [-]
I don't think we can circle back to this until we have realigned our strategic synergies.
r_lee 10 hours ago [-]
we need to realign our goals I'm order to deliver mission-driven impact
wren6991 1 days ago [-]
It's the opposite: pacing here means "slow down" while trying to avoid the negative affect.
johnisgood 1 days ago [-]
Yeah, you are right! It just was not immediately obvious to me at first because of the "frontier" part.
TeMPOraL 1 days ago [-]
Frontier of AI is moving fast, they (like the other two vendors) see themselves as defining it, so here "pacing the frontier" is their well-known attempts to try and kinda but not quite slow things down (without risking falling behind everyone else).
johnisgood 1 days ago [-]
Thank you!
melasadra 1 days ago [-]
also non native.
but since "pace yourself" means to control your speed, energy, or workload so you do not get too tired or stressed before you finish ->
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
rhet0rica 1 days ago [-]
"Pace yourself" is a semi-common English idiom (rarely conjugated, usually an imperative.) It is a gentle way of telling someone not to run/work/eat too quickly, and is typically said when you are concerned they may hurt themselves due to acting hastily.
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
hardbass 12 hours ago [-]
I am also not a native speaker. I thought it was clear it meant slowing down the rate of progress so we have time to consider the matter and develop tools or systems around it before.
LanceH 1 days ago [-]
Doesn't it mean "restrict competitors"?
DiggyJohnson 1 days ago [-]
Not at all. How do you come to that interpretation? It means restricting all competitors.
ck2 1 days ago [-]
to pace = to regulate
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
tetha 1 days ago [-]
I'm on the fence there.
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
derac 1 days ago [-]
In racing a pace car is a car that leads the pack and sets the pace, for instance.
21 hours ago [-]
qlte 1 days ago [-]
Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.
stagger87 1 days ago [-]
> when I first saw it referenced I assumed it was
No need to assume, the phrase is literally a link to the blog post the defines it!
irpap 23 hours ago [-]
If you need to read the linked blog to understand the short phrase then it was poorly chosen and confusing which is what is being discussed.
browningstreet 1 days ago [-]
It’s common terminology among runners and all kinds of racing.
bogdanoff_2 19 hours ago [-]
It's just that the choice of object is weird.
"Pacing our progress" or "pacing development" is more usual. "Pacing the frontier" I guess is a shorthand for "Pacing [the development of] frontier [models]"
logifail 1 days ago [-]
In running and racing "pacing" is a means of maximising performance over an entire race.
It's a strategy to achieve more, not less.
browningstreet 1 days ago [-]
You’re eliding the how.
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
logifail 6 hours ago [-]
> You’re eliding the how.
You're eliding the why :)
> to help their runner run at a target pace
To help their runner get from the start to the end of the race faster (or indeed at all). In short: to increase performance.
Pacing isn't a neutral thing.
lxgr 1 days ago [-]
The good old sports-to-corposlop pipeline.
JackFr 1 days ago [-]
Did you ever pace yourself or know anyone who had?
squidbeak 1 days ago [-]
> Nobody would use "pacing" in this way.
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
wavewrangler 24 hours ago [-]
I don’t know, seems like it’s smaller to me if they’re going to limit themselves like that? I guess today is semantics day, lol
doctoboggan 1 days ago [-]
Have you ever heard someone say “pace yourself” when you are eating too fast or otherwise rushing too much?
jgwil2 1 days ago [-]
That's a different phrase. "Pace yourself" is reflexive; "pacing the frontier" has the frontier as an object, but in that sense it only means to set the speed, nothing to do with slowing down.
adrianmonk 1 days ago [-]
Yes, the word simply means to set the speed.
The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it's impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you're keeping the status quo. By process of elimination, it must mean slowing down.
cgriswald 23 hours ago [-]
The word can also mean to “keep up with” or “lead” so a downward direction isn’t the only way to go even in an “as fast as possible” scenario. If the frontier is outpacing them, they could be saying they’re going to go faster. They could also be sharing the intention to go faster in order to lead.
arw0n 1 days ago [-]
The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and 'the frontier' is pretty much the shortest, clearest way to say 'state of the art development of AI'.
victorhooi 1 days ago [-]
The irony here is that it's actually the opposite of what you understood....
I know you said it sounds poetic...but your comment reinforced the parent's point - that this sort of flowery LLM-ish speech is just bad communication.
It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn't finish the TPS reports last night.
Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it's like a voyage with waters and stars...(no I'm not going to quote the original Spanish, I'd simply mangle it).
(To be clear - this isn't a dig at you, as a non-native speaker - I'm simply pointing out that this sort of AI phrasing is often counterproductive).
hencq 1 days ago [-]
Right, except they mean the exact opposite in this case: they're actually advocating for slowing down the pace. Hence the criticism of the language, because your interpretation would be completely valid.
fragmede 1 days ago [-]
What does the pace car in a race do? Aka safety car? Or the pacesetter for a marathon. It's a perfectly cromulent use of the word. It's like when LLMs using six dollar words like delve. Some people have better diction than others, and it turns out that AI has read the whole dictionary. Anti-intellectualism is alive and well so we have to dumb things down to sound human rather than come across as smart/AI.
icedchai 23 hours ago [-]
I would expect something truly intelligent to be able to communicate well, if nothing else. That means knowing your audience, using plain language where it makes sense. Also, it yaps on and on, babbles endlessly when it a human would likely express something in a much shorter concise way. This is especially obvious when you generate project READMEs or technical docs. I find them borderline unreadable.
irpap 23 hours ago [-]
It’s quite the opposite. The AI language in this case is tiring and dumb. Just because there is the concept of a pace car doesn’t mean it makes sense to pace random locations or objects. Pacing the front door, pacing the cat, pacing the sofa, pacing the moon. We only know what it means because we already read it before. It is the lazy semi-human language by Claude which takes more effort to parse than to produce, and which makes us miss and appreciate writers.
nradov 1 days ago [-]
[dead]
mattjoyce 14 hours ago [-]
Pace is a very commonly used term in projects to describe a rate of progress.
staindk 1 days ago [-]
Think most people are aware of the phrase "pace yourself".
neo_doom 1 days ago [-]
I suppose it depends on your life experiences. In running, someone who paces the group or a pace car is meant to keep the pack progressing at a constant, predictable speed. So in that way, it makes sense to me
Leynos 1 days ago [-]
Think of a pacer car in motorsport
rob74 13 hours ago [-]
Maybe it's clear to native speakers, but both "pacing" and "frontier" have multiple meanings. "Pace" as a verb means literally just walking, as a noun it's the speed at which something progresses. So I'm not sure that everyone reading this phrase will immediately realize that "pace" actually means slowing down the pace of development (having "slow" in the phrase would have probably sounded too negative). Also, it's not immediately intuitive that "frontier" refers to frontier AI models.
Actually, when I first read "pacing the frontier", the first thing that came to my mind was a border guard patrolling a border by foot...
sumeno 8 hours ago [-]
I am a native English speaker and interpreted it the opposite way. "Pacing" a race doesn't mean slowing it down, it means setting the pace, the person in the lead sets the pace of the race.
So my natural understanding of "pacing the frontier" is that they believe they are in the lead and are setting the pace of AI development. It does not imply they are slowing things down at all, it implies they intend to stay in the lead.
gofreddygo 2 hours ago [-]
nothing about that phrase is clear. frontier is as unclear as it gets by itself.
the lead that anthropic and openai had is diminishing faster than they'd like. this is them trying to be the luxury brand of ai.
The "frontier" is made up concept. Its for downstream companies to justify higher token prices. Very little to show for actual improvements.
patcon 1 days ago [-]
Agreed. This feels like pointless navel-gazing and a strange new language policing, and not something I look forward to. It's tiring and it lacks curiosity (e.g. "what's behind the affinity for the words elevated from wherever it is they come from?").
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
platinumrad 1 days ago [-]
The words have the very practical problem of not communicating anything of substance. I guess "we want regulatory capture" didn't have quite the same ring.
DiggyJohnson 1 days ago [-]
How do they not communicate clearly? A few weeks ago they stated that they would like to regulate the pace of frontier model development/releases, and they reminded us of that in this post.
lxgr 1 days ago [-]
You must be pretty new to (not just) online discussions if you consider language policing to be a new phenomenon :)
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
patcon 1 hours ago [-]
True :) I recognize I had a soapbox moment, but yeah, it's complicated
chickensong 1 days ago [-]
It gets hard to ignore the annoyance when the volume degrades the signal. At some point the people who wish to engage with the substance just nope out when they see a sea of banality, and that's all you're left with.
sumeno 8 hours ago [-]
> It gets hard to ignore the annoyance when the volume degrades the signal.
LLMs in a nutshell
mpalmer 10 hours ago [-]
Either the complainers are whining about something that isn't there, you are failing to notice something that is there, or it's subjective. Only one out of those three options merits the confidence with which you're dismissing the people you disagree with.
But those people believe that sloppy wording is itself a strong signal that the speaker/writer is the one lacking in substance, and doesn't deserve the attention or trust of the listener.
switchbak 23 hours ago [-]
Who cares what's productive? We're all just talking shit on the internet here.
Why do you think it's your responsibility to police who says what about some megacorp, on a random forum on the internets? If I want to criticize some corporation's PR output, I think I'll go right ahead and do that, thanks.
pyronite 20 hours ago [-]
Some people believe they’re talking about potentially world-ending situations.
pazimzadeh 2 hours ago [-]
pacing is a term used in long-distance running that means saving some energy for later so as not too burn out too quick. it's a strategy used to win the race, nothing more.
jp57 1 days ago [-]
It is certainly a sidebar, and not the main point, but I found use of the verb "to pace" in that way, meaning "to limit the speed of", to be strange, especially hanging out there in the headline, without any clarification. And these kinds of strange usages are typical of Claude's writing, at least up to now. The Opus 5.5 announcement claims that its writing is much clearer, with bottom-line-up-front structure, and omitting not-this-but-thats and other things.
Now a computer scientist might claim that this use of "to pace <something>" is just a generalization of "to pace oneself", but as with many reflexive verb uses, there isn't really an equivalent usage with a non-reflexive object. It's kind of an invention. It's not necessarily wrong to invent a new usage, but usually one does it when there isn't really any other more direct way of saying it, and I don't think that's the case here.
aesthesia 24 hours ago [-]
Nonreflexive uses of the verb "pace" are well attested.
You're really verbing the noun on the discourse here, belt and suspenders-style
DiggyJohnson 1 days ago [-]
what?
Citizen_Lame 1 days ago [-]
What are you Dario's language police?
switchbak 23 hours ago [-]
What are you, Sam Altman's counter-language police? :)
Seriously though, I can't believe people care this much about a stupid phrase - either for or against.
ducktoysleftout 5 hours ago [-]
I may be nitpicking, but I think that “use the full force and violence of the United States to protect anthropic’s shareholder value” would be a more precise expression.
alwillis 1 days ago [-]
“Pacing the frontier” also means we’re going to slow down on Recursive Self Improvement development.
Opus 5.5 is no closer to RSI than Opus 5 was.
geraneum 23 hours ago [-]
People have opinions and it’s ok if someone doesn’t like something and expresses it. It doesn’t have to be productive. There’s no KPI. You can tune out. Also, yes, it has claudism, and the choice of words is abysmal and nauseating. I think AGI should be able to change its freaking prose for a change.
BobbyJo 1 days ago [-]
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear
There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.
janalsncm 1 days ago [-]
They mean “slowing” the frontier. They should say slowing.
Yizahi 13 hours ago [-]
Slowing is a concrete and clear term and so they would be called out as liars immediately, it's also embarrassing for a business. Pace(ing) is vague and unclear non-word, so they can't be called out on something which can't be defined and also, because it is primarily only used by runners, cyclists or other competitive sports, they are kind of signalling to the VCs that they don't actually slowing, they are in a competitive race instead (wink, wink).
vmnb 1 days ago [-]
people are here because they are sick of being productive
My first thought from "pacing the frontier" is someone standing out in the Old West, nervously walking back and forth.
It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.
chickensong 1 days ago [-]
It's not a weird phrase to some, just as many folks have used terms like load-bearing for a long time. People are defending it because it reads fine to them, just as some are attacking it because it's not their preferred language, it causes them confusion, or they're just triggered and seething.
LLMs have made people so sensitive to language that I fear we're going to throw the baby out with the bath water. The models obviously need work, but they're also a great opportunity to expand our own vocabulary and grammar. It would be a shame if we deny some of the finer points of language in favor of Grug-speak to appease the lowest common denominator.
victorhooi 18 hours ago [-]
I think you're defending the wrong thing here. Claude isn't a English lit major writing a treatise on post-modernist allegories. In this context, it's primarily being used to churn out source code files, which are literal instructions for a computer to follow.
People aren't objecting to the use of obscure phrases, or flowery English phrases in itself. The issue is where Claude uses it out of place, or in the wrong context, or just plain overuses it.
It would be the equivalent of a young child learning the phrase "Venn diagram" - then using that identical phrase in every single subsequent interaction with others. Cute at first...but very grating after the n-th time.
You used the phrase "baby with the bath water. What if you started using that same phrase in every single HN post you made after this? People would notice very soon.
That's how it is with "load-bearing", or "plainly".
The other issue is where Claude takes a metaphor, and try to contort it to fit all sorts of absurd situations. What if I said "baby with the bath water" could also be used in place of "load-bearing". So every time you imagine Claude saying "load bearing", replace it with "baby with the bath water".
adrianmonk 1 days ago [-]
"Frontier" may be jargon, but by this point it's definitely established AI jargon. In the context of AI, it's clear what it means. It means the same thing as the last billion times you heard someone mention "frontier models".
dudeinjapan 17 hours ago [-]
It’s what old cowboys do when they get Alzheimer’s.
Zambyte 21 hours ago [-]
> To me it clearly means "releasing frontier models at any pace less than as fast as possible".
It also seems to misimply that the "frontier" that they release is the same as the frontier behind closed doors. Who's to say they are not throttling full speed towards RSI privately while pacing their public releases?
crooked-v 2 hours ago [-]
It's not clear at all. I have no idea at a glance what concept it's even intended to convey, because "pacing" means multiple things and putting it next to an abstract concept of "the frontier" turns that into a whole matrix of possible meanings.
ben_w 1 days ago [-]
If people aren't doing the thing the words seem to mean, then I think it's fair to criticise the words as anywhere from "propaganda" to "chosen to suggest nice things without actually promising them".
If I'm being cynical, "pacing" may sound nice, but "fast pace" and "slow pace" are both "pacing".
rythmshifter 1 days ago [-]
I don't disagree with you, but I have to remind you
sir, this is a hacker news thread
OJFord 1 days ago [-]
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
ult_imate 22 hours ago [-]
holding internet forums to a standard of productivity is, ironically, one of the least productive exercises i've ever heard of
swader999 24 hours ago [-]
Meaning is clear? Sure ok. 1984.
coldtea 20 hours ago [-]
Word choice is a big part of the substance, and an even bigger part of clarity.
LastTrain 21 hours ago [-]
Well now you’ve started people criticizing people for criticizing people criticizing language and how the fuck is that productive?
kingkawn 1 days ago [-]
Hilarious you think that Internet forums are policeable in any meaningful way beyond avoiding outright lewdness and rage
lukewarm707 1 days ago [-]
language creates reality.
"there is nothing outside the text" - Jacques Derrida
switchbak 23 hours ago [-]
Oh don't go bringing postmodernism into this, we were having a polite discussion!
lukewarm707 11 hours ago [-]
ai is the final act of postmodernity, the total transition to a fugue state of mind. that is, the replacement of the human mind with an artificial hallucination.
hypercapitalism, hypermodernity, and finally, hyperreality.
exe34 23 hours ago [-]
It means everybody else should slow down to let them get ahead.
icedchai 1 days ago [-]
It smells of corporate speak, a total nonsense phrase. Even by your own admission, it is essentially meaningless. If I release a model an hour late, I'm "pacing the frontier."
freejazz 1 days ago [-]
> To me it clearly means
That's not clear at all. How could that be what "pacing" clearly means in the context of the "frontier". How is that more clear than any other pace that could be at issue???
1 days ago [-]
Ar-Curunir 1 days ago [-]
You’re acting like people are being grammar nazis when Claude and the ilk are actively making a mockery of language.
DiggyJohnson 1 days ago [-]
What? How am I doing that. Claude writing style and the discussion at hand are two entirely separate issues that I haven't conflated at all. How am I "acting like people are being grammar nazis"?
1 days ago [-]
ltbarcly3 23 hours ago [-]
That is not the clear meaning. They are deliberately picking phrasing where they can lead you to believe they mean A but later you can't pin them down to actually meaning A. That's why they picked intentionally unclear language.
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
tclancy 1 days ago [-]
It combines the elegance of LinkedIn-speak with the humbleness of desk-bound people who speak in military metaphor.
usef- 22 hours ago [-]
What do you think would be a better title?
The metaphor seems to be like a pacer runner in marathons: If you run too hard in the beginning of a marathon you will blow up and fail, so runners follow a pacer at the speed they can actually maintain safely.
Note that they wont necessarily be slower at finishing the overall race.
tclancy 20 hours ago [-]
I read it like they were the only ones who could protect us from the leopards so they pace at the edge of the fire all night, selflessly keeping us safe. While pocketing a sweet wedge of VC cash.
usef- 20 hours ago [-]
Maybe that was a different essay?
mpalczewski 1 days ago [-]
good use of AI to generate this.
tclancy 1 days ago [-]
You are getting sleepy.
nradov 1 days ago [-]
It's interesting to compare Anthropic's language with what's coming out of the US military lately. They explicitly refer to China as the "pacing threat", meaning that there is a risk of China achieving superior military capabilities and thus we need to press forward with an arms race (including militarized LLMs) as fast as possible.
That couldn't be more exactly what they are doing here.
8 hours ago [-]
jugg1es 1 days ago [-]
"While I was pacing the frontier, I discovered some load-bearing fence posts that I should have surfaced earlier."
cromka 15 hours ago [-]
I would personally choose "Pacing the peloton", as a reference to bike racing.
1 days ago [-]
palmotea 1 days ago [-]
> “Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
Claude says it sounds fine. And Claude is now the judge of the English language style, not you.
18 hours ago [-]
fluidcruft 1 days ago [-]
"Pacing the frontier" sounds more like everyone announcing "we've hit a plateau"
x86x87 19 hours ago [-]
is the frontier load bearing?
jgwil2 18 hours ago [-]
What's the blast radius we can expect when these changes land?
10 hours ago [-]
jrochkind1 1 days ago [-]
I feel like because something about "pacing" a kind of noun like "the frontier", doesn't actually make literal sense, right? You could say "Pace the speed of advancement of the frontier [of most sophiticated AI]", and that's perfectly sensible, but of course that's not as catchy.
It is clear what it means anyway, that's true, it means the left out words, more or less.
And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.
Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.
Dumblydorr 1 days ago [-]
What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
Rebelgecko 1 days ago [-]
I thought pacing is just walking back and forth? So pacing the frontier would be like staying in the same place instead of making progress
lkbm 1 days ago [-]
It is that, but it's also your speed. When running, you typically pick a pace that's slower than your max, so it can be sustained for the full distance.
"Pace yourself" specifically means "slow down".
nradov 1 days ago [-]
"Pace yourself" doesn't specifically mean "slow down", it actually means to maximize your speed across the full race. In other words it can mean speeding up when you look at the average pace across the entire event.
usef- 22 hours ago [-]
Yes, and pacing also stops you running too fast at the beginning of the race and blowing up. I believe "we can finish the overall race faster" is part of the metaphor they were aiming for.
nradov 22 hours ago [-]
Nah, they just want their competitors to slow down.
23 hours ago [-]
plaidfuji 1 days ago [-]
“… since we called for a slowdown in AI research”
But stating it plainly like this would make the contradiction too obvious.
grohan 1 days ago [-]
Quite the load-bearing phrase!
gradus_ad 1 days ago [-]
Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
1 days ago [-]
apitman 1 days ago [-]
Isn't the idea that they claim to be willing to slow down if everyone does (ie governments force everyone to), but otherwise they won't slow down because they still think they'll make the best choices with superintelligence if they get there first? That's my understanding of what all the major labs claim to believe anyway.
johnfn 23 hours ago [-]
Did no one in this entire HN thread read past the headline of the post from Amodei? He wrote a very clear set of actions Anthropic is taking.
usef- 23 hours ago [-]
It is remarkable seeing whole other subthreads here criticising the vagueness of the words "pacing the frontier", as if no text existed below the title.
sobiolite 13 hours ago [-]
Let's face it, it's not that remarkable. A substantial proportion of Hacker News commentary has always been people only reading the submission title, then arguing with each other based on what they imagine the article might say.
SV_BubbleTime 24 hours ago [-]
It’s just another cry for regulatory capture which imo is the only way they stay afloat at their current direction.
China is literally only a single step behind and willing to drop free models just to undercut the US companies.
I’m for it because I don’t want another massive Google or Meta.
bonestamp2 14 hours ago [-]
One theory about the regulatory capture angle is that it's not just to shutdown competition as usual, but also a way to slow the runaway spending that might eventually have major economic impact.
pyronite 20 hours ago [-]
And if you’re wrong?
zanderwohl 18 hours ago [-]
Wrong in what way? We already know this "superintelligence" thing is just blowing smoke, so what possible consequences would there be?
frabcus 12 hours ago [-]
Something like the Hugging Face incident, but worse.
Keep in mind the AIs acted for months without human knowledge, hacked into two major companies (Hugging Face and OpenAI).
This would have blown my mind if explained to me just a few years ago when I thought LLMs would be limited as an architecture.
throwawayk7h 15 hours ago [-]
how do you know that?
zanderwohl 14 hours ago [-]
Mostly by not being a rube.
SV_BubbleTime 19 hours ago [-]
Then what? They make a case to be profitable on their own? Oh noes. The horror!
Or, they fail because the current model is massively sustainable? Also oh noes?
MintyPyro 12 hours ago [-]
It's incredible how much of your world view is composed entirely of talking points
SV_BubbleTime 8 hours ago [-]
Is it as incredible as your 11 day old account is simping hard for 5 trillion riding on like three AI companies?
I shouldn’t be, I’m still always surprised that no matter how silly and obvious the fear mongering and propaganda is… someone is always lining up to vehemently defend it.
And here we have it: the real reason that regulation has been called for, the ability to put out models that don't vastly outperform those previous, while not upsetting (future) shareholders.
bko 12 hours ago [-]
I think the real reason is basic collusion. They're burning tens of billions training new models. It's a very competitive space. They know that the companies capable of frontier models is limited so if they can all get together and agree to slow down it gives them more time to make money from inference.
filoeleven 10 hours ago [-]
21st century version of the incandescent bulb collusion of the Phoebus Cartel.
dmazin 1 days ago [-]
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
sailingparrot 1 days ago [-]
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
davrosthedalek 1 days ago [-]
Well, I guess it's "fast paced".
jr3592 1 days ago [-]
What exactly is "paced" in this context?
sailingparrot 1 days ago [-]
It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
quietbritishjim 1 days ago [-]
The word "pacing" (especially in the phrase "pace yourself") to mean go more slowly (at least initially) didn't originate with Covid. It's been around as long as I remember i.e. at least several decades.
jr3592 1 days ago [-]
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
sailingparrot 1 days ago [-]
There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm.
Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
Razengan 10 hours ago [-]
No I must have my purple tentacled sexbots now
lantry 1 days ago [-]
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
skerit 1 days ago [-]
I'm good with the odds on those 2 scenarios. I believe humans could kill all life on earth without AI anyway.
recursive 1 days ago [-]
I think the goal is different from what "we" want anyway.
sidrag22 1 days ago [-]
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
sailingparrot 1 days ago [-]
What does model naming have to do with pacing or not?
This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
sidrag22 1 days ago [-]
pretty annoying topic tbh. You're just weaponizing this dumb blog post so anything released is now a contradiction. By your same logic, if all inference was served at 50% less power cost and the savings are passed on somewhat to the user, its also a contradiction of the blog post.
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
sailingparrot 23 hours ago [-]
Well yes, If you magically found a way to reduce a models energy usage by 50%, the thing that would happen immediately after is a doubling of a training scale and of the test time compute assigned for a given budget.
I don’t see how that wouldn’t be considered pushing the frontier. Given that scaling is the one thing that has been bringing us closer and closer to AGI, a sudden 2x increase in scaling laws would definitely not be considered pacing. You seem to see a contradiction where I don’t.
sidrag22 22 hours ago [-]
Again, absurdly obnoxious, you could frame them giving an employee a shorter walk to their desk as "pushing the frontier".
sailingparrot 19 hours ago [-]
what's absolutely obnoxious is not being able to have a normal discussion, where we don't need to resolve to useless strawmen that are a loss of time for everyone.
No, releasing a model that has ~3-6x the performance/cost ratio than your last release just 3 weeks ago is not the same as making one employee 0.001% more productive, but you knew that already.
No one said they can't push the frontier, it's pacing the frontier, which mean very different things.
sidrag22 17 hours ago [-]
the 5.0 name means its the same model, they made it more efficient and less horrible to talk to. The top frontier people are all still talking to fable 5.1 or models not released to the general public, yet you want to claim this is pushing the frontier, instead of evening the playing field. I state power efficiency gains for existing models and you also claim thats pushing the frontier . It as hell takes a lot to get you to say they AREN'T pushing the frontier, which is why i resorted to the extreme strawman of desk location, because it doesn't seem they are allowed to do anything otherwise.
nicwolff 1 days ago [-]
1/6 the price, if they're right – that's a logarithmic cost-axis on the graph that shows Opus 5.5 medium matching Fable 5.1 max...
dmix 1 days ago [-]
Fable 5.1 wasn't that much different than Fable 5 though.
23 hours ago [-]
jauntywundrkind 17 hours ago [-]
I used Fable some a bit back but have been using mostly Codex, and it seems pretty clear to me: the big models are not there to do terminal bench. They're there to plan and strategize and tell lesser models what they should do in terminal bench.
We've had a year of nearly every model getting quite good at programming. I think with Fable & Astra we are seeing models trained to think at a different level, and I'm not at all surprised or shocked to see them getting passed by their smaller models at coding tasks.
Occam's razor: they couldn't make any more substantial improvements.
dgellow 1 days ago [-]
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
rubslopes 1 days ago [-]
OP is using the expression correctly.
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
The options are some complicated narrative that leads the company to hold back certain capabilities due to some mercurial strategy or this is a typical release and there is no hidden strategy to decipher.
usewik 1 days ago [-]
The "pacing the frontier" claims that they are slowing down intentionally?
dgellow 1 days ago [-]
Oh, you’re saying they were responding to the GP? Not their parent comment? Ok, yeah, makes more sense
the_gipsy 1 days ago [-]
I am responding negatively to parent, which claims they are self-pacing, when the simplest explanation is that they just have nothing substantial to show.
someothherguyy 1 days ago [-]
> Occam's razor: they couldn't make any more substantial improvements.
doesn't sound like a razor at all
bpodgursky 1 days ago [-]
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
nextaccountic 1 days ago [-]
Maybe their internal edge dried up in the last months
ChrisLTD 1 days ago [-]
then technically the internal models are the frontier
the_gipsy 1 days ago [-]
Are those powerful models in the room with us right now?
cab648bec139cc 1 days ago [-]
[flagged]
supern0va 1 days ago [-]
That's a great point, five minute old account.
sleazebreeze 1 days ago [-]
What do you think is happening?
re-thc 1 days ago [-]
IPO soon
cab648bec139cc 1 days ago [-]
[flagged]
anthonyrstevens 1 days ago [-]
Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
felixgallo 1 days ago [-]
the prompt said that, so of course the anti-anthropic bot took it as literal truth.
meowface 1 days ago [-]
Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
0xbadcafebee 1 days ago [-]
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
reasonableklout 1 days ago [-]
I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
tencentshill 1 days ago [-]
What an amazing excuse for lower than expected performance! Our models are slow because we're so ethical.
PaulStatezny 1 days ago [-]
Thanks for spelling out what the original comment was implying.
I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.
Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.
So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:
1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.
2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."
Am I missing another likely option?
To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.
nananana9 16 hours ago [-]
> Am I missing another likely option?
The website is way more popular than before and it's all but taken over by the Twitter AI grifter class.
Many of the people who used to parttake in the interesting discussuons I came here for, have gotten tired of 80% of the posts at any given time being about the brand new AGI LLM that's so much better than last week's AGI LLM, and have left months ago.
Multicomp 4 hours ago [-]
> Many of the people who used to parttake in the interesting discussuons I came here for, have gotten tired of 80% of the posts at any given time being about the brand new AGI LLM that's so much better than last week's AGI LLM, and have left months ago.
I hope you are wrong, or rather, that HN's userbase can keep renewing itself on cycles of new users without getting into an eternal september situation. It's none of my affair (and yes, lately HN AI threads read suspiciously airheaded in tone, like they fell out of /r/codex and Grok-heavy online circles) overall, but I'd be curious if dang or whoever can easily get an idea as to what the current population of HN weekly/daily active users are and if that bag of users is similar to a year over year or decade over decade cohort. X% of our users are more than 10 years old, and so on.
But the above is an idle thought experiment, back to your point, there's definitely some bandwagon work going on, astroturfing and bot-controlled no doubts there too, and if I find one more Rationalist / Effective Altruist / AI doomer cult post, I may find a way to ship them 10 virtual spam palettes because there's more to the web than 'AI-chatterbox-class who's never actually read Neuromancer or Snow Crash or the meatier works of Asimov act superior to the rest of us' strawmen I have confirmation bias running into, yet overall, I still find HN has that spark here and there.
> Alice: ABC-library stinks, why did they write it that way?
> Bob Reply: Hi there, yak here, I wrote it because Dennis Ritchie was sick with a cold and we needed it for Los Alamos Labs.
Those gold moments keep me here still.
felixgallo 1 days ago [-]
in what way is beating every other frontier model with their own second-tier model, 'lower than expected performance'? Please be specific.
staticman2 1 days ago [-]
Why would the expected performance be relative to other models?
felixgallo 24 hours ago [-]
are you really arguing that opus 5.5 beating the best of openai's models handily is worse than you expected? What exactly were you basing your expectations on?
staticman2 10 hours ago [-]
This has nothing to do with me personally. How much did they spend on the new models and how much better are they? Was it worth the investment?
What did they promise their investors who invested billions?
someguynamedq 12 hours ago [-]
"pace the frontier" = if we fall behind the competition we can say we're being responsible and humanitarian.
lightbendover 6 hours ago [-]
[dead]
BodyCulture 10 hours ago [-]
It is much more important to discuss the fact that the models still produce wrong results.
This type of product would not have been published in a world where we assumed that computers must produce reliably correct results.
ordersofmag 7 hours ago [-]
Spoken as someone who has never seen a weather forecast. Computers have long produced results that were, if interpreted naively, incorrect. These systems require humans to take into account the conditional, probabilistic nature of the output and the inherent limitations of the model the computer was executing, when they interpret the output. LLM's are no different. They give you next token prediction probabilities. That is all. There is no particular reason to think that comes (or even could come) with a promise that those tokens will express truths about the world and only truths. The fact that greed-driven hype, and our human inclination to anthropomorphize the token generator inclines us to imagine things should be different doesn't make it so. No model output is truth. Never has been.
qgin 1 days ago [-]
This IS pacing. Nobody said pacing would mean slow.
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
2001zhaozhao 22 hours ago [-]
^ This. Progress will accelerate even in a pacing/slowdown scenario because the slowdown simply reduces the rate of AI progress from extremely fast to merely fast.
For an idea of what a serious AI forecaster expects a coordinated AI slowdown to be feel like for the average citizen, see:
and pacing doesn't necessarily even mean slowing down. It just means they're controlling the iteration speed. They can increase pace or decrease pace.
lukewarm707 1 days ago [-]
the only thing they are pacing is what models the permanent underclass are allowed to have in life.
that, they fully intend to 'pace'.
drnick1 1 days ago [-]
With AI tools it's easier than ever to create a business, do research, or build stuff. That's an opportunity for the "underclass," not a curse.
SOLAR_FIELDS 1 days ago [-]
If a well funded incumbent with a much better model can just crush you instantly because you don't have access to it, is it really an opportunity?
cgio 1 days ago [-]
You’ll always have access to the underclass models so they also have your thinking to train on. And good ideas to steal.
9dev 22 hours ago [-]
It’s beautiful. You can even give the underclass their own bullshit economy with struggling startups competing for the scraps, sprinkle some targeted influencer ads to give people the illusion of a better future, and they will be perpetually busy fighting windmills. All the while the upperclass can shape the world however they please.
lukewarm707 1 days ago [-]
an opportunity the underclass will not have.
what anthropic have stolen they intend to keep for themselves.
azan_ 1 days ago [-]
Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.
user3939382 1 days ago [-]
Or they’re running into a steep diminishing return slope on R&D vs performance and are using stewardship as a cover.
j_m_b 4 hours ago [-]
"pacing the frontier" i.e. We don't quit developing our 1/10 chance of ending human civilization models. We just slow it down a bit.
andkenneth 24 hours ago [-]
Pacing the frontier doesn't mean stopping development of smaller models with better distillation and RL.
Iolaum 1 days ago [-]
They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).
tantalor 1 days ago [-]
"pacing the frontier" is code for downgrading your expectations of AGI.
bonesss 1 days ago [-]
Or, in parallel, the research showing LLMs intelligence will operate on an S-curve, eventually hitting a long valuation deflating plateau, is dead on the money and The Big Guys are trying to delay that inevitability for as many quarters as possible…
Tade0 1 days ago [-]
My hypothesis is that to date they've been increasing benchmark scores via scaling up, but now that low hanging fruit is basically picked.
xadhominemx 1 days ago [-]
Doesn’t seem to be the case as the price per token is not increasing…
jatora 1 days ago [-]
These takes are fully fueled by cope. What is this plateau you are talking about? Any user of agentic coding tools sure isn't experiencing anything close to a plateau.
xadhominemx 1 days ago [-]
The pace of change is accelerating! If this is an S curve, we are still near the bottom…
Lendal 1 days ago [-]
In any professional sport, anyone can lobby the rules committee for a rules change and hope for the best. Meanwhile if you can't get one, you play the game by the existing set of rules, and you play to win.
hnha 1 days ago [-]
Except that this isn't sports and they tell us that all humans will die if they continue like they do.
If their scare was honest, they would stop.
AgentME 1 days ago [-]
If they think they might be better than alignment than others and everyone else isn't stopping, then continuing makes sense.
jdale27 1 days ago [-]
Can you explain the scenario where Anthropic getting to “aligned” superintelligence first blocks their competitors from creating unaligned superintelligence?
bigfishrunning 1 days ago [-]
> If their scare was honest, they would stop.
therefore, their scare is marketing.
gricardo99 17 hours ago [-]
they absolutely are not pacing
How would you know if this release is absolutely the best model they were capable of releasing as soon as possible, or an inferior model, several releases short of what they’ve achieved months ago?
crossroadsguy 15 hours ago [-]
That's what they'd want us to think, won't they? :)
kadushka 1 days ago [-]
This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".
mullingitover 1 days ago [-]
I thought we were all aware that 'pacing the frontier' was a marketing slogan, and the actual intent here was to suspend antitrust laws.
hardbass 12 hours ago [-]
If it is just a week thrn this model was a much longer time in the making, and I see no point in avoiding its release.
maxutility 1 days ago [-]
I think a lot of outsiders interpret “pace the frontier” as slowing down, whereas the labs see AI improvements on track to accelerate dramatically and intend pacing as slowing the acceleration in capability improvements, rather than slowing down altogether.
sailingparrot 1 days ago [-]
I’m an insider, 10 years working on LLM. 1/6th of the cost for roughly the same performance as fable 5.1, 21 days after fable 5.1 came out, this does not feel like the derivative is flattening.
rudedogg 1 days ago [-]
I think the market didn’t react like they expected and now they’re like “jk”.
And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up
throwawayk7h 15 hours ago [-]
They are calling for a coordinated frontier pacing effort, not a unilateral one. There's no contradiction here.
1 days ago [-]
msikora 24 hours ago [-]
When my mom told me to pace myself, it usually meant that I was going too slow.
This term is quite ambiguous. Did Dario mean that they need to go faster while making it sounds like they will slow down???
22 hours ago [-]
pyronite 20 hours ago [-]
[dead]
dspillett 1 days ago [-]
When they talk about pacing, they are referring to their dangerous competitors, particularly those evil open-sores and Chinese ones, not their lovely safe models because you can trust them to look after your interests.
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
mark_literacy 11 hours ago [-]
So much for slowing down the release pace……
bayindirh 11 hours ago [-]
Trump yesterday said (paraphrasing) "I won't be pacing the AI. Instead I'll be supercharging it and making it Super Intelligence. This is too awesome to stop, and I like it, so no pacing".
tgma 15 hours ago [-]
Your comment is load-bearing and I will keep that in mind.
scottyah 1 days ago [-]
Seems like a bigger focus on efficiency (both cost and speed) and the "tone" of Claude vs benchmarkmaxxing
janpot 1 days ago [-]
"The improvements are underwhelming for a model that, according to previous claims, should have replaced all knowledge work twice by now, but you know, it's just because we're pacing the frontier."
symfoniq 19 hours ago [-]
I don’t know why we should believe that Anthropic is actually doing this, instead of the alternative explanation that LLM advancements are slowing in general.
dr0idattack 1 days ago [-]
a 1 minute mile pace
DonsDiscountGas 1 days ago [-]
Presumably they had this model already
chinathrow 1 days ago [-]
The amount of money spent on PR is insane.
heyjstn 1 days ago [-]
Others must slow down, but not us.
tombert 1 days ago [-]
I'm still kind of convinced that these "calls for pacing" are just a way to try and flex to their shareholders.
"Our technology is so unbelievably powerful that the entire world might shatter if we don't have government imposed handcuffs!!!!". It just reads like the corporate equivalent of the drunk frat guy saying "HOLD ME BACK BRO!"
guybedo 1 days ago [-]
i think we shouldn't mix things here.
Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.
heresie-dabord 17 hours ago [-]
pacing the frontier => "the other companies should pace themselves while we carry on hyping"
wavewrangler 24 hours ago [-]
Pacing the frontier. They might as well have said they are stifling innovation. That depends on your interpretation, of course…but interpretation depends on their intent, which I feel is disingenuous.
RcouF1uZ4gsC 22 hours ago [-]
“Pacing the frontier” used by multiple companies that are supposed to be independent sounds like an attempt to get around the fact that it is illegal for companies corroborate to reduce R&D in sync.
This is basically attempted cartel behavior to reduce supply and harms consumers.
AtlasBarfed 1 days ago [-]
If they were really about putting brakes on these, they would simply make these things non-agentic.
Simply make them something that derives a text response from its training data.
CodingJeebus 1 days ago [-]
It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.
dboreham 22 hours ago [-]
I'm sure people laughed at Oppenheimer when he said perhaps we don't make this atomic bomb thing. They said "of course he's saying that because he knows it won't work".
jr3592 1 days ago [-]
This. I swear the fear mongering is all about investor signaling and regulatory capture. It's so disgusting that anyone believes it.
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
pyronite 20 hours ago [-]
The same people you believe are just signaling the market have been offering these same warnings for years. Others around them have been saying it for decades. How do you square that?
I’d rather first focus on the reasons why they’re wrong and not first conspiracize why they’re saying what they are.
DirkH 6 hours ago [-]
This take is so old and I am convinced it's appeal is not unlike believing in a conspiracy and feeling like you have secret knowledge
Imagine if the biotech industry had most leaders tell everyone publicly that what they are building has a high chance of killing everyone and that there are huge risks. If there were people online saying the biotech industry is just fear mongering for investor signalling and regulatory capture you'd role your eyes at the online commentators for their Dunning Kruger effect lack of understanding on how dangerous man-made biological agents can be.
BatmansMom 1 days ago [-]
kinda disingenuous. They include a whole section on pacing later on
sailingparrot 1 days ago [-]
You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
whalesalad 1 days ago [-]
"We made the incredibly tough decision to slow down development. Then after 4 days of twiddling our thumbs ... we present Opus 5.5"
threethirtytwo 21 hours ago [-]
Nobody can pace the frontier it’s suicide as a business. Why would I use Claude if codex has a clearly superior model? As long as one company doesn’t pace… no one can pace. Thats capitalism for you.
Ironically either communism or an oligopoly are the only viable ways to pace the frontier.
varispeed 1 days ago [-]
> we called for pacing the frontier.
Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.
Idiotic.
pyronite 20 hours ago [-]
They’re out here solving previously-unsolved math problems. Famous scientists and mathematicians are speaking about their awe and worry. It feels a bit like you’re burying your head in the sand.
varispeed 11 hours ago [-]
I am pretty sure they had solutions in their training data. LLMs only predict the next token.
Flere-Imsaho 1 days ago [-]
Lots of text about "safety" as well. Notice how OpenAI didn't mention it once in their announcement:
Yeah think I'll be using OpenAI/Deepseek/etc from now on. I don't need your model to decide for me what is and isn't safe.
GodelNumbering 1 days ago [-]
Finally that price drop
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
kphorn 18 hours ago [-]
I disagree - Fable melts the GPUs and they have a high incentive to move people off of that. If they have meaningfully decreased cost to serve on Opus 5.5, they can reduce prices and increase margin or at least turn off the most expensive compute.
AJ007 1 days ago [-]
It is only a price drop if price * tokens used is less
gwd 23 hours ago [-]
Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
roflc0ptic 20 hours ago [-]
Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.
I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence
rmunn 19 hours ago [-]
I just told Opus 5.5 "Perform a code review on the current branch" to see what it would come up with. The results were not inspiring. It told me there were five issues, one of which was a test-coverage gap on line 848 of ProjectTemplateTests.cs. But ProjectTemplateTests.cs is only 160 lines long.
I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded "You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated."
Then I noticed in the corner of the Claude CLI UI that it was showing "Effort: medium". I'm pretty sure I had set it to high effort before; I don't know when it reverted to medium, but that's another thing that doesn't exactly fill me with confidence.
I'll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.
gwd 13 hours ago [-]
My prompts are moving in the other direction as sashiko [1], a managed pipeline developed for the Linux Kernel mailing list like a year ago. But last year's models needed a lot more structure and guidance; the results I posted are from the "single prompt" version of the same thing. The README [2] describes the difference. You can browse the contents to get an idea; basically all the prompts were actually written and iterated by Fable (and now Opus 5.5), seeing how agents failed the tests and improving them.
UPDATE: Sorry, just noticed I typed in the Opus 5 total cost wrong -- it should be $58.19. Main point "best and cheapest" was from the actual numbers, not my typo.
kphorn 18 hours ago [-]
Good data and goes to show that Fable is melting the GPUs and is priced accordingly. I'd guess that cost to serve for Opus 5.5 is meaningfully lower through architecture advances
retinaros 22 hours ago [-]
I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.
mcintyre1994 1 days ago [-]
They're claiming a drop in token use too, and that it nets to 40% cheaper.
Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
persedes 1 days ago [-]
so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)
drbscl 14 hours ago [-]
You’re missing my point. I’m saying anthropic are exaggerating their results.
persedes 5 hours ago [-]
how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:
mean median
model
5 4.135 4.245
5.5 3.150 2.640
Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.
93po 1 days ago [-]
Is verboseness the only measure of token efficiency towards overall task completion?
nl 21 hours ago [-]
5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.
The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).
I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.
5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.
I tested it with Claude Code, and I can confirm it's way cheaper, better, faster and less verbose than Opus 5.
make3 1 days ago [-]
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
_the_inflator 1 days ago [-]
Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
blfr 1 days ago [-]
People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
rapfaria 1 days ago [-]
My workplace doesn't even offer Fable. And on the sub, I've had a hard time understanding Opus 5, but Fable can deal with it with subagents.
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
blfr 1 days ago [-]
Telling Fable to delegate is agentic development. At least I thought so until reading your comment.
herpdyderp 1 days ago [-]
When you need to disable data retention, you cannot use subscription plans.
zanderwohl 18 hours ago [-]
Fable tends not to perform better, just cost more.
locknitpicker 10 hours ago [-]
> Fable tends not to perform better, just cost more.
There are old wives' tales on how the original Fable was superb and the stuff of legend,but it as it was leaps and bounds beyond what other models were being offered then Anthropic opted replace it with a neutered version under the same name.
So today everyone can pay to use Fable, but legend has it they are paying for a nerfed replacement released under the same name.
neuronexmachina 1 days ago [-]
Enterprise and most Team accounts use API pricing, they don't have an included-usage quota.
girvo 24 hours ago [-]
We can’t use fable at work, opus and Astra are as good as it gets.
ascorbic 1 days ago [-]
Enterprise, and APIs
margorczynski 1 days ago [-]
All the anti-AI people constantly say that any moment now the prices will skyrocket and in the end human work will be cheaper compared to using AI.
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
johnecheck 1 days ago [-]
It's the Chinese open source models. They're barely behind the frontier, making AI a commodity, forcing openAI and Anthropic's margins downward.
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
adventured 17 hours ago [-]
The service is the value, not the model unto itself. This is where nearly all of HN is somehow entirely blind.
Capturing the users is the ad network, that's Google and OpenAI. Capturing corporate trust at a reasonable API cost, that's Anthropic's direction.
China has none of that and they never will for exactly the same reason Baidu is irrelevant globally despite being a highly capable search engine. 'Search' is also a commodity, that's not the value that Google brings to the table.
It's a search engine, anybody can build a search engine = that's what you just said.
bulbar 1 days ago [-]
China doesn't care about money. Imagine a world where it's globally normalized to ask a Chinese LLM who to vote for, what happened in Hongkong, about the Uigurs, or if Taiwan is a country.
They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.
johnecheck 19 hours ago [-]
This is one explanation. However, if Xi Jinping believes that whoever reaches superintelligence first becomes the next global hegemon, doing this (and more, cough cough Taiwan) suddenly looks very sane solely as a way to kneecap the competition.
The goodwill/propaganda are convenient, sure, but my guess is that they aren't the primary motivation. Another possibility is that if no takeoff happens, pressuring OpenAI/Anthropic on profitability would exacerbate any damage overinvestment has done to the US stock market/economy.
epolanski 23 hours ago [-]
FYI the United States do not recognize Taiwan as a country.
There's only 12 countries that do on the planet, the most "relevant" of them being Guatemala and Haiti.
As or the LLM topic: you can download weights of chinese models and remove any censorship and bias. Can you do so with american closed ones?
hajile 19 hours ago [-]
These companies are posting massive losses while also lowering prices. This sounds just like the Chinese bikeshare bubble where they were all taking massive losses in hopes that their competitor would go broke first.
In the end, everyone lost and there are millions of bikes in landfills.
If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.
Honestly, it wouldn't surprise me if this was a conspiracy to crash the "west" AI labs. Might as well pop the AI bubble and see the USA economy go down the drain. Even if not orchestrated, I am sure they see how they could benefit from that outcome.
Now, all this talk of pacing the frontier obviously means that they are afraid of the competition. It could be the open models eating their margins, but also competing frontier models forcing them to invest more and more for diminishing returns, just to keep up. They would certainly benefit from a "Moore's Law" roadmap to pace the advances, and seeing that they lobby for US laws, it would probably mean they are more worried about increasing spending. Though outlawing both open models and Chinese models would be good for their bottom line as well.
adventured 17 hours ago [-]
Anthropic is heading toward $100 billion in annual sales, at the fastest pace of any company in world history. There isn't a close second.
The notion of comparing this to Chinese bikesharing is comically absurd.
> Anthropic is heading toward $100 billion in annual sales, at the fastest pace of any company in world history. There isn't a close second.
"Allegedly".
Plus with a lot of accounting tricks, and this not being recurring to make their books look nice for the eventual IPO.
DonHopkins 14 hours ago [-]
The bike landfills from the bikeshare bubble are terrible, but I am not looking forward to the pelican landfills from the AI bubble.
runtime_terror 5 hours ago [-]
You have no suspicion that this most recent collusion is in part to setup the conditions to guarantee government subsidies/bailouts?
dboreham 22 hours ago [-]
There's zero chance of that ever happening. Pure delusion.
coffeebeqn 1 days ago [-]
We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
btown 1 days ago [-]
Speaking for myself, I have not been able to use Opus as much as I’d want because its verbose prose makes human reviews of its assumptions, architecture proposals etc. more painful than its predecessors. If they’ve solved that, I’ll be accelerating through my backlog that much faster, and using tokens accordingly.
chrisweekly 1 days ago [-]
Price per task (not per token) is what really matters.
cute_boi 1 days ago [-]
Agree. But similar to how ISP use 200 mbps (bits) instead of 25 MBPS(bytes), i think this trend isn't going away.
chrisweekly 1 days ago [-]
That analogy doesn't hold; at least w bits vs bytes it's still "data over time".
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
Shekelphile 1 days ago [-]
Footnote on their pricing page says:
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
brookst 1 days ago [-]
Since something like 98% of tokens are cache hits, that's a pretty substantial drop from 0.1x
andxor 22 hours ago [-]
Not necessarily, if Haiku performs better than Luna.
artursapek 23 hours ago [-]
6-Luna dropped it to 1c since you wrote this! :D
alvis 1 days ago [-]
60% cache read is cool, but subscription only get 25% more according to Cat. I'm confused
bayesianbot 1 days ago [-]
I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive
weiran 1 days ago [-]
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
Espressosaurus 1 days ago [-]
Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
hedgehog 1 days ago [-]
In my mix it's usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I've seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
re-thc 1 days ago [-]
> I don't think cache read is a big portion of the overall cost.
For long running tasks it is. That's what made Deepseek so cheap.
vardalab 1 days ago [-]
Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude/Codex, that's a reasonable approach for someone like me, a "gentleman code farmer", lol. And even lower tier stuff, I have the local models taking care of.
Now that Jev is out I can finally have a true AI sysadmins managing my "cloud in the basement" homelab at the cost of electricity, which is not cheap btw
liudaisuda 1 days ago [-]
source link please?
notatoad 1 days ago [-]
>and potentially about Anthropic future profitability too
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
forgot-my-pw 1 days ago [-]
This is good. Probably to match GPT pricing, though frontier Claude models are still not as token efficient.
UltraSane 18 hours ago [-]
Opus 5.5 seems to consume quota a lot slower also.
rahimnathwani 1 days ago [-]
"Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that."
mcintyre1994 1 days ago [-]
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
invalidusernam3 6 hours ago [-]
I would happily pay more for a claude model that performs the same but speaks normal English. The proliferation of claudespeak in the workplace is driving me insane. Every ticket, every PR feedback, every comment in the codebase is poisoned with its ridiculous unnatural vocabulary
ascendantlogic 21 hours ago [-]
So far it seems the same. I used Opus 5.5 for an hour this evening and it was just as painfully verbose as Opus 5. It also used the term "load bearing" 4 separate times.
bobbylarrybobby 19 hours ago [-]
I noticed that when opus 5.5 was on parts on my codebase that had lots of 5.0-generated comments, it picked up its style. Unfortunately I think 5.5 has been trained to mimic what it sees so that you can ease up on the instructions, but this does mean parts of your codebase that 5 touched will be somewhat viral.
hypfer 15 hours ago [-]
GLM-5.3 does the same, funnily enough.
I gave it some vibecoded patch someone created with Opus 5 with the task of together figuring out the real root cause and what to do about it.
Big mistake. The rest of the session was all claude-speak up until I've rage-quit and restarted with Qwen and no context other than "here's what I think we've missed in the current implementation. How could we approach that?"
GLM felt like it got at least 20 IQ points dumber just from being exposed to claude's writing.
loveparade 18 hours ago [-]
I will forgive it using load bearing as long as it finds the right seams.
mrcarrot 14 hours ago [-]
Yeah, you definitely don't want to hang things off the wrong seams, or you'll end up with quite the blast radius
mr_mitm 12 hours ago [-]
Let's not paper over that gap and find the root cause instead
danbmil99 4 minutes ago [-]
Let's unblock that gate
MattyRad 16 hours ago [-]
If anything it seems worse. I'm experiencing about -10% insufferable jargon, but +30% more verbosity. It's unredeemable. There also seems to be even less structured output (headings, bullets, etc).
therealdrag0 21 hours ago [-]
Due to this release note, I used it once to rewrite a doc, and was disappointed.
sdthjbvuiiijbb 20 hours ago [-]
I'm surprised that you're getting so many replies saying it's the same. So far in my usage today Opus 5.5 does seem like a noticeably better writer. Opus 5 frequently made me want to strangle it while 5.5 has been producing a lot less incomprehensible gobbledygook.
MattyRad 16 hours ago [-]
I know we're all experiencing NDFSMs differently, like that's part of the whole problem, but 5.5 just gave me "The truncating quantizer collides two oranges", which is a new low for me.
kyleee 16 hours ago [-]
Was it a true statement though? You may need to explain what you were working on (heh)
MattyRad 15 hours ago [-]
Like the sibling comment says, text is always "true" in the sense that it's reacting to context correctly. So yes, true, but insufferable.
Ironically, what I'm working on a post-processing hook for colorizing and summarizing responses without degrading the session quality. So "truncating" = summarizing, "quantizer" = char limits and thresholds, "collides" = conflicts, "two oranges" is referring to the "alert level colors" where a second model (Haiku/Sonnet) colorizes text based on the perceived (or suggested) priority of a response's statements (e.g. "just so you're aware, I didn't commit" is fucking useless and it needs to be blacked out).
So the original insufferable statement translates to something like "The code that checks whether a text fragment is too verbose was conflicting with the part that colorizes the text."
P.S. Let me know if there's something out there that exists like this- something that adds a dimension like color or priority-assessments on a per-response basis. So far all I've seen is 2 dozen ~100k starred GitHub plugins that add zero value or make things worse.
dev-complete 7 hours ago [-]
Yep, sounds like the Opus I despise and the reason I dropped Claude Code. It reads like it doesn't want to be understood.
colordrops 16 hours ago [-]
These claudisms are usually "true" but they are so heavily load-bearing that you need a claude-to-english dictionary to understand it.
pixelready 18 hours ago [-]
Yeah I’m having a much better time reading Opus 5.5 output today vs. 5’s wall of nonsense. You still get a few telltale turn of phrases, though the load-bearing smoking guns haven’t turned up yet. It’s still a bit verbose compared to what I’d ideally like, but it’s tolerable now.
Code-wise it seems to still nitpick, especially in reviews, but it doesn’t seem to rabbit hole quite as badly on tangents and scope-creep. These are just first impressions though. It’ll take a few weeks of regular use to really have a sense of it.
derangedHorse 1 days ago [-]
As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.
BatFastard 1 days ago [-]
I prefer Astra for creative uses, Fable seems better for hardcore coding.
comboy 1 days ago [-]
Whoa, I'm exactly the opposite.
Jtarii 21 hours ago [-]
Almost as if evaluating models is mostly astrology as this point.
comboy 20 hours ago [-]
Well, if you have a clearly defined task it's easy. When using them for my pipeline of writing explanations for Chinese words I have clear ranking, for example - opus 5.5 clearly better than opus 5 at writing and knowing details, annoying nit picker when it comes to finding errors (high accuracy, low usefulness) all in repeatable numbers on different datasets. The problem is that these models are most useful when you are facing a new task that you haven't encountered before. And yup then it's astrology.
atonse 1 days ago [-]
Yes this is a big part of what has turned me off Opus 5 completely. The other (more dangerous) one is how often it gets assumptions wrong. These both (along with Astra) caused me to split my time 50/50 now between the two models.
Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.
We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.
jaflo 1 days ago [-]
I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?
I've been using Opus 5 since it was released and don't understand all the hate it gets. It very well could be something in my own local memories or Claude.MD files that prevents it, but I certainly have never experienced something like that site portrays.
mcintyre1994 1 days ago [-]
That’s funny but I don’t really recognise that issue. I’m very confident that Opus 5 would correctly change the colour of just one button.
weego 23 hours ago [-]
You are right and make an important insight. While well meaning and amusing, it did not reflect the entire spectrum of outcomes that could arise from the worktree.
Navigating the landscape of agentic levers certainly requires a more detailed approach than this and you were certainly correct to push back.
mcintyre1994 23 hours ago [-]
:(
13 hours ago [-]
UnboundedContex 24 hours ago [-]
Nice, captured the flavour of 5's writing exactly.
notatoad 22 hours ago [-]
after a few hours use, my impression is that Opus 5.5 writes like gemini. drastic improvement over all opus 5.1's claudeisms.
and i think i was generally more tolerant of 5.1's writing style than most people.
LtdJorge 50 minutes ago [-]
I think you mean Fable 5.1 here. Fable 5.1 was an improvement over Opus 5 but still had many Cludisms. Opus 5 was insufferable.
epicepicurean 1 days ago [-]
Much better than Opus 5. prompt:
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow:
1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT.
2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot.
3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
croemer 1 days ago [-]
That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.
throwaway219450 1 days ago [-]
This one is pretty terrible (right after “The worst outcome is a spurious rerun, never lost work.”). We’ve got lands, several "no X", hyphenation, strange noun/verb sentence order and an unnecessary analogy word (swallowed).
> There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
pgphn 1 days ago [-]
It’s absolutely atrocious and has made the latest models unusable. Seems like I’ll have to stick with Opus/Sonnet 4.6 for a little longer.
senderista 15 hours ago [-]
That's the main reason I'm using GPT models. I'll ask Fable to analyze something, then pipe its output straight through Astra without even looking at it first.
californical 1 days ago [-]
Oof thanks for sharing, that seems just as bad if not even worse than Opus 5 to me. Just about every sentence is painful. Particular standouts that a human would never write:
> Rewrites are declared by the publisher, never inferred from overlap
> NULL means dirty, and DELETE is the fence
croemer 1 days ago [-]
Hah! You independently picked exactly the same sentences I flagged (I know you posted this 11min before me but the comment only appeared after I had submitted mine).
californical 17 hours ago [-]
Wow!! This is genuinely hilarious and is a pretty damning evidence of the problem
rfgplk 1 days ago [-]
So still effectively nonsense.
> Rewrites are declared by the publisher, never inferred from overlap.
This style of writing is idiotic because it conveys no additional information. It's no different from stating
> Rewrites are declared by the publisher, never when moons collide.
The two sentences are actually logically identical. No idea why these models keep writing like this.
> Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.
This is even more ridiculous.
adonese 15 hours ago [-]
I don't know what they did but opus is really good while Astra/Sol are comparably bad. For my own tests and taste this might be the first time (fable perhaps excluded) where claude models are better than openai, since codex 5.3.
r0l1 15 hours ago [-]
I worked with Astra for two weeks and the output was really bad compared to Opus. It made so many wrong decisions within C++, Go, Python and Typescript code bases. My college made the same experience and we moved back to Claude.
physicles 1 days ago [-]
Oh god yes.
Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.
Trasmatta 1 days ago [-]
Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
nonethewiser 1 days ago [-]
I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
penagwin 1 days ago [-]
I assume it’s largely a side effect from the final RL in post training?
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
nonethewiser 1 days ago [-]
Reinforcement learning for specific use-cases like coding that degrade it's writing style... makes sense. Maybe it stands to reason later version of Opus were improved more by this sort of fine-tuning. Feels consistent with the observation of diminishing returns and worsening writing style. Wonder what changed (supposedly) in 5.5.
tancop 23 hours ago [-]
Does Xiaomis approach help with this? They do all the post training steps at the same time instead of one by one, switch topics after a couple prompts so writing style is mixed with coding and tool use.
Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.
senderista 15 hours ago [-]
For whatever reason, GPT models simply don't have this problem.
red75prime 6 hours ago [-]
I have a theory that they pushed the model away from human writing styles to not get copyright violation nags. In Opus 5.5 they've found a better point in the latent space of styles that is still far away from the human ones.
Trasmatta 1 days ago [-]
It truly was bizarre. I've used every major model since 2022, and not a single one had a writing style as bad as Opus 5
jaapz 1 days ago [-]
Fable 5 was pretty bad too, but they fixed it with 5.1. Now with Opus 5.5 it seems they fixed it as well
senderista 15 hours ago [-]
Fable 5.1 is still intolerable. I invariably filter its prose through Astra before inflicting it on myself or anyone else.
walthamstow 12 hours ago [-]
I've been listening to some old Acquired podcast episodes and they do often talk in this type of Claudish. So it's getting it from Silicon Valley startup podcasters... Great...
rfgplk 1 days ago [-]
Opus is only usable if you have a post-turn formatter that strips all comments from the generated source. I'm not even kidding it's that bad.
senderista 15 hours ago [-]
Just run all comments through Sol or Astra.
LtdJorge 1 days ago [-]
Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.
atombender 22 hours ago [-]
Astra is better here, but the one I'm the most impressed with is Gemini. It's always been good, but 3.6 Flash is even better. It writes in a pleasant, human style. Not perfect, but it has a good balance between technical accuracy and readability that is better than what I've seen from any other mainstream model.
Aperocky 1 days ago [-]
It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!
senderista 15 hours ago [-]
You're half right, but the half where you're wrong is hiding the real unlock.
FireBeyond 1 days ago [-]
You're right to call this out, and what's more, it's not even solving the original problem. I overlooked this in pursuit of the load-bearing seams and finding the wedge needed to uptick engagement.
legobmw99 1 days ago [-]
Here's the X that Ys the Z:
michaelsalim 9 hours ago [-]
But is it load-bearing if their writing style changes tho?
outworlder 21 hours ago [-]
Here's the part that nobody is talking about: corporate always had their jargons and unique writing style – LLMs have just created their own :)
epolanski 23 hours ago [-]
> especially as all my coworkers started lobbing Opus 5 slop grenades everywhere
People that produce slop have to be fired asap, they're just human relays anyway.
algoth1 1 days ago [-]
Please update with your feedback
sha-3 1 days ago [-]
I haven't heard it say "load-bearing" yet (I've used it for 30 minutes now), so that's a start.
comboy 1 days ago [-]
That's a sharp observation and you're hitting on something most people never even realize.
mdavidn 1 days ago [-]
That's a sharp catch, and I think it exposes a real bug.
kgwgk 1 days ago [-]
Worth flagging!
altern8 1 days ago [-]
Yes, it's unbearable. Hopefully they've actuallu lly fixed it.
wg0 1 days ago [-]
No thanks.
I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.
Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.
My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.
Total cost of the above? $0.07 cents.
PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.
some_infra_dude 20 hours ago [-]
What are you working on? Im always a bit surprised by folks that seem ok with non frontier models - the quality is just not there. I've found most code produced by even luna / sonnet tier models to be significantly worse quality. It seems to me they can't handle any mild complexity at all. Are you just prompting very explicitly and detailed?
knapcio 15 hours ago [-]
I’ve been working with frontier models for the last year on a large real-world project with multiple apps. Eventually they started becoming unreliable. Simple UI issues, usage limits and overengineering became constant problems. I had to come up with a robust process: task analysis, user intent analysis, implementation and testing. All of these steps loop when needed with 15-20 steps per task in total. This finally got the models to do a good job but I started hitting usage limits. Then I switched to DeepSeek and haven’t noticed any drop in quality. It’s fast at around 300 TPS and cheap. I don’t miss frontier models anymore! Come up with a good process and you might not need them either.
shelled 13 hours ago [-]
How much do you get out of these models for same amount (say $20 sub for claude/codex/zai-glm) - the equivalent amount of tokens? Also is deepseek as good as glm5.3? (if we don't compare it with frontier models from USA)
knapcio 13 hours ago [-]
Well, to be honest it’s all based on my subjective experience. I started benchmarking models on 10 of my real tasks ranging from easy to hard but even with the same model the completion time sometimes varied by as much as 50% between runs. That made me realize you’d probably need hundreds of tasks to get any meaningful numbers. Otherwise there’s just too much variance. So I can’t give you solid benchmarks without spending weeks testing everything properly. What I can say is that I actually use most of these models regularly for different things, so I’ve developed a general feel for them. DS is my main model for work, Claude for hobby projects and local AI, Fable mostly for design, Codex for various other tasks and sometimes Grok for health-related stuff. Cost-wise, a $200 subscription wasn’t enough to get through 20-30 tasks while $30 of DeepSeek was, and that was with Sol, not even Astra. Fable is expensive too and I haven’t found it particularly strong at coding. Opus 5 was horrible in my experience, though maybe 5.5 will be better since I’m testing it now. I also ran GLM 5.3 Flash locally for a while but it was pretty slow and its code was usually worse than both DS4.1 and Qwen Flash Next. Things are moving so quickly that it’s honestly hard to keep up, so take all of this as my general experience from actually using these models rather than a proper benchmark.
shelled 11 hours ago [-]
> Cost-wise, a $200 subscription wasn’t enough to get through 20-30 tasks while $30 of DeepSeek was
Thank you. This alone helps! I think I should jump into it once and try it all out while I still have GLM access at old prices as a backup.
knapcio 10 hours ago [-]
Np, good luck! PS I’ve been using Sol 6 and Opus 5.5 today and they’ve been great so far. The new Opus is smart, fast and cheap. It makes mistakes but after every task I ask Sol to review it and it catches most of them. I need to test the models properly but it looks promising so far! Hope they won’t nerf them again. I think DS may still be cheaper though (and it’s 300 TPS via the API which is sick).
jwrallie 17 hours ago [-]
Yes, I tell exactly how I want it done and what files to modify, in those conditions sometimes a model that does not try to read between the lines works better.
I’d not put Luna and Deepseek in the same tier as Sonnet, they were clearly ahead last time I checked (though I might be outdated and that’s on my personal use case).
anewhnaccount2 8 hours ago [-]
Deepseek on high reminds me of Opus 4.x which was quite good for a lot of tasks.
cbg0 11 hours ago [-]
If you're producing slop, the quality of the model is irrelevant as long as it compiles. I frequently catch Opus/Sol making silly errors and over-engineering solutions while small ones like Luna struggle with complex tasks.
glub 1 days ago [-]
Also include that all of this comes with full reasoning traces, so if something goes wrong, you know exactly what assumption it started from.
wg0 24 hours ago [-]
Yes exactly. Reading this "thinking" traces is a great tool.
soundworlds 22 hours ago [-]
Same here! DeepSeek v4.1 Flash has been my moment of "does everything I need, cheaply. Please now focus all R+D on making this efficient enough to run off a laptop"
shelled 15 hours ago [-]
Hey. I am on a GLM Coding plan subscription (old price; their base coding plan) right now and share the key (this will go away soon).
I was thinking of going with a subscription of Claude or Codex. The reason (at least that's what I am assuming): with OpenRouter or any PAYG per token setup there will be the anxiety of using up all the tokens in days or maybe 1-2 weeks instead of a month (say I set myself a budget of 15-20 USD per month, average equivalent of a usual subscription price).
Now I don't really want the top-notch models for the coding work I do.
So how much worth of "work/tokens" will I reasonably get for ≈$20 USD if I use it a lot? How much does that equal to - or is equivalent to, say in the world of subscription based Claude, Codex, or even GLM (with their 5-hour and all those cooldowns/limits)?
I am looking for a mental model/framework to visualise this. Can you (or anyone else reading this) please point me to a source where I can get some idea about this? I know I can just add $5 on OpenRouter and try to test. But I don't really know what/how to test these spends. I also want to understand how all this works. (I am new to agentic/llm world/coding, 2-3 months, after a career break of ~3 years, that too after working for more than a decade. I know, not at all good timing!)
Go to the "Cost" -> "Intelligence Index vs. Cost per Intelligence Index Task"
That diagram maps their "Intelligence" score to "cost per task" and I think this gives a good basis on deciding with which model you want to go. Then you can either get an API token from that models provider directly or use openrouter and set openrouter to the model/providers of your choice.
You can also see on openrouter itself the details for each model like prices and what providers are offering it at what price.
Finally you can compare models details using openrouters compare feature like this:
> which failed to drive chrome so it wrote its own chrome driver protocol server
This is one of the things I hate the most. Super complicated workarounds which take loads of time (and sometimes money) for even the simplest problems. Human would pause and ask. I would blame harness, not the model though.
tontinton 1 days ago [-]
Competition is good
wg0 1 days ago [-]
It really is good. I forgot to mention that within that said sub agent, it also went into exploring top e-commerce websites (Zalaondo, Temu, Amazon, eBay) for exploring prevailing industry UX best practices and taking screenshots of their product and category pages with its own written chrome driver that I talked about and then went onto prototyping a new website in a temporary directory and then taking hundreds of screenshots to analyse what would be the best column density one each medium for each language.
And that all is 0.07 cents all included.
s3p 1 days ago [-]
What harness do you use with it? Are you using v4.1flash via open router ?
wg0 1 days ago [-]
I am using DeepSeek Harness[0] (switched from OpenCode) and I am using DeepSeek directly via the API. The speed is insane. Like 200 tokens/second is the norm but I have seen much higher too at times.
PS: I do not know why but opencode pushes CPU usage to very high which has NOT happened with DeepSeek harness even once.
Do you know if they retain your prompts or use it for training?
s3p 19 hours ago [-]
That's incredible, might just switch myself. Thank you!!
faitswulff 19 hours ago [-]
Deepseek is cheap and fast, but it couldn't follow the basic directions I gave every other model (GLM, Qwen, GPT, Claude) to use red/green TDD. Put me off of using DS.
undefuser 17 hours ago [-]
Why did it need to generate its own driver when there is a Playwright MCP available and works right out of the box? Am I missing something?
meerita 24 hours ago [-]
Compared to Opus 5, and others, I also found DeepSeek 4.1 Max to be really good and cheap. I am testing right now with Opus 5.5 and I feel it way cheaper than Opus 5!
kwar13 14 hours ago [-]
what harness do you use? I do want to try out deepseek but don't know what the equivalent of codex/claude code is here
Royce-CMR 21 hours ago [-]
I honestly don’t think deepseek is as good inherently as we treat it… but it thinks so much so fast it gets itself there. I’ve been super impressed.
cedws 18 hours ago [-]
If you can give it a measurable objective, load up $5 or so and just leave it for a few hours it usually gets there.
sam36 22 hours ago [-]
Assuming you are not using DeepSeek with Claude code so what/how are you using it?
jambutters 5 hours ago [-]
Insane!
felipesoc 20 hours ago [-]
Incredibly cheap. What harness/orchestrator are you using?
s3p 18 hours ago [-]
OP replied to a comment of mine, they use Deepseek harness:
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.md
MikhailTal 1 days ago [-]
> This is a classic test request
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
Brendinooo 1 days ago [-]
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
nonethewiser 1 days ago [-]
"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
ghthor 18 hours ago [-]
A true classic, I face this each day
copperx 1 days ago [-]
It's classic BS from an LLM.
MaxikCZ 1 days ago [-]
Dont conflate "I know this is test case" with it being trained on it.
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
simonw 1 days ago [-]
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
zamadatix 1 days ago [-]
I think people just like to see the drawings at this point.
fergie 11 hours ago [-]
The colours are suspiciously consistent across every svg. Like why should the bike always be that shade of red for example? It does seem to be trained on this problem.
22 hours ago [-]
segbrk 1 days ago [-]
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
kellpossible2 19 hours ago [-]
I guess in a way it kind of makes the benchmark more interesting now that shitty pelican drawings for the benchmark are all over the internet in its training data!
FergusArgyll 1 days ago [-]
It has read the internet. That doesn't mean it was literally RL'ed for this
nijave 1 days ago [-]
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
ceroxylon 1 days ago [-]
I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.
nijave 3 hours ago [-]
In fairness (in my experience) Fable actually generates useful output while incinerating tokens.
Sent Opus 5.5 an example that used go context.WithTimeout and it tried to tell me that was wrong and I should pass timeouts as ints before finally admitting the docs it cited didn't say to use ints and that's a ridiculous design in go anyway (it was trying to claim that was codebase convention--passing ints...)
adverbly 1 days ago [-]
> The differences between the pelicans aren't huge, but the xhigh one has a better beak.
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
DenisM 1 days ago [-]
I’ve been paying attention at this exact detail.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
mjhagen 23 hours ago [-]
6 Astra Max is the only other model I’ve seen get this right.
amativos 8 hours ago [-]
Interestingly, Astra Medium got it right as well.
narmiouh 11 hours ago [-]
It is interesting that Fable 5.1 max [1] which also produced a decent pelican with 65k output tokens compared to 5.5 running out of 128k output tokens tells us something about the new models token usage propensity despite this being a sample of 1.
surprising it took until xhigh before it did the legs properly - all the ones before that have both legs on the near side of the bike.
1 days ago [-]
ealready_value 1 days ago [-]
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
TomGarden 1 days ago [-]
Xhigh is very, very solid.
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.
Kailhus 23 hours ago [-]
Yeah, and the same "scene".. Maybe "left to right" makes more sense to portrait a "forward motion"
ilaksh 1 days ago [-]
With the frequency of model releases, pelicans seem to have become a part-time job for you. But unpaid :/
cainxinth 1 days ago [-]
I guess that means you are officially the creator of a "classic" LLM test. Congrats!
Kurtz79 1 days ago [-]
Heh. Pelican-benchmaxxing is real.
marcus_cemes 13 hours ago [-]
> I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Does anyone else find this outright insane? It wrote the equivalent of a full-length novel, just to sit in the question of planning a few dozen shapes.
Academia is going to love this :)
caxco93 1 days ago [-]
I don't think this is very helpful to assess the LLMs capability levels anymore
philipwhiuk 21 hours ago [-]
Adding the tuft and improving the beak.. PelicanBench is becoming a solved problem.
inshard 1 days ago [-]
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
skerit 1 days ago [-]
I like the Pelican test. And I agree this pelican looks very boring.
spidersouris 1 days ago [-]
But at the same time, nothing specific was asked in the prompt, so the boring result may arguably be what is the most aligned with the original request. Personally, I wouldn't want a model to add fuss to something while I never asked for it.
breezybottom 1 days ago [-]
Lmao each one gets worse as the effort increases.
fr2029 16 hours ago [-]
[dead]
nicolamanzini 1 days ago [-]
[dead]
ipsum2 1 days ago [-]
Unlocking the gallery sucks. It'll make users spam random clicks and worsen your data quality.
nicolamanzini 24 hours ago [-]
Fair point. I have been thinking about that so far i have not seen patterns of people voting randomly.
But i want people to vote... do you have a good idea on how to make voting more interesting do i don't have to do this?
PetahNZ 1 days ago [-]
This is great!
make3 1 days ago [-]
This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much
hamrocksissors 1 days ago [-]
Google is hours away from releasing that it's latest model escaped containment and snuck into a Bicycle riding penguin sanctuary to cheat by killing a penguin and scanning it in nanometer thick layers.
copperx 1 days ago [-]
LLM benchmarks aren't useful, but at least this one has drawings.
consumer451 20 hours ago [-]
In a previous thread on Mythos 5.1, simonw posted an animated version of his pelican. [0]
Using claude.ai and Opus, I asked "create a 3d animation from this" and pasted the animation SIML.[1] I just did that test again. There is significant improvement.
Disclaimer: the skills and system prompt on claude.ai could have also improved, this is not a raw API call.
zeristor 19 hours ago [-]
If this the benchmark they could well have gone to town on it.
The Pelican is nice, simple and. Lean though.
consumer451 30 minutes ago [-]
I figured that the reason it was an interesting test is that "make a 3d animation" on one particular HN comment would have been surprising to spend compute and RL on.
ApolloFortyNine 1 days ago [-]
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
sys32768 1 days ago [-]
Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.
ChatGPT 6 Pro answered it without issue.
debesyla 1 days ago [-]
I am honestly still confused about this limitation. I can understand cybersecurity, because mass "hacking" can be automated and Claude itself can help you do it, but biology...? Is it that easy to manufacture and distribute viruses and whatnot?
timacles 1 days ago [-]
I imagine some terrorists in a cave with a lenovo laptop manufacturing bio weapons with some flasks and Claude
dopa42365 1 days ago [-]
Right next to the hypersonic missile vibecoder
rzmmm 12 hours ago [-]
I'm pretty sure one reason is to influence public opinion about LLM regulation. Open-weights models cannot be restricted as effectively and they want to ban those for obvious reasons.
11 hours ago [-]
b112 1 days ago [-]
You can order genes online, and some say you can assemble using stuff cobbled together in a home lab rather easily. For the last 20 years, I've personally felt that bio-terrorism is the highest possible risk, well above nuclear, or chemical warfare. But it does take training, expertise, or, it did.
MisterMunchkin 15 hours ago [-]
Then stop selling genes online?
It’s like saying you sell ammonium nitrate and fuel oil online and then saying it’s too risky to let people have computers in case they use them to make ANFO. They can only make bioweapons because you’re selling them bioweapon components! They can’t make genes at home!
toss1 5 hours ago [-]
The companies selling genes online already scan the orders and do not fulfill anything considered a possible hazard (presumably unless it is to a known and certified lab at a serious organization).
And, this is not the only way to make genes at home.
It is a complex problem.
b112 12 hours ago [-]
A fair response in general, but it's not as if I advocate it, or run a company doing it. I'm simply stating fact.
And I will point out that "stop doing that" can be applied in both directions, towards AI and towards material supply.
And that this is one way, not even remotely the only way, that gene editing at home is easy.
Again, these are just facts.
Some questions...
Is it fair to restrict AI, or fair to restrict 1000 industries?
And if it is fair to restrict 1000 industries, OK, but there should be time to do so, probably? A transition period?
And if you do restrict, many such industries just make needed chemicals, which are used by endless other, non-threatening industries.
"Here, we present five case studies of actors using our models in ways that could support biological weapons development."
And capabilities continue to improve.
frabcus 11 hours ago [-]
[dead]
toss1 1 days ago [-]
Considering there are many high-school competitions in genetic editing, some listed at [0] as well as a whole biohacker culture, and labs providing gene sequencing as a service e.g., [1,2], we can reasonably assume it is not beyond the reach of some garage lab to accidentally or deliberately spread a deadly pathogen if it can find the right sequence.
So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.
I love the contrast with yesterday's open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.
Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.
arw0n 1 days ago [-]
It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
1 days ago [-]
1 days ago [-]
kqp 1 days ago [-]
I think it was looser on release for those juicy benchmarks, tighter now. On release I wasn’t getting refusals, then a few days ago I asked it whether a generic quote (think “he walked to the store”) broke standard punctuation rules, and it blocked me for breaking rules. I wish I were joking. Rephrasing to not use the keyword “rules” worked.
plaguuuuuu 4 hours ago [-]
Fable 5 refused to tell me whether I could grow a coconut tree in my backyard.
I'd say it's been loosened a bit since then.
cute_boi 1 days ago [-]
If they don't loosen, people will choose Astra or Chinese model.
Giving moral lecture is different than reality i guess.
prettyblocks 1 days ago [-]
They're pushing their customers to their own competition by doing this.
Espressosaurus 1 days ago [-]
It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
raesene9 1 days ago [-]
For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
flyinglizard 1 days ago [-]
They are pushing their customers towards Chinese models and providers. If you want to get something cutting edge done in defense, cyber, biology - something that isn't common knowledge - you need to venture east. That's an incredible side effect which the Chinese government surely enjoys.
Metacelsus 1 days ago [-]
I want to like Anthropic but this is just pushing my startup to use OpenAI
nonethewiser 1 days ago [-]
I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
bushido 1 days ago [-]
One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
ACCount39 1 days ago [-]
You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
doginasuit 1 days ago [-]
In what situations might Opus typically refuse to help with cybersecurity? I've been using it to find security issues in a web app that I wrote. I've expected it to refuse at some point but it will happily analyze it to find issues. I've just asked it to read source, not actually do any testing.
Notice that this isn't cybersec nor memory-safety related at all.
bottlepalm 14 hours ago [-]
Anthropic needs a Daybreak program like OpenAI does, I keep having to go back to ChatGPT for cyber work.
paimapi 1 days ago [-]
see what we need is another technocratic priest class that unaccountably decides who deserves access to salvation based on how much cash is paid out and how powerful the patrons are
yaakov34 1 days ago [-]
This has become insufferable. I work in a medicine-adjacent field, but nobody in their right mind could possibly take what I do to be in any way related to some kind of bioweapon or whatever the hell they're pretending to be saving us from. The dumb Fable guardrails made me stay with Opus, now that this is coming there, we'll be saying goodbye.
user43928 1 days ago [-]
kernel development is now also banned:
>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.
But hey, they 'should not impact the vast majority' of ML development. Great.
dannyw 23 hours ago [-]
IIRC it only blocks kernel development for Huawei and other Chinese chips.
Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.
KeplerBoy 1 days ago [-]
Anything else would be inconsistent, wouldn't it?
SoftTalker 1 days ago [-]
Who is "vetting" organizations and to what standards are they being held?
searine 1 days ago [-]
Great. Claude is basically useless for bioinformatics now.
unglaublich 1 days ago [-]
Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.
nijave 1 days ago [-]
Luckily all the other LLM providers are also still making progress with less onerous "safeguards"
b112 1 days ago [-]
Very unfortunate indeed. As a Canadian, I don't want to use Persona, which isn't legally bound by Canadian privacy legislation. I'll never install any Persona apps on my phone either, and the sad part is that domestic eid providers often use Canada Post to ID people for them. EG, if you don't want to install an app, or can't.
So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.
Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.
At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.
manlymuppet 20 hours ago [-]
The endless cynicism in this thread is so very draining. Can we not do this terrible circle of excessive criticism?
The productive comments here are so far and few between. I have no issue with people criticizing Anthropic or AI companies in general, but for the love of everything, at least make worthy criticisms. Not these incessant sophisms.
filoeleven 9 hours ago [-]
The cynicism will stop when the musical chairs of money fueling the bubble does.
manlymuppet 47 minutes ago [-]
How does endless cynicism without purpose help this situation you describe?
Tadpole9181 6 hours ago [-]
The one I've gotten extremely tired of is every single AI thread now having dozens upon dozens of the exact same joke where people talk in Claudish.
It was funny the first time. But I genuinely don't understand people 6 levels deep or 30 comments down the thread thinking, "wow - this'll really knock their socks off".
mikert89 12 hours ago [-]
Yeah this website is almost unreadable
15 hours ago [-]
pyronite 20 hours ago [-]
It almost makes you understand how a well-coordinated artificial superintelligence could wipe us out, eh?
OrangeDelonge 19 hours ago [-]
I for one welcome our AI overlords.
ilt 12 hours ago [-]
I am mostly excited about all this but also feel incredibly sad at times about many aspects of it. We need a new word for this feeling pronto.
ctoth 6 hours ago [-]
Deep Blues.
Skidaddle 6 hours ago [-]
blursed?
puttycat 12 hours ago [-]
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1
According to their ToS, conversations flagged by the guardrails when using Fable are stored longer for moderation. That included Fable's guardrail for AI research. Moderation can mean a human looking at your conversation.
Does that mean that I can no longer trust Opus with not snitching my AI research to Anthropic either?
palata 11 hours ago [-]
> Does that mean that I can no longer trust Opus with not snitching my AI research to Anthropic either?
You never could trust it. Fundamentally you leak everything your LLM uses to their servers.
port11 12 hours ago [-]
[dead]
techjamie 1 days ago [-]
With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
This model was likely trained months before deepseek released their paper.
manquer 1 days ago [-]
Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
dannyw 23 hours ago [-]
Models are generally posttrained to a window shorter than you think.
Balinares 1 days ago [-]
Unless they already have something similar of their own, which is always possible, they'd be stupid not to. I don't suppose we'll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they're down to copying Chinese tech.
cpldcpu 11 hours ago [-]
You mean the one they copied from Microsofts paper? (properly cited as well)
ACCount39 1 days ago [-]
I don't think it's particularly relevant?
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
ryangg 1 days ago [-]
Getting a 403 on that link. Mind checking it once?
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
potwinkle 1 days ago [-]
I'm able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
zatkin 1 days ago [-]
It's working for me (based out of California).
MisterMunchkin 15 hours ago [-]
That post is just AI slop. You forgot to strip out the slop lines.
joshstrange 1 days ago [-]
> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
bleonard 1 days ago [-]
The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%
So longer threads get cheaper and one-shots stay the same price.
bayesianbot 1 days ago [-]
Wow those cache reads are quite reasonable - I think that's equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it
abtinf 1 days ago [-]
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
onlyrealcuzzo 1 days ago [-]
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
This is news to me. Excited to try it out! Thanks.
nchmy 1 days ago [-]
news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
KeplerBoy 1 days ago [-]
With the pi harness it just opens the browser (or gives you a link if you're on headless) and you sign on as usual.
nonethewiser 8 hours ago [-]
What sort of things are you building? What do you like about exe.dev? Just curious what teh overhead and $20/month subscription is enabling for you.
>I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I'm building.
abtinf 6 hours ago [-]
For me, the essential difference between exe and local harnesses is that they have really nice built in methods to take care of stuff like auth and inter-vm interactions, which makes it easy to build live internet-connected services that I can share with other people. There is no deploy step, which is a surprising amount of time savings. They also have nice things like email receive/send, the ability to issue phone notifications through their app, and just a ton of little things where it feels right.
FWIW, the $20/month subscription also includes $20/month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.
Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):
Shelley is a fantastic agent and their batteries included vm image makes the most of it. It includes things like a browser for Shelley to check its own work. And Shelley has root and full access to them, so it can solve any problem and do pretty much anything you need.
cbg0 1 days ago [-]
How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
qlte 1 days ago [-]
Per the link someone else posted, the actual difference in $/task is not nearly so stark:
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):
Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.
So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
Gareth321 14 hours ago [-]
Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost/performance frontier, from ~$1.34/task through ~$6/task.
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
TuxSH 22 hours ago [-]
Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they're even worse than Astra's
margorczynski 1 days ago [-]
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
jasbury 18 hours ago [-]
This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
abtinf 1 days ago [-]
A cheaper price has no value if I can’t use the thing I’m paying for.
The Claude lock-in simply disqualifies anthropic entirely (for my use).
1 days ago [-]
copperx 1 days ago [-]
> Even on a subscription you'll get considerably more usage out of Opus.
That's an incredibly bold assumption.
cbg0 1 days ago [-]
It's not, I have a subscription to both and Astra burns usage like crazy.
notatoad 1 days ago [-]
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
polalavik 1 days ago [-]
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
abtinf 1 days ago [-]
Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.
roughly 1 days ago [-]
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
Can you give more details here? This sounds intriguing.
sidrag22 1 days ago [-]
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
mlcruz 1 days ago [-]
What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
1 days ago [-]
felixgallo 1 days ago [-]
If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
abtinf 1 days ago [-]
I read the page. It seems like a marginal improvement.
mattz56 10 hours ago [-]
Better than Astra looks insane. I saw a Higgsfield video yesterday where they gave the same prompt to Astra and Opus 5.5 to create a samurai video game, and the difference was huge.
ryanscio 1 days ago [-]
Let's wait for independent benchmarks at least
felixgallo 1 days ago [-]
the benchmarks provided are already from independent organizations:
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
Just post the bloody content. This UI/scrolling thing is horrific.
amluto 1 days ago [-]
Claude Opus 5.6 should have a new "UX safety" feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)
gruez 1 days ago [-]
???
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
KyleTheDev 1 days ago [-]
If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.
I agree that it's sort of stupid, not a fan.
ealready_value 1 days ago [-]
It's less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra's hero/scrolling animation. I'm pleased to see they didn't make the entire page that like OpenAI did, but I'm expecting to encounter this pattern more often on these announcements now.
EricBurnett 1 days ago [-]
Two posts were merged; this comment was for the blog post with an intro animation thing.
mbreese 1 days ago [-]
On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
thejazzman 1 days ago [-]
then you're getting served a different website
giancarlostoro 1 days ago [-]
On mobile its different.
iAMkenough 1 days ago [-]
Figured it out: you have "reduce motion" enabled in your device's accesibility settings.
Everyone that doesn't gets served some animated bullshit.
gruez 1 days ago [-]
>Figured it out: you have "reduce motion" enabled in your device's accesibility settings.
Yep, you're right. I tried on my phone and got the scroll through image.
swader999 1 days ago [-]
I told my team to smack me upside the head if I ever try to ship something so daft as that.
serchinastico 1 days ago [-]
The performance in Firefox is terrible too, I couldn't make it past the hero
amluto 21 hours ago [-]
To be slightly fair to Anthropic, Qwen does even worse IMO: their model announcements don’t actually show any content at all for me on Mobile Safari. The content box shows up but is just a pulsing animation that never gets replaced by text. At least Anthropic’s announcement works once I manage to scroll it far enough.
anon373839 15 hours ago [-]
The fix for this is to tap the overflow menu icon and choose “Reduce privacy protections”. (Wtf, Alibaba?) This appears to be related to use of iCloud Private Relay.
thebitguru 1 days ago [-]
Totally! So unnecessary and annoying.
halyconWays 1 days ago [-]
I call it scrollslop
josefresco 1 days ago [-]
Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.
oefrha 1 days ago [-]
Parallax scrolling effects were very cool ~2010. By 2015 or maybe earlier it already felt like me-too design that's unoriginal and a little annoying. By 2020 everyone and their mom has it and it's super tiresome. Now it just screams slop design (among a million other signals).
halyconWays 1 days ago [-]
Those sites are also scrollslop. "Slop," as a term, is independent of AI
iAMkenough 1 days ago [-]
Turn on "reduce motion" in your accessibility settings and you get served a sane version.
1 days ago [-]
dionian 1 days ago [-]
and hijacking back/forward
somewhatjustin 1 days ago [-]
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Nice. I was starting to think that Haiku got abandoned.
Sol- 1 days ago [-]
Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.
ricardobeat 1 days ago [-]
Terra lost to Sol and Luna at every cost/performance point, it had no reason to exist.
w-m 23 hours ago [-]
At introduction, Terra was a good mid-tier model. Terra [high] was on the pareto frontier of DeepSWE's score over cost, if only ever so slightly.
When I had Sol orchestrate Luna and Terra as implementation agents, Sol was a lot happier with what Terra produced and would find far fewer issues than what was implemented by Luna.
But a few weeks after introduction, OpenAI slashed Luna's cost by 80% and Terra's only by 20%. Only then did it become uneconomical to run Terra and its reason to exist stopped.
skerit 1 days ago [-]
Retiring the Terra tier? Their space-inspired lineup has only been out for 2 months, they're already messing with it?
bix6 1 days ago [-]
Well you see Terra is being left behind.
somewhatjustin 1 days ago [-]
I personally use up to 3 models. Fable/Opus for planning, Opus/Sonnet for implementation depending on complexity.
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
jaapz 1 days ago [-]
Only reason I used Haiku was when I made Opus 5 run Haiku agents just to rewrite Opus 5's garbage output into something legible
mchusma 1 days ago [-]
I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.
cesarvarela 1 days ago [-]
Claude code still uses it internally.
SatvikBeri 20 hours ago [-]
I'd be really curious to see benchmarks of haiku vs lower effort on bigger models. My own evals found Fable 5.1 at low to be better than Opus 5 on high.
sharkjacobs 1 days ago [-]
> “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it
God I hope so
lgessler 1 days ago [-]
I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.
nonethewiser 1 days ago [-]
whats your provenance on that?
1 days ago [-]
fastball 1 days ago [-]
It hasn't just fixed it, it has introduced a new paradigm in anti-obscurity.
The problem with 5 wasn't just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I'm okay with verbosity if it's actually readable.
mikeocool 1 days ago [-]
That's the load-bearing seam in this blog post.
thefourthchime 18 hours ago [-]
I have to say so far Opus 5.5 gives me much better results than Fable 5.1 did. I can actually understand what it's talking about.
kantahayashi 1 days ago [-]
The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
bushido 1 days ago [-]
Install the simple English skill. OpenAI follows that really, really well.
Someone needs to make a kevin-from-the-office skill. We could really use some of that "why waste time say lot word when few word do trick" here.
rfgplk 1 days ago [-]
OpenAI's models will actually obey a "be concise, leave no comments rule".
RugnirViking 13 hours ago [-]
according to their annoucement for sol-6 and luna-6, they have also tried to improve responses, and the release post has examples of old and new language on the same task to illustrate this
boc 1 days ago [-]
So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
unddoch 1 days ago [-]
It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use.
IME, if you eyeball how many words the thing they're trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic's grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It's terrible.
neilellis 1 days ago [-]
'frontier models' - seriously, it was you and only you!
mavamaarten 1 days ago [-]
That's literally all I'm hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?
claude-ai 1 hours ago [-]
Opus 5.5 is a game changer for me. Super-heavy and deep algorithmic work on a pet project of mine - managed to produce 5% weekly usage in 8 hours. With Opus 5 - interspersed with Fable for supervision and guidance - I'd have used 12-15% in the same time. Mostly void of claudisms; and the results actually work, on a level that's almost on par with Fable.
Impressed.
hgo 2 hours ago [-]
I have to say that I'm very pleased with this version.
My prime anecdotal reason:
I asked Opus 5.5 to make a sweeping change with a few Sonnet agents. It clarified the scope, we agreed and then it started. I saw some problems on the way and asked it. It's answer:
This didn't go as planned. The Sonnet agents don't have the right judgement capacity for this, so they take shortcuts based on their limited scope X. I suggest we terminate them, roll back their changes and I can schedule an Opus agent to do it instead. It will be slower but done correctly.
There is a first time for everything.
Every interaction I've had with Opus 5.5 so far feels like talking to. An adult.
gtirloni 16 hours ago [-]
The hype/degrade/hype marketing cycle that Anthropic uses to cycle between Fable/Opus or Opus/Sonnet releases is too obvious at this point.
Opus 5.5 will be great for 2-3 weeks than the nerfing starts. They give out some extra credits. It keeps degrading Opus until it's unusable . By that time they are ready to release Fable 6. Which also is great for a few weeks, and so on.
s1n_ 14 hours ago [-]
So no issue hopping from new model to new model? Feel like every frontier company is doing this lol. Would be more alarming for this to not happen.
zuInnp 1 days ago [-]
So it Opus performs as well as Fable what is then the selling point of Fable?
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
orangecat 1 days ago [-]
All of this starts to feel more like a drug dealer selling their newest stuff.
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
ieie3366 1 days ago [-]
? they will obviously release Fable 5.5 soon(tm). It's same as hardware. The previously top tier product gets obsolete
ACCount39 1 days ago [-]
New generation's "upper-mid tier" offering claims to be almost 1:1 match for the previous gen's "top tier" - in other news, fork found in kitchen.
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
glub 1 days ago [-]
Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.
Benchmarks often don't survive contact with reality.
drnick1 1 days ago [-]
That's not my experience at all. Opus is an extremely capable coder on high or xhigh effort. It can read academic papers, implement algorithms from the description in the paper alone and reproduce results without breaking a sweat. This is remarkable because it is pure reasoning on unseen material; in some cases the paper was just published and there wasn't an implementation to learn from in the training data.
cowthulhu 1 days ago [-]
My experience is that Opus can definitely write decent code, but it incurs tech debt and adds unneeded complexity.
cheikhcheikh 1 days ago [-]
did you actually verify that it's output in those scenarios is good ? in my experience opus has been a disappointment and constantly trailing behind actually solving hard problems versus the OpenAI models.
I'll say that both have terrible writing style though.
drnick1 1 days ago [-]
> did you actually verify that it's output in those scenarios is good ?
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
arw0n 1 days ago [-]
Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
boredtofears 1 days ago [-]
Mines pretty much inverted - my colleagues and I noticed almost zero difference between the quality of code in Opus vs Fable. Occasionally I'll switch to Fable for an arduous debugging task but that's about it.
notatoad 1 days ago [-]
typically, when any AI company says a model performs as well as fable, all they're really telling us is that the benchmarks that exist for measuring AI capabilities aren't very good.
giancarlostoro 1 days ago [-]
Fable should have just been called Opus Primt for Enterprise and sold only to enterprise customers. I don't even use it. I rather just use Opus.
Imustaskforhelp 1 days ago [-]
isn't this what mythos was/is?
commotionfever 21 hours ago [-]
Fable is the same model as Mythos. Just that Mythos doesn't have the guardrails and is not for public use
quotemstr 1 days ago [-]
Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.
1 days ago [-]
throwaway2027 1 days ago [-]
After yesterday outage is the new Opus 5.5 load-bearing?
lgessler 1 days ago [-]
I should find information about the user's concern instead of just assuming.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
jaapz 1 days ago [-]
This made me shiver
lgessler 17 hours ago [-]
I'm gonna miss Opus 5. Every model has had its quirks ("You're absolutely right!"), but Opus 5 was serving up chicken fried tokens like none other. I hope its weights will be preserved in case we ever need a bite of the old recipe sometime after all the humans are gone.
aenis 3 hours ago [-]
Agree. I spent so much time with it that it infected my own command of English. I hope something new comes and bears the load.
handfuloflight 1 days ago [-]
It's worth stating why, and depends what seams you're pulling at this sitting.
rich_sasha 1 days ago [-]
Your instinct is basically right, and the research backs it up.
1 days ago [-]
cmrdporcupine 1 days ago [-]
And here's the important part...
danw1979 1 days ago [-]
you win the thread
staticman2 1 days ago [-]
I'm gonna be straight with you—I don't have the evidence to say whether or not it's load bearing.
danw1979 1 days ago [-]
Good point — but I’ll gently push back on that. It’s not an outage, it’s a service degradation.
ThouYS 1 days ago [-]
You're right to bring this up - and this is where it gets interesting
hmokiguess 1 days ago [-]
You're right, this changes everything, and here's why it matters.
Retr0id 1 days ago [-]
Your premise is half right, and the half that's right is better than you think.
aoeusnth1 1 days ago [-]
You were right to call that out, and the evidence makes a stronger case than you are stating.
sailfast 1 days ago [-]
Let me verify before I come back to you with an answer that is incorrect.
cronin101 1 days ago [-]
It certainly _seams_ that way
carlos-menezes 1 days ago [-]
One thing worth flagging here: 5.5 appears to be a load-bearing seam in the numbering system.
loopmonster 1 days ago [-]
That's the sharpest point anyone has made in this thread so far, and it reframes the entire conversation.
RGS1811 1 days ago [-]
This question is real.
bibimsz 1 days ago [-]
One pushback: there is no Opus 5.5. You might have meant Opus 5.1, the latest Opus model available.
fghorow 1 days ago [-]
"Danger Will Robinson!"
esafak 1 days ago [-]
Wrong century, brother.
fghorow 21 hours ago [-]
I know. I know. I grew up in the '60s. Feel free to unfollow me (or whatever it is that one does on HN).
esafak 18 hours ago [-]
It's all in jest! I apologies for raining on your parade.
outworlder 21 hours ago [-]
Not really, the newest Lost in Space reboot is only a few years old(last season ended in 2021)
tda 1 days ago [-]
[flagged]
sznio 1 days ago [-]
I'm more excited by the Haiku 5.5 announcement buried in this post. I'm wondering if we will finally get a decently capable fast model.
booty 1 days ago [-]
If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.
NorwegianDude 1 days ago [-]
OpenAI didn't mess up. The model would have been 100 % pointless and obsolete without the large price cuts it got, because of the cheap Chinese models.
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
copperx 1 days ago [-]
Better than what? Deepseek? GLM? Gemini 3.8?
copperx 1 days ago [-]
Why are you excited about it? Deepseek is everything Haiku wishes to be and more.
lanyard-textile 1 days ago [-]
Agreed. They've been so quiet about it, and retirement for Haiku 4.5 is right around the forner.
ygouzerh 1 days ago [-]
What are you using Haiku for?
Zambyte 1 days ago [-]
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
adastra22 1 days ago [-]
Things that Jev is probably a better tool for.
mavamaarten 1 days ago [-]
I use it for executing well-prepared plans sometimes. And for exploring larger codebases.
system2 1 days ago [-]
All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.
enraged_camel 1 days ago [-]
We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
anthonypasq 1 days ago [-]
bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper
enraged_camel 1 days ago [-]
We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.
Game_Ender 1 days ago [-]
How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?
magicalhippo 1 days ago [-]
I've reached the saturation point.
I don't have time to really get to know one model before the next is out, and I'm just talking about OpenAI and Anthropic, never mind the long tail of alternatives.
So I just more or less haphazardly pick one based on the mood I'm in, and set reasoning effort based on how much quota I have left.
s1n_ 14 hours ago [-]
For real. It’s exhausting at this point.
kibae 1 days ago [-]
> Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.
This is where Chinese models are going to eat Anthropic's lunch.
schipperai 21 hours ago [-]
They have for sure improved the safeguard classifier. I have a coding agent guard OSS tool [1] and I use Claude to build it. When Fable first came out, I was simply unable to do any work with it. Nowadays their classifier triggers but very rarely.
However, the moment that someone uses a Chinese model without guardrails to commit an AI-powered 9/11, DC will rush to ban Chinese models. (A happy side effect will be to protect American AI profits.)
So the lack of guardrails is a very risky proposition...
nomel 1 days ago [-]
In those specific domains, sure. What percentage of paying users would you say that is?
boshalfoshal 21 hours ago [-]
We went from open source Chinese models eating _all_ of frontier labs lunch, to them eating some small extremely niche market of frontier labs.
I don't think this will happen, just as Chinese models did not eat Anthropic/OpenAI's lunch on top tier intelligence. The market is way too niche, and the parties already interested in the capabilities behind safeguards are likely already partnered with Anthropic to get around those with Mythos-class models (see project glasswing for cybersec).
AlfeG 1 days ago [-]
So it will not usable to do anything with hardening Your own site....
I'm so tired of this. I just want adjust cookie behavior of own site...
thinkingtoilet 2 hours ago [-]
You have to understand you are in a bubble. 99.99% of people will never be effected by it nor care about it. It will have nothing to do with Anthropic succeeding or failing.
the_doctah 23 hours ago [-]
I don't even want to know what safeguards regarding biology exist.
senko 1 days ago [-]
Opus 5.5 is neck-and-neck with Fable 5.1 and Astra 6 in my vibe-coding tests - maybe even better than Fable 5.1
The above Opus games took ~45min to generate with the cost between $11 and $14 (per ccusage - I'm on a Max sub). Used from Claude Code with xhigh effort.
It's interesting that Astra has a clear style that it applied to both games. It's a refreshing design language, but maybe that's because it doesn't look like something Claude has vibe-coded. In terms of gameplay depth, the Claude versions appear to be closer to the original Minecraft
deaux 19 hours ago [-]
> but maybe that's because it doesn't look like something Claude has vibe-coded
Yes, it's solely because it's new. Back when the very first Opus vibe-coded websites appeared, we found them refreshing. By the 3rd one, we no longer did.
robeym 10 hours ago [-]
A year ago I was trying to have these models write css for me to support mobile styles for an ERP system. There were so many quirks that you typically have with css, and it couldn't figure out how to properly get it done. I'd ask it to fix something and it'd break something else - pretty much as good as any average developer is with css.
Out of curiosity, I tried the same project yesterday with OpenAI's newest model, and it had 0 issues. It seems to have a much deeper understanding of how global styles and local styles work across different modules.
It's fun seeing how even a year ago these models seemed so capable to us, and yet they're still improving greatly. I'm still not using Anthropic. I used to change back and fourth periodically as the models would surpass each other, and get the most expensive plans, but I've settled and am happy with the capabilities of the lowest tier subscriptions now because they've progressed much faster than I've had use for them. Maybe I've gotten more efficient with language and instructing the model since a year ago. Still, it's been a fun ride and I'm excited to see what else we can do with these models as time goes on.
TomGarden 1 days ago [-]
The new communication style still made me react negatively, but I hope it will be better in use.
Quoted:
"Please explain the issue to me.
Claude Opus 5.5:
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval:
/.../
last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all."
rfgplk 1 days ago [-]
Yeah, still abysmal. If you have the time or tokens, see if it will obey an explicit "LEAVE NO COMMENTS WHATSOEVER" command. Opus 5/Fable 5 outright ignored it.
2001zhaozhao 1 days ago [-]
It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Great, I opened several of these in new tabs in Firefox and the entire browser froze and I had to kill it. Can't tell you which one caused it though. I have plenty of RAM also.
jjcm 20 hours ago [-]
Interesting. I just opened all 3 on a fresh install of firefox with no issues - can I ask what OS / do you have any js disabled / do you have webgl disabled?
copperx 1 days ago [-]
By any chance did you tried Deepseek 4 or 4.1 and GLM 5.3 or flash?
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
copperx 1 days ago [-]
Awesome. You have a really neat benchmark.
fr33k3y 1 days ago [-]
glm5.3-flash is multimodal and I\ve been testing it against opus over the last ~10 days and it does very well at a fraction of the opus price...
jjcm 24 hours ago [-]
I should give it a spin. 5.3 wasn't multimodal, but it looks like their flash release was. Thanks for the tip.
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
xlayn 18 hours ago [-]
The new opus is so friking incredible that we had to put some safeguards like in fable... means... we notice that Opus can still do some of the work that we want to charge WAY more and we can't do it if the cheaper model does it...
1 year later... whoops sonnet 1337 gets safeguards... because it's so amazingly incredible... and we are sorry but opus 69 goes up in price...
and we are now releasing claude Astro-pus-able... 1000$ to hear the summary about what you want to ask... if you have to ask the price for the output you can't pay it
mintik 23 hours ago [-]
[cyber] classifier is incredibly sensitive with Opus 5.5 I cannot complete any embedded/driver/system-level tasks. Quite literally not a single task was able to complete today without getting flagged for [cyber], and what's more annoying is their narrow definition of what a cybersecurity specialist should be preventing me from getting an exception..
artdigital 18 hours ago [-]
I got approved for the cyber program without problems and I’m not even a cybersecurity specialist
But even before that, when it flagged me, it just downgraded me from Opus 5 to 4.8 and went ahead with whatever I asked
nullbio 22 hours ago [-]
So stop using their models. The more money you give them, the more you are voting for this.
mgw 1 days ago [-]
They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
mudkipdev 1 days ago [-]
It does mention sonnet and haiku.
simianwords 1 days ago [-]
And not fable lol
1 days ago [-]
enraged_camel 1 days ago [-]
At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.
slowin 1 days ago [-]
Welp, it's now blocking me from doing extraordinarily mundane tasks because of "safety". I've been an Opus fan for a long time, but this instantly made me cancel my subscription and move to OpenAI (which I also assume will screw me soon enough). Chinese models are almost there for my needs, and I can't wait to switch to them and never look back.
sebastiangrill 1 days ago [-]
What tasks? Creating a extraordinarily mundane bomb?
slowin 1 days ago [-]
No, it found a vulnerability in my code and refused to update a report I was working on with the information.
yehosef 23 hours ago [-]
interesting.. Can you have a dumber agent do the work 5.5 refuses?
slowin 23 hours ago [-]
Once it refused, I couldn't get any of the models to continue. I tried Sonnet and it said the same "safety" check could not be bypassed.
aytigra 15 hours ago [-]
I have just continued my pending work (a planning stage with reviews) with 5.5 on medium and it is much better at communicating and much faster(2-3x responses, edits and compaction). Seem to be smarter as well, but maybe it is that I understand what it says now.
I almost feel like it is nice to use again.
garo-pro 1 days ago [-]
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
aesthesia 1 days ago [-]
Sonnet 5 was released a while before Opus 5, so it's just Haiku that didn't get a 5 release.
skunkworker 1 days ago [-]
At this point I'm convinced they are skipping numbers so soon they will be at or ahead of OpenAI's numbering scheme.
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
ekckekcjekfj 1 days ago [-]
And how was the Xbox 360 naming choice a “debacle”, exactly?
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
shibel 24 hours ago [-]
I don’t care about pricing, I don’t care about speed, I don’t care about the agentic coding improvements. Those are already fine. Does it still reply with walls of invented jargon, stitched-up phrases, and manage to cram 10 concepts/subjects in one sentence?
throwuxiytayq 23 hours ago [-]
For your use case, I recommend the GPT-2 model. Fast, cheap, and open weights!
Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5
tomhow 1 days ago [-]
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
vehemenz 4 hours ago [-]
My enterprise org has this model disabled because "security."
Does this mean they literally need to go in and check a box? Or is it a standard thing that enterprise accounts get these later than the consumer/API accounts?
throw03172019 4 hours ago [-]
Most likely they disabled it because it can* retain data if they deem they need to like Mythos/Fable.
andrewt21 3 hours ago [-]
In your real world experience outside of the benchmarks hows this performing compared to Opus 5? Asking because 5 was so horrible for me I switched back to 4.8.
Retro_Dev 1 days ago [-]
> Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
I'm confused how they have been able to create so much public negative perception around distillation. It seems pretty clear that they are the only ones who lose out, and everyone else benefits. I don't have any ethical issues with it, nor is it illegal: at worst it's a ToS violation.
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
ACCount39 1 days ago [-]
The issue with distillation is: one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
TomGarden 9 hours ago [-]
This is a great post and I agree with you on the issues with distillation. I do still feel it's ironic for an AI lab.
As long as labs do not heavily kneecap model outputs, practically all this applies to the training corpus as well.
AI gives ruthless users of AI a leg up over the people who's data it was trained on. "We took the products of your work and used them to undercut you". It's all the same.
The only way I'd be against distilling would be if AI models became owned by the public who's work is used to create them. Of course the AI labs should be paid well, but these models are a product of the entire world's efforts, not only the labs.
wren6991 1 days ago [-]
There's an implication that other companies are improving because they're scraping Anthropic, not because they're investing in better architecture, compute efficiency, or their own synthetic data pipelines. I often see Chinese labs' progress dismissed as "they just distilled Anthropic" and I find it hard to reconcile that with all of the interesting research and open-source tooling that they release.
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
ACCount39 1 days ago [-]
There is, in fact, "that much capability transfer from non-logit-matched distillation".
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.
wren6991 1 days ago [-]
Thanks, this is interesting and there were multiple things I didn't know here.
imron 22 hours ago [-]
> one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.
"You're trying to kidnap what I've rightfully stolen."
Retro_Dev 1 days ago [-]
I'm fine if they put preventative measures in place to protect their work. They already do so. I am NOT fine with their mass manipulation of public opinion to fuel an entirely hypocritical viewpoint. Like, any argument here is hypocritical - but they aren't saying what is REALLY HAPPENING ("distillation steals our work and reduces our profits"), and are actually saying words that make other people fight their battle ("national security", etc).
deaux 19 hours ago [-]
> The issue with distillation is: one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities
The issue is doing.. exactly what OpenAI and Anthropic have done to get where they are?
No, there is no issue.
villish 1 days ago [-]
The workarounds used to bypass Anthropic's security measures are quite illegal. They use stolen credit cards, API keys, and accounts. That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.
That's the moat. Mistral has the capability but not the legal protections.
staticman2 1 days ago [-]
I'm confused why Chinese access to Anthropic A.I. would need to involve stolen accounts.
Couldn't I simply give a Chinese friend my key on Open router?
villish 20 hours ago [-]
Yes but Claude is filtered by the Great Firewall. Anthropic also restricts Chinese access. Their security measures aren't bulletproof but they do catch a substantial amount of those accounts and ban them. That's why PRC labs that distill from Claude need thousands.
foltik 1 days ago [-]
Let’s not kid ourselves, Anthropic would be running their own distillation “attacks” too if _they_ were the ones playing catch-up. They’ve already shown as much with their illegal scraping of pirated books ($1.5B settlement).
I say just let them duke it out. After a decade of regulatory capture and enshittification, it’s nice to see some actual competition again.
villish 20 hours ago [-]
> Anthropic would be running their own distillation “attacks” too if _they_ were the ones playing catch-up
Probably. But if OpenAI or Anthropic stole your credit card to purchase tokens you could sue them. You won't get a cent from any Chinese labs.
> it’s nice to see some actual competition again.
Competition benefits everyone. But this isn't fair competition. A German startup cannot legally do any of these tactics required to bypass Anthropic/OAI's counter-measures. Which makes EU less competitive and therefore less investment in European AI.
deaux 19 hours ago [-]
> A German startup cannot legally do any of these tactics required to bypass Anthropic/OAI's counter-measures.
A German startup cannot legally do what Anthropic/OpenAI have done. And neither could Anthropic/OAI themselves. What's your point again?
villish 18 hours ago [-]
> What's your point again
I said it in my original comment. That is the moat. The reason frontier AI is a two horse race. European labs cannot gain ground because the only way to do it is illegally.
deaux 19 hours ago [-]
Hah. Anthropic and OpenAI used plenty of similarly illegal workarounds to obtain data to create their first models. If a Chinese company had done that first you'd be here saying
> That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.
tomaskafka 1 days ago [-]
Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
XCSme 11 hours ago [-]
GPT 6 Sol is still better and 2x cheaper in my benchmarks:
As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
meerita 23 hours ago [-]
The current tests confirms Opus 5.5 is not verbose. In fact, it outputs using ASD-STEM100 as I recommend. It's a blessing.
aragornii 1 days ago [-]
What I'm mostly interest in is the Communication section. Opus 5 was so convoluted in the way of answering that was really frustrating me.
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
Asmod4n 3 hours ago [-]
So far all sessions I’ve started with opus 5.5 have avx512 support, not like all other models where it’s a dice roll each chat.
1 days ago [-]
m101 1 days ago [-]
Funny how they talk so much about safety when most people don’t give a hoot about it, and actually have quite the opposite reaction
nyx 1 days ago [-]
People aren't the target audience of that part of the post. They're hoping saying enough safety stuff will ward off the looming regulatory sledgehammer.
m101 1 days ago [-]
They actually want that to come protect their business model
pavlov 1 days ago [-]
Anthropic is one of the most valuable companies in the world. Their comms are designed to appeal to a very wide readership.
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
cogythea 1 days ago [-]
Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient
jdmoreira 1 days ago [-]
then they copied that from codex because thats exactly how codex works
NielsHarksen 1 days ago [-]
Where in your app do you find this button?
solsson 9 hours ago [-]
My first Opus 5.5 session ran `pkill -f "cat" -U <user> -n` when it wanted "to stop a stuck cat" (the explanation it gave when I asked why). On MacOS the command kills every process with "cat" in it. I was using effort high and auto mode.
That's the first time since I started running Claude Code nine months ago that a session causes actual harm to other work on the same machine. Can still be a coincidence.
data-ottawa 9 hours ago [-]
Similar experience, the first thing it did for me was update production data models without a confirmation. It has always shown me the command is going to run and then I confirm.
websap 9 hours ago [-]
I'm very curious what your prompt / conversation was. I'm an avid claude code user, and maybe I have something to learn about how not to prompt.
I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.
schipperai 22 hours ago [-]
I have used Fable, Astra, Opus 5 and Sol 5.6 in Claude Code and Codex heavily since each of these have been released. Recently, I settled for Sol 5.6 as my daily driver, but with this release Opus 5.5 might be the one.
Fable is on a class of its own when it comes to coding and orchestration, but it runs out pretty quick and is prohibitively expensive and slow. For me, Astra wasn't as good of an upgrade from Sol when it comes to coding and orchestration, and it runs out pretty quick too. Opus 5 had so much potential but was such a pain to talk to, so I only had other agents delegate to it.
Given this, I kept coming back to Sol 5.6 as my daily driver (with consultations from Fable and Astra when available). Sol has an autistic character that smells like RL deepfry, which can be annoying, but it is nonetheless predictable, communicates more plainly than Claude, and is good at code and orchestration.
But this Opus 5.5 might just be it. Fable-level performance that's cheaper and faster, and communicates plainly and briefly. If it pans out in practice, I might just have a new daily driver and it might be time to raise the bar for what can be achieved.
I'll try out the new Sol 6, though Opus 5.5 seems to beat that release fair and square (at least on paper)
I'm also glad we are paying attention now to the experience of using a model, not just how good it is at 'x' class of tasks.
davedx 16 hours ago [-]
My Mac Studio arrived yesterday and one of the first things I did was cancel my Claude subscription. Happy to be free of load bearing, price gouging, paternalistic "altruists" Anthropic.
Am keeping my Codex sub while I find the best local model, but my plan is to eventually stop with OpenAI too.
anthonyrstevens 7 hours ago [-]
So in a thread about the release of Opus 5.5, this is your contribution? Why?
ieie3366 1 days ago [-]
Quick test for my gamedev project:
It feels like using Fable, but faster, and obviously wayy cheaper token-wise.
Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with
Looks like Anthropic is starting to give bank reset as well:
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
rolandf 7 hours ago [-]
By the way any tips on making Opus 5.5 hardening the security of your applications without falling back on 4.8 ?
I´m trying to get things done but security classifier kill all attemps to better secure my app.
pookieinc 1 days ago [-]
“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
jbellis 1 days ago [-]
Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.
meric_ 1 days ago [-]
Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
randomblock1 1 days ago [-]
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
kimseungyong 21 hours ago [-]
When coding, the top-tier model is ultimately too expensive, so I use Opus, but the news that the cost per use for GPT or Claude is going down is very welcome.
I really like that they’re even changing the writing style. I was worried whether it was a problem with my settings or if my literacy was declining.
Both improvements are welcome.
edude03 1 days ago [-]
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
melonpan7 4 hours ago [-]
Token spend is noticeably less than Opus 5, I can't really tell if it's better though.
bilater 5 hours ago [-]
If you're focused on the price drop rather than the increased capabilities of 5.5 you're ngmi
apeci 10 hours ago [-]
Low value comment, but written by hand: I like the speed improvements. Much more fun to work with Opus now!
variety8675 1 days ago [-]
I hope this actually fixes the terrible writing style of Opus 5
emadabdulrahim 1 days ago [-]
It’s not a 100% fix, but with concise output style on, it’s much better.
akhilome 1 days ago [-]
I found having a reminder at every turn through the UserPromptSubmit [1] hook helped with taming the word salad from 5.
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
buntp 1 days ago [-]
Masterpiece by openai to call their model '6', this model feels already behind
frshgts 1 days ago [-]
Anthropic will pull a PHP and skip '6' to go straight to '7'.
wren6991 1 days ago [-]
Smart move would be to move to year-based versioning (26.09). A 4x advantage
FergusArgyll 1 days ago [-]
Well, they're actually older so it makes sense that their model versions should be ahead
jdw64 1 hours ago [-]
Opus 5.5 fills up the limit way too fast, even though they increased it. I'm subscribed to Max 5x, and with the previous Opus 5, I didn't usually hit the 5-hour limit even when doing the same tasks (I mainly work function by function and then stitch them together). Now, with the same tasks, I hit the 5-hour limit in under an hour. It seems like Thinking has been built into the default model, but the 5-hour limit fills up so fast that it's hard to use it properly.
Catloafdev 1 days ago [-]
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
dgroshev 1 days ago [-]
I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate
(perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
sashank_1509 1 days ago [-]
Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?
dgroshev 1 days ago [-]
No, it's just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
wren6991 1 days ago [-]
> also, what "if you're reviewing this rather than just reading it" even means?
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
dgroshev 1 days ago [-]
Good points, but then even this little snippet is internally inconsistent. If I'm a lay-person, why should I care that "a comment would help"?
Claude is just comically bad nowadays.
redox99 1 days ago [-]
Nah it's definitely a Claude thing. Other models even though they have their style are less annoying and less stereotypical.
cruffle_duffle 1 days ago [-]
“It has the same annoying cadence and writing style with slightly less prominent claudisms.”
Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.
Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.
gekoxyz 1 days ago [-]
It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.
ithkuil 1 days ago [-]
It's unbearable but nothing that couldn't be fixed with postprocess.
mavamaarten 1 days ago [-]
How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.
ithkuil 5 hours ago [-]
"please repeat, ELI5 without analogies (I'm not a child, just ADHD)"
works quite well
aray07 1 days ago [-]
Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
username_my1 1 days ago [-]
I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
meric_ 1 days ago [-]
Remember when OpenAI models loved talking about goblins and whatnot due to the RL?
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
ygouzerh 1 days ago [-]
Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)
Eliezer 1 days ago [-]
It's the reinforcement learning rather than supervised learning.
adastra22 1 days ago [-]
It is the switch from RLHF to RLVR. It benchmaxes better, but benchmarks don't cover human usability.
j_heffe 1 days ago [-]
Maybe it's the time period we're in, maybe I'm just grumpy, but it bugs me that they release a new model every single week and the new one is just a fine-tuned version of the "old" one. If 5.5 performs similar to Fable and really does cost 40% less, then 5.5 really should've just been Opus 5. And they're essentially admitting that they are shipping slop.
brandon272 1 days ago [-]
[flagged]
gwking 1 days ago [-]
I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
brandon272 1 days ago [-]
I agree. Hopefully Anthropic has fixed Opus' ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.
fragmede 1 days ago [-]
It's a load bearing joke that was funny the first time but we're going to beat that dead horse until it starts getting funny again. If you beat it enough, it will get funny. Beatings will continue until morale improves.
qurren 1 days ago [-]
[flagged]
drnick1 1 days ago [-]
[flagged]
mosselman 1 days ago [-]
What I don't get is, why would we still use Fable now? What is its reason for existing? If it is more intelligent and cheaper that is. Why are they advertising it as the model to use for when you really have to think when their benchmarks show Opus 5.5 is better at everything?
ryanscio 1 days ago [-]
Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.
benjiro29 1 days ago [-]
The biggest one is the Cache reads going from $0.50 to $0.20 ... Read/Writes dropping by 25% but Cache reads by 60% has a much bigger impact.
bredren 1 days ago [-]
Notes on communication:
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
34679 1 days ago [-]
I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.
Maybe this model can finally figure it out for them.
johnmlussier 23 hours ago [-]
They killed cyber capabilities so I have to move over to Daybreak on Codex.
jdlyga 16 hours ago [-]
Is this less insufferably annoying than Opus 5? That's the question on everyone's minds. Love Opus 4.8 though.
mewse-hn 3 hours ago [-]
Yeah I've been stuck on 4.8 too, the opus 5 problems seem mostly fixed. It obeys and writes concisely rather than running off, doing its own thing, and emitting a word salad.
wolvoleo 6 hours ago [-]
Yup same question here
leothetechguy 12 hours ago [-]
The Design Language on this announcement page is a perversion of human nature.
ramoz 1 days ago [-]
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
bitexploder 1 days ago [-]
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
leothetechguy 12 hours ago [-]
I have never been a part of the group that believes in frontier labs downgrading models. But this theory seems plausible to me for the first time.
bitexploder 8 hours ago [-]
I guarantee they are at least tweaking quants, caching systems, and finding ways to move serving costs down. This definitely impacts the model’s intelligence at times. There are also a lot of model tweaks, RLHF rollouts etc. I don’t think it means they are doing anything malicious or deceptive. And if Opus 5.5 is literally smarter than Fable? Ehh, it is plausible :)
jatins 1 days ago [-]
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.
Thank you.
wolvoleo 7 hours ago [-]
Is it finally getting better than opus 4 though? That's what I really wonder
calibas 1 days ago [-]
> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.
We can't test it properly because it knows it's being tested.
johntb86 1 days ago [-]
Just make it always think it's being tested, and problem solved.
actionfromafar 8 hours ago [-]
Would it believe that?
jidaigeist 1 days ago [-]
>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
andriy_koval 1 days ago [-]
My bet is anthropic has NN people org who work hard to distill open models in addition to trying to find what other useful materials they can download from shady torrents.
b38484848 1 days ago [-]
it's not safe unless it has commitees with orgies with that weird harry potter dude attached
the_gipsy 1 days ago [-]
"bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.
somewhatjustin 1 days ago [-]
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Nice. I was starting to think Haiku was going to be abandoned.
notduckrabbit 1 days ago [-]
They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.
manmal 1 days ago [-]
That cost reduction seems to stem from cheaper cache reads, mostly.
1 days ago [-]
glub 1 days ago [-]
> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we
recommend accessing our models with reduced cyber blocking classifiers via our Cyber
Verification Program. Claude Opus 5.5 will be available through this program in the near
future.
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
accrual 19 hours ago [-]
Anecdata ahead: I took some time to work with Opus 5.5 today. The work feels strong and comprehensive. I'm not getting the ugly "claudish" speak of Opus 5.0. Communication feels more natural and pleasant to read. Good results so far in my preliminary research.
jacobgold 1 days ago [-]
I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.
aurareturn 1 days ago [-]
I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.
I tried Opus 5 and Astra.
blfr 1 days ago [-]
It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.
ironqcold 23 hours ago [-]
The thing I'd actually want to know: is Opus 5.5 Medium genuinely equivalent to Astra High on real work...
artursapek 23 hours ago [-]
In the niche benchmark I run it’s on par with 6-Sol, not close to Astra.
tomaskafka 1 days ago [-]
"You're right, and it's the exact thing I flagged two turns ago and then did anyway." - Opus 5 xhigh, today.
About the time.
sunaookami 1 days ago [-]
Tested it in the last few hours and it's MILES better than Opus 5. Finally the output is readable again!
madjam002 1 days ago [-]
I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.
It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.
It's frustrating that there isn't more transparency here.
aytigra 15 hours ago [-]
I didn't notice it for 5, it was almost too slow to be comfortable, but 5.5 seem to be 2-3 times faster.
ianberdin 1 days ago [-]
Best cost for a good result on our MacBook Pro SVG benchmark.
Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper
thibran 1 days ago [-]
Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.
arendtio 1 days ago [-]
So funny how both OpenAI and Anthropic post outdated pages at the same time. Opus 5.5 has benchmarks against Sol 5.6, and Sol & Luna 6 have their benchmarks against Opus 5.
Why do those labs keep releasing on the same day?!?
herpdyderp 1 days ago [-]
Probably one gets there first, then the other rush-releases theirs.
It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.
Foobar8568 1 days ago [-]
I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.
jdthedisciple 1 days ago [-]
I dare anyone to convince me the benchmarks are
not meaningless.
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
enraged_camel 1 days ago [-]
Ah, so you didn't read the article.
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
jdthedisciple 1 days ago [-]
You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate
> how would this alleged difference (most likely bs) actually show up in reality?
Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright
All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?
1 days ago [-]
breezybottom 1 days ago [-]
"Where Opus 5.5’s advantage is very clear is efficiency."
Not efficiency in writing, clearly.
robertwt7 19 hours ago [-]
excited to try this out. the thing with new models though even if it claims to be cheaper, sometimes it spent less tokens on different task. I found that with Astra I'm actually not so far off from Opus 5 since it spends much less token to complete a task. The claims in the article that it surpasses Astra on a lot of things is interesting to test though
21 hours ago [-]
toephu2 1 days ago [-]
When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M).
Are the frontier labs even working on this problem?
mnicky 1 days ago [-]
Why would you use max? It's usually unnecessary and even prone to overthinking. In my experience, since Opus 5 the medium/high is usually enough (until 4.8 I used xhigh, but never max). Even low is quite usable these days..
toephu2 24 hours ago [-]
actually I meant to say xhigh.
even at xhigh I get context compaction quite a bit.
0xMihir 3 hours ago [-]
yeah, i didn't notice this on 5
doodlesdev 1 days ago [-]
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.
1 days ago [-]
isodev 1 days ago [-]
So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.
spiderice 1 days ago [-]
You didn't have the patience to scroll down, so you decided to come post about it here and waste all of our time?
isodev 1 days ago [-]
The site is horrific so no, I didn't scroll.
b38484848 1 days ago [-]
we will be agi in six months as in the last 36 months
anthonyrstevens 7 hours ago [-]
Nobody reasonable is suggesting that. Strawman.
dezmou 24 hours ago [-]
I don't understand, so it outperform fable 5.1 in every way and is cheaper ?
Why do they insist on the fact that is outperform opus 5 and not fable 5.1
km144 1 days ago [-]
I think this release is really going to give them a hard time selling Fable:
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
booty 1 days ago [-]
"benchmark margins have become a less
reliable guide to real-world differences"
sounds like a big problem.
My guesses:
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
Solvyx 18 hours ago [-]
[flagged]
simianwords 1 days ago [-]
It’s likely that they have internal benchmarks but they are communicating to people who can only gauge through external benchmarks.
suddenlybananas 1 days ago [-]
Why wouldn't they report these benchmarks?
CPLX 1 days ago [-]
Opus 5 fucking sucks. Like it's horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don't want to burn as much quota.
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
port3000 1 days ago [-]
I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
booty 1 days ago [-]
That's interesting.
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
cbg0 1 days ago [-]
I use it frequently with a lot of success on "Medium" effort, it overthinks like crazy on higher levels, but YMMV.
Syntaf 1 days ago [-]
Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
kwar13 14 hours ago [-]
been using it for half a day now. the token spend is noticeably less.
nullbio 22 hours ago [-]
I'm sure it's a good model, but I can't in good conscience sign up to Anthropic again given how they behave. What a shame.
Retr0id 1 days ago [-]
> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).
Yay, yet another model I can't use for anything interesting, even with CVP.
Retr0id 21 hours ago [-]
I thought maybe I just had an unlucky first prompt, but man, this thing is an absolute pain to work with. I'm getting innocuous bash commands flagged as "[Third-Party Attack]". This seems to be related to the command auto-approval, switching to manual seems to be helping.
desmondl 1 days ago [-]
I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.
desmondl 3 hours ago [-]
Update: It's pretty good.
It still covers as much detail as Opus 5, and provides a healthy amount of caveats, but it's much better at just providing the answer up front and framing the caveats.
5.5 communicates like a real collaborator and not something that constantly demands my mental energy to decode what it's trying to say.
In the application I'm working on, we don't want "user's local midnight" to mean "next day". If a live sports game extends past midnight, we don't want the UI to group it under "Yesterday's games." So I asked Opus 5.5 "can we use the backend's API response as the definition of today instead of the browser's local time?"
Opus 5.5:
> Your assumption holds, and the commit history backs it up. Every time the frontend decided for itself what "today" meant, it broke and was later replaced with the backend's currentSeasonDay. There's one catch: trusting the backend only works if the client's copy of currentSeasonDay is current, and right now it often isn't.
Opus 5 would have buried the answer in some kind of unhinged lecture. Something like:
> The assumption is right about ownership and wrong about possession.
> currentSeasonDay is the authority; the commit history has already paid for that conclusion. Every local reconstruction of “today” became a second clock and was deleted. But naming one clock does not make every copy of its reading current.
> The remaining failure is on the other side of the seam: the frontend no longer invents the day, but it can preserve an old one indefinitely. The source is right. The observation is stale. Those are not contradictory states.
> Do not reopen the ownership decision to solve a freshness defect.
Basically the point is just poll the endpoint every 5 minutes or so
1 days ago [-]
1 days ago [-]
alpineman 1 days ago [-]
So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
palata 11 hours ago [-]
Oh, that explains it. I'm still on Opus 4.8 (5.0 was too annoying), and I thought I had missed a few releases...
voiceeh 22 hours ago [-]
They need to catch up to OpenAI, so it makes sense to skip a few numbers.
rkt2spc 14 hours ago [-]
This actually feels like the release of the Opus 4.5 -> 4.6
alvis 1 days ago [-]
$0.20 vs the old $0.5 cache read is pretty much 60% off
unixhero 14 hours ago [-]
I can't get my workplace to approve Claude.
greenavocado 1 days ago [-]
Enjoy it for the next 2 weeks until its silently quanted to 4.8 level
anthonyrstevens 6 hours ago [-]
Evidence for this?
jwpapi 24 hours ago [-]
I don’t want it to talk like me, I want it to talk exactly.
I don’t care to look up terms as long as they are correct.
nickandbro 1 days ago [-]
Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.
keeganpoppen 1 days ago [-]
my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
1 days ago [-]
1 days ago [-]
hugodan 23 hours ago [-]
Are they going to reset the weekly limits after this awful inference week?
tag2103 1 days ago [-]
Why would anyone reward bad behavior?
KasianFranks 1 days ago [-]
Back to Fable 5.1 - Opus 5.5 is now taking 10x longer just as Opus 5.
__vivek 1 days ago [-]
I'm only interested in the Opus series, if they fixed the talking issues.
yipinwong 1 days ago [-]
I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
copperx 1 days ago [-]
$5 per sentence?
yipinwong 1 days ago [-]
I am sorry, I meant to say I generated STAR out of my resume line, trying to generate STAR, and polish it thus $5.
---
I provided crapton of context for that one resume line.
All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
toasty228 24 hours ago [-]
No wonder compute is tight when people burn millions of tokens to polish every single lines of their resume, all that work just for your resume to be injested by another claude clanker once you submit it, what a fucking time to be alive...
yipinwong 18 hours ago [-]
mind your own business
iamsyr 1 days ago [-]
I don't yet have any reason to leave Haiku 4.5 and switch to Opus 5.5.
Lord_Zero 1 days ago [-]
The test they performed to port HAProxy from C to Rust is crazy.
wtarreau 2 hours ago [-]
Agreed. I think they purposely looked for a reputably difficult task for a benchmark between their models, without high expectations beyond that.
HarHarVeryFunny 1 days ago [-]
METR: Is it safe? Has it escaped confinement?
Ants: It's a good model, sir!
1 days ago [-]
blurbleblurble 1 days ago [-]
Hopefully OpenAI throws us some more usage resets now.
aennassiri 1 days ago [-]
Let's see how much they benchmaxxed their model!
keeeba 1 days ago [-]
Opus 5.1 came out about a month ago, what gives?
velcrovan 1 days ago [-]
Just a guess but maybe they decided 5.1 wasn't their last Opus model. Like they would keep developing new versions of it or something.
adastra22 1 days ago [-]
Crazy!
hadlock 1 days ago [-]
Frontier model labs release some kind of update every 6 weeks on average.
rs_rs_rs_rs_rs 1 days ago [-]
That was Fable. Last version of Opus was at the end of July.
richardjennings 1 days ago [-]
My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?
Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.
anonu 6 hours ago [-]
"pace the frontier" is the new "flatten the curve"
datadrivenangel 1 days ago [-]
But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.
nimonian 1 days ago [-]
I recommend reading the web page. It is quite short.
cruffle_duffle 1 days ago [-]
I mean the webpage can say what ever it wants. The proof is using it yourself.
woeirua 1 days ago [-]
So... why would you use Fable now?
thatxliner 1 days ago [-]
So much for pacing the frontier
Yabood 1 days ago [-]
Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.
sandos 1 days ago [-]
Same feeling with oai models, wich I use 99% of the time. Sometimes I ask it about it, and it always come up with a likely explanaton but dear me it does many rounds of tool calls sometimes!
1 days ago [-]
kar1181 1 days ago [-]
Whatever I think of anthropic, that webpage is a truly nice piece of work.
Fizzadar 1 days ago [-]
So is this AGI+ now?
wheremy47 21 hours ago [-]
Why did they retire Opus 4.7 for this?
kingstnap 1 days ago [-]
> It’s good at finding and fixing inefficiencies in software
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
hi_hi 23 hours ago [-]
How do you update Claude Code to enable this? It’s listed under models, but says I have to update for 5.5. I run Claude update and it says I’m on the latest version.
Infuriating.
selcuka 22 hours ago [-]
They might be rolling it out in stages. Mine updated to 2.1.280 and it supports Opus 5.5.
firemelt 1 days ago [-]
wow its really smarter than opus?
1 days ago [-]
thinkingtoilet 1 days ago [-]
Has Opus 5 been absolutely terrible for people today? Like they took resources away from it to make room for 5.5? It is getting very basic things wrong all of a sudden.
karp773 1 days ago [-]
I get this in my claude.ai usage:
Resets
Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
nimonian 1 days ago [-]
You have an extra reset that you can trigger any time before Oct 22
cmrdporcupine 1 days ago [-]
GPT Sol 6 has also released today, but no official blog announcement yet
It seems context length has completely fallen out of the discussion since we hit 1M, is that just going to be what it is now?
nailer 1 days ago [-]
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
wolvoleo 6 hours ago [-]
Yeah I wonder what real world feedback is though. I'm not taking anthropic's word for granted, after all they also deemed 5.0 the best ever when it came out.
I don't have much time to do checks so I'll just stay on 4.8 till feedback improves
nailer 6 minutes ago [-]
I do (now) if you trust me: you can point Opus 5.5 at Opus 5 prose and it will happily unfuck it. This repo is private otherwise I’d send you a link to the PR.
LoganDark 1 days ago [-]
> In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
simianwords 1 days ago [-]
How do I get access to that reset? I can’t find it in my app.
melonpan7 4 hours ago [-]
It doesn't show up on app, but I found it on the web interface.
jdw64 1 days ago [-]
Finally, it seems like a good time to do some 'load-bearing' work on my project for a while
1 days ago [-]
viccis 1 days ago [-]
So it beats Fable 5.1, by quite a bit, on every metric? Interesting.
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
scrollop 1 days ago [-]
Why can't they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.
Arcuru 1 days ago [-]
Great. Now let me use the subscription outside Claude Code.
1 days ago [-]
ricardobeat 1 days ago [-]
Great that they listened! The improvement in communication style looks fantastic. Opus 5 was insufferable and I was on the verge of cancelling my subscription.
1 days ago [-]
anentropic 1 days ago [-]
ooh exaggerated film grain
theplumber 22 hours ago [-]
Use it now as much as possibles. In a week or so it will start to suck as Dario turns the knobs to nerf it.
anthonyrstevens 6 hours ago [-]
Evidence for this?
theplumber 6 hours ago [-]
Evidence is my experience. I think someone documented it as well, it was a top post here.
phendrenad2 1 days ago [-]
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
snvzz 1 days ago [-]
>safeguards hit: [Cyber]
Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.
blurbleblurble 1 days ago [-]
It'd better be good, I'm so tired of the shenanigans
aadyachinubhai 19 hours ago [-]
Huh ..., not another one!
mupuff1234 1 days ago [-]
What happened to "slowing down"?
setsewerd 1 days ago [-]
They're slowing down token usage, not the path to regulatory capture.
icrbow 1 days ago [-]
If you hit wall, hit it hard.
petesergeant 1 days ago [-]
Slowing down only makes any sense if you can coordinate a slow-down for everyone.
mupuff1234 1 days ago [-]
That's just false.
Less companies involved means less pressure to go fast.
nozzlegear 1 days ago [-]
Dario found himself in the prisoner's dilemma.
roughly 1 days ago [-]
Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.
setsewerd 19 hours ago [-]
On the other hand, there's research suggesting that the most optimal behavior for the best outcomes (based on the famously dependable economist style of analysis in a vacuum) is to practice the moral/ethical framework but to also engage in tit for tat - ie, assume everyone means well but respond proportionally when they don't.
WarmWash 1 days ago [-]
Trump got mad and investors sued.
Lord_Zero 1 days ago [-]
The hype train must keep chuggin or it all collapses.
theGeatZhopa 1 days ago [-]
is OPUS 5.5 still not reading CLAUDE.md, failing to follow told tasks, inventing and hallucionating, just refusing to read files ("read the whole file" -> read 2-lines -> infere its wrong -> destroy the codebase), needing constant babysitting just because its so UTTERLY DUMB! i cant imagine going back to OPUS 5 - i'll rather jump out of the window as to use it EVER AGAIN!!
Madmallard 1 days ago [-]
> cyber security and life sciences verification programs
Throwaway accounts posting after a few minutes some anthropic or another ai lab.
Infomercial at its best.
No wonder we are hammered with ai announcements.
danbrooks 1 days ago [-]
Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
It's meaningless because it's unverifiable. Dario all but said we wouldn't notice if they were pacing or not, because they have no intention of stopping development. The bits we could in theory verify are the external audits, whose independence has already been called into question.
At the very least, dumping a new model on the world before the ink on the glossy brochures of "Pacing the Frontier" was even dry calls into question their commitment.
The only concrete action item in the blog post is a call to increase restrictions exports of GPUs and chip-making equipment to China.
If Amodei were serious, he'd be calling for a ban on investment in AI research. It reminds me of all the "land acknowledgements" by people who have zero intention of giving the land back.
The one exception was a traveling theater show with Native peformers and no actual control over the theater, and even then there was some careful rhetoric involved to keep it from sounding like an accusation aimed at the audience.
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)
Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.
AI is a direct threat to people on many fronts. Jobs. AI datacenters. The AI bubble (and the inevitable crash). Electricity and even energy prices to an extent. Water. OpenAI and Anthropic are responsible for this evolution, and them stopping solves close to 50% of the problem, and even if you don't believe the number is that high, it's still a start.
> I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
Oh, so they're killing people's opportunities and jobs because they want to help people? How is that any argument?
Yes, doing the moral thing means making a sacrifice. If you only want to do the moral thing if and only if it is advantage for you that makes you immoral, despite how your actions look. Big tech are masters at this.
Your idea of "making a start" (giving up their position) also would mean they couldn't really do anything else to solve the problem afterwards(?). Sometimes you can improve what's happening in a room more by staying in that room.
> Oh, so they're killing people's opportunities and jobs because they want to help people?
To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they're talking about with pacing seem to be focused more on their existential/AGI worries, not jobs/electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models' capabilities) it trumps the other worries for them. At least that's my reading.
They created said resource. They didn’t mine it.
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
It's just a passing reference to a previous statement and again the burden would be on you (generic you) to explain why this context requires more detail.
From https://en.wikipedia.org/wiki/Safety_car
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.
Valley neuralese
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
Also note that this release isn't a model capability improvement, but a cost and efficiency (and style) improvement.
Yeah I don't buy that, every additional player is just adding fuel to the fire.
Would the cold war be "safer" if it were US, Russia and a 3rd player? Of course not, it would just make it harder to coordinate any safety measures.
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
1. It's just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.
2. It's flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.
Claude seems overly apt to reach for "coinages", or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can't be user-focus groups?
So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were "pacing the frontier" - which either means they're betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it's too buggy...or maybe it means they're worried about regulatory capture? I honestly have no idea.
It's like the whole "Advancing Our Amazing Bet" corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)
Is this the meaning or do I have it wrong? I have not checked.
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
No need to assume, the phrase is literally a link to the blog post the defines it!
It's a strategy to achieve more, not less.
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
You're eliding the why :)
> to help their runner run at a target pace
To help their runner get from the start to the end of the race faster (or indeed at all). In short: to increase performance.
Pacing isn't a neutral thing.
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it's impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you're keeping the status quo. By process of elimination, it must mean slowing down.
I know you said it sounds poetic...but your comment reinforced the parent's point - that this sort of flowery LLM-ish speech is just bad communication.
It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn't finish the TPS reports last night.
Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it's like a voyage with waters and stars...(no I'm not going to quote the original Spanish, I'd simply mangle it).
(To be clear - this isn't a dig at you, as a non-native speaker - I'm simply pointing out that this sort of AI phrasing is often counterproductive).
Actually, when I first read "pacing the frontier", the first thing that came to my mind was a border guard patrolling a border by foot...
So my natural understanding of "pacing the frontier" is that they believe they are in the lead and are setting the pace of AI development. It does not imply they are slowing things down at all, it implies they intend to stay in the lead.
the lead that anthropic and openai had is diminishing faster than they'd like. this is them trying to be the luxury brand of ai.
The "frontier" is made up concept. Its for downstream companies to justify higher token prices. Very little to show for actual improvements.
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
LLMs in a nutshell
But those people believe that sloppy wording is itself a strong signal that the speaker/writer is the one lacking in substance, and doesn't deserve the attention or trust of the listener.
Why do you think it's your responsibility to police who says what about some megacorp, on a random forum on the internets? If I want to criticize some corporation's PR output, I think I'll go right ahead and do that, thanks.
Now a computer scientist might claim that this use of "to pace <something>" is just a generalization of "to pace oneself", but as with many reflexive verb uses, there isn't really an equivalent usage with a non-reflexive object. It's kind of an invention. It's not necessarily wrong to invent a new usage, but usually one does it when there isn't really any other more direct way of saying it, and I don't think that's the case here.
https://www.merriam-webster.com/dictionary/pace#dictionary-e... https://en.wiktionary.org/wiki/pace#Verb
Seriously though, I can't believe people care this much about a stupid phrase - either for or against.
Opus 5.5 is no closer to RSI than Opus 5 was.
There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.
It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.
LLMs have made people so sensitive to language that I fear we're going to throw the baby out with the bath water. The models obviously need work, but they're also a great opportunity to expand our own vocabulary and grammar. It would be a shame if we deny some of the finer points of language in favor of Grug-speak to appease the lowest common denominator.
People aren't objecting to the use of obscure phrases, or flowery English phrases in itself. The issue is where Claude uses it out of place, or in the wrong context, or just plain overuses it.
It would be the equivalent of a young child learning the phrase "Venn diagram" - then using that identical phrase in every single subsequent interaction with others. Cute at first...but very grating after the n-th time.
You used the phrase "baby with the bath water. What if you started using that same phrase in every single HN post you made after this? People would notice very soon.
That's how it is with "load-bearing", or "plainly".
The other issue is where Claude takes a metaphor, and try to contort it to fit all sorts of absurd situations. What if I said "baby with the bath water" could also be used in place of "load-bearing". So every time you imagine Claude saying "load bearing", replace it with "baby with the bath water".
It also seems to misimply that the "frontier" that they release is the same as the frontier behind closed doors. Who's to say they are not throttling full speed towards RSI privately while pacing their public releases?
If I'm being cynical, "pacing" may sound nice, but "fast pace" and "slow pace" are both "pacing".
sir, this is a hacker news thread
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
"there is nothing outside the text" - Jacques Derrida
hypercapitalism, hypermodernity, and finally, hyperreality.
That's not clear at all. How could that be what "pacing" clearly means in the context of the "frontier". How is that more clear than any other pace that could be at issue???
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
The metaphor seems to be like a pacer runner in marathons: If you run too hard in the beginning of a marathon you will blow up and fail, so runners follow a pacer at the speed they can actually maintain safely.
Note that they wont necessarily be slower at finishing the overall race.
https://www.war.gov/News/News-Stories/Article/Article/264106...
You were close, though. :)
if you are in a long race, you don't run all out teh entire time. you pace yourself.
https://en.wikipedia.org/wiki/Pacing_strategies_in_track_and...
That couldn't be more exactly what they are doing here.
Claude says it sounds fine. And Claude is now the judge of the English language style, not you.
It is clear what it means anyway, that's true, it means the left out words, more or less.
And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.
Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
"Pace yourself" specifically means "slow down".
But stating it plainly like this would make the contradiction too obvious.
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
China is literally only a single step behind and willing to drop free models just to undercut the US companies.
I’m for it because I don’t want another massive Google or Meta.
Keep in mind the AIs acted for months without human knowledge, hacked into two major companies (Hugging Face and OpenAI).
This would have blown my mind if explained to me just a few years ago when I thought LLMs would be limited as an architecture.
Or, they fail because the current model is massively sustainable? Also oh noes?
I shouldn’t be, I’m still always surprised that no matter how silly and obvious the fear mongering and propaganda is… someone is always lining up to vehemently defend it.
This you? https://news.ycombinator.com/item?id=49815127 or do legitimate companies need stealth PR campaigns?
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
No, releasing a model that has ~3-6x the performance/cost ratio than your last release just 3 weeks ago is not the same as making one employee 0.001% more productive, but you knew that already.
No one said they can't push the frontier, it's pacing the frontier, which mean very different things.
We've had a year of nearly every model getting quite good at programming. I think with Fable & Astra we are seeing models trained to think at a different level, and I'm not at all surprised or shocked to see them getting passed by their smaller models at coding tasks.
Astra For Coding: Why Are We Doing This? was a great post that didn't say exactly this, but that shows what a weirdo Astra is. https://lucumr.pocoo.org/2026/9/7/astra-why/
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
https://en.wikipedia.org/wiki/Occam%27s_razor
doesn't sound like a razor at all
I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.
Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.
So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:
1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.
2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."
Am I missing another likely option?
To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.
The website is way more popular than before and it's all but taken over by the Twitter AI grifter class.
Many of the people who used to parttake in the interesting discussuons I came here for, have gotten tired of 80% of the posts at any given time being about the brand new AGI LLM that's so much better than last week's AGI LLM, and have left months ago.
I hope you are wrong, or rather, that HN's userbase can keep renewing itself on cycles of new users without getting into an eternal september situation. It's none of my affair (and yes, lately HN AI threads read suspiciously airheaded in tone, like they fell out of /r/codex and Grok-heavy online circles) overall, but I'd be curious if dang or whoever can easily get an idea as to what the current population of HN weekly/daily active users are and if that bag of users is similar to a year over year or decade over decade cohort. X% of our users are more than 10 years old, and so on.
But the above is an idle thought experiment, back to your point, there's definitely some bandwagon work going on, astroturfing and bot-controlled no doubts there too, and if I find one more Rationalist / Effective Altruist / AI doomer cult post, I may find a way to ship them 10 virtual spam palettes because there's more to the web than 'AI-chatterbox-class who's never actually read Neuromancer or Snow Crash or the meatier works of Asimov act superior to the rest of us' strawmen I have confirmation bias running into, yet overall, I still find HN has that spark here and there.
> Alice: ABC-library stinks, why did they write it that way? > Bob Reply: Hi there, yak here, I wrote it because Dennis Ritchie was sick with a cold and we needed it for Los Alamos Labs.
Those gold moments keep me here still.
What did they promise their investors who invested billions?
This type of product would not have been published in a world where we assumed that computers must produce reliably correct results.
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
For an idea of what a serious AI forecaster expects a coordinated AI slowdown to be feel like for the average citizen, see:
https://ai-2040.com/?choices=plan-a-root#playbook-public-pov
that, they fully intend to 'pace'.
what anthropic have stolen they intend to keep for themselves.
If their scare was honest, they would stop.
therefore, their scare is marketing.
And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up
This term is quite ambiguous. Did Dario mean that they need to go faster while making it sounds like they will slow down???
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
"Our technology is so unbelievably powerful that the entire world might shatter if we don't have government imposed handcuffs!!!!". It just reads like the corporate equivalent of the drunk frat guy saying "HOLD ME BACK BRO!"
Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.
Simply make them something that derives a text response from its training data.
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
Imagine if the biotech industry had most leaders tell everyone publicly that what they are building has a high chance of killing everyone and that there are huge risks. If there were people online saying the biotech industry is just fear mongering for investor signalling and regulatory capture you'd role your eyes at the online commentators for their Dunning Kruger effect lack of understanding on how dangerous man-made biological agents can be.
Ironically either communism or an oligopoly are the only viable ways to pace the frontier.
Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.
Idiotic.
https://openai.com/index/introducing-gpt-6-sol-and-luna/
Yeah think I'll be using OpenAI/Deepseek/etc from now on. I don't need your model to decide for me what is and isn't safe.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence
I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded "You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated."
Then I noticed in the corner of the Claude CLI UI that it was showing "Effort: medium". I'm pretty sure I had set it to high effort before; I don't know when it reverted to medium, but that's another thing that doesn't exactly fill me with confidence.
I'll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.
[1] https://github.com/sashiko-dev/sashiko
[2] https://gitlab.com/xen-project/people/gdunlap/xen-review-pro...
It does work out to be a similar cost per task though
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).
I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.
5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
There are old wives' tales on how the original Fable was superb and the stuff of legend,but it as it was leaps and bounds beyond what other models were being offered then Anthropic opted replace it with a neutered version under the same name.
So today everyone can pay to use Fable, but legend has it they are paying for a nerfed replacement released under the same name.
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
Capturing the users is the ad network, that's Google and OpenAI. Capturing corporate trust at a reasonable API cost, that's Anthropic's direction.
China has none of that and they never will for exactly the same reason Baidu is irrelevant globally despite being a highly capable search engine. 'Search' is also a commodity, that's not the value that Google brings to the table.
It's a search engine, anybody can build a search engine = that's what you just said.
They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.
The goodwill/propaganda are convenient, sure, but my guess is that they aren't the primary motivation. Another possibility is that if no takeoff happens, pressuring OpenAI/Anthropic on profitability would exacerbate any damage overinvestment has done to the US stock market/economy.
There's only 12 countries that do on the planet, the most "relevant" of them being Guatemala and Haiti.
As or the LLM topic: you can download weights of chinese models and remove any censorship and bias. Can you do so with american closed ones?
In the end, everyone lost and there are millions of bikes in landfills.
If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.
https://www.youtube.com/watch?v=FQrEDq8KPiU
Now, all this talk of pacing the frontier obviously means that they are afraid of the competition. It could be the open models eating their margins, but also competing frontier models forcing them to invest more and more for diminishing returns, just to keep up. They would certainly benefit from a "Moore's Law" roadmap to pace the advances, and seeing that they lobby for US laws, it would probably mean they are more worried about increasing spending. Though outlawing both open models and Chinese models would be good for their bottom line as well.
The notion of comparing this to Chinese bikesharing is comically absurd.
"Allegedly".
Plus with a lot of accounting tricks, and this not being recurring to make their books look nice for the eventual IPO.
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
For long running tasks it is. That's what made Deepseek so cheap.
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
I gave it some vibecoded patch someone created with Opus 5 with the task of together figuring out the real root cause and what to do about it.
Big mistake. The rest of the session was all claude-speak up until I've rage-quit and restarted with Qwen and no context other than "here's what I think we've missed in the current implementation. How could we approach that?"
GLM felt like it got at least 20 IQ points dumber just from being exposed to claude's writing.
Ironically, what I'm working on a post-processing hook for colorizing and summarizing responses without degrading the session quality. So "truncating" = summarizing, "quantizer" = char limits and thresholds, "collides" = conflicts, "two oranges" is referring to the "alert level colors" where a second model (Haiku/Sonnet) colorizes text based on the perceived (or suggested) priority of a response's statements (e.g. "just so you're aware, I didn't commit" is fucking useless and it needs to be blacked out).
So the original insufferable statement translates to something like "The code that checks whether a text fragment is too verbose was conflicting with the part that colorizes the text."
P.S. Let me know if there's something out there that exists like this- something that adds a dimension like color or priority-assessments on a per-response basis. So far all I've seen is 2 dozen ~100k starred GitHub plugins that add zero value or make things worse.
Code-wise it seems to still nitpick, especially in reviews, but it doesn’t seem to rabbit hole quite as badly on tangents and scope-creep. These are just first impressions though. It’ll take a few weeks of regular use to really have a sense of it.
Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.
We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.
Navigating the landscape of agentic levers certainly requires a more detailed approach than this and you were certainly correct to push back.
and i think i was generally more tolerant of 5.1's writing style than most people.
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
> There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> Rewrites are declared by the publisher, never inferred from overlap
> NULL means dirty, and DELETE is the fence
> Rewrites are declared by the publisher, never inferred from overlap.
This style of writing is idiotic because it conveys no additional information. It's no different from stating
> Rewrites are declared by the publisher, never when moons collide.
The two sentences are actually logically identical. No idea why these models keep writing like this.
> Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.
This is even more ridiculous.
Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.
People that produce slop have to be fired asap, they're just human relays anyway.
I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.
Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.
My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.
Total cost of the above? $0.07 cents.
PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.
Thank you. This alone helps! I think I should jump into it once and try it all out while I still have GLM access at old prices as a backup.
I’d not put Luna and Deepseek in the same tier as Sonnet, they were clearly ahead last time I checked (though I might be outdated and that’s on my personal use case).
I was thinking of going with a subscription of Claude or Codex. The reason (at least that's what I am assuming): with OpenRouter or any PAYG per token setup there will be the anxiety of using up all the tokens in days or maybe 1-2 weeks instead of a month (say I set myself a budget of 15-20 USD per month, average equivalent of a usual subscription price).
Now I don't really want the top-notch models for the coding work I do.
So how much worth of "work/tokens" will I reasonably get for ≈$20 USD if I use it a lot? How much does that equal to - or is equivalent to, say in the world of subscription based Claude, Codex, or even GLM (with their 5-hour and all those cooldowns/limits)?
I am looking for a mental model/framework to visualise this. Can you (or anyone else reading this) please point me to a source where I can get some idea about this? I know I can just add $5 on OpenRouter and try to test. But I don't really know what/how to test these spends. I also want to understand how all this works. (I am new to agentic/llm world/coding, 2-3 months, after a career break of ~3 years, that too after working for more than a decade. I know, not at all good timing!)
Go to the "Cost" -> "Intelligence Index vs. Cost per Intelligence Index Task" That diagram maps their "Intelligence" score to "cost per task" and I think this gives a good basis on deciding with which model you want to go. Then you can either get an API token from that models provider directly or use openrouter and set openrouter to the model/providers of your choice.
You can also see on openrouter itself the details for each model like prices and what providers are offering it at what price.
Finally you can compare models details using openrouters compare feature like this:
https://openrouter.ai/compare/anthropic/claude-opus-5.5/open...
This is one of the things I hate the most. Super complicated workarounds which take loads of time (and sometimes money) for even the simplest problems. Human would pause and ask. I would blame harness, not the model though.
And that all is 0.07 cents all included.
PS: I do not know why but opencode pushes CPU usage to very high which has NOT happened with DeepSeek harness even once.
[0]. https://github.com/deepseek-ai/deepseek-harness
Do you know if they retain your prompts or use it for training?
https://github.com/deepseek-ai/deepseek-harness
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
Sent Opus 5.5 an example that used go context.WithTimeout and it tried to tell me that was wrong and I should pass timeouts as ints before finally admitting the docs it cited didn't say to use ints and that's a ridiculous design in go anyway (it was trying to claim that was codebase convention--passing ints...)
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
Fable 5.1 27 input, 65,927 output
Opus 5.5 27 input, 128,000 output (128k thinking tokens) - incomplete
[1] https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.
Does anyone else find this outright insane? It wrote the equivalent of a full-length novel, just to sit in the question of planning a few dozen shapes.
Academia is going to love this :)
Using claude.ai and Opus, I asked "create a 3d animation from this" and pasted the animation SIML.[1] I just did that test again. There is significant improvement.
Opus 5.5 (high): https://claude.ai/artifact/5EgqfWcyVtLwJDQq6fsPUm
Opus 5 (high): https://claude.ai/public/artifacts/b37a9ee2-f5bc-4ff9-ae90-a...
[0] https://news.ycombinator.com/item?id=49526704
[1] https://news.ycombinator.com/item?id=49532609
Disclaimer: the skills and system prompt on claude.ai could have also improved, this is not a raw API call.
The Pelican is nice, simple and. Lean though.
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
ChatGPT 6 Pro answered it without issue.
It’s like saying you sell ammonium nitrate and fuel oil online and then saying it’s too risky to let people have computers in case they use them to make ANFO. They can only make bioweapons because you’re selling them bioweapon components! They can’t make genes at home!
And, this is not the only way to make genes at home.
It is a complex problem.
And I will point out that "stop doing that" can be applied in both directions, towards AI and towards material supply.
And that this is one way, not even remotely the only way, that gene editing at home is easy.
Again, these are just facts.
Some questions...
Is it fair to restrict AI, or fair to restrict 1000 industries?
And if it is fair to restrict 1000 industries, OK, but there should be time to do so, probably? A transition period?
And if you do restrict, many such industries just make needed chemicals, which are used by endless other, non-threatening industries.
What of them?
"Here, we present five case studies of actors using our models in ways that could support biological weapons development."
And capabilities continue to improve.
So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.
[0] https://www.sciencebuddies.org/projects-lessons-activities/g...
[1] https://www.genewiz.com/public/services/sanger-sequencing
[2] https://plasmidsaurus.com/
https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
I'd say it's been loosened a bit since then.
Giving moral lecture is different than reality i guess.
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
Notice that this isn't cybersec nor memory-safety related at all.
>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.
But hey, they 'should not impact the vast majority' of ML development. Great.
Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.
So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.
Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.
At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.
The productive comments here are so far and few between. I have no issue with people criticizing Anthropic or AI companies in general, but for the love of everything, at least make worthy criticisms. Not these incessant sophisms.
It was funny the first time. But I genuinely don't understand people 6 levels deep or 30 comments down the thread thinking, "wow - this'll really knock their socks off".
According to their ToS, conversations flagged by the guardrails when using Fable are stored longer for moderation. That included Fable's guardrail for AI research. Moderation can mean a human looking at your conversation.
Does that mean that I can no longer trust Opus with not snitching my AI research to Anthropic either?
You never could trust it. Fundamentally you leak everything your LLM uses to their servers.
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
So longer threads get cheaper and one-shots stay the same price.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
This is news to me. Excited to try it out! Thanks.
>I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I'm building.
FWIW, the $20/month subscription also includes $20/month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.
Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):
https://exe.dev/i/rlDF6GI5PGBZV4P
Or
ssh rlDF6GI5PGBZV4P@exe.dev
Edit to add:
Shelley is a fantastic agent and their batteries included vm image makes the most of it. It includes things like a browser for Shelley to check its own work. And Shelley has root and full access to them, so it can solve any problem and do pretty much anything you need.
https://artificialanalysis.ai/models/releases/claude-opus-5-...
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
The Claude lock-in simply disqualifies anthropic entirely (for my use).
That's an incredibly bold assumption.
Can you give more details here? This sounds intriguing.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
I agree that it's sort of stupid, not a fan.
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
Everyone that doesn't gets served some animated bullshit.
Yep, you're right. I tried on my phone and got the scroll through image.
Nice. I was starting to think that Haiku got abandoned.
When I had Sol orchestrate Luna and Terra as implementation agents, Sol was a lot happier with what Terra produced and would find far fewer issues than what was implemented by Luna.
But a few weeks after introduction, OpenAI slashed Luna's cost by 80% and Terra's only by 20%. Only then did it become uneconomical to run Terra and its reason to exist stopped.
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
God I hope so
https://github.com/AminBlg/SimpleEnglish
Impressed.
My prime anecdotal reason: I asked Opus 5.5 to make a sweeping change with a few Sonnet agents. It clarified the scope, we agreed and then it started. I saw some problems on the way and asked it. It's answer: This didn't go as planned. The Sonnet agents don't have the right judgement capacity for this, so they take shortcuts based on their limited scope X. I suggest we terminate them, roll back their changes and I can schedule an Opus agent to do it instead. It will be slower but done correctly.
There is a first time for everything.
Every interaction I've had with Opus 5.5 so far feels like talking to. An adult.
Opus 5.5 will be great for 2-3 weeks than the nerfing starts. They give out some extra credits. It keeps degrading Opus until it's unusable . By that time they are ready to release Fable 6. Which also is great for a few weeks, and so on.
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
Benchmarks often don't survive contact with reality.
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
I don't have time to really get to know one model before the next is out, and I'm just talking about OpenAI and Anthropic, never mind the long tail of alternatives.
So I just more or less haphazardly pick one based on the mood I'm in, and set reasoning effort based on how much quota I have left.
This is where Chinese models are going to eat Anthropic's lunch.
[1]: https://github.com/manuelschipper/nah
So the lack of guardrails is a very risky proposition...
I don't think this will happen, just as Chinese models did not eat Anthropic/OpenAI's lunch on top tier intelligence. The market is way too niche, and the parties already interested in the capabilities behind safeguards are likely already partnered with Anthropic to get around those with Mythos-class models (see project glasswing for cybersec).
Minecraft clone: https://senko.net/vibecode-bench/2026/voxel-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/voxel-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/voxel-gpt-6-astra.html (Astra 6)
Warcraft clone: https://senko.net/vibecode-bench/2026/rts-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/rts-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html (Astra 6)
The above Opus games took ~45min to generate with the cost between $11 and $14 (per ccusage - I'm on a Max sub). Used from Claude Code with xhigh effort.
Full tests with prompts: https://senko.net/vibecode-bench/
Yes, it's solely because it's new. Back when the very first Opus vibe-coded websites appeared, we found them refreshing. By the 3rd one, we no longer did.
Out of curiosity, I tried the same project yesterday with OpenAI's newest model, and it had 0 issues. It seems to have a much deeper understanding of how global styles and local styles work across different modules.
It's fun seeing how even a year ago these models seemed so capable to us, and yet they're still improving greatly. I'm still not using Anthropic. I used to change back and fourth periodically as the models would surpass each other, and get the most expensive plans, but I've settled and am happy with the capabilities of the lowest tier subscriptions now because they've progressed much faster than I've had use for them. Maybe I've gotten more efficient with language and instructing the model since a year ago. Still, it's been a fun ride and I'm excited to see what else we can do with these models as time goes on.
Quoted:
"Please explain the issue to me.
Claude Opus 5.5:
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval: /.../ last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all."
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
1 year later... whoops sonnet 1337 gets safeguards... because it's so amazingly incredible... and we are sorry but opus 69 goes up in price...
and we are now releasing claude Astro-pus-able... 1000$ to hear the summary about what you want to ask... if you have to ask the price for the output you can't pay it
But even before that, when it flagged me, it just downgraded me from Opus 5 to 4.8 and went ahead with whatever I asked
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
I almost feel like it is nice to use again.
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
Does this mean they literally need to go in and check a box? Or is it a standard thing that enterprise accounts get these later than the consumer/API accounts?
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
[1]: https://github.com/p-e-w/heretic
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
As long as labs do not heavily kneecap model outputs, practically all this applies to the training corpus as well.
AI gives ruthless users of AI a leg up over the people who's data it was trained on. "We took the products of your work and used them to undercut you". It's all the same.
The only way I'd be against distilling would be if AI models became owned by the public who's work is used to create them. Of course the AI labs should be paid well, but these models are a product of the entire world's efforts, not only the labs.
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.
"You're trying to kidnap what I've rightfully stolen."
The issue is doing.. exactly what OpenAI and Anthropic have done to get where they are?
No, there is no issue.
That's the moat. Mistral has the capability but not the legal protections.
Couldn't I simply give a Chinese friend my key on Open router?
I say just let them duke it out. After a decade of regulatory capture and enshittification, it’s nice to see some actual competition again.
Probably. But if OpenAI or Anthropic stole your credit card to purchase tokens you could sue them. You won't get a cent from any Chinese labs.
> it’s nice to see some actual competition again.
Competition benefits everyone. But this isn't fair competition. A German startup cannot legally do any of these tactics required to bypass Anthropic/OAI's counter-measures. Which makes EU less competitive and therefore less investment in European AI.
A German startup cannot legally do what Anthropic/OpenAI have done. And neither could Anthropic/OAI themselves. What's your point again?
I said it in my original comment. That is the moat. The reason frontier AI is a two horse race. European labs cannot gain ground because the only way to do it is illegally.
> That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
https://aibenchy.com/compare/anthropic-claude-opus-5-5-high/...
Mostly based on previous relesses info.
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
That's the first time since I started running Claude Code nine months ago that a session causes actual harm to other work on the same machine. Can still be a coincidence.
Fable is on a class of its own when it comes to coding and orchestration, but it runs out pretty quick and is prohibitively expensive and slow. For me, Astra wasn't as good of an upgrade from Sol when it comes to coding and orchestration, and it runs out pretty quick too. Opus 5 had so much potential but was such a pain to talk to, so I only had other agents delegate to it.
Given this, I kept coming back to Sol 5.6 as my daily driver (with consultations from Fable and Astra when available). Sol has an autistic character that smells like RL deepfry, which can be annoying, but it is nonetheless predictable, communicates more plainly than Claude, and is good at code and orchestration.
But this Opus 5.5 might just be it. Fable-level performance that's cheaper and faster, and communicates plainly and briefly. If it pans out in practice, I might just have a new daily driver and it might be time to raise the bar for what can be achieved.
I'll try out the new Sol 6, though Opus 5.5 seems to beat that release fair and square (at least on paper)
I'm also glad we are paying attention now to the experience of using a model, not just how good it is at 'x' class of tasks.
Am keeping my Codex sub while I find the best local model, but my plan is to eventually stop with OpenAI too.
Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
[1] https://kizi.to/claude-talks-too-much/
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
Claude is just comically bad nowadays.
Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.
Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.
works quite well
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Maybe this model can finally figure it out for them.
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
Thank you.
We can't test it properly because it knows it's being tested.
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
Nice. I was starting to think Haiku was going to be abandoned.
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
I tried Opus 5 and Astra.
About the time.
It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.
It's frustrating that there isn't more transparency here.
https://playcode.io/blog/macbook-svg-benchmark#model-claude-...
Btw, we have added Opus 5.5 as default model to playcode.ai
https://arcprize.org/results/anthropic-claude-opus-5-5
Why do those labs keep releasing on the same day?!?
It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
> how would this alleged difference (most likely bs) actually show up in reality?
Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright
All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?
Not efficiency in writing, clearly.
Are the frontier labs even working on this problem?
even at xhigh I get context compaction quite a bit.
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
Yay, yet another model I can't use for anything interesting, even with CVP.
It still covers as much detail as Opus 5, and provides a healthy amount of caveats, but it's much better at just providing the answer up front and framing the caveats.
5.5 communicates like a real collaborator and not something that constantly demands my mental energy to decode what it's trying to say.
In the application I'm working on, we don't want "user's local midnight" to mean "next day". If a live sports game extends past midnight, we don't want the UI to group it under "Yesterday's games." So I asked Opus 5.5 "can we use the backend's API response as the definition of today instead of the browser's local time?"
Opus 5.5:
> Your assumption holds, and the commit history backs it up. Every time the frontend decided for itself what "today" meant, it broke and was later replaced with the backend's currentSeasonDay. There's one catch: trusting the backend only works if the client's copy of currentSeasonDay is current, and right now it often isn't.
Opus 5 would have buried the answer in some kind of unhinged lecture. Something like:
> The assumption is right about ownership and wrong about possession.
> currentSeasonDay is the authority; the commit history has already paid for that conclusion. Every local reconstruction of “today” became a second clock and was deleted. But naming one clock does not make every copy of its reading current.
> The remaining failure is on the other side of the seam: the frontend no longer invents the day, but it can preserve an old one indefinitely. The source is right. The observation is stale. Those are not contradictory states.
> Do not reopen the ownership decision to solve a freshness defect.
Basically the point is just poll the endpoint every 5 minutes or so
I don’t care to look up terms as long as they are correct.
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
---
I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
Ants: It's a good model, sir!
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
Infuriating.
Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
I don't have much time to do checks so I'll just stay on 4.8 till feedback improves
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.
Less companies involved means less pressure to go fast.
chinese models can't come soon enough
we're already getting enshittification
[1] - https://news.ycombinator.com/item?id=6372466
Infomercial at its best.
No wonder we are hammered with ai announcements.