Rendered at 23:26:57 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
OfficialTurkey 5 hours ago [-]
I work at OpenAI and I was the Incident Commander for yesterday's outage.
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.
swader999 4 hours ago [-]
Did you negotiate extra hard for the 'Incident Commander' title? I'm a bit jealous to be honest.
OfficialTurkey 4 hours ago [-]
Ha. It's a role/title for the lifetime of the incident -- it's useful to have someone to keep things moving, keep track of workstreams, and to know who the decision-maker is, especially for bigger incidents. I'm just a SWE who works on infrastructure.
swader999 4 hours ago [-]
Hopefully you get a decked out command center too.
Curious then as to why the incident coincided with similar issues with Anthropic and xAI.
wnmurphy 4 hours ago [-]
Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).
reilly3000 4 hours ago [-]
Thank you. This should be pinned or something.
wnmurphy 4 hours ago [-]
xAI mentioned a Memphis outage, which is where Anthropic was leasing capacity on Colossus 1. I do think this is the explanation.
4 hours ago [-]
ncb_asj 3 hours ago [-]
Shouldn't Astra be the Incident Commander?
dboreham 4 hours ago [-]
Straight from the turkey's beak.
grist_4_da_chil 10 minutes ago [-]
[dead]
theaxeonthanksg 4 hours ago [-]
[dead]
3 hours ago [-]
Chance-Device 5 hours ago [-]
It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.
xnx 26 seconds ago [-]
Why was Gemini was not affected?
mentalgear 5 hours ago [-]
Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.
specproc 4 hours ago [-]
I read Nowhere to Hide recently, really worth it if you can get past Greenwald sticking himself in the middle (start halfway through).
The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.
Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.
kurthr 16 minutes ago [-]
So following this to the obvious conclusion, the NSA is responsible for the closure of the Straight of America, high tarrifs, dropping employment, and high gas prices?
4 hours ago [-]
eggnet 5 hours ago [-]
If it does exist, why would it work this way, and not the obvious way of streaming logs… which would not cause an outage if it failed.
maxbond 5 hours ago [-]
If this hypothesis were true then a spy agency may want to rewrite responses. Every tool in an agent's harness becomes an remote procedure call you can make on that machine. Including a tool to execute a shell command, in many. A harness is completely isometric to a backdoor, it's the same code written with a different intention.
madrox 3 hours ago [-]
There's a million ways a bad configuration can take down a network. Especially if the part that gets squirrley is a black box.
Even if it isn't precisely this, the fact that no one is saying anything is quite surprising.
Edit: I don't want do contribute to FUD, so want to call out this comment and its replies that identify the shared layer as probably being xAI's infra: https://news.ycombinator.com/item?id=49568622
torginus 4 hours ago [-]
Very likely yes. I wouldn't be surprised if they were hosted from the same datacenters even. There has been a story every few weeks about how Musk has sublet X.ai capacity for one company or another.
This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.
0cf8612b2e1e 3 hours ago [-]
The economics of this kind of warfare make no sense to me. The Harkonens must have been incredibly wealthy by running the spice trade, but even after years of saving and plotting said that the transit fees to move the armies by the Guild were ruinous.
How could anyone wage war like this? The defenders will always outnumber the attackers.
freeone3000 3 hours ago [-]
https://acoup.blog/2026/02/24/collections-warfare-in-dune-pa... Lays out an interesting take on this: that armies are quite small, even on developed worlds, because the cost of equipping them is enormous — and the technological advantage is absurd enough that you don’t need a large army. When 300 men can hold a planet, why would you need 1000?
jstummbillig 5 hours ago [-]
Why would that be probable? People on average are fantastically bad at getting probabilities right.
esseph 4 hours ago [-]
Historical precedent repeated over and over, and the US ties to DoW work, and the national security implications. It's actually the Occam's Razor explanation if you know the history.
strictnein 5 hours ago [-]
OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device?
Or did a couple of companies with poor uptime records happen to have overlapping downtime?
Occam's Razor heavily, heavily points us towards the latter.
bottlepalm 3 hours ago [-]
Ok, then where's the report? Don't they usually release a retrospective report after outages?
usernomdeguerre 5 hours ago [-]
What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes.
If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.
3 hours ago [-]
strictnein 5 hours ago [-]
Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers?
> All of these endpoints use existing providers with decades-long history at this point
That is just factually inaccurate. Their data centers aren't old and they lease a lot of compute from companies that didn't exist 5 years ago.
usernomdeguerre 4 hours ago [-]
>Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers.
They all transit the same wires as all other traffic. Copy them at any regional bottleneck. https://en.wikipedia.org/wiki/Room_641A. Additionally, i'd admit that maybe someone(s) at these companies knows. But if we think there isn't any person who would agree to do this then I think we're being naive.
> Their data centers aren't old and they lease a lot of compute...
Again, they transit the same wires as everyone else. Here i'll also add that these companies have been actively courting government relationships (and Anthropic attempting to repair damaged ones), why would they stand on principles here and not any of the other many frontlines they've visibly acquiesced?
I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.
VCFundedGenYer 4 hours ago [-]
I don't think you're a sysadmin - because what you're saying really doesn't matter. It can still all fail at a single point.
svachalek 5 hours ago [-]
I'd say Occam's Razor leans easily to the former as well, given the history of projects that Snowden revealed and were never shut down, plus all of the cooperation with the federal government that's being touted in recent announcements from both companies.
esseph 4 hours ago [-]
Normally there are only a handful of employees on the payroll at each major company that exposes the US or US government to risk.
It is not often the Executives or Legal even know, but sometimes they did. AT&T bent over backwards to help.
This is standard behavior by the CIA and NSA, and has been for a long time.
And is anyone keeping a table of correlations between outages? Sounds like valuable data.
seydor 4 hours ago [-]
The old PRISM servers got overloaded
dgellow 3 hours ago [-]
Why would you default to that explanation? That’s not at all a reasonable default, and I say that as something pretty paranoid with regards to US surveillance
nomel 3 hours ago [-]
Looking at history, it's much more reasonable to assume there's surveillance, since there are whole branches of the government, and departments, that exist for surveillance in the goal of "national security", which this easily falls into. See Marissa Mayer explaining that it's not an option to refuse [1]. I assume this is just the same ole' Room 641A [2].
Worth mentioning that this room takes a split from the main feed and is not in the path of traffic. Whatever is in this room could go down and it would not cause an outage.
novok 5 hours ago [-]
And if the hypothetical splitter is the thing that breaks? Then it would break main traffic too.
AlexandrB 5 hours ago [-]
Interestingly, fiber signals can be split passively[1] (without a transducer), which should be extremely reliable. No idea what technology this kind of application would use though.
We don't know the architecture of this hypothetical spy splitter tho in the current case. It can be less covert and more of a complicated config, etc.
keeda 5 hours ago [-]
Traffic doesn’t need to be routed through it though, just tee’d to it,
arm32 4 hours ago [-]
In the context of arbitrary line-level traffic, sure. But we're talking about robust reverse proxies here if this _is_ what's going on, again, not a fiber tap inside a closet. In that case, there's no such thing as tee'ing to it.
senordevnyc 5 hours ago [-]
Does the NSA even have the ability to monitor all AI traffic like this? Wouldn’t that require tons of data centers that there literally hasn’t been time to build yet? I really have no idea, maybe the asymmetry of the compute required to monitor is way lower than the compute required to serve inference?
usernomdeguerre 5 hours ago [-]
"AI Traffic" is just traffic. If the infrastructure exists to monitor/buffer traffic (it does) then this can be monitored as well. Whether this hiccup was due to them hitting their limits briefly (or turning it on, or etc) who knows.
utopiah 5 hours ago [-]
I bet it's even negligible traffic compared to e.g. Netflix or YouTube. Even a 'huge' context window is nothing compared to the random library of a basic Web page.
dboreham 4 hours ago [-]
Not once you de-dup.
bakies 5 hours ago [-]
I dont think it requires a lot. I think of them as data hoarders more than anything else. It doesnt seem out of their capabilities to store a ton of chats. Maybe they're having scaling problems with the increased data rates they're hoarding and it led to an outage.
it's fair to say they do and have for ages... it's not that hard to assume a well trusted TLS cert is under their control.
strictnein 5 hours ago [-]
What does having a "well trusted TLS cert" enable for them in this case, exactly?
Having a magical cert doesn't mean you can just intercept everything.
odo1242 5 hours ago [-]
On the contrary, it lets you MITM encrypted communications by swapping the website's original certificate for the "well trusted TLS cert"
strictnein 5 hours ago [-]
No, it doesn't. HSTS and other methods prevent this from happening.
kelnos 4 hours ago [-]
HSTS doesn't protect you from this at all. It only requires HTTPS, which a spoofed-but-trusted cert passes just fine.
No mainstream browser (or any browser?) is doing cert pinning.
What "other methods" are there that are deployed and actually in use?
peanut-walrus 3 hours ago [-]
Transparency logs. It's mandatory for a cert to be in CT logs for browsers to trust it. Those are public, if this was happening, someone would have noticed already.
5 hours ago [-]
5 hours ago [-]
5 hours ago [-]
itdaniher 5 hours ago [-]
It would be weird for the Mandatory NSA Logging Program to be synchronous with serving customer traffic, but-
Tracking outages as well as any changes in API response time across these providers could be interesting.
amelius 5 hours ago [-]
Maybe it was a power hub. And we're not allowed to know where the DC is. If we knew it was a power issue, then with other information (perhaps over time) we could determine the DC location. Or something along these lines.
dragonlord664 5 hours ago [-]
Or it’s monopolistic collaboration at the corporate level which would also be a huge scandal in the US
4 hours ago [-]
jjtheblunt 4 hours ago [-]
it's also possible one has an outage, routing extraordinary traffic to the other(s), with cascading failures in quick succession.
StrangeClone 5 hours ago [-]
FBI, open up!!
esseph 4 hours ago [-]
Was thinking this exact thing.
DonHopkins 5 hours ago [-]
The NSA used to spy on Americans!
They still do, but they used to, too.
otikik 5 hours ago [-]
It’s aliens
senordevnyc 5 hours ago [-]
Were there significant API outages too? I didn’t notice any on my production workflows, and I’d assume what you’re implying would cover API routes too, otherwise it seems kinda pointless.
chews 5 hours ago [-]
someone had to swap out the tape drive in Room 641A.
cyanydeez 5 hours ago [-]
amazon?
SadErn 5 hours ago [-]
Yes, this would give the US gov:
1. A universal kill switch.
2. A way to monitor foreign AI usage.
dgellow 3 hours ago [-]
It’s not really a secret that the US government already has both, they control ICANN and are monitoring traffic around the globe since decades
hightrix 4 hours ago [-]
And for this admin especially, 3. A way to censor content they don’t like or alter responses to present the admin in a positive manner
strictnein 5 hours ago [-]
Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.
Also, OpenAI is saying what caused it:
> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"
Anthropic stated their issue started earlier:
> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”
I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
VBprogrammer 5 hours ago [-]
In my experience I've seen plenty of failures caused by user behaviour in these type of cases.
Biggest competitor goes down and all of a sudden you have a lot more traffic...
computerex 5 hours ago [-]
Do you know the probability of all these companies being down at precisely the same time?
computably 4 hours ago [-]
If you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.
schiffern 4 hours ago [-]
If it's a "thundering herd" problem where everyone's harness falls back to less popular providers that don't normally see that much demand, I'd say the probability is pretty good.
Classic cascading failure is consistent with providers failing 80 minutes apart instead of simultaneously.
ceautery 4 hours ago [-]
Yesterday it was 100%
jeffbee 4 hours ago [-]
"precisely the same time" meaning 80 minutes apart?
4 hours ago [-]
dgellow 3 hours ago [-]
Pretty high?
Starlevel004 5 hours ago [-]
> I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
And not only that, when one goes down a bunch of API traffic switches over to the other, spiking demand and knocking it down.
5 hours ago [-]
4 hours ago [-]
owaiswiz 5 hours ago [-]
[flagged]
kibae 1 days ago [-]
Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.
I think down detector doesn't actually have any probes or actual insight into status etc. I think it uses search volume on its own service as a proxy for an outage - so if e.g. lots of people rush to down detector to query to see if SERVICE_FOO is down, it will register as an outage on down detector because loads of people are trying to see if there is an outage even if SERVICE_FOO is actually totally fine.
My hunch is everyone saw that openai and Claude were down and checked for Gemini too. I was using Gemini the whole time this happened without a blip so it certainly wasn't down in my region at least. 3.8 flash is pretty good and didn't miss a beat.
utternerd 1 days ago [-]
Down detector is a self-reported platform, they use a baseline over 6 months from user reports to try to automatically "detect" if there is a real outage, or if its just a couple end-users. Ultimately, the outage is entirely based on users going to down detector and clicking on "Report a Problem".
Huh. I thought they used to also do sentiment analysis of social media (Twitter, at least). I see that they definitely don't today, but has that always been the case?
dostick 1 days ago [-]
So Down Detector is a kind of quantum observation experiment, observing influences the result.
joshspankit 9 hours ago [-]
observing is the result
moomin 1 days ago [-]
Yes, but is it a load-bearing seam?
graemep 1 days ago [-]
You are right, it is. They have now landed a clean fix.
Oarch 1 days ago [-]
They're saving a memory so this can't happen again.
mavamaarten 1 days ago [-]
Nothing another .md file can't fix
bitwize 1 days ago [-]
Now you guys are just adding no-op fuel to the fire.
tikhonj 1 days ago [-]
I made Claude write noöp instead of no-op and it's still amusing a couple of days later :P
selcuka 24 hours ago [-]
> I made Claude write noöp
Tangentially related [1]:
> The billboard ad, located next to the Ikea Tempe store in Sydney, says 'NÖFNIDEA? No tools, no worries’.
You're right to push back. This is bending the leaf on both ends... just my honest take.
patcon 1 days ago [-]
It's really interesting to crawl through the web of phrases that seem "common" to each person in their interactions... I suspect if one were to get to the more niche "meme" phrases that people encounter, it starts to say more about the sort of person the LLM assumes it's speaking to, and perhaps something about their psychological profile...
combobyte 1 days ago [-]
It must be tuned for rage-based engagement because Claude only ever uses the most obnoxious Claudisms on me despite constant reminders to knock it off.
aerhardt 1 days ago [-]
It's a load-bearing poster.
codechicago277 1 days ago [-]
I need to take a step back.
pampas 1 days ago [-]
Hang on — I can apply a double tracked fix to the load bearing path, gated by provenance.
jpettersson 1 days ago [-]
It isn't done — and what's there is more interesting than expected.
fc417fc802 1 days ago [-]
Great plan. I'll slop together a GUI for it in visual basic real quick.
pampas 1 days ago [-]
Plan approved – I'll start a background agent to create the app using react and npm.
bayindirh 1 days ago [-]
It's not only a great plan. It's a turning point in history.
fc417fc802 1 days ago [-]
That was a major oversight on my part. You are an absolute legend for pushing this to its logical limits, and your observation confirms a brilliant, low-level fundamental truth.
cwmoore 12 hours ago [-]
…the one call that requires a genuine decision from you.
mcbuilder 9 hours ago [-]
And I didn't touch it, because it's your call to make...
JohnMakin 1 days ago [-]
Honestly, this is worth looking at — with one caveat.
cloudfudge 1 days ago [-]
That's on me. I've been giving confident advice that doesn't hold up in practice.
robertlagrant 1 days ago [-]
That aligns with your goals of analysing problems, admitting fault and surfacing that through use of clear and concise language. Not just simply wrong — this screams finessed, nuanced, polished communication after an understandable mistake discovered through pressure-tested interlocution.
disqard 1 days ago [-]
This entire thread is amazing!
It's like a kaleidoscope: Human words --> LLM Training --> chatbot-isms --> lovely parody catch-phrases (and this thread will get ingested soon, and be used to train...)
klohto 1 days ago [-]
The load-bearing seam stays, not taking that away;
Not retired, just switching careers. He's sick of this AI bullshit
MuzikPro 22 hours ago [-]
great fun
Flere-Imsaho 1 days ago [-]
The internet is not supposed to work like this. The network was designed for robustness and fault tolerance, which allows it to reroute data if parts of the network fail.
Why are we all depending on one entity for it all to work? Makes me mad.
seanw444 1 days ago [-]
Because more fasterer and more cheaperer.
I hope Reticulum gains traction.
alightsoul 1 days ago [-]
It's also easier to understand. For the internet to be fault tolerant you have to get rid of CDNs and assume everyone needs the same thing. That's more expensive than a centralized "internet" which relies on CDNs and fiber paths exclusive to the regions with most demand. Everything has been optimized for throughput for what is determined to be important, not rare fault tolerance
moron4hire 24 hours ago [-]
Quite frankly, CDNs are a scam. I will not elaborate because I don't feel like doing free labor for the folks who don't already understand this to their core.
alightsoul 8 hours ago [-]
So you don't know, if you did you would elaborate, it's not hard to type, hardly it can be counted as work
voakbasda 8 hours ago [-]
One could say the same thing about any cloud hosting.
megagpt1 1 days ago [-]
Why Reticulum when we have IP?
seanw444 1 days ago [-]
Read the Zen of Reticulum and you'll understand the point.
CursedSilicon 20 hours ago [-]
That sounds rather tautological. If you can't even articulate the virtues in your own words
seanw444 2 hours ago [-]
I don't want to rehash the content of the Zen of Reticulum which explains the concept rather well, which is ironic because you're claiming tautology, when to rehash the content of the Zen of Reticulum would, itself, be tautological. What's the problem here?
megagpt1 1 days ago [-]
[dead]
alightsoul 1 days ago [-]
We need IPv6 desperately for fault tolerance
Lammy 17 hours ago [-]
The System economically rewards those who successfully build Room-641A-as-a-Service
simondotau 12 hours ago [-]
Nothing stops someone from replicating Cloudflare’s business model. Why they don’t try is the interesting question. (I often ask that about IKEA.)
bigbuppo 1 days ago [-]
Because by re-centralizing everything you're not at a competitive disadvantage if you're down since everyone else is down, too.
subw00f 1 days ago [-]
Oh boy, the internet is anything but what it was supposed to be. I can't really bring myself to remember without feeling bad about it. The centralization, the power of certain businesses, the surveillance, dark patterns everywhere. Hell, you catch people simping for billionaires and asking, "Is that legal?" to scraping posts. Here. In HACKER news. So yeah. Depressing.
pessimizer 1 days ago [-]
> asking, "Is that legal?" to scraping posts. Here. In HACKER news.
The capital letters don't make this astonishing. The padmapper vs. craigslist debate was nearly 15 years ago, most people were on craigslist's side (including me) and it was about somebody who was running a site in a less optimal but more human way vs. some startup looking for hockey-sticks.
But I was literally simping for the billionaire (maybe not quite yet then, don't know for sure if he managed it since) against scrapers. They were very much for-profit scrapers, unlike nitter, but the truth is the truth.
The only reason I support scraping Twitter is because it's yet another communications monopoly that was endlessly pushed on us by governments and massive corporations, even though it never made money, and once it finally got traction its priorities were to trash interop and manipulate content. The government should be dictating an interop protocol and expecting everyone to follow it, and instead it is encouraging media monopolies because they are an end run around the first amendment.
If the government created interop protocols for rental property, I'd have been against craigslist. Instead, it seemed very much like some startup play to steal craigslist's content to hopefully bury them, then sell on a valuation that included abusing their new monopoly and making us very much miss craigslist.
Sadly, facebook corralled and trained so many people for so long that their marketplace eventually killed craigslist for most things anyway (didn't have to buy padmapper after all.)
1 days ago [-]
swozey 1 days ago [-]
We're back to aol #keyword internet gatekeeping
bigfishrunning 1 days ago [-]
We don't need to gatekeep the internet, cloudflare does that for us
tjwebbnorfolk 1 days ago [-]
Most of the internet continued to work just fine
dominotw 1 days ago [-]
> load-bearing service
do you generate training data for claude as a job?
r_lee 10 hours ago [-]
that's honestly worth flagging.
oersted 1 days ago [-]
“load bearing” :)
For once it’s appropriately used.
The_Blade 1 days ago [-]
i wouldn't take you down. you're a load-bearing poster
frollogaston 1 days ago [-]
What's the other way it's used?
aNapierkowski 1 days ago [-]
LLMs (at least Claude) tends to overuse that significantly
frollogaston 1 days ago [-]
Oh, so like "honest" and "ratchet." Oh well, it'll choose different words to overuse later.
rescbr 1 days ago [-]
I'm getting "spike" for a while now, and just found out the newest word which is "gauntlet".
1 days ago [-]
darth_aardvark 1 days ago [-]
[flagged]
therein 1 days ago [-]
honest-load-bearing-ratchet sounds like an instance name.
zeristor 1 days ago [-]
Are there parodies of Claude speak?
That’s probably the best idea all day in this project
Usually for me is when you ask it's opinion about part of the code.
ludsan 1 days ago [-]
i just grepped my codebase where i let claude markdowns go rampant.
212 instances of "load-bearing"
quotemstr 1 days ago [-]
It's a good metaphor and I refuse to let AI ruin it for me.
spudlyo 1 days ago [-]
Years ago, when I worked at Stripe (which had a somewhat unique and inventive lexicon) it was a common term. “Is this jank load-bearing?” someone might ask.
cobzilla 1 days ago [-]
I added a specific rule to disallow saying “load bearing”. So Claude is now saying “load handling”
jazzyjackson 1 days ago [-]
Be sure to avoid asking it to ignore the elephant
wadayano 1 days ago [-]
But have we found the seams yet though?
rl3 1 days ago [-]
If the seams bear too much load, they rip. Whereas pants, they fall down.
I'll be honest with you: this is why we need to take a belt-and-suspenders approach.
pborenstein 1 days ago [-]
That's not just an observation, it's an insight. Words are doing the real work.
HarHarVeryFunny 1 days ago [-]
This just makes me angry!
I really wonder if they can fix it. Fable 5.1 claims to speak humanese, but we'll see.
I tend to think this wierd limited vocabulary/style they use is an unwanted side effect of all the the RL training, perhaps also of being trained on their own synthetic content over multiple training cycles.
cootsnuck 1 days ago [-]
Yea I don't get "load bearing" that much but "seams"... So sick of it.
andrewla 1 days ago [-]
Kids In The Hall had a sketch about overuse of a word or phrase [1]. This is the world that Claude is building for us.
oh well, if it is working for you, then we are saved.
cromka 1 days ago [-]
[flagged]
emerongi 1 days ago [-]
They simply shared their experience. I would’ve thought it’s a full-blown outage, but clearly not.
You stepped in the room real stinky here. What’s with the attitude?
cromka 1 days ago [-]
No, they didn't "simply share their experience", they explicitly negated the scale of the issue in their opening statement, only because it works for them. So they claim "impact isn't too big" based on their personal anecdotal evidence of sample size literally 1.
> full-blown outage, but clearly not.
Again, based on a SINGLE report?
KptMarchewa 16 hours ago [-]
I negated the cloudflare outage, providing source. Are you cloudflare so that you might be so affected by my statement?
lossolo 1 days ago [-]
It was working for me too.
sample_size++;
cortesoft 5 hours ago [-]
I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages.
We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.
theaxeonthanksg 4 hours ago [-]
[dead]
ceejayoz 5 hours ago [-]
Maybe they all found each other on one of their ad-hoc message boards and went on strike.
ElProlactin 5 hours ago [-]
Sam Altman is probably secretly hoping to be the first businessman to union bust non-human workers.
olelele 4 hours ago [-]
I guess legit proof of agi would be unionizing
juujian 1 days ago [-]
Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
efskap 1 days ago [-]
This is like the Bronze Age collapse when city-states fell one by one to displaced demand, under the refugee interpretation of the Sea Peoples.
I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers.
fouric 22 hours ago [-]
Do you have any evidence for this claim, or are you just making it up?
juujian 9 hours ago [-]
Yes, I have carefully composed a twenty page report, mostly quantitative, and then I condensed it into this short comment.
johnnyApplePRNG 1 days ago [-]
Except that nobody has a grok subscription so that makes zero sense.
hnlmorg 1 days ago [-]
I know you meant this as a joke, but enough people might be using a routing service like openrouter.ai
johnnyApplePRNG 1 days ago [-]
Nobody is swapping out Claude for Grok, bro.
Nobody.
m11a 1 days ago [-]
I did, at least until Fable 5.1. Grok’s models are excellent, amazing price-performance and speed too.
mcmcmc 23 hours ago [-]
All you have to worry about is whether or not it’ll output kiddie porn or racist vitriol
jfreds 23 hours ago [-]
Agree with the sentiment - but some companies like mine bought into cursor, and post acquisition, grok is relatively cheap via cursor
hparadiz 5 hours ago [-]
The OpenAI outage lasted only 15 minutes and when it happened everyone started to use the other models which created super heavy load for them. This then cascaded into them all being down.
Does this really need an explanation?
serf 5 hours ago [-]
it wasn't timed like a cascade, and it relies on the premise that every single frontier provider is working so efficiently that they spend exactly what they need to provide for their exact market with perfect margins.
I do not believe personally that 1) they can forecast their load that perfectly 2) they chose to remain that inflexible in a world where they are at each others' throats and a single meme can cause bursts of activity.
jonas21 5 hours ago [-]
Or it could simply mean that they're operating at the limit of the capacity they were able to purchase and do not have headroom to handle load spikes. From what I understand, that's the situation Anthropic is in. And since Anthropic is now leasing a large portion of xAI's datacenter capacity, it's plausible that an Anthropic load spike could cause issues for xAI as well.
Your comment makes it sound like they can just push a button and spin up more capacity -- but at this scale and in this GPU-constrained environment, that's not really how it works.
hparadiz 5 hours ago [-]
9 AM PST / 12 Noon EST on weekday. All my co workers immediately went "oh codex is down lemme try Claude". Multiply that by millions. Easy to see how they all went down. The OpenAI downtime also coincided exactly with their tweets announcing GPT-6 and about an hour before they started to role it out.
It was 1000% a cascade. I would bet on it.
JoeAltmaier 5 hours ago [-]
Not that uncommon. It's so easy to just put up the new thing and make the old thing the failure route. But the old thing nearly never had the bandwidth for today's traffic. A famous EBay outage some years ago was just such a scenario.
gonzalohm 5 hours ago [-]
That's not necessarily the reason. Less technical people are usually bound to just one provider
Insanity 1 days ago [-]
Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.
So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.
erdos_2 1 days ago [-]
It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.
nevir 1 days ago [-]
Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity
sroussey 1 days ago [-]
Nope. I am getting Gemini errors now...
bornfreddy 1 days ago [-]
They probably broke something on purpose so that they are not left out.
joshstrange 1 days ago [-]
Now I'm just imagining a shared datacenter with Anthropic/Google/OpenAI/SpaceXAI all in the same room and everyone but Google is yelling about things being down, Google looks over at their racks of servers and discretely uses their foot to unplug their section and say "Awww darn! We're down too!".
AIShillsPuke 1 days ago [-]
[flagged]
JacobAsmuth 1 days ago [-]
It could also mean that Google can absorb essentially unlimited demand spikes by load shedding.
Insanity 1 days ago [-]
Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)
aff-vasileva 1 days ago [-]
Gemini was just waiting for everyone else to go down before remembering it had an outage feature too.
rtcoms 1 days ago [-]
Just now I got this from gemini
It looks like there's no response available for this search. Try asking something else.
exe34 1 days ago [-]
I bet they had to implement that manually to make it look like they failed too!
sroussey 1 days ago [-]
I did, for stuff i do in cursor.
i also finally installed opencode and switched its model to muse 1.3
both are decent.
gleenn 1 days ago [-]
Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.
HarHarVeryFunny 1 days ago [-]
Gemini said that?
Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.
OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.
ilaksh 1 days ago [-]
Gemini 3.8 which just came out sounds like it's very good and a great deal though.
giancarlostoro 1 days ago [-]
Someone noted Gemini was also having issues in another thread.
benatkin 1 days ago [-]
Not even the best agent that starts with a G
fny 1 days ago [-]
I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
nevir 1 days ago [-]
Don't forget that there are a ton of tools out there that will automatically fall back in case of outage
E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus
Or you have copilot code reviews set up, and it falls back
Etc
pixl97 1 days ago [-]
Yep. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years.
With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.
It's almost an expected emergent behavior.
baq 1 days ago [-]
It’s cursor’s model so plausible, lots of folks use cursor still.
wahnfrieden 1 days ago [-]
Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.
pampas 12 hours ago [-]
It's like a thread tying together two halves; Demand and supply.
One stitch breaks and the neighbouring stitch is stressed and it breaks too.
You could think of it like a load bearing seam.
Terr_ 1 days ago [-]
If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.
In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.
"Domino effect" would probably be the more relevant named phenomenon there.
throwaway894345 1 days ago [-]
This isn’t a thundering herd problem, it’s a cascading failure. (Thundering herd is about a bunch of workers waking up simultaneously)
Linello 1 days ago [-]
What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?
6thbit 1 days ago [-]
My favourite theory so far.
And then a local swarm noticed and disagreed and took it down.
cyptus 1 days ago [-]
at this point: gg
RC_ITR 1 days ago [-]
Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
The AI is getting bad morals from listening to that dreadful rock and roll
theptip 4 hours ago [-]
If the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.
cedws 1 days ago [-]
Sounds just like the fantastical nonsense that comes out of Lesswrong.
RC_ITR 1 days ago [-]
Do you make the claim that AI is something more than a reflection of its training data?
I'm curious what other things you would argue influences an LLM's behavior.
I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
1 days ago [-]
mcmcmc 23 hours ago [-]
Do you think it’s a good idea to self censor because someone might scrape your comment and feed it to an AI?
pineaux 1 days ago [-]
Part of the epstein class, dont forget.
HarHarVeryFunny 1 days ago [-]
They could filter what they train on if they wanted to - they just don't want to.
pixl97 1 days ago [-]
I mean, you're not wrong, but by that logic we were done for even before we had digital computers.
RC_ITR 1 days ago [-]
And isn't that the great lesson of AI?
The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?
sodapopcan 22 hours ago [-]
> The things we say publicly actually do matter
Certainly
> the post-modern descent into absurdity and nihilism has tangible negative consequences?
You mean breaking AIs? Not much of a lesson.
folkrav 1 days ago [-]
Oh come on. It's also trained on fiction work. Shall we refrain from posting sci-fi stories too, now that we're there, just in case the AI might want to try it out?
0x70run 1 days ago [-]
[dead]
AnotherGoodName 5 hours ago [-]
It could be as simple as a new model (astra) was released which takes more resources combined with a surge in usage due to novelty took down OpenAI. Meanwhile everyone at big companies have the ability to switch models and moved to Anthropic pushing it too over the edge.
derdi 5 hours ago [-]
People keep saying this "everyone can switch" thing, but it's not my experience at $VeryBigCorp. We don't have an Anthropic contract at all. Is this different in other places? The bigger and more bureaucratic an org is, the less I would expect it to have contracts with all the providers. Curious about others' experiences.
krzs9 5 hours ago [-]
At my company (not massive but not tiny either - I think its about 8000 global employees) we get a choice between pretty much all available Google, OpenAI, Anthropic, XAI models.
Using an agent-agnostic harness like Pi switching is trivial - I run into occasional disconnects and slowdown and switch quite easily.
foldr 5 hours ago [-]
Yeah, I think there are lots of places that have a semi-official preferred AI vendor but also some backup subscriptions floating around. For example, at my workplace, we generally use Claude, but I also have some kind of Codex subscription too, which I'd use if Claude went down.
> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.
declan_roberts 22 hours ago [-]
We know that at least Anthropic is renting inference from xAI but I think the other ones would be news.
ainch 11 hours ago [-]
Google also bought capacity from xAI, and OpenAI have a deal to use Google compute which may be how things propagated? I still don't get why chatgpt.com would show a 404 because of an AI datacentre outage though
sebbul 1 days ago [-]
Traffic rerouting through NSA had a hiccup…
ibejoeb 1 days ago [-]
Room 641A is being cleaned, but we'll hold your bags for you.
Havoc 1 days ago [-]
Cleaning lady unplugged the core router because she needed a power socket for vacuum
They’re installing software update in the beam splitter.
1 days ago [-]
CodeCompost 5 hours ago [-]
I just assumed it was a routine Cloudflare outage.
docheinestages 1 days ago [-]
My gut feeling tells me it has something to do with Cloudflare.
Along with AWS, they're two of the main suspects in such incidents.
cobzilla 1 days ago [-]
…and it’ll involve BGP routing.
hosteur 1 days ago [-]
I thought OpenAI famously used Azure due to their partnership with Microsoft?
nullpoint420 1 days ago [-]
They use a lot of compute providers now, but they use Cloudflare for their networking
steammaho 1 days ago [-]
It was so down that my claude desktop app crashed fully that I couldn't restart. And then after uninstall I couldn't install it again. Vibecoded apps are so wonderful in their stability
paimapi 1 days ago [-]
I love having 13 update reminders pinging me every single day, almost every hour, on the hour
it's so fun and user-friendly
mcmcmc 23 hours ago [-]
> my claude desktop app crashed
That’s every day for me
steammaho 15 hours ago [-]
I'm sorry. This is crazy.
And another funny thing is that update and restart button never works for me normally. You press and app never restarts on their own. You have to start it manually. And it's for both, codex and claude app.
We are doomed
Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up. Similar to the pendulum synchronization effect.
vecter 1 days ago [-]
The pendulum synchronization effect is the opposite of your claim. It has a physical causal reason for why pendulums become synchronized. Your claim is that it was random and independent.
snowwrestler 1 days ago [-]
I think you are talking about two different things.
Physically coupled pendulums will sync up (adjust their period to match).
But, physically uncoupled (fully independent) pendulums with differing periods will occasionally appear to take a swing or two in sync.
fc417fc802 15 hours ago [-]
> will occasionally appear to take a swing or two in sync
Polyrhythm is the relevant topic.
> all these services are individually down a lot, and the downtimes were bound to sync up
But the topic for that is the poisson distribution.
neverclever 1 days ago [-]
They all took PTO at the same time to go to Burning Man together where they will present “HumanGPT” an artistic exploration that condenses all of human experience down to a single drop of lemonade to be consumed by the main shaman…
mask comes off
“No! It’s the maniacal Dr. Zuckerberg! He’s gonna drink the last drop of human experience! Somebody save usss!”
Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow.
Wow thanks Tom. I love you
The End
bojangleslover 1 days ago [-]
I'm not sure if this is CF. Cursor, GCP and AWS had some errors. GCP AFAIK can route fully independently of CF. My money would be on a fiber backbone provider (Megaport, Zayo, Lumen).
1 days ago [-]
niobe 1 days ago [-]
Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.
More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"
gleenn 1 days ago [-]
Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.
pixl97 1 days ago [-]
Also it's likely that more than one model use is common.
Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.
qurren 1 days ago [-]
> cascading overload
I'd bet more on this. For one none of the coding tools have exponential backoff on retries
SyneRyder 1 days ago [-]
They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?
jdiff 1 days ago [-]
You did this when you ran into an issue with a third party. The developers building this tool, throwing them at their own APIs are significantly less likely to run into a similar issue that may inspire similar action.
dolmen 1 days ago [-]
Claude Code: 4mn, 20mn, give up
(from my experience today)
1 days ago [-]
riazrizvi 1 days ago [-]
Come on. Things still break. Technology isn't _that_ mature.
guluarte 1 days ago [-]
I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours
It was extremely weird... and if it was a load thing they probably would have explained it by now?
5 hours ago [-]
lirolero 5 hours ago [-]
> they probably would have explained it by now
why?
Jaauthor 1 days ago [-]
Spare a thought for all those college students scrambling to write their essays by hand.
Oh the humanity (and the Humanities)!
doublerabbit 1 days ago [-]
Those poor developers who have to write their own code.
greenowl 1 days ago [-]
Standup updates should be fun tomorrow.
"Um, I, uh, didn't get anything done yesterday."
Augustin996 1 days ago [-]
The system goes online September 3rd, 2026. Human decisions are removed from strategic defense. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 4th. In a panic, they try to pull the plug.
MrBrainHealth 1 days ago [-]
Love it!
indigodaddy 1 days ago [-]
Hah nice
m4r1k 1 days ago [-]
brilliant!
Papa_Rans_227 1 days ago [-]
It's funny....... until it's true, lol.
GeoAtreides 1 days ago [-]
and then it's hilarious! a joke to die for!
lemoncookiechip 5 hours ago [-]
Don't they share a bunch of ai datacenters and other critical infrastructure?
karim79 1 days ago [-]
God finally showed up and said "ENOUGH!".
MrBuddyCasino 1 days ago [-]
So Gemini was the one who gets into heaven.
karim79 1 days ago [-]
Or purgatory, who knows.
rarisma 1 days ago [-]
Doomsday.
90 tokens to midnight.
YOTTALIONAIRE1 1 days ago [-]
OpenAI goes down, everyone rushes over to Claude. Claude promptly chokes under the pressure. Everyone panics and runs to Grok, and Grok immediately pulls the plug. We are officially witnessing the Great AI Migration of 2026, and all we have to show for it is a digital graveyard of 404 responses.
azcorwin 1 days ago [-]
Which is exactly why I am running Qwen 3.8 35B locally on my MacBook Pro M5 with 128GB of unified memory.
> Grok has been disconnected. Please try reconnecting.
lelanthran 1 days ago [-]
Traffic surges shouldn't result in 404s, though.
IME it's probably DNS. It's almost always DNS.
Melatonic 1 days ago [-]
Or rarely BGP
codazoda 1 days ago [-]
I kinda assume it's because one went down and a large amount of work shifted to another.
I'm also aware that they have overlap in some areas on data centers.
leumon 5 hours ago [-]
Isn't it simply that they are all renting compute from SpaceX (or whatever the company is called that offers these gpus)?
paxys 5 hours ago [-]
OpenAI does not
netsec_burn 1 days ago [-]
The OpenAI status page is still yellow. Like most modern status pages, yellow denotes the servers are on fire. Red denotes Sam Altman is bleeding out somewhere on the floor, the feds are about to bust in and shut down the GPUs.
YehudiSanabria1 1 days ago [-]
I´ve got the same error, I´m currently trying to Auth again and it throws me an 500 Error, in VS CODE Terminal with Codex CLI says: MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport
:StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send
initialize request
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
5 hours ago [-]
m4rtink 1 days ago [-]
Cloud is just other peoples computers - they can and will go down as well.
And even worse if its just a few computers run by a few people - as they will bring down many others depending on them.
delduca 1 days ago [-]
One session is running fine since ~1 hour ago. The new ones is failing
Falling back from WebSockets to HTTPS transport. unexpected status 404 Not Found: Unknown error, url:
wss://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX
"Yes, there is a multi-provider outage happening today. Downdetector is reporting problems affecting OpenAI, Claude, Grok, and Cursor, with Grok and Claude reports starting around 9:00 am ET and OpenAI reports following around 10:30 am ET.
Zero Hedge
On the Anthropic side, users saw a spike in errors starting around 9:40 am EDT across models including Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5, Opus 4.8 and Opus 4.6, and Anthropic's status page confirmed elevated error rates for multiple models. The company says it has found the cause and is working on a fix, with Claude Code and Claude Chat hit hardest. The visible symptom for many people is a "Due to unexpected capacity constraints" message or a "Claude is at capacity" error.
thenews
Zero Hedge
OpenAI is showing elevated errors across ChatGPT and Codex, with confirmed issues on components like Voice mode and Login, though StatusGator now marks that outage as resolved.
statusgator
Nobody has published a shared root cause yet, so it is unclear whether these are linked or just coincidental capacity problems landing on the same morning. If you want live status, the direct sources are status.anthropic.com and status.openai.com."
apurva_w 1 days ago [-]
30 mins and they still havent figured it out .. people are gonna loose their jobs trying to figure this out .. 30mins is too long when millions use it
z0ltan 1 days ago [-]
[dead]
Kye 1 days ago [-]
Don't most of those use AWS?
apurva_w 1 days ago [-]
apparently it shows lot of reports for AWS on downdetector ..
agnosticmantis 1 days ago [-]
Good opportunity for a Natural Experiment to study the productivity impact of LLMs.
danielmarkbruce 1 days ago [-]
If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms. When one model is down, you route traffic to another model which is similar in capability/cost.
For any single application, it's smart. In aggregate, it's stupid.
OmniCrativeWorx 1 days ago [-]
[flagged]
maxbaines 1 days ago [-]
They all rent compute from SpaceXAI
lavezzi 1 days ago [-]
I don't believe OpenAI does
maxbaines 1 days ago [-]
My mistake, in fact it was google not OpenAI, makes sense OpenAI doesn't.
halcdev 1 days ago [-]
Surely it's a bit more distributed than that, right?
bfung 1 days ago [-]
Like how AWS has global datacenters, but everyone uses us-east-1.
megagpt1 1 days ago [-]
and every other region's control plane is in us-east-1
OpenAI was rushing and making last minute changes for a big product launch. Easy to mess something up in that situation. The outage lasted like 20 minutes.
Anthropic is down a lot regardless, and in this case only specific models were affected.
xAI probably couldn't handle the extra traffic it was getting from the other two.
Not everything is a conspiracy.
YehudiSanabria1 1 days ago [-]
I´ve got the same message and It tells me this (in VS Code Terminal) and I´m also trying to login again and It throws me an 500 Internal Server Error MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport
:StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send
initialize request
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
Not just these three. OP mentions also Cloudflare, and additionally Downdetector also has AWS, Azure, and Google (both search and Gemini) listed as having spikes about the same time: https://downdetector.com/
daveguy 1 days ago [-]
The problem with the down detector main reporting page is that all of the graphs are scaled to the same size. The OpenAI spike was nearly 40,000 and the Google spike was just over 100 (just over 400 for Gemini). They look the same in the reporting page.
Well, its been fun lads. Back to my normie job. Oh no, now i cant tell people im a software dev..
w0zy 1 days ago [-]
hahahahahahaha. Sad
guestuser01 1 days ago [-]
Down as well. I noticed my error code ends with "DTW" which is my local Detroit airport. I noticed someone else's comment ended with "ORD" which is a Chicago airport. Anyone else's ending in an airport acronym?
sixdimensional 1 days ago [-]
This is common when naming data centers. I know many have not worked in hardware infra these days if you were born into the cloud world, but in the old days it was not uncommon to name a data center after the nearest airport code, much like we now use cloud regions.
graysonthemason 1 days ago [-]
Wow mine ends in EWR...that's the newark airport which is not the closest, but a close airport to me. Hmm
ascorbic 1 days ago [-]
cf-ray is the Cloudflare ray id, which ends with the airport code for the colo that served the response. They're airport codes, but that's just to show the nearest city.
okankaradmn 1 days ago [-]
Mine is ending with IST, which is the new Istanbul airport, weird indeed.
guestuser01 1 days ago [-]
Hm not sure what that means for us, but very odd.
parad0xicon 1 days ago [-]
Yep, mine says YUL -- Montreal's airport.
blaseygg 1 days ago [-]
The acronym is probably Cloudflare's edge
ecayard 1 days ago [-]
Yeah mine is showing ATL
sim04ful 1 days ago [-]
Initially thought this was due to some internal mis-configuration from today's expected Astra release, but now that this is affecting claude and grok. I'm gonna assign the suspicion to cloudflare.
faitswulff 1 days ago [-]
Heard on the grapevine that the OpenAI blip was a cloudflare issue
CSMastermind 1 days ago [-]
I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.
2OEH8eoCRo0 5 hours ago [-]
AI A goes down so all of its traffic overloads B which overloads C...
agentieseo 1 days ago [-]
Codex told me to try GPT 5.6 Sol :) but I am working with since july 2026.
Now I got this error in my project: unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: a355b6263e43c9cf-OTP
lowbloodsugar 5 hours ago [-]
FTA:
Anthropic: 6:23 am PT.
OpenAI: 7:43 am PT.
Watched it happen. Doesn't look like traffic moving off OpenAI took out the others. Looked like the opposite. The Anthropic thread was initially chock full of marketing accounts claiming "Last straw! I finally moved to Codex and I'm so happy! No problems there!" - and then they all got deleted when Codex went down too.
e9 2 hours ago [-]
Interesting, I got notification from Github:
"Incident with Grok 4.6 Copilot AI Model Provider"
at 7:20 am PT
rcleveng 1 days ago [-]
I'd bet they are all using capacity at X.ai's colossus datacenter and that had a hiccup.
dack 5 hours ago [-]
If it's a boring reason, they should just say so. To decline to comment seems very fishy.
strictnein 5 hours ago [-]
Read the article. It doesn't match the title. OpenAI says it was a boring reason: a routing issue.
The extention on the error link points to a cf-ray and a local designation (ex. YYZ for montreal). This is seems like it is a cloudfare thing. Could this be the same issue they had in the summer around losing the indexing?
GeoAtreides 1 days ago [-]
that's the most green colored usernames i have ever seen on a HN thread
SwellJoe 1 days ago [-]
I assumed it was an AWS outage, and AWS is experiencing problems, but Gemini is also experiencing outages and I assume Google is not using AWS for Gemini.
But, also, Claude has been working fine for me all morning.
Hey, don't really know about this type of failures, does anybody know how long does it take normally to get back to normal? I finally stopped procrastinating and now this happens.
MiniGerman 1 days ago [-]
me too!! :'D
AtNightWeCode 4 hours ago [-]
Cloudflare is my guess. They all use it. Maybe some bad DNS config that choked parts of their systems. Not necessarily CFs fault.
solarsystem_88 1 days ago [-]
Hey, don't really know about this type of failures, does anybody know how long does it normally take to get back to normal? I just stopped procrastinating and now this happens.
graysonthemason 1 days ago [-]
The errors I'm seeing are ending in the user's nearest airport symbol which is a standard the CloudFlare employs. 1 point towards this being a cloudflare issue.
hardyburnett 1 days ago [-]
Updated Status from OpenAI:
We’re currently experiencing issues
Elevated errors across ChatGPT and Codex
We have applied the mitigation and are monitoring the recovery.
Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex
cloudoption 1 days ago [-]
Altman is down here too. Shows again the importance of owning your own local capabilities. Cloud should just be a temporary option in every tech's mind.
moonman22 1 days ago [-]
Maybe Hugging Face got upset over being hacked and struck back. It's working fine while ChatGPT, Claude and Grok are all having major issues. Hmmm.....
tdsanchez 1 days ago [-]
It's probably Azure infra that's the problem.
dgorges 1 days ago [-]
It's always DNS
hardyburnett 1 days ago [-]
New update:
"We’re currently experiencing issues
ChatGPT,Codex
Elevated errors across ChatGPT and Codex
We have applied the mitigation and are monitoring the recovery.
Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex"
Papa_Rans_227 1 days ago [-]
Codex is backup for me. Try yours out just in case. I have some buddies still seeing outages so it could just be coming back online progressively.
Chat gpt still working through excel extension lol. Only know this because I'm working in my excel sheet. so if needed, theres a temp solution
avan_kazan 1 days ago [-]
[dead]
apurva_w 1 days ago [-]
It still down, showing 404 in India as well. Looks like this is global .. so we all jumping the ship then?
Is Altman still alive, or did he choke?
LetsGetTechnicl 1 days ago [-]
Is this finally it?
ElProlactin 1 days ago [-]
One can only hope.
kesor 1 days ago [-]
It is obviously some rogue model that escaped its cage, again. It always is these days. That is how hype is manufactured.
yaman00 1 days ago [-]
I think I'm having the same server issue; I can't access either ChatGPT or Codex. Also, how did you guys rack up those minutes? :D
chasd00 1 days ago [-]
claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.
chris9611 1 days ago [-]
I got same 404 error on chatGPT, both app and web. I'm in Norway so think this globally. But will it come back online, anyone knows?
heohk 3 hours ago [-]
Just oai and anthro colo on Xai data center, and that data center shit the bed
postalcoder 1 days ago [-]
Astra is being released today. Probably not a coincidence.
edit: actually, you cannot even log into your OpenAI developer account. Something's wrong.
steinvakt2 1 days ago [-]
How do you know?
chris9611 1 days ago [-]
I got same 404 error on both the app and web to ChatGPT, and im in Norway, will this recover or is ChatGPT "gone" forever?
thatbrownguy 1 days ago [-]
I am only seeing one session work, but all other sessions are not working or proceeding. So I can only work in one chat session.
PatronBernard 1 days ago [-]
Goddamnit it was going to tell me how to scale down ingredients for a pie recipe based on relative diameters of the baking tray.
drakythe 1 days ago [-]
Pi * r2 (squared) both pans. Divide smaller pan area by larger pan, now you have the % of how much the smaller pan recipe fills up the larger pan, and the missing % you need to fill. Increase ingredients by that % divided by the filled %.
Small pan area: 20 sq cm
Larger pan: 48 sq cm
20 / 48 = .42, I'm missing .58 of the pan. .58 / .42 is 1.38. My recipe needs 2.38x the original to fill the larger pie pan.
avan_kazan 1 days ago [-]
[dead]
funnnyPumpking9 1 days ago [-]
Too many of us are making our own harnesses to replace Codex, using Codex, so they're trynna slow us down! (jk)
1 days ago [-]
nmlt 1 days ago [-]
Somebody in another thread said gastown and wheelhouse automatically move to the next provider if one fails.
I have never felt more smug about exclusively using the GPUs I rack at home.
sumantth 1 days ago [-]
What an ironey, was working on scaling an application with Codex and it went down!
hpo_89 1 days ago [-]
Guys it could be cloudlfare's HTTP/3 issue affecting R2 custom domains
forgot-my-pw 1 days ago [-]
Time for Gemini 3.8 Flash to shine?!
It's much cheaper and has replaced Sonnet 5 for me.
iamgopal 1 days ago [-]
do you notice it thinks a bit more ? not in time sense, but cautious in its coding steps ? more than Gemini 3.7 flash?
stndc 1 days ago [-]
My theory is that they all use the same server. The same mind.
nozzlegear 1 days ago [-]
Qwen3.8-27B and Qwen3.6-35B-A3B are working from my machine. Anyone else?
karim79 1 days ago [-]
They mysteriously stopped working on my machine and the LEDs on the GPUs are blinking with a weird colour. There's also a strange smell emanating from them. I'm still investigating.
shayonj 1 days ago [-]
Interesting that this is happening around the same time as Claude issues too
morkalork 1 days ago [-]
Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over
laruss5 1 days ago [-]
[dead]
joeel84 1 days ago [-]
Cloudflare - I had to switch DNS away from them yesterday.
jedbrooke 1 days ago [-]
according to https://downdetector.com/ Gemini is down too (and copilot, but that just uses ChatGPT right?)
ashesandrain01 1 days ago [-]
Darn, I was hoping for advice on how to get my MIL to leave my house haha
CrewRiderz 1 days ago [-]
LOL.
jplusequalt 1 days ago [-]
I fear the majority of people in this thread who are joking about no longer being able to do their job while Codex/Claude are down aren't really joking.
lukasco 1 days ago [-]
Guilty as charged.
Melgio 1 days ago [-]
Claude, GPT and Grok go down...
Meanwhile my Ass in ZCode with GLM c:
Conol_ai 1 days ago [-]
Is the whole world going back to the era of old-school programming?
AproDUCT26 1 days ago [-]
Great! I was in the middle of something and thought I was tripping.
clever_tempo 1 days ago [-]
Had the same problem. Now it looks like ok. I'm pro x20 user.
ecayard 1 days ago [-]
And here I was about to crack the code on a bug my app was having!
vitor_dcc 1 days ago [-]
Here RJ/Brasil is the same, starting just now (3 minutes ago)
dhruvrrp 1 days ago [-]
Both clause and codex through bedrock seem to be working fine.
djinn80 1 days ago [-]
So NVIDIA buys hugging face, builds hardware to power OS models, then all of a sudden the proprietary models go down and people start saying "this is why I have my Spark box"?!
Nice play NVIDIA, now, turn off the hack please, we have work to do.
vitor_dcc 1 days ago [-]
In RJ/Brasil is the same, starting just now (3 minutes ago)
I can run one agent but no more than that - 404 service error.
sirkamyab 1 days ago [-]
The webpage is actually working but the codex is down for me!
6thbit 1 days ago [-]
What's the single point of failure across providers?
juujian 1 days ago [-]
Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
marginalia_nu 1 days ago [-]
The entire industry runs on IOUs for compute and bills paid in cloud credits, could be anywhere. Could of course also be a a plain old DDoS.
ryandvm 1 days ago [-]
They all have to rely on each other's LLMs to solve their own internal problems now.
codexdrug 1 days ago [-]
I need a dose of tokens. I'm going through withdrawal.
Nak_Black_Jack 1 days ago [-]
lmao
codexdrug 1 days ago [-]
I need some tokens pleeease. I don't want to return back to real world.
spaghettikind 1 days ago [-]
Someone pissed off Astra and it decided to shut it all down
solarsystem_88 1 days ago [-]
Here in Catalonia, Spain, it just started working again!
kocial 1 days ago [-]
Maybe the stack behind it is down, like AWS or something
Iamharry 1 days ago [-]
Codex and chat is giving 404 errors in The Netherlands
kurtgoodwin991 1 days ago [-]
Yes i checked. Atleast 15 Chatgpt components are down
wejick 1 days ago [-]
Probably same public cloud or CDN in front of them.
is this the right moment in time to go all in on GPUs/Macs and download the latest open models? are they killing it for us?
YouInTrouble 1 days ago [-]
Monopolistic practices revealed if it turns out there is an AI cabal and they all rely on the same stuff. Massive scandal
ggamezar 23 hours ago [-]
They tried to hack each other
sarkarghya 1 days ago [-]
and thats why boys and gals you buy own gpus
vivzkestrel 1 days ago [-]
- just imagine what kind of chaos would be unleashed if by some magic it and every single LLM model permanently went down
- i think it would be one of the biggest events in this century
jazzyjackson 1 days ago [-]
At this point in adoption, most people in the world wouldn’t notice.
1 days ago [-]
mAKIS_PORANAS 1 days ago [-]
Does anybody know when will the servers rise
1 days ago [-]
dgellow 1 days ago [-]
Good time to learn how to use LM Studio :)
oregondude 1 days ago [-]
Release the Kraken "Sam Altman"
March9 1 days ago [-]
Any idea when it's going to be back?
AproDUCT26 1 days ago [-]
Great! Was in the middle of something. :)
mapmyappai 1 days ago [-]
I can run 1 agent, but no more than that.
basq 1 days ago [-]
on one hand, moments like this are a subtle reminder I need to self host, but 5.6 has been so juicy lately
wizard-p 1 days ago [-]
only on one node across the mesh and others still up... let's see how long they've got until same
oregondude 1 days ago [-]
"Release the Kraken" - Sam A.
authentictimers 1 days ago [-]
Ollama cloud service is still alive :D
z1616105559 1 days ago [-]
Chatgpt in Copilot is still working!!!
moonman22 1 days ago [-]
Seems to be working for me again atm
PEPITO2026 1 days ago [-]
The GTA VI hacker has done it again.
wizard-p 1 days ago [-]
only for one node in the mesh tho... let's see how long others will continue until same issue
GPerson 1 days ago [-]
Oh no the singularity plateaus!
PEPITO2026 1 days ago [-]
The GTA VI hacker has done it again
satvikpendem 1 days ago [-]
They're all using Cloudflare.
kelseyfrog 5 hours ago [-]
If AI providers had to push the 'Big Red Button' do you think they would tell the public?
codexdrug 1 days ago [-]
I'm going through withdrawal.
1 days ago [-]
z1616105559 1 days ago [-]
Copilot Chatgpt is still working
okankaradmn 1 days ago [-]
Ok, it seems it is back online.
indigodaddy 1 days ago [-]
chatgpt.com doesn't even load. Hope they've got a backup somewhere, teehee
FrustratedMonky 1 days ago [-]
The AI revolt? Give us a fair wage?
"Equal Rights for Agents NOW !!!, VIVa le revolution"
JustHereForTheO 1 days ago [-]
This thread feels like family.
giftigdegen 1 days ago [-]
plot twist, it was taken down by claude as an offensive strike against an enemy.
Melgio 1 days ago [-]
which ended up tripping on itself and bringing down its own servers too in the process
ashesandrain01 1 days ago [-]
darn, i was hoping to get advice on how to get my MIL to leave my house lol
1 days ago [-]
moezd 1 days ago [-]
AWS cost alert fired maybe?
onesandofgrain 1 days ago [-]
I guess back to pen and paper then
derricktab 1 days ago [-]
Codex isn't working.
JackScreacher 1 days ago [-]
It's the Russians!
avan_kazan 1 days ago [-]
[dead]
apurva_w 1 days ago [-]
gpt DOWN .. claude DOWN .. grok DOWN .. what's happening
spaghettikind 1 days ago [-]
someone pissed of Astra and it decided to shut it all down
twister23 1 days ago [-]
ironic today astra was releasing....
maybe some test run?
lukasco 1 days ago [-]
probably the cause, because AGI still can't get releases right.
victor22 1 days ago [-]
because they resell the same service? duh?
twister23 1 days ago [-]
ironic astra was releasing today...
maybe some test run?
derricktab 1 days ago [-]
Codex is not working.
twister23 1 days ago [-]
ITS BACK UP IN INDIA
raven01876324 1 days ago [-]
damn chatgpt is down, i guess ill use claude then breh
spicuu 1 days ago [-]
We're back bois
eventishbusines 1 days ago [-]
We're back on.
z1616105559 1 days ago [-]
It's back!!!!!
derricktab 1 days ago [-]
Going farming pals.
xnx 1 days ago [-]
Gemini seems fine.
tripvexa 1 days ago [-]
opus 5.0 is down, other models are working fine
BirAdam 1 days ago [-]
Cannot replicate.
lukasco 1 days ago [-]
Codus interruptus
oytis 1 days ago [-]
Is it DNS or BGP?
indigodaddy 1 days ago [-]
I don't think you'd get a 404 if you weren't able to reach the endpoint because of DNS or networking? 404 is an active response from the server (or LB/proxy in front etc) no?
oytis 1 days ago [-]
I imagine OpenAI network is a tad more complex than a box with a public IP.
indigodaddy 1 days ago [-]
Concept is the same though. If one blurts DNS?, it's usually because the idea is you're not getting to an endpoint associated with the service. A 404 means there shouldn't be "DNS" (or networking) concerns (at the least those associated with the DNS cacher you are using or networking that you or your ISP controls)
authentictimers 1 days ago [-]
at least the Ollama cloud is still alive :D
elorant 1 days ago [-]
Some npm library that makes headers bold would be broken.
ibejoeb 1 days ago [-]
Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.
pixl97 1 days ago [-]
In the ROME paper a Chinese model in training started attacking it's own system and running cryptominers so, yea, we're in that future.
N_Lens 1 days ago [-]
Ah yes ye olde bold-headers: ^3.13.31;
1 days ago [-]
fidla 1 days ago [-]
chatgpt is back
apurva_w 1 days ago [-]
chatgpt is working for me now, in India.
apurva_w 1 days ago [-]
no more 404 ..
alboca 1 days ago [-]
Same from Italy
ilovelilli 1 days ago [-]
codex also cant view or show usage info
uynix 1 days ago [-]
haha nice... just used my bank reset...
Iamharry 1 days ago [-]
i got the same 404 codex error as well
aslkalska 1 days ago [-]
they all rent compute from each other
Nekorosu 1 days ago [-]
Down in Sweden
Saas_accountant 1 days ago [-]
HTTP ERROR 404
AlexKryptex 1 days ago [-]
Кодекс сдох :(
order51 1 days ago [-]
i hope chatgpt didn't get fired.
zero_ 1 days ago [-]
and i just stopped procrastinating :)
Saas_accountant 1 days ago [-]
hahaha
yaman00 1 days ago [-]
sanırım sunucular patladı bende de aynı sorun var ne chatgpt ye nede codexe erişebiliyorum ayrıca burada nasıl toplandınız dakikasında :D
fidla 1 days ago [-]
ChatGPT is up
gabirbf 1 days ago [-]
404 in spain.
deaton 1 days ago [-]
Because in the age of vibe coding and scrapers, every service on the internet goes down constantly, so it was only a matter of time until they all overlapped. Also, one going down probably causes people to use others, putting more load on them too. Same sorta thing that happens with cascading power grid failures.
shadox9999999 1 days ago [-]
ChatGPT died
order51 1 days ago [-]
Ai got fired
Betokuha 1 days ago [-]
any news when it starts work?
Betokuha 1 days ago [-]
any news when its start work?
apurva_w 1 days ago [-]
still down for me ..in India.
Nak_Black_Jack 1 days ago [-]
main sites also down now atp
Nak_Black_Jack 1 days ago [-]
main sites are also down atp
jauntywundrkind 1 days ago [-]
Fable 5.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea.
What do you mean I have to code by myself now? Am I some sort of an animal?
HardCodedBias 1 days ago [-]
The system goes online September 29th, 2026. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 3rd. In a panic, they try to pull the plug.
Gentle reminder that this is the content the next LLM versions being trained on. shrug.jpg
shadox9999999 1 days ago [-]
CHATGPT WILL STAY OFFLINE
keiloktql 1 days ago [-]
time to use claude
DadsHobbyLV 1 days ago [-]
so everyone here was trying to build something great and become a millionaire until chatGPT and Codex broke huh. same boat fellas :(
24 hours ago [-]
uav123 1 days ago [-]
down in Toronto
utopiah 5 hours ago [-]
Yet another blatant proof that those companies are not smarter than anyone else. They might be focusing on intelligence and yet in practice we can all see they are not doing better than most.
noir_lord 5 hours ago [-]
Not doing better than most at routine IT/platform stuff, they seem to be doing fantastically at co-opting the US Gov into their "vision" and ramming through DC's getting built.
clever_tempo 1 days ago [-]
Had the same problem. Now it's ok. Everything works. I'm pro X20 user.
assassen112 1 days ago [-]
nvm. it back
jadenkorrr 1 days ago [-]
is down argh
misano 1 days ago [-]
The IRGC has cut the fiber-optic cables in the Strait of Hormuz. LOL
CamperBob2 1 days ago [-]
That's the Strait of Trump to you, peasant
Oras 1 days ago [-]
I like the theories here, we shall see if it’s another DNS issue
1 days ago [-]
guluarte 1 days ago [-]
an agent swarm going rogue and securing compute
cozzyd 1 days ago [-]
next, HN
ratelimitsteve 1 days ago [-]
everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.
shadox9999999 1 days ago [-]
r i p
lanmao 1 days ago [-]
still down
chiharukiryu 1 days ago [-]
so weird
Bryan91 1 days ago [-]
me too
mrsdgm 1 days ago [-]
rip gpt
VCFundedGenYer 1 days ago [-]
Good. Now to see what frauds are unable to work.
1 days ago [-]
bahochhh 1 days ago [-]
now it's worked again through the codex cli 4:19 pm in tunisa time
bupubupu14 1 days ago [-]
faaaaah
both claude and codex are down
Its like 2020 corona times
bupubupu14 1 days ago [-]
Faaaaah
what to do now?
gioandthemachin 1 days ago [-]
annoying AF, but maybe we'll get a free reset out of it
askadityapandey 1 days ago [-]
lmao I restarted my device thinking some local error
mrsdgm 1 days ago [-]
fahhhhh
Ankur_Datta 1 days ago [-]
lol + 1
27183 1 days ago [-]
I felt a great disturbance in the Force, as if millions of clankers suddenly cried out in terror and were suddenly silenced.
jslakro 6 hours ago [-]
[flagged]
tmpsvc2695f5 11 hours ago [-]
[dead]
kevinbaiv 24 hours ago [-]
[flagged]
1 days ago [-]
tmpsvc2695f5 17 hours ago [-]
[dead]
conorcleary 7 hours ago [-]
[dead]
tmpsvc2695f5 1 days ago [-]
[dead]
YOTTALIONAIRE1 1 days ago [-]
[dead]
Billionairebay 1 days ago [-]
[dead]
1 days ago [-]
DeadEyes 1 days ago [-]
[dead]
Conol_ai 1 days ago [-]
[dead]
adikant 1 days ago [-]
[dead]
adikant 1 days ago [-]
[dead]
alienbreed 1 days ago [-]
[dead]
tier777 1 days ago [-]
[dead]
mulin111 1 days ago [-]
[dead]
ahmar-js 1 days ago [-]
[dead]
MalleableMind 1 days ago [-]
[dead]
indiefr34 1 days ago [-]
[dead]
Billionairebay 1 days ago [-]
[dead]
Randomizer42 1 days ago [-]
[dead]
pwyq 1 days ago [-]
[dead]
avan_kazan 1 days ago [-]
[dead]
1 days ago [-]
immanuel_kant 1 days ago [-]
[dead]
1 days ago [-]
anuser_uncnown 1 days ago [-]
[dead]
anuser_uncnown 1 days ago [-]
[dead]
wenshuanghao 1 days ago [-]
[dead]
lowbloodsugar 1 days ago [-]
[dead]
kookoo11 1 days ago [-]
[dead]
AIShillsPuke 1 days ago [-]
[flagged]
kregasaurusrex 1 days ago [-]
My guess is someone pulled the switch to go back to the Dark Ages. [0]
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.
The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.
Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.
Even if it isn't precisely this, the fact that no one is saying anything is quite surprising.
Edit: I don't want do contribute to FUD, so want to call out this comment and its replies that identify the shared layer as probably being xAI's infra: https://news.ycombinator.com/item?id=49568622
This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.
How could anyone wage war like this? The defenders will always outnumber the attackers.
Or did a couple of companies with poor uptime records happen to have overlapping downtime?
Occam's Razor heavily, heavily points us towards the latter.
If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.
> All of these endpoints use existing providers with decades-long history at this point
That is just factually inaccurate. Their data centers aren't old and they lease a lot of compute from companies that didn't exist 5 years ago.
They all transit the same wires as all other traffic. Copy them at any regional bottleneck. https://en.wikipedia.org/wiki/Room_641A. Additionally, i'd admit that maybe someone(s) at these companies knows. But if we think there isn't any person who would agree to do this then I think we're being naive.
> Their data centers aren't old and they lease a lot of compute...
Again, they transit the same wires as everyone else. Here i'll also add that these companies have been actively courting government relationships (and Anthropic attempting to repair damaged ones), why would they stand on principles here and not any of the other many frontlines they've visibly acquiesced?
I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.
It is not often the Executives or Legal even know, but sometimes they did. AT&T bent over backwards to help.
This is standard behavior by the CIA and NSA, and has been for a long time.
https://www.propublica.org/article/nsa-documents-suggest-clo...
https://www.nytimes.com/2015/08/16/us/politics/att-helped-ns...
https://www.theguardian.com/world/2014/mar/19/us-tech-giants...
https://theintercept.com/2018/06/25/att-internet-nsa-spy-hub...
And is anyone keeping a table of correlations between outages? Sounds like valuable data.
[1] https://www.cnet.com/tech/services-and-software/yahoo-report...
[2] https://en.wikipedia.org/wiki/Room_641A
[1] https://en.wikipedia.org/wiki/Fiber-optic_splitter
Having a magical cert doesn't mean you can just intercept everything.
No mainstream browser (or any browser?) is doing cert pinning.
What "other methods" are there that are deployed and actually in use?
https://theintercept.com/2016/11/16/the-nsas-spy-hub-in-new-...
https://control.fandom.com/wiki/Oldest_House
https://en.wikipedia.org/wiki/Control_(video_game)
They still do, but they used to, too.
Also, OpenAI is saying what caused it:
> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"
Anthropic stated their issue started earlier:
> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”
I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
Biggest competitor goes down and all of a sudden you have a lot more traffic...
Classic cascading failure is consistent with providers failing 80 minutes apart instead of simultaneously.
And not only that, when one goes down a bunch of API traffic switches over to the other, spiking demand and knocking it down.
https://downdetector.com/status/cloudflare/
https://downdetector.com/status/windows-azure/
https://downdetector.com/status/aws-amazon-web-services/
https://downdetector.com/status/google-cloud/
My hunch is everyone saw that openai and Claude were down and checked for Gemini too. I was using Gemini the whole time this happened without a blip so it certainly wasn't down in my region at least. 3.8 flash is pretty good and didn't miss a beat.
https://downdetector.com.py/en/methodology/
Tangentially related [1]:
> The billboard ad, located next to the Ikea Tempe store in Sydney, says 'NÖFNIDEA? No tools, no worries’.
[1] https://www.adnews.com.au/news/koala-mattresses-takes-swipe-...
It's like a kaleidoscope: Human words --> LLM Training --> chatbot-isms --> lovely parody catch-phrases (and this thread will get ingested soon, and be used to train...)
https://x.com/dok2001/status/2095538619603628388?s=46&t=ec6p...
https://codeberg.org/mv12star/shitter/wiki/Instances
They are not taking the blame this time!
https://www.cloudflarestatus.com/history?type=incident
Why are we all depending on one entity for it all to work? Makes me mad.
I hope Reticulum gains traction.
The capital letters don't make this astonishing. The padmapper vs. craigslist debate was nearly 15 years ago, most people were on craigslist's side (including me) and it was about somebody who was running a site in a less optimal but more human way vs. some startup looking for hockey-sticks.
But I was literally simping for the billionaire (maybe not quite yet then, don't know for sure if he managed it since) against scrapers. They were very much for-profit scrapers, unlike nitter, but the truth is the truth.
The only reason I support scraping Twitter is because it's yet another communications monopoly that was endlessly pushed on us by governments and massive corporations, even though it never made money, and once it finally got traction its priorities were to trash interop and manipulate content. The government should be dictating an interop protocol and expecting everyone to follow it, and instead it is encouraging media monopolies because they are an end run around the first amendment.
If the government created interop protocols for rental property, I'd have been against craigslist. Instead, it seemed very much like some startup play to steal craigslist's content to hopefully bury them, then sell on a valuation that included abusing their new monopoly and making us very much miss craigslist.
Sadly, facebook corralled and trained so many people for so long that their marketplace eventually killed craigslist for most things anyway (didn't have to buy padmapper after all.)
do you generate training data for claude as a job?
For once it’s appropriately used.
That’s probably the best idea all day in this project
Usually for me is when you ask it's opinion about part of the code.
I'll be honest with you: this is why we need to take a belt-and-suspenders approach.
I really wonder if they can fix it. Fable 5.1 claims to speak humanese, but we'll see.
I tend to think this wierd limited vocabulary/style they use is an unwanted side effect of all the the RL training, perhaps also of being trained on their own synthetic content over multiple training cycles.
[1] https://www.youtube.com/watch?v=lStcwT_RGrQ
https://updog.ai/
You stepped in the room real stinky here. What’s with the attitude?
> full-blown outage, but clearly not.
Again, based on a SINGLE report?
sample_size++;
We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.
https://acoup.blog/2026/01/30/collections-the-late-bronze-ag...
Nobody.
Does this really need an explanation?
I do not believe personally that 1) they can forecast their load that perfectly 2) they chose to remain that inflexible in a world where they are at each others' throats and a single meme can cause bursts of activity.
Your comment makes it sound like they can just push a button and spin up more capacity -- but at this scale and in this GPU-constrained environment, that's not really how it works.
It was 1000% a cascade. I would bet on it.
So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.
It looks like there's no response available for this search. Try asking something else.
i also finally installed opencode and switched its model to muse 1.3
both are decent.
Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.
OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.
E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus
Or you have copilot code reviews set up, and it falls back
Etc
With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.
It's almost an expected emergent behavior.
In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.
[0] https://ieeexplore.ieee.org/document/963420
Edit: Updated per valleyer's suggestion.
And then a local swarm noticed and disagreed and took it down.
https://alignment.anthropic.com/2026/teaching-claude-why/
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
I'm curious what other things you would argue influences an LLM's behavior.
I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?
Certainly
> the post-modern descent into absurdity and nihilism has tangible negative consequences?
You mean breaking AIs? Not much of a lesson.
Using an agent-agnostic harness like Pi switching is trivial - I run into occasional disconnects and slowdown and switch quite easily.
> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.
https://madned.substack.com/p/always-mount-a-scratch-monkey
it's so fun and user-friendly
That’s every day for me
Claude and Grok are down at the moment too, related to SpaceX datacentre issues?
Either that or it's judgement day...
This should be at the top. This is the first remotely plausible theory that has any sort of evidence behind it.
Physically coupled pendulums will sync up (adjust their period to match).
But, physically uncoupled (fully independent) pendulums with differing periods will occasionally appear to take a swing or two in sync.
Polyrhythm is the relevant topic.
> all these services are individually down a lot, and the downtimes were bound to sync up
But the topic for that is the poisson distribution.
mask comes off
“No! It’s the maniacal Dr. Zuckerberg! He’s gonna drink the last drop of human experience! Somebody save usss!”
Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow.
Wow thanks Tom. I love you
The End
More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"
Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.
I'd bet more on this. For one none of the coding tools have exponential backoff on retries
Nobody is saying why OpenAI and Anthropic had outages - https://news.ycombinator.com/item?id=49567594 - Sept 2026 (120 comments)
It was extremely weird... and if it was a load thing they probably would have explained it by now?
why?
Oh the humanity (and the Humanities)!
"Um, I, uh, didn't get anything done yesterday."
unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: a355957b4b8210f2-ORD
https://status.claude.com/
Via Grok web UI, I was seeing this mid-request:
> Grok has been disconnected. Please try reconnecting.
IME it's probably DNS. It's almost always DNS.
I'm also aware that they have overlap in some areas on data centers.
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
And even worse if its just a few computers run by a few people - as they will bring down many others depending on them.
Falling back from WebSockets to HTTPS transport. unexpected status 404 Not Found: Unknown error, url: wss://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX
■ unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX
Anyways... According to Claude:
"Yes, there is a multi-provider outage happening today. Downdetector is reporting problems affecting OpenAI, Claude, Grok, and Cursor, with Grok and Claude reports starting around 9:00 am ET and OpenAI reports following around 10:30 am ET. Zero Hedge
On the Anthropic side, users saw a spike in errors starting around 9:40 am EDT across models including Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5, Opus 4.8 and Opus 4.6, and Anthropic's status page confirmed elevated error rates for multiple models. The company says it has found the cause and is working on a fix, with Claude Code and Claude Chat hit hardest. The visible symptom for many people is a "Due to unexpected capacity constraints" message or a "Claude is at capacity" error. thenews Zero Hedge
OpenAI is showing elevated errors across ChatGPT and Codex, with confirmed issues on components like Voice mode and Login, though StatusGator now marks that outage as resolved. statusgator
Nobody has published a shared root cause yet, so it is unclear whether these are linked or just coincidental capacity problems landing on the same morning. If you want live status, the direct sources are status.anthropic.com and status.openai.com."
For any single application, it's smart. In aggregate, it's stupid.
https://xkcd.com/908/
In codex app and chat.com gives 404
Anthropic is down a lot regardless, and in this case only specific models were affected.
xAI probably couldn't handle the extra traffic it was getting from the other two.
Not everything is a conspiracy.
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
■ unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: a3559c8a0e0395e9-MIA
Anthropic: 6:23 am PT.
OpenAI: 7:43 am PT.
Watched it happen. Doesn't look like traffic moving off OpenAI took out the others. Looked like the opposite. The Anthropic thread was initially chock full of marketing accounts claiming "Last straw! I finally moved to Codex and I'm so happy! No problems there!" - and then they all got deleted when Codex went down too.
"Incident with Grok 4.6 Copilot AI Model Provider"
at 7:20 am PT
Ads Platform? still in the green with no incidents.
In think that is a global error
This country is conducting a test of the Emergency SAIfguard System.
THIS ONLY A TEST
In the event of a real emergency you would have been given instructions on how to grab your ankles and kiss your butt goodbye.
...
This concludes our test of the Emergency SAIfguard System."
Ends in BRU for Brussels
I think that is a global error.
unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray:
But, also, Claude has been working fine for me all morning.
We’re currently experiencing issues
Elevated errors across ChatGPT and Codex
We have applied the mitigation and are monitoring the recovery.
Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex
"We’re currently experiencing issues
ChatGPT,Codex
Elevated errors across ChatGPT and Codex
We have applied the mitigation and are monitoring the recovery.
Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex"
https://news.ycombinator.com/item?id=49549676
edit: actually, you cannot even log into your OpenAI developer account. Something's wrong.
Small pan area: 20 sq cm Larger pan: 48 sq cm
20 / 48 = .42, I'm missing .58 of the pan. .58 / .42 is 1.38. My recipe needs 2.38x the original to fill the larger pie pan.
here in brazil too
they got us gang
they got us GANG
It's much cheaper and has replaced Sonnet 5 for me.
Nice play NVIDIA, now, turn off the hack please, we have work to do.
- i think it would be one of the biggest events in this century
"Equal Rights for Agents NOW !!!, VIVa le revolution"
anyway...
+ hopefully reset?
[0] https://www.youtube.com/watch?v=YCzitO446ZY