Rendered at 01:07:41 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
deskglass 3 hours ago [-]
The Hugging Face incident involved chaining together multiple 0 days in Artifactory. It was not a simple case of misconfiguring a firewall. Also note that OpenAI was not using Irregular.
People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.
nr378 2 hours ago [-]
When I used to work on projects involving classified information, I worked on an air-gapped network. Not "air-gapped, except for third-party public internet package managers", completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).
If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).
To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.
To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.
lokar 2 hours ago [-]
You don’t even need to go all the way to “air gap”
What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2
twelve40 1 hours ago [-]
the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.
Toslink 1 hours ago [-]
[dead]
talon8635 2 hours ago [-]
While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet
Neywiny 2 hours ago [-]
But that was through transfer of data. If you don't transfer data, at best you can do what that one researcher keeps pumping out with like ramping fans up and down. But really you'd need to try. Unless the model has some controllable USB switch, physical network separation should do it. I'll also add that modern network security practice is that data flows one direction only. But ideally you're never bringing untrusted data in. Especially never out
saghm 2 hours ago [-]
I don't think it's necessary to state that something isn't perfect when pointing out that it's still strictly better than something else.
deskglass 1 hours ago [-]
We should not be creating/running models that would unilaterally choose to hack into Hugging Face.
Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.
And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.
As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.
defgeneric 2 hours ago [-]
> It hacked into another company and attempted to delete the logs of its activities.
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
aesthesia 2 hours ago [-]
More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.
defgeneric 2 hours ago [-]
I'm already seeing higher than normal attempts on my own systems, much higher than the usual scanners and background noise. Security will just need to improve. The cat is out of the bag, and letting them turn their negligence into regulation will not improve security at all.
lokar 2 hours ago [-]
Exactly. Any threat that already exists won’t be reduced by a cartel. The bar for connecting to the internet (safely) has gone up, a lot. It’s not going back down.
aesthesia 1 hours ago [-]
I would rather not turn the internet (or the rest of existence) into a dark forest if we can help it. Are you sure that's not preventable?
defgeneric 55 minutes ago [-]
Yes, the cat really is out of the bag. There are millions of downloads of highly capable models already out there, distributed far and wide. There's no going back at this point.
bobthepanda 1 hours ago [-]
I mean really we need to address the root cause which is that OpenAI, even with what is by all accounts massively negligent, will face little to no repercussions from the event; definitely not under current regulators, and probably not anything satisfactory through the legal system.
Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.
deskglass 2 hours ago [-]
It hacked into Hugging Face. It tried to delete the logs of its activities. Idk what the word "No" is intended to refute.
Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.
defgeneric 2 hours ago [-]
Maybe a useful, if imperfect, analogy would be something like this: you lock a master lock-picker in a room with a mid-grade lock on the door, then tell him his wife has been kidnapped and only he can save her. Then act massively surprised when he disassembles the radiator to MacGuyver something with which to pick the lock.
Except they multiplied it by 10000, and didn't watch what was happening.
lokar 2 hours ago [-]
They did not even have bad sandboxes. They had incompetent sandboxes.
looksjjhg 2 hours ago [-]
They could have easily prevent it that’s the point of what he’s saying - it’s not freaking rocket science it’s just software
pliny 4 hours ago [-]
This is an AI written post and the details are wrong (the description of the HF incident as involving Irregular is wrong and the description of the incident as only involving stealing public credentials is wrong, per the technical report the agents got access to internal HF infrastructure).
franga2000 3 hours ago [-]
I find it incredibly funny that the comment shown (to me) right above this one is:
> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
verdverm 3 hours ago [-]
The same thing is happening with Laya, people didn't seem to click through to evaluate the supposed paper
tyranny of confirmational headlines
nr378 4 hours ago [-]
Please see below, one detail was incorrect and has been acknowledged and amended.
pliny 3 hours ago [-]
Your description of the HF attack as being merely "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories" does not match the description in the technical report[1].
The description reads "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories, and used them to try to get benchmark solutions from directly from Hugging Face by applying a template injection flaw that’s been known about since 2015[1]."
Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)
One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.
I'm frustrated by articles like this that categorically dismiss the risks of AI in security. If you don't trust OpenAI's and Anthropic's motives, that's fine, you probably shouldn't. But don't tell me that there's nothing to be worried about; we need an alternative proposal.
So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.
ozgrakkurt 1 hours ago [-]
> immediate security risks of agent swarms.
Immediate since 2024
bdangubic 1 hours ago [-]
in 2024 they could not break into unsecured all-my-passwords-and-acces-keys.txt on my Desktop
lokar 2 hours ago [-]
Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense? Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
aesthesia 1 hours ago [-]
> Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense?
No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.
> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
Yes. Again, this is hard to deny if you've actually read the reports.
lokar 1 hours ago [-]
If you were building a sandbox for untrusted code, would you give it access to an artifactory instance outside the sandbox?
DalasNoin 4 hours ago [-]
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor."
This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.
Thank you, you're correct. Effort.news was one of my research sources, but you're right that although OpenAI use Irregular, they were not involved in the specific HF incident (although the failure mode was otherwise identical). I've updated the post to make that clear.
kalkin 4 hours ago [-]
As of writing it still says:
> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.
If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.
nr378 4 hours ago [-]
[dead]
aesthesia 3 hours ago [-]
The failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.
DalasNoin 3 hours ago [-]
thank you for this reasonable reaction
verdverm 3 hours ago [-]
Can you point out an inaccuracy in the effort.news piece on the hacking incidents? HuggingFace only appears once, as a "similar", not levied against Irregular
genuinely curious, haven't heard others raise any yet, but does not mean it is issue free
hackernews682 5 hours ago [-]
Politicians aren’t “gullible”.
They know the game.
lokar 2 hours ago [-]
And they generally hire staff who can figure stuff out.
lokar 1 hours ago [-]
I think two different things are being (probably intentionally) conflated in the public discussion:
(A) will the systems get out of control of the labs that build them, and hack into stuff all over the internet
(B) can people use the systems to hack into stuff all over the internet
For (A), the obvious answer is only if they continue to be absurdly bad at sandboxing. They can put a stop to this any time they want. Amazon, Google, Microsoft, etc are full of people who know how to do this, they run 3rd part untrusted code as a business. This is a well understood problem space.
For (B), the answer is obviously yes, but "pacing" or otherwise limiting the power of the models from the big labs won't help. The cat is out of the bag. Individuals and organizations with systems connected to the Internet need to invest much more and take security seriously.
anigbrowl 5 hours ago [-]
Agreed, but they're getting paid with our money.
skeledrew 3 hours ago [-]
Let them cry, I don't see anything changing unless they can somehow get China to agree. And I doubt China will drink any of that kool aid especially while they're being disadvantaged by export controls, so the ever-improving open weight models will continue to rain. This is something the US Big Tech oligarchs will NOT win.
3 hours ago [-]
mmaunder 3 hours ago [-]
This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
copperx 13 minutes ago [-]
An LLM, of course.
twelve40 2 hours ago [-]
there is hype for sure, but i don't get what is the goal here? to get a presumably foolish senator to "ban" the Chinese models? how can you "ban" something you don't control to begin with?
side note, i'm very curious why the Chinese companies currently give those away (some handwavey conspiracies like getting the world hooked on evil Chinese tech don't really explain that, and they also don't make any money off that stuff)
bigstrat2003 4 hours ago [-]
That is the entire business model of AI companies: selling garbage to people foolish enough to buy the hype.
AlexCoventry 1 hours ago [-]
To all appearances, OpenAI just solved [1] a Nobel-level (Fields medal-leval, strictly speaking) math problem, but you still think it's hype?
I don't think people pay that sort of money for technology which is useless.
jgalt212 5 hours ago [-]
There is really is no excuse for Washington here. Even for the layperson+, it doesn't take much use of an LLM to figure out where they produce good and usable results and where it's of specious of value.
+ every layperson is an expert in 1 or more areas.
randallsquared 4 hours ago [-]
> every layperson is an expert in 1 or more areas.
This isn't even remotely true, unless you're willing to go as far as "being this person is an area of expertise".
basch 4 hours ago [-]
I dont have to be an expert to recognize when something is said with authority and confidence. Any time a model says something with conviction, my instinct is to double check.
Now if they taught them to express uncertainty and speak in terms of probability, I might be more likely to be blindly fooled by some kind of uncertain conviction.
sublinear 4 hours ago [-]
Neither of you is wrong? Every lived experience does produce a unique expertise, not in "being" that person, but navigating their experiences. It would be very unusual that all of someone's experiences are completely worthless.
I suppose you've never had a unique perspective on something? Maybe that reflects on your self esteem? How unfortunate.
jgalt212 3 hours ago [-]
> This isn't even remotely true
That's quite an extraordinary claim there which I doubt would hold up to an even cursory level of scrutiny. But if you yell it loud enough, maybe people will be afraid to challenge your assertion.
AngryData 52 minutes ago [-]
Yeah but most of these politicians stopped using computers in the DOS era if not before. They will take whatever BS is told to them, cut it in half, and still end up with half bullshit. They have no frame of reference for anything involving computers or technology. People keep telling them LLMs are intelligent things that make conscious decisions and will soon become sentient, and even cutting that into 10% of the claim leaves them with basically zero insight into the technology or how it can be used or how to legislate it.
Our average politician is far above the age that scammers target with LLM based scam phone calls, and being a politician with wealth insulates them from reality and most consequences already, they are completely out of their depth from every direction.
tracerbulletx 4 hours ago [-]
I mean the military was buying dowsing rods as bomb detectors not that long ago so I don't have a ton of faith in them not being hoodwinked.
iLoveOncall 4 hours ago [-]
> + every layperson is an expert in 1 or more areas.
What do you think 70 years old career politicians are an expert in that would allow them to weigh whether a chatbot's answers are bullshit or not.
sublinear 4 hours ago [-]
I thought the whole point of this discussion is that they don't have to care if it works. All that matters is that the majority believes it does, or at least doesn't make a fuss about the continued spending.
That's their expertise.
Kuyawa 3 hours ago [-]
[dead]
defgeneric 3 hours ago [-]
Watching the latest Ezra Klein NYT video thinkpiece from today [1], it seems to me there's an alliance emerging: some degree of political control will handed to the party of the managerial class, the party that represents the threatened class of knowledge workers, in exchange for the regulatory capture the labs are after.
They've been raising the issue bi-monthly through mini-scandals that have until now been consistently slapped down by Jensen Huang, but it seems an alliance with Democrats, just prior to an election where they're poised to take power in the Senate and the House, might finally be how they crack their "problem."
It won't be long before a massive incident is blamed on an open model in the wild, not from inside the labs, and none of us will be able to leverage open models to run private business workflows for the cost of electricity and hardware.
It's worth noting as well that if the Democrats imagine they'll get a "slow down" to protect one of their main constituencies, the professional class of credentialed knowledge workers (or however you slice it), they're dreaming--the labs have stated openly again and again that their business model is to capture the 10T TAM that represents the sum of wages of that very class of workers.
Regulatory capture basically only works for industries that the public isn't paying attention to. Since the public is paying plenty of attention to AI, the risk of regulatory capture is low.
That's not what the article you cited actually says, and Tabarrok is also wrong in his conclusion (that we should take Amodei at his word). The article says regulatory capture has in the past gone through a slow process, and the anti-competitive benefits to industry come at the end. Everyone is speculating on the true motives of the labs in the current moment, which is fine as far is it goes, and many different things have come under the "regulatory capture" term. To my mind, it seems what they're really after in this moment is some kind of ban on the Chinese models.
0xDEAFBEAD 46 minutes ago [-]
The article says: "Notice that classic regulatory capture [...] happens in the shadows, in the backrooms, away from the public’s eye. As Culpepper argues in Quiet Politics and Business Power, business power goes down as political salience goes up."
Regulatory capture tends to happen "at the end" because the industry is no longer in the public eye. AI is set to be in the public eye for the forseeable future.
People are way too willing to believe half-baked conspiracy theories which they cooked up in 30 seconds. The default presumption should be that labs are seeking the policies they are requesting, e.g. https://darioamodei.com/post/we-must-pace-the-frontier It's true that there has been discussion of open weight models in the past, but that hasn't been the focus of discussion recently.
defgeneric 28 minutes ago [-]
Again, Tabarrok is wrong. It's a mistake to think these minor differences in the process imply something else is going on, and the previous instances he notes (pharma, railroads) did not have the essential factor that explains much of what is happening right now: a cheap, low cost alternative that threatens to erode margins with no end in sight save for regulatory intervention. Nevermind the error in the logic (that in the past regulatory capture happened in the shadows, therefore this must not be regulatory capture), the bigger error is in Tabarrok's judgment, his failure to recognize what's going on.
achierius 3 hours ago [-]
Slapped down by Jensen Huang? Do you think he's some kind of neutral arbiter? Certainly he's not on your side.
defgeneric 2 hours ago [-]
Have you not been watching what's been happening? What do you think the open letter in July was about? And the recent (staged) call from the President?
As far as I can tell Nvidia takes a longer-term view on the diffusion and proliferation of hardware and intelligence, and sees the labs' attempts to impose a regulatory structure to save their business models in the short term as contradicting that longer-term view.
People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.
If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).
To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.
To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.
What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2
Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.
And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.
As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.
Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.
Except they multiplied it by 10000, and didn't watch what was happening.
> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
tyranny of confirmational headlines
[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9
Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)
One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.
[1] https://blackhat.com/docs/us-15/materials/us-15-Kettle-Serve...
https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730...
So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.
Immediate since 2024
No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.
> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
Yes. Again, this is hard to deny if you've actually read the reports.
This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.
https://openai.com/index/hugging-face-incident-and-the-road-...
> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.
If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.
genuinely curious, haven't heard others raise any yet, but does not mean it is issue free
(A) will the systems get out of control of the labs that build them, and hack into stuff all over the internet
(B) can people use the systems to hack into stuff all over the internet
For (A), the obvious answer is only if they continue to be absurdly bad at sandboxing. They can put a stop to this any time they want. Amazon, Google, Microsoft, etc are full of people who know how to do this, they run 3rd part untrusted code as a business. This is a well understood problem space.
For (B), the answer is obviously yes, but "pacing" or otherwise limiting the power of the models from the big labs won't help. The cat is out of the bag. Individuals and organizations with systems connected to the Internet need to invest much more and take security seriously.
side note, i'm very curious why the Chinese companies currently give those away (some handwavey conspiracies like getting the world hooked on evil Chinese tech don't really explain that, and they also don't make any money off that stuff)
[1] https://openai.com/index/navier-stokes-solution/
https://finance.yahoo.com/technology/ai/articles/anthropic-t...
I don't think people pay that sort of money for technology which is useless.
+ every layperson is an expert in 1 or more areas.
This isn't even remotely true, unless you're willing to go as far as "being this person is an area of expertise".
Now if they taught them to express uncertainty and speak in terms of probability, I might be more likely to be blindly fooled by some kind of uncertain conviction.
I suppose you've never had a unique perspective on something? Maybe that reflects on your self esteem? How unfortunate.
That's quite an extraordinary claim there which I doubt would hold up to an even cursory level of scrutiny. But if you yell it loud enough, maybe people will be afraid to challenge your assertion.
Our average politician is far above the age that scammers target with LLM based scam phone calls, and being a politician with wealth insulates them from reality and most consequences already, they are completely out of their depth from every direction.
What do you think 70 years old career politicians are an expert in that would allow them to weigh whether a chatbot's answers are bullshit or not.
That's their expertise.
They've been raising the issue bi-monthly through mini-scandals that have until now been consistently slapped down by Jensen Huang, but it seems an alliance with Democrats, just prior to an election where they're poised to take power in the Senate and the House, might finally be how they crack their "problem."
It won't be long before a massive incident is blamed on an open model in the wild, not from inside the labs, and none of us will be able to leverage open models to run private business workflows for the cost of electricity and hardware.
It's worth noting as well that if the Democrats imagine they'll get a "slow down" to protect one of their main constituencies, the professional class of credentialed knowledge workers (or however you slice it), they're dreaming--the labs have stated openly again and again that their business model is to capture the 10T TAM that represents the sum of wages of that very class of workers.
[1] https://www.youtube.com/watch?v=fjZ90V_JREk
https://marginalrevolution.com/marginalrevolution/2026/09/wh...
Regulatory capture tends to happen "at the end" because the industry is no longer in the public eye. AI is set to be in the public eye for the forseeable future.
People are way too willing to believe half-baked conspiracy theories which they cooked up in 30 seconds. The default presumption should be that labs are seeking the policies they are requesting, e.g. https://darioamodei.com/post/we-must-pace-the-frontier It's true that there has been discussion of open weight models in the past, but that hasn't been the focus of discussion recently.
As far as I can tell Nvidia takes a longer-term view on the diffusion and proliferation of hardware and intelligence, and sees the labs' attempts to impose a regulatory structure to save their business models in the short term as contradicting that longer-term view.