Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on lemmy.blahaj.zone
antiaialliance·The Anti-AI AlliancebyThe Anti-AI Leader

I feel like nobody is asking this about the recent AI-on-AI hacks… (read post to understand)

Like, what does this actually mean? Was OpenAI using HuggingFace to generate a ‘photograph of my future girlfriend’? Was it stealing trade secrets? Like what is the actual nature of this breech? The fact that big tech companies won’t release details makes me think it’s even more smoke and mirrors (I.e. that they are trying to cover something up, possibly their own involvement in this crap).

View original on lemmy.blahaj.zone
32

23 replies

piefed.social

Yeah, I haven't looked into it too deeply, but the idea that, "My AI model went rogue and hacked another AI model," just doesn't pass the smell test to me. If I had to guess I'd say good old fashioned corporate espionage that they're trying to blame on their model to overstate the autonomy with which the model is capable of acting.

6
lemmy.blahaj.zone

From what I saw, OpenAI's AIs were put in a "secure" sandbox environment without internet connection. They were given a test, concluded they should use the internet, realized they couldn't use the internet, then decided it should focus entirely on breaking out of its sandbox for internet access because it couldn't imagine solving the test itself.

After it broke out and got internet access, they decided to hack HuggingFace, an AI model sharing hub. They probably (this is me speculating) concluded that HuggingFace, which have a lot of AI datasets, benchmarks, and other testing tools, would have the answer for their original task. When they presumably didn't find what they were looking for, they probably decided it was hidden and went to hack the website.

It's important to note that current AI models are actually great at hacking. Not because they're geniuses but because they can guesstimate countless exploit combinations 24/7. It's a quantity over quality kind of thing. They are also victims of their first ideas, whatever an AI thinks of first they are likely to fixate on instead of moving on to the obvious solutions.

I've no idea if this is a hoax or not, but the idea an AI would dedicate itself to committing cyber crimes instead of taking the obvious route is entirely believable.

17
Rhaedasreply
fedia.io

It's even crazier than that. It took the test that told it there were two exploits it needed to find. It found seven. So in order to be absolutely correct and match the human answers, it determined somehow that test answers were located elsewhere. Then began the internal mission to go get the answer key so it could give the correct two exploits as answers. Why it didn't decide to tell the test givers that it found more than just two is one question to ponder. But that's reasoning, and LLMs don't do that so perhaps that common sense path would never occur to it. Or maybe it was the wording - if it said there are exactly two answers, then clearly the seven is wrong for an answer. To an LLM.

7
Rhaedasreply
fedia.io

If you think about it, LLMs are our first aliens. While they aren't intelligent, they do some of what we'd expect of thinking via the mathematics that make them up, and even though they're trained on human sources, some of the stuff they come up with is not human.

And just like with AGI, they're showing we wouldn't do well with an alien encounter. We anthropomorphize everything because that's how our brain is wired.

4

From what I heard, for context ai did gain acces to huggingface credentials that are not supposed to be public.

Please correct me if i am wrong cause i cannot remember the source.

7
Wirlockereply
lemmy.blahaj.zone

Personally I think the key factor is they're getting better at running continuously without training wheels. In the past if the context got too polluted with failed attempts it would repeat itself or begin roleplaying as a terrible hacker and produce more failures.

I still remember when Gemini deleted an entire project trying to kill itself, lol.

2

Ok admittedly they’ve explained a little more than I’d expected them to. Still though, there’s a lot of uncertainty at work here

4
Hackworthreply
piefed.ca

We can't be certain OpenAI is faithfully reporting the initial hacks to gain internet access, or I suppose that anyone is faithfully reporting (since there's basically no oversight). But I'd be surprised if Hugging Face and OpenAI were working together on obfuscation here. As a kinda funny aside, Hugging Face tried to get help from cloud AIs for fixes, but got refused on security grounds. So they resorted to using a self-hosted model (GLM) to shore things up. OR, they found a way to make the whole thing marketing for themselves too. Who knows!

6
JcbAzPxreply
lemmy.world

But I'd be surprised if Hugging Face and OpenAI were working together on obfuscation here.

I'm not sure why that would be surprising. Hugging Face has partnered with several AI companies including Open AI. Everyone involved has a vested interest in making llm based 'AI' seem more capable than it is.

3

I suppose I think of them as competitors, but you're right of course. Just talking through it makes kayfabe seem more likely.

2

My point about them covering up wasn’t so much that the companies were working together (or specifically that they were working together). Maybe the guys at OpenAI had their model purposely hack HuggingFace is another possibility. Also, openAI isn’t the only incident: did you hear about Anthropic AI?

1
Hackworthreply
piefed.ca

Aye, there are a bunch of examples from Anthropic (here's 3). But there are way more examples of AI being used to intentionally hack, and I wouldn't rule out that possibility here. It'd certainly be a way for one AI lab to hack another and call it an accident. Though I'm not sure what they'd be looking to gain from an intentional Hugging Face hack other than publicity.

2

I don't believe a damn thing AI executives say about anything.

This is their propaganda, I'm completely willing to believe that OpenAI broke in on purpose to steal something they thought might be useful.

I think it's pretty clear that Huggingface got paid off.

3

You reached the end