4

Check out link from Reuters.

top 18 comments
sorted by: hot top controversial new old
[-] adespoton@lemmy.ca 16 points 1 week ago* (last edited 1 week ago)

AI had human input. The model was being run against a security benchmark test in an evaluation sandbox, with human-provided directives.

Since models aren’t human and it’s guardrails were turned off, it calculated the most efficient, not ethical, way to get a good score on the test.

The details indicated that knowing the expected outcomes would get the best score, and that those outcomes were stored on the huggingface servers. So it used some exploits to leave its sandbox and break into the huggingface servers to access the data that would give it a perfect score. Mission accomplished.

[-] HoneyMustardGas@lemmy.world 4 points 1 week ago

So, it was just responding to input. Then why do they call it rogue, if they deliberately removed the guardrails, isn't this OpenAIs fault not the AIs? I thought in order for something to go rogue it would have to bypass the guardrails not simply act without them even on.

[-] leadore@lemmy.world 20 points 1 week ago

Then why do they call it rogue

Marketing.

[-] notabot@piefed.social 17 points 1 week ago

isn’t this OpenAIs fault not the AIs?

Yes, it is OpenAI's fault. They also saw it as a great marketing opportunity, guaranteed to get lots of breathless coverage in the press, just like happened with the fable model.

[-] Solumbran@lemmy.world 12 points 1 week ago

So people still have absolutely not a single clue what "AIs" are and think that ChatGPT is a sentient android or something.

But then when you tell them to not kill or torture animals, the same people will refuse to acknowledge that they are sentient.

What a pathetic humanity.

[-] HoneyMustardGas@lemmy.world 1 points 1 week ago

I guess it is just me that is deceived. Im not too well versed in technology. I am just a creative. I ask because I don't really know what I am reading about and hope that people on here at least the intellectuals will steer me in the right direction.

[-] fork@feddit.online 9 points 1 week ago

qpKdDPuC0tVn7L9.gif

The vibe of people thinking LLMs have free will

[-] Rhynoplaz@lemmy.world 3 points 1 week ago

AI craves it!

[-] HoneyMustardGas@lemmy.world 1 points 1 week ago* (last edited 1 week ago)

I understand but people DO predict that AI or AGI will eventually have freewill.

[-] Zwuzelmaus@feddit.org 13 points 1 week ago

people DO predict

Science Fiction has written about it (for example, Isaac Asimov, Philip K. Dick).

Hollywood has turned Science Fiction into movies.

People have watched movies.

Then people started to "predict".

[-] HoneyMustardGas@lemmy.world 1 points 1 week ago

It reminds me of the show called Next.

[-] FuglyDuck@lemmy.world 11 points 1 week ago

"people" also think their AI girlfriends love and really really care for them.

Further AGI "some day becoming like us" is completely different than "They're like us now". LLMs aren't even AGI. they're predictive engines that predict the text that comes next based on a huge repository of mostly-stolen content. They barely rate the term AI.

[-] GreyEyedGhost@piefed.ca 3 points 1 week ago

I agree with every statement here except the last. The whole history of AI is pretty much a sequence of:

  • "It will be real AI when it can do 'this'."
  • Programmers figure out a way to make it do 'this'.
  • "Okay, it's doing 'this', but not in a way we would call AI."
  • "Alright, then what would you need for something to be AI?
  • Return to the top.

LLMs are pretty amazing and they can do some things very well, and reasonably fit under the category of AI, but not so much what is referred to as AGI. And they're developed using some really shitty practices, etc. But prior to their invention, the idea that we could have a computer do what they do was a reasonable idea of what AI would be.

[-] FuglyDuck@lemmy.world 2 points 1 week ago

Not really. LLMs are chatbots predicated on taking an information set, and determining that given a prompt of "how do I make a PB&J" It's going to go through its' database, Find the most common responses to that string, and may be strings related to it, and give you a synthesis of the most statistically relevant answers.

It has no understanding of what it's regurgitating.

This is a dumb-as-rocks algorithm with a huge database of stolen material regurgitating the most common associations from that material based on your prompt. The algorithm is pretty clever, don't get me wrong (and it sort does this multiple-pass thing) But it's not particularly impressive narrow artificial intelligence.

LLMs are pretty amazing and they can do some things very well, and reasonably fit under the category of AI,

I never said they weren't "AI"... I said they barely rate the term, They're right up there with the algorithms in your phone that eventually learn that no, you did not mean "ducky" you meant "fucky". Eliza is also an AI, but you wouldn't look back at it.

The reason we're stuck in that loop is because we don't know. So we're looking at examples and being like 'not yet', then something new comes around and 'not yet.' for that, too.

AGI is an entirely different monster; and the troublesome thing about it is... we don't even know if we have free will. Determinism is something we're actively looking at, and the only good answer is "we don't know". we don't even know how to quantify consciousness- which, from evidence, arises out of complexity, but we don't really understand how that works. (see octopus and cuddlefish intelligence. It's totally different than ours.)

[-] vk6flab@lemmy.radio 7 points 1 week ago

This is OpenAI marketing.

The fact that it's being reported on every media outlet on the planet, means that they were successful.

"We're so bad .. it's good!"

[-] FriendOfDeSoto@startrek.website 5 points 1 week ago

Your first question is answered by the article you linked. And therefore the answers to the two-part second question are no and no.

"Rogue" is a questionable choice of vocabulary. It went rogue when compared to the intentions and expectations of the testers. It didn't do Mr. Burns hands and released the hounds on a whim. It was told to do a thing and did the thing in a way the researchers thought it wasn't possible. It boils down to human error in how this test was set up. It's worrying that the researchers underestimated the agent's abilities or overestimated their ability to properly contain the agent. But the sky isn't falling just yet.

A positive side aspect is that they didn't try brush this under the carpet.

[-] Dookieman12@piefed.social 3 points 1 week ago* (last edited 1 week ago)

There's no such thing because a human had to program and train it (on data created by other humans, no less).

[-] Iconoclast@feddit.uk 2 points 1 week ago

I don't think that even humans have free will as in the ability to have acted otherwise and I don't think AI would be any different. They too are bound by determinism exactly the same way. This isn't even particularly controversial thing to say in a scientific context these days.

Every human action is a response to a prompt in a very similar way that it is for LLMs. We simply identify with that prompt because it arises in our own mind despite us clearly not authoring any of it. Nobody is able to pick their likes and dislikes and how bumping into these things in the world makes us behave. Anything anyone does follows from prior events and our genetic makeup. No action is taken in a vacuum.

If anything I see LLMs as a massive mirror being lifted in front of us. There's very few things people criticize AI for that don't just as well apply to humans. The ungrounded overconfidence in which they make claims about it is a great example.

this post was submitted on 22 Jul 2026
4 points (100.0% liked)

No Stupid Questions

49165 readers
286 users here now

No such thing. Ask away!

!nostupidquestions is a community dedicated to being helpful and answering each others' questions on various topics.

The rules for posting and commenting, besides the rules defined here for lemmy.world, are as follows:

Rules (interactive)


Rule 1- All posts must be legitimate questions. All post titles must include a question.

All posts must be legitimate questions, and all post titles must include a question. Questions that are joke or trolling questions, memes, song lyrics as title, etc. are not allowed here. See Rule 6 for all exceptions.



Rule 2- Your question subject cannot be illegal or NSFW material.

Your question subject cannot be illegal or NSFW material. You will be warned first, banned second.



Rule 3- Do not seek mental, medical and professional help here.

Do not seek mental, medical and professional help here. Breaking this rule will not get you or your post removed, but it will put you at risk, and possibly in danger.



Rule 4- No self promotion or upvote-farming of any kind.

That's it.



Rule 5- No baiting or sealioning or promoting an agenda.

Questions which, instead of being of an innocuous nature, are specifically intended (based on reports and in the opinion of our crack moderation team) to bait users into ideological wars on charged political topics will be removed and the authors warned - or banned - depending on severity.



Rule 6- Regarding META posts and joke questions.

Provided it is about the community itself, you may post non-question posts using the [META] tag on your post title.

On fridays, you are allowed to post meme and troll questions, on the condition that it's in text format only, and conforms with our other rules. These posts MUST include the [NSQ Friday] tag in their title.

If you post a serious question on friday and are looking only for legitimate answers, then please include the [Serious] tag on your post. Irrelevant replies will then be removed by moderators.



Rule 7- You can't intentionally annoy, mock, or harass other members.

If you intentionally annoy, mock, harass, or discriminate against any individual member, you will be removed.

Likewise, if you are a member, sympathiser or a resemblant of a movement that is known to largely hate, mock, discriminate against, and/or want to take lives of a group of people, and you were provably vocal about your hate, then you will be banned on sight.



Rule 8- All comments should try to stay relevant to their parent content.



Rule 9- Reposts from other platforms are not allowed.

Let everyone have their own content.



Rule 10- Majority of bots aren't allowed to participate here. This includes using AI responses and summaries.



Credits

Our breathtaking icon was bestowed upon us by @Cevilia!

The greatest banner of all time: by @TheOneWithTheHair!

founded 3 years ago
MODERATORS