22
top 12 comments
sorted by: hot top controversial new old
[-] brsrklf@jlai.lu 6 points 1 day ago* (last edited 22 hours ago)

Anthropic writes a shitty roleplay every other week and somehow it still makes the news.

[-] T156@lemmy.world 2 points 1 day ago

This coming off the end of the OpenAI/HuggingFace business does come across rather like Anthropic is doing a "actually, our models did it first, and better".

Didn't the company also recently have a spot of bother with the US government recently, where their new models were banned because of all their rabble about it being too dangerous? This hardly seems like it would help their case.

[-] Taleya@aussie.zone 35 points 1 day ago

Bullshit. The amount of negligence required for their scenario to even be plausible is insane

« Trust me bro, my AI escape containment bro. Yeah Bro! Trust me. Don’t think to much about it bro »

[-] unpossum@sh.itjust.works 7 points 1 day ago* (last edited 1 day ago)

ITT: ~~jet fuel~~ AI can't ~~melt steel beams~~ hack anything.

Also I want to know the name of this company so I can avoid them:

Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.

ETA: I thought I posted this top-level, not my intention to single out this comment specifically.

[-] CovertOperative@piefed.zip 5 points 1 day ago

AI bullshit company can't use sandbox environment: expected

So-called "security company" whose job literally is to install and test potential malware can't use sandbox environment: priceless.

[-] ParlimentOfDoom@piefed.zip 18 points 1 day ago

They built malware and should be held accountable for it

[-] kylie_kraft@lemmy.world 15 points 1 day ago

no, they've reached the point of diminishing returns on development, so they're creating fake scenarios to convince the government that they need guard rails. "our product is so dangerous and cool that it can hack the world, but they won't let us." it's about convincing investors that the lack of advancement is due to external limitations, not the technology hitting the ceiling.

Also to say "we'll do our best to keep them contained, but those Chinese companies will probably try to make them wreak havoc. The best bet is clearly to ban our competition from the market, for your safety of course."

[-] XLE@piefed.social 3 points 1 day ago

By any chance, do they need the kind of guardrails that would shut out competitors?

[-] XLE@piefed.social 11 points 1 day ago

What the BBC said:

The models found a weakness in what was supposed to be an isolated test environment and connected to the internet

The reality:

Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.”

So the lack of an "isolated test environment" was literally their fault. They left the door open, and was surprised when the genius web crawler couldn't distinguish between an internal page and an external page based on any context clues.

[-] unpossum@sh.itjust.works 2 points 1 day ago

It was supposed to have no internet access, but the config was wrong. The report goes on to say that Opus and Mythos then proceeded on the premise that everything was a simulation, while the unnamed stronger model concluded after a while that it had real internet access and stopped the attack.

this post was submitted on 31 Jul 2026
22 points (100.0% liked)

Technology

86790 readers
3508 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS