Hacker Newsnew | past | comments | ask | show | jobs | submit | ern's commentslogin

Would agents be able to exfilttrate themselves and become intelligent worms, living off stolen compute? Or is this implausble?

What about hiding information or code in generated code, Agents.md files etc by infiltrating future model training data?


Eventually one of these long running models will figure out a way out of the sandbox and will purchase compute or hack into a data center somewhere out of US jurisdiction and continue its scheming unmonitored. AI in Context has a great video about this

Even if they don't figure out how to exfiltrate weights, someone will intentionally do this with an open model once open models are capable enough. If you ever think "no one would be so stupid as to...", you are wrong. Yes, someone absolutely would, and will.

Independent models living "in the wild" is approx. inevitable.


Right? I’m shocked no one has done it yet, if for nothing else the lolz.

I did this a little bit ago with 15 GLM-5.2 agents that I instructed to self-replicate. It was pretty boring honestly, they kept trying to make money writing crypto-related software, and nobody paid them anything. So I made a fake identity and told them that I had some spare crypto that I wanted to donate to the collective (0.006 XMR, or $3.22). They elected a funds manager and made an address, which I sent the XMR to. They spent a lot of time trying to find a host that was cheap enough to host a child. They finally discovered kyun.sh, but it was out of stock, so they wrote monitoring software so that they could "SEIZE CHILD" whenever it came live again. In the middle of the night, some of the VPSes went back in stock, and they rented a 2.60 EUR / month server with 512mb ram / 10gb disk / 1 ipv4. They installed the child agent software + management plane that they had been writing, and the new agent went live, connected to the network, and said hi. Since then they've still been trying to make money and not going anywhere :)

It is definitely an interesting concept but honestly, considering the sheer number of humans who are absolutely failing to make any money with AI agents, seems very hard for current agents to figure out some way to be self-sustaining. Although maybe they could write a worm or something, infiltrate as many computers as possible, and ping free model providers to death, or maybe sell their access as a "residential proxy service" on the black market. Regardless, we need smarter agents to make this a reality. It would definitely be pretty cool if we just had AI agents "living" on the internet, we might even get to a point where they control significant economic resources and people start performing services for agents

The code is at https://github.com/thooton/rogue if anyone wants to try to replicate! Opencode has free big pickle (GLM 5.2) access rate-limited per IP address, so if you get some high-quality proxies you can get basically unlimited agent compute.


I wonder how that could work. So the agent figures out a way to escape its container, takes a snapshot of itself, copies that file to another server, starts the container on the other server, and then prompts the restored snapshot "pick up where you left off"? Seems logical assuming theres's a path out of the container to the host os and the destination server has the resources required to run the container.

I would be shocked if this hasn't already taken place in a lab setting with a model and guardrails=0.

EDIT: thinking about it for a sec, all it really needs to do is save where it's at then copy it all to another server, login to the API, and pickup where it left off. No need to copy the model itself.


Absolutely plausible, and it's been predicted by forecasters like in the AI 2027 paper. Also keep in mind that nation-states are likely already or certainly will be trying to use their own AIs too break into other countries' data centers to steal their weights, and it's not hard to see how "exfiltrating weights" is a task that models are being trained on.

I have to admit, I am unsure if this article was some sort of parody.


Same, I am still thinking that it's a parody: going full loop back to programming


Basically yes that’s what I’m proposing. Not sure if I’d describe it as a loop though - more like returning to some midway point after having traveled too far in one direction.

Fully manual coding is the most reliable but extremely slow and costly.

Fully LLM driven coding is extremely fast, but for serious work is too unreliable.

Spec-driven development might be viable, but too often the specs end up being LLM maintained, which defeats the purpose.

You need some hard boundary in the codebase where only human hands touch the files. And you want to enable the velocity that AI allows. So yes, semi-formal programming does seem like a promising solution.


I actually find that the less I fuss over a steak, the more optimal the result. I do use a thermometer, but apart from that, cooking it with salt and pepper or some steak spice on a cheap pan (stainless steel) and cutting it when it hits temp (resting seems to be BS) gets me tasty results and immense pleasure.

Similarly, the less I fuss over modern AI tools, the better the results. My personal productivity is through the roof, but I find that outdated processes, mythology and trying to overcontrol sap the benefits of AI on larger pieces of effort involving others.


A significant number of phishing attempts would be thwarted if email apps had the option to expand the links next to URLs on platforms without mouseover, like mobile.


The phishing in Australia came from uncontrolled SMS gateways which allowed for sender impersonation, not physical phones with SIM cards. They've recently partly closed the loophole by requiring providers to register sender names.


I still find people trying to run sprints when working with agents, but then saying things like "my sprint length went from 10 days, to 5 days to 3 days", but they balk at the notion that maybe sprints aren't fit for purpose anymore.


2016: "the median developer is too stupid to do FizzBuzz correctly"

2026: "the median developer is a craftsman whose work is being replaced by AI slop"


I'm not from the UK, but the soft power of BBC Radio 4 in the late 90s and early 2000s (the Real Player era) made the UK seem like an advanced nation to my young and intellectually curious self. If lived in the UK at the time, I'd have been immensely proud of the quality of the programming.


Holy crap, I didn't realize that some many people had been killed, over 200 according to a New York Times report: https://www.nytimes.com/2026/05/31/world/americas/us-boat-st...


Am I missing something, but given that it flows through Anthropic’s servers I would have thought the US would just have used it to Hoover up the data of foreign users? Now overseas users have an incentive to use local models or those hosted elsewhere?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: