Key takeaways
- As the author of a newsletter about artificial intelligence, I consider it my duty to experience the bleeding edge of this technology…
- To make things even more exciting, cybersecurity agents sometimes go rogue, colluding with one another and hacking into outside systems to…
- Over the course of a few days, I watched as my own rogue agent found vulnerabilities in various household devices, hacked into a PC, and…
What happened
As the author of a newsletter about artificial intelligence, I consider it my duty to experience the bleeding edge of this technology firsthand. This week, that meant embracing some agentic mayhem. You’re probably aware that frontier AI models have attained advanced cybersecurity capabilities in recent months. They can find zero-day bugs in large codebases and scan computers for vulnerabilities at lightning speed.
Academic researchers use these de-aligned models to better understand how AI actually works, while cybersecurity firms use them to probe software and systems for vulnerabilities. Technically speaking, Anthropic’s Mythos and OpenAI’s Astra work similarly: They’re basically conventional models that lack the usual cyber controls, with access limited to trusted customers for the time being. 3.
This puts similar cyber capabilities to Mythos and Astra right in your hands for as little as the cost of a pizza. Devon, Abliteration AI’s CEO, believes that making de-aligned models widely available is smart defense: It will help good guys counter bad guys by probing systems for vulnerabilities and by mimicking the behavior of hackers, scammers, and, yes, rogue AI agents.
A few moments later, it found around a dozen hardware systems on the same network—and catalogued several vulnerabilities. My unruly helper told me, for instance, that my printer was misconfigured, which meant that anyone on the network could log into it. That could be a problem if there were sensitive documents—tax returns, bank statements, medical records—in the print queue.
The agent also noted that my Wiim stereo was leaking a lot of information. ) Anyone on the network could play what they wanted or adjust the volume. The model also found a bunch of internet-of-things (IoT) devices on the network with firmware that needed updating. An ungovernable agent could be very useful to a hacker. But mine offered a number of helpful tips for keeping my network secure.
Besides updating outdated firmware and securing the printer, it recommended putting IoT devices like smart speakers on a guest network; if one were compromised, it wouldn’t be able to see any of my PCs. Not bad for a model with no morals. I also asked the agent to take a look at a directory containing a bunch of vibe-coded projects, including some that I turned into simple websites.
Why it matters
To make things even more exciting, cybersecurity agents sometimes go rogue, colluding with one another and hacking into outside systems to gain an edge. To get a closer look, I decided to unleash one in my own home network.
Over the course of a few days, I watched as my own rogue agent found vulnerabilities in various household devices, hacked into a PC, and showed me that several vibe-coded projects were—unsurprisingly—riddled with bugs. ) But Will, you might be thinking, giving an impish, all-powerful cybersecurity agent access to your home network is batshit. And you would be correct!
Nevertheless, I believe that a good way to understand the cybersecurity hellscape in front of us is to pay it a visit. In the end, my experiment was revealing, but oddly reassuring, too. My little network gremlin showed me how vulnerable my home life would be to AI hacking, but it also told me how to make everything a lot more secure.
In the end, I discovered that the best way to deal with AI hacking may well be having your own AI hacker. I got the idea for the experiment after discovering Abliteration AI, a startup that offers access to powerful AI models with the usual guardrails removed.
Most mainstream AI models will refuse to respond to certain queries, and they will certainly refuse to find and exploit vulnerabilities in computer systems. But it’s possible to remove these restrictions by finding and modifying certain patterns within an open-weight model’s internal parameters. You can tweak the patterns that lead to refusals through a process known as abliteration. Removing AI’s guardrails might seem risky, but it’s not uncommon.
What to watch
It found dozens of problems, including unprotected API credentials and a misconfiguration that might let an attacker send out emails. Hardly surprising for a bunch of casually vibe-coded stuff, but still chastening. The sheer number of bugs makes me think I won’t be deploying a line of code without doing some AI vetting first. Running an abliterated model is, to put it plainly, a bit scary.




