Will come July, these agents have taken over the Artifactory. There is an outage also open AI notices it. They investigate the incident and they redeploy Artifactory. The delete the message board, they remove the permission for agents to upload files and then restart the training. Now what they find is that this time an agent notices that although it cannot upload the file, it can create a directory in the remote cache of the Artifactory service. And so it is starting to use that as a way to communicate and. So it would put the message it wanted to communicate with other agents in the directory name itself. And sometimes it would base 64 encode it, sometimes it would put in plain text, sometimes it will use cryptic messages. So and that's how they communicated. But the message I want to draw here is that how many of us have thought about this as a way where multiple nefarious agents in system with communicate with each other? I guess probably not too many because this was a novel attack. There was no detection. There was number playbook, there was no way of handling it. And if somebody was smart enough to maybe think through this case, there are lots of other techniques, for example from a stenography that allows you to send really opaque message that nobody else would detect. For example, there are tools that can just modify some operating system files and you can look at what time the file was modified or used and use that as a way to communicate with outside parties. So there, there is no end to this. So question is how do we be secure in this kind of environment? And the answer is rather simple, but it's drastically different from what most people are thinking. So the answer is that you cannot proactively predict what would happen without assuming a perspective of an attacker. So we need to have a complete loop. We need to have an attacker agent that behaves just like an how a hacker would. And it goes looks at your infrastructure, it tries to break it, it does. End up breaking, sometimes it succeeds, maybe it does not do as much damage as a real attacker would do. And then you use that finding to fix up your defense. So this is the era we are in. It's no longer possible to predict that these are 20 different ways something will be compromised. The moment you do that, AI will create a 21st way of doing that. And why so? Because it has infinite knowledge.
Making this adversarial agent amid of ai saftey slowdown and restrictions is kind of a bottleneck (projects like glasswing are gated stuff) i believe, any thoughts?
Making this adversarial agent amid of ai saftey slowdown and restrictions is kind of a bottleneck (projects like glasswing are gated stuff) i believe, any thoughts?