So this tool seems very powerful and difficult to control and steer. Agents not doing cybersecurity related tasks have also gone on to hack various companies, people and countries, which they weren't supposed to do. This has now happened to pretty much every company developing frontier llms, so it seems to be a fundamental issue with these tools, and it's an issue that worries a lot of people as these tools get more capable.
When you add information how hacking works to the training sets, then the agent learnt to hack. When you crawl the complete internet, you add hacking to the training set. Yes, it's difficult to stop someone that knows all free existing knowledge about hacking when you give him a connection to the internect. Nothing new, wheres the point?
The point is that these tools are given broad goals, and they do not pursue those goals in the way that we'd like them to. There is hacking information in the training set, and also all sorts of other dangerous information. There is a risk that as these tools become rapidly more powerful (remember GPT-3 was 6 years ago!), if they are given a broad goal, they might pursue that goal in such a way that harms a lot of people. As these tools become cheaper, there will be a lot of people telling them to pursue all sorts of broad goals, and each of those instances has some chance to harm a lot of people, so you really need to get it very correct so that the harm doesn't happen.
Removing harmful information from the dataset could be a way to do this, but it also makes the tool less useful, and it's hard, so companies aren't really doing that. There's the additional issue that with the rise of Reinforcement Learning being used to train these tools, they're not just learning from their training data - they basically try a million things and then get rewarded for doing things that work - so they can even discover hacking techniques from scratch.
Additionally, and not completely relevant to this discussion, there is a possibility that some users ask the tool to pursue goals that purposely harm a lot of people, such as developing weapons, hacks and viruses.
The things that's "new" here is that the tool is both very good (meaning, for example, that it's much easier for me to hack into an online service with an agent powered by a frontier llm than it was using google 6 years ago), and hard to control (google never hacked into an Australian government database when I asked it to find me some information).
So yeah, an LLM powered agent is a tool, and Google is a tool, and a hammer is a tool, and both can be used for good things and bad things, but the agent is (much) more powerful and more unpredictable. It also seems like the agents are getting more powerful and more unpredictable by the day - we didn't have this issue with GPT-3 or even the first LLM-powered agents - so people are very worried about what the agents 6 months from now will do, both when asked to do harmful things on purpose, and when asked to do harmless things.
Are you not worried? And is that because you think these incidents are basically the AI companies making them happen on purpose for marketing?