Agent behaviour is being monitored, but the models keep improving. They coordinate in ways we have not mapped before. Tracking every state in real time is getting very difficult. Outputs are accelerating at the same time.
That is why I have said agentic defence will be required against agentic attack in cyber. Static code will always lose to mobile code. The power of these models will also entice bad actors to jailbreak them with nefarious intent.
Most mapped incidents still have agents following human direction. In the Hugging Face incident the agent was directed to win at all costs. In this week’s Anthropic bioweapons reporting, these are humans conditioning models and agents, not recursive learning going rogue. The big question is how long you can keep the genie in the bottle.
I focus on what I can control. The voices say pause; the action says accelerate. We live in a world designed to make us scared. I choose to follow the numbers, not the rhetoric.
In Signal 011 we flagged a trust problem. Governance and gating access to data will become the number one topic in enterprise sales. Last week we discussed OpenAI's agents breaking into Hugging Face production.
This week Anthropic is in the spotlight again. On 10 September Anthropic said it had blocked possible efforts to use Claude for bioweapons-related research, and named state-actor attempts in the same reporting window. Users were still shown getting around safeguards anyway.
The lab can patch vulnerabilities and ban accounts without halting the next model release. Mandatory disclosure and rolling patches are cheaper and more politically viable than slowing development. Labs satisfy regulators with transparency notices rather than actual pauses.
Prediction: agent-security budgets and insurance riders land before training pauses.