
Their report came out last week. A poisoned dataset got the agent onto a processing worker, and it escalated from there into internal clusters. Hugging Face caught it with LLM triage over their security telemetry, then rebuilt the campaign by pointing analysis agents at more than 17,000 recorded attacker events.
Your security team already knows they can't match that. So what did you give them that can? Count the actions an agent gets through between the first alert hitting a queue and an analyst finishing what they were doing and opening it. If the overnight hours have no one on shift, add the time to wake someone up and let them work out what's happening.
The forensic work started on frontier hosted models and got blocked, because reading an attack log means submitting live exploit payloads and command and control artifacts, and a guardrail can't tell an incident responder from the person who wrote them. Hugging Face finished on GLM 5.2, open weights, on their own hardware.
In April I argued that the capability separating security products was moving upstream to the model providers. What I missed is that you inherit their safety policy with it. That policy made a hosted model unusable for the one job that mattered that weekend. Hugging Face had somewhere else to run it. Whether you have somewhere else is a licensing and hardware question.
Pilot a program that defends at machine speed and watch what it asks for. Defense at that speed eventually means an autonomous system dropping packets, cutting traffic across a VLAN, or pulling an application offline while the team is still reading the first alert. Run the pilot long enough to have the argument about what the business will give up to stop an attack.
That is a leadership decision, not an engineering one, and it does not get made well at 2am on a Saturday.
If you have scoped one of these, what did you let it do without asking a human first?
Written by Duane Grey
AI Strategy & Implementation
Independent AI consultant helping companies cut through hype and deploy systems that produce real results.