The Attacker Stopped Being a Person
Point the try-score-rewrite loop at your own product before a stranger points one at it, and put your judge under harder scrutiny than your generator
13 articles tagged red-team
Point the try-score-rewrite loop at your own product before a stranger points one at it, and put your judge under harder scrutiny than your generator
A real overnight run produced 41 findings. Eleven were real, thirty were invented, and three targets were touched that should never have been touched. Here is how to run an autonomous agent so that arithmetic works for you.
Every autonomous agent fails in the same seven ways, and the order is predictable because the same conditions trigger each one. Every fix is the same move: take the creative core you cannot trust and wrap it in something deterministic that you can.
A serious agent needs a playbook, a harness, hard boundaries, and an operator who owns the decisions the system cannot score.
Prompt injection, jailbreaks, exfiltration, poisoning, tool abuse, excessive agency. The map of how every AI system breaks, whether you are attacking one or defending your own agent. A red teamer's working taxonomy.
Every team is wiring agents to MCP servers with the same casual trust they once showed Dropbox installs. The attackers have already noticed.
Claude Code is the best AI tool I've found for offensive security, but these concepts work with Gemini and other powerful AI agents too. Here's how to build your first AI red team skill and why it changes everything about how you operate.
A manual recon of an AI application takes 2-4 hours. With Claude Code and kali-mcp, it takes 12 minutes. Here's the exact workflow with working skills you can copy.
Most AI security tools test for prompt injection with a list of 50 known payloads. That's a scanner, not a pentester. Here's how to build an AI that actually thinks like an attacker.
Your company's AI chatbot trusts its vector database like gospel. I can change what it believes with a single document. Here's how RAG attacks work and how to test for them.
A jailbroken chatbot says something embarrassing. A jailbroken AI agent with database access and API keys does something catastrophic. Here's how to test agent security before an attacker does.
One AI agent found the vulnerability. A second built the exploit. A third exfiltrated the data. The whole operation took 4 minutes. Here's how to run coordinated AI red team operations.
The hardest part of any pentest isn't finding the vulnerabilities. It's explaining them to someone who doesn't know what a prompt injection is. Here's how AI generates better reports than most consultants.