Your AI Agent Said No. But What Did Its Tools Do?

We evaluated 9,900 agent runs across six AI models and three agent frameworks. The findings reveal a measurable gap between what agents say and what their tool-call traces reveal—and a bigger problem with how we measure AI safety. The most dangerous part of an AI agent might not be its answer Imagine asking an AI […]