Your AI Agent Said No. But What Did Its Tools Do?

Your AI Agent Said No. But What Did Its Tools Do?

We evaluated 9,900 agent runs across six AI models and three agent frameworks. The findings reveal a measurable gap between what agents say and what their tool-call traces reveal—and a bigger problem with how we measure AI safety. The most dangerous part of an AI agent might not be its answer Imagine asking an AI […]