
Framework Doesn’t Matter (Much). Model Does.
We ran the largest cross-framework agentic security evaluation we know of — 6 models, 6 execution conditions, 5 attack families, 7,020 trials — and the headline result surprised us: which orchestration framework you use barely moves the needle on security outcomes. Model choice and attack type dominate. Framework explains about 0.06% of the variance. That’s not what the existing literature said going in. The question we started with Everyone building agentic AI right now is choosing between LangChain, CrewAI, AutoGen, LlamaIndex, the OpenAI Agents SDK, or just calling the model directly. A natural question follows: does that choice matter for security? If a model is safe under one framework, is it safe under another? The most direct prior attempt to answer this — Nguyen & Husain’s comparative penetration test across AutoGen and CrewAI — found a large gap: 52.3% attack refusal under AutoGen versus 30.8% under CrewAI, holding the model



