I have been spending a lot of time lately running stress tests on human facing AI agents across sales, support, e-commerce, healthcare, finance, and legal.
150+ tests in. Average score is still < 75/100.
The patterns we keep seeing are interesting enough that we want to keep building on it. More industries, more agents, more data.
We are also working on publishing our first report on the state of AI agent performance. The data is starting to tell a real story and I want to get it out there properly.
If you are building or deploying an AI agent that talks to real people, I want to hear about it. Not to pitch you anything. I genuinely want to understand what you built, what problem it is solving, and how you thought about the conversation design.
Drop a comment or send me a DM. Happy to run a free test and share the scorecard back with you. No strings attached.
The more agents we test the better the data gets for everyone.
Brought to you by Sourcee
We find journo requests from across the web and deliver them directly to your inbox.