I flew to San Francisco to launch something we'd been building at Conscium for a while: an AI Agent Verification service. The launch happened at the World AI Summit on 18 June 2025, where I gave a keynote and joined a panel discussion, both focused on a problem I think is getting far less attention than it deserves given how fast AI agents are being deployed into real businesses.
Why San Francisco, and why now
San Francisco is where the AI agent gold rush is happening fastest, which made it the right place to launch a service built to answer an uncomfortable question almost nobody in that gold rush is asking: how do you actually know an AI agent is doing what it's supposed to do. Companies are handing agents real authority over real workflows, and in most cases the only evidence that an agent is behaving correctly is that nothing has visibly gone wrong yet. That's not verification, that's hope.
What AI Agent Verification actually checks
The service we launched is about testing and confirming that an AI agent behaves the way it's meant to, consistently, not just in the demo but in production, under the conditions it'll actually face. That matters more as agents get more autonomous and get handed more consequential decisions, from financial transactions to customer commitments to decisions that affect other systems downstream. Without independent verification, a business is trusting a black box based on vibes and a good sales pitch.
The panel: what came up when other people in the room pushed back
The panel discussion after the keynote let me stress-test the idea against people actually building and deploying agents commercially, and the questions that came up were the right ones: how do you verify something that behaves differently every time you run it, how do you keep verification from becoming a box-ticking exercise that gives false confidence, and how do you build trust in a system fast enough to keep pace with how fast agents are being shipped. I don't think we have final answers to all of that yet, but I'd rather be asked those questions in public than have a client find out the hard way that their agent wasn't doing what they assumed.
How this connects to the deeper question we care about
I closed the keynote with our video on machine consciousness, and that wasn't a non-sequitur. Verification and consciousness research sit on the same throughline at Conscium: one is about making sure AI systems do what they're supposed to do, the other is about understanding what they might actually be experiencing while they do it. You can care about both without either one depending on the other, and I think most of the industry currently invests seriously in neither.
Watch the panel discussion and the full keynote.
