← Recent activity
Panel· San Francisco, CA

Hosting a panel on enterprise AI agents in San Francisco

I moderated and gave the keynote for a panel with leaders from ServiceNow, TikTok, Tencent, and VeloDB on enterprise AI agents, agent safety, governance, and how frontier labs and traditional enterprises differ on AI.

Placeholder for a photo of the enterprise AI agents panel in San Francisco

On September 19th I hosted a panel in San Francisco on enterprise AI agents. I opened with a keynote and then moderated the discussion. The panelists brought very different vantage points:

The keynote

Before the panel I gave a short keynote on evaluating AI agents. The core argument: you cannot govern what you cannot measure. An agent that calls tools and takes actions needs the same rigor we apply to any production system, which means tracing every step, scoring outcomes against a golden set before launch, and watching real traffic after. I walked through how we do that at AgentX, including LLM-as-judge scoring and the human feedback loop that keeps the judges honest.

Robin on stage giving the keynote on AI agent evaluation
Opening keynote on AI agent evaluation.

What we talked about

Enterprise AI agents in practice. Where agents are actually running inside large companies today, what separates a demo from a system a business unit will depend on, and why the hard part is rarely the model.

Agent safety. What "safe" means once an agent can call tools and take actions on real systems, how teams think about blast radius, and where human review still belongs in the loop.

Governance of agents. Why governance has to be designed in rather than bolted on: who owns an agent, how its behavior is evaluated before and after it ships, and how you audit what it did and why. This is the part I care most about, and it is the reason AgentX exists.

Frontier labs versus traditional enterprises. The labs and the enterprises deploying their models often have different stances on risk, speed, and control. We compared how each side thinks about it and where the two are converging.

My takeaways

  1. Everyone on stage agreed that evaluation and observability are the foundation for trusting an agent, not a nice-to-have.
  2. Governance is becoming an engineering problem, not just a policy document.
  3. The most interesting enterprise work is happening in the gap between what frontier labs ship and what a regulated company can actually adopt.

Thanks to the panelists and to everyone who came out. If you were in the room and want to keep the conversation going, reach out.