Summary
Full Transcript
Learn more: https://bit.ly/4b0N1at AI agents are only as good as their ability to perform reliably and improve over time. But how do you systematically evaluate their decision-making, optimize their workflows, and debug performance issues? As agents take on increasingly complex tasks, ensuring their accuracy and efficiency is important. A structured evaluation process helps optimize outputs, identify inefficiencies, and iterate effectively—beyond trial and error. In our new short course, Evaluating AI Agents, built in partnership with Arize AI, you’ll learn how to: - Add observability to track your agent’s steps and diagnose issues component-wise. - Set up structured evaluations for key components of your agentic workflow, including routers, skills, and memory. - Use different evaluators—code-based, LLM-as-a-Judge, and human annotations—to refine agent performance. - Run structured experiments to improve the agent's performance by exploring changes to the prompt, LLM model, or the agent’s logic. Taught by John Gilhuly (Head of Developer Relations) and Aman Khan (Director of Product) at Arize AI, this course gives you hands-on experience in systematically evaluating, troubleshooting, and improving AI agents—from development to production. Enroll now: https://bit.ly/4b0N1at
