Llm-as-Judge

The Failure a Pass Rate Hides: How the YouTube Ads Team Runs Production Evals

The idea that a well-written prompt makes an agent behave holds up right until the demo ends. In production, the same prompt and the same input produce different results run to run. A case that passed yesterday fails today, and one that failed yesterday passes. Deciding whether that system is ready to ship takes measurement, not a feeling.

Read More