About this event
Most teams pick an AI model the same way they pick a coffee machine - the demo looks good or they hear about it from someone, so they assume it works. But in production, evaluation decisions directly impact cost, latency, accuracy, and user trust. Without a structured eval practice, you're flying blind when a model regresses, or a new release claims to be better.
This session gives you a framework for evaluating LLMs against *your* workloads and operationalizing evaluation as a continuous practice, not a one-time exercise.
**What we'll cover:**
* Why standard benchmarks are often misleading for enterprise use cases
* Building a golden dataset: how to sample, label, and version evaluation sets from real production traffic
* Key metrics beyond accuracy: hallucination rates, precision/recall tradeoffs, cost-per-correct-answer, and latency percentiles
* Running evals at scale: open-source tools (Ragas, PromptFoo, LangSmith) and using Amazon Bedrock Model Evaluation
* Setting a model evaluation gate: how to enforce a promotion threshold before any model goes to production
* Demo: running a side-by-side eval across two models on a real classification task
**High level schedule:**
* 4:30 - 4:45: Arrivals
* 4:45 - 5:45: Presentation & Demo
* 5:45 - 6:00: Q&A
* 6:00 - 6:30: Networking
**Important instructions**
* Event starts at 4:45 pm EST, but please allow at least 15 minutes for security to process registration
* Please ensure that your meetup profile has your **full name.** Both first and last name are required and we will not be able to register attendees with just abbreviations or incomplete names.
* There is an optional networking and Q&A event at the end of the meetup
Group: AWS New York | Official Meetup
Group page: https://www.meetup.com/aws-nyc
Hosts: Avichal Chum, William Torrealba
Going: 137
Group members: 7,592
Official listing: https://www.meetup.com/aws-nyc/events/316007124/