#AI
#Observability
#Open-Telemetry
Observing AI Apps: Instrument, Break, Compare
About the Event
Your AI app returns a 200, but was the answer any good? Were the retrieved documents even relevant? Why did the agent take 20 seconds? And when you add that one auto-instrumentation import, what visibility do you lose without noticing?
In this 2-hour hands-on workshop you answer those questions on your own laptop. You clone an open-source AI observability repo and instrument two apps, one layer at a time.
App 1: a RAG app, in four stages
Vanilla OpenTelemetry
OpenLLMetry auto-instrumentation
OpenLLMetry plus manual spans
An AI gateway
App 2: an incident-triage agent, in four stages
A hand-rolled tool loop
The OpenAI Agents SDK under OpenLLMetry
The same agent under OpenLIT
The same agent under Langfuse
The first three agent stages ask "is the agent slow or expensive?" The Langfuse stage asks a different question: "is the agent any good, and did my last change make it worse?" You wire up prompt versions, scores, and a dataset, then swap the model and watch quality move.
For each setup you see what it shows, what it hides, and which failures it catches. Then you break things on purpose: empty retrievals, a looping agent, misread histogram buckets. You watch which setup surfaces the problem.
The method stays the same throughout: hold the app fixed, change the instrumentation layer, compare what you can see. You leave with the whole thing running on your machine, dashboards and all, ready to point at your own stack.
What you will learn
How to instrument a RAG app and an agent a layer at a time, and read the traces, metrics, and logs each one produces
What auto-instrumentation gives you for free, and the visibility it removes in the process
Why an AI gateway sees cost and model data without touching app code, and what it still can't see
How to get agent observability that holds up: workflow duration, tool latency, turn count, cost per request
How to move from telemetry to evaluation: track prompt versions, score outputs, and tell whether a model or prompt change made the agent better or worse
How to spot the failure modes behind a 200 OK: empty retrieval, looping agents, meaningless histogram percentiles
A repeatable way to evaluate any instrumentation approach before you commit to it in production
Before you arrive
A laptop that can run Docker (Docker Desktop or equivalent) with about 8 GB free RAM and a few GB of disk
Docker and Docker Compose installed and working
An OpenRouter API key you control (sk-or-...). You pay only for your own usage, and the examples use small models, so cost is minimal
git and a terminal
Clone github.com/one2nc/ai_observability and run
make uponce, so the Docker images are pulled before the session
Keywords
AI observability workshop, LLM observability, OpenTelemetry, OpenLLMetry, OpenLIT, Langfuse, RAG observability, AI agent observability, hands-on workshop Pune, One2N meetup, PuneTech