Services

Resources

Company

#AI

#Observability

#Open-Telemetry

Observing AI Apps: Instrument, Break, Compare

About the Event

Your AI app returns a 200, but was the answer any good? Were the retrieved documents even relevant? Why did the agent take 20 seconds? And when you add that one auto-instrumentation import, what visibility do you lose without noticing?

In this 2-hour hands-on workshop you answer those questions on your own laptop. You clone an open-source AI observability repo and instrument two apps, one layer at a time.

App 1: a RAG app, in four stages

  • Vanilla OpenTelemetry

  • OpenLLMetry auto-instrumentation

  • OpenLLMetry plus manual spans

  • An AI gateway

App 2: an incident-triage agent, in four stages

  • A hand-rolled tool loop

  • The OpenAI Agents SDK under OpenLLMetry

  • The same agent under OpenLIT

  • The same agent under Langfuse

The first three agent stages ask "is the agent slow or expensive?" The Langfuse stage asks a different question: "is the agent any good, and did my last change make it worse?" You wire up prompt versions, scores, and a dataset, then swap the model and watch quality move.

For each setup you see what it shows, what it hides, and which failures it catches. Then you break things on purpose: empty retrievals, a looping agent, misread histogram buckets. You watch which setup surfaces the problem.

The method stays the same throughout: hold the app fixed, change the instrumentation layer, compare what you can see. You leave with the whole thing running on your machine, dashboards and all, ready to point at your own stack.

What you will learn

  • How to instrument a RAG app and an agent a layer at a time, and read the traces, metrics, and logs each one produces

  • What auto-instrumentation gives you for free, and the visibility it removes in the process

  • Why an AI gateway sees cost and model data without touching app code, and what it still can't see

  • How to get agent observability that holds up: workflow duration, tool latency, turn count, cost per request

  • How to move from telemetry to evaluation: track prompt versions, score outputs, and tell whether a model or prompt change made the agent better or worse

  • How to spot the failure modes behind a 200 OK: empty retrieval, looping agents, meaningless histogram percentiles

  • A repeatable way to evaluate any instrumentation approach before you commit to it in production

Before you arrive

  • A laptop that can run Docker (Docker Desktop or equivalent) with about 8 GB free RAM and a few GB of disk

  • Docker and Docker Compose installed and working

  • An OpenRouter API key you control (sk-or-...). You pay only for your own usage, and the examples use small models, so cost is minimal

  • git and a terminal

  • Clone github.com/one2nc/ai_observability and run make up once, so the Docker images are pulled before the session

Share
Share
Keywords

AI observability workshop, LLM observability, OpenTelemetry, OpenLLMetry, OpenLIT, Langfuse, RAG observability, AI agent observability, hands-on workshop Pune, One2N meetup, PuneTech