
Token Economics, Paisa, Prompts & Production | Entain Tech Brews
In this Tech Brews at Entain talk, Srivatsa RV and Biswa dive deep into token economics and the art of token optimization for LLM applications. They explain the fundamental concepts of tokens, the critical differences between prefill and decode phases, and how caching strategies can drastically reduce AI inference costs. By breaking down prompts into cache reads and un-cached tokens, they reveal practical techniques for maintaining cache stability, structuring context, and optimizing dynamic data to achieve up to 90% cache hit rates. The session also covers LLM observability, defining metrics like Time to First Token (TTFT) and Cost Per Successful Task, and how to govern AI agents in enterprise and API-driven workloads.









