Monitor

This chapter covers logging and tracing for inference services, resource monitoring and dashboards, and model bias and drift monitoring. Together they give observability across the full lifecycle of an inference service: real-time replica logs, and multi-dimensional dashboards for infrastructure, GPU resources, token throughput, and API traffic. Logging and resource monitoring serve inference service users, while configuring monitoring dashboards requires the administrator view.

Bias and drift monitoring is provided by TrustyAI from the Evaluate & Safety chapter; deploying the TrustyAIService it needs is covered there, in Deploy TrustyAI Service.

Logging & Tracing

Resource Monitoring

Bias and Drift Monitoring