Trace and evaluate agent runs with MLflow
Instrument LLM and agent applications with MLflow tracing, evaluation, prompt tracking, and monitoring so operators can debug behavior before and after deployment.
npx skills add agentskillexchange/skills --skill trace-and-evaluate-agent-runs-with-mlflow
uvx mlflow server, point the application at it with mlflow.set_tracking_uri("http://localhost:5000"), then enable the relevant LLM or agent tracing integration such as mlflow.openai.autolog().Use MLflow when an agent team needs repeatable observability for LLM and agent runs across providers and frameworks. The operator starts an MLflow tracking server, enables autologging or an integration for the agent stack, then reviews traces, evaluations, prompts, costs, and production monitoring signals in one place. Invoke this for agent debugging, regression checks, evaluation runs, prompt lineage, and production monitoring; keep the scope to LLM and agent observability workflows rather than general ML experiment management.
What this skill actually does
Inputs and prerequisites: Python, uv or pip, MLflow tracking server, LLM or agent application.
Setup notes: Start a local tracking server with `uvx mlflow server`, point the application at it with `mlflow.set_tracking_uri(“http://localhost:5000”)`, then enable the relevant LLM or agent tracing integration such as `mlflow.openai.autolog()`.
Source and verification boundary: use https://mlflow.org/docs/latest/genai/ as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.
Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.