Skill Detail

Trace and evaluate agent runs with MLflow

Instrument LLM and agent applications with MLflow tracing, evaluation, prompt tracking, and monitoring so operators can debug behavior before and after deployment.

Monitoring & AlertsMulti-Framework
Monitoring & Alerts Multi-Framework Security Reviewed
⭐ 26k GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill trace-and-evaluate-agent-runs-with-mlflow Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python, uv or pip, MLflow tracking server, LLM or agent application
Install & setup
Start a local tracking server with uvx mlflow server, point the application at it with mlflow.set_tracking_uri("http://localhost:5000"), then enable the relevant LLM or agent tracing integration such as mlflow.openai.autolog().
Author
MLflow
Publisher
Organization
Last updated
May 20, 2026
Quick brief

Use MLflow when an agent team needs repeatable observability for LLM and agent runs across providers and frameworks. The operator starts an MLflow tracking server, enables autologging or an integration for the agent stack, then reviews traces, evaluations, prompts, costs, and production monitoring signals in one place. Invoke this for agent debugging, regression checks, evaluation runs, prompt lineage, and production monitoring; keep the scope to LLM and agent observability workflows rather than general ML experiment management.

How it works

What this skill actually does

Inputs and prerequisites: Python, uv or pip, MLflow tracking server, LLM or agent application.

Setup notes: Start a local tracking server with `uvx mlflow server`, point the application at it with `mlflow.set_tracking_uri(“http://localhost:5000”)`, then enable the relevant LLM or agent tracing integration such as `mlflow.openai.autolog()`.

Source and verification boundary: use https://mlflow.org/docs/latest/genai/ as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.

Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.