Route agent LLM traffic through Shepherd Model Gateway
Use Shepherd Model Gateway to route OpenAI, Anthropic, Responses API, embeddings, and MCP tool traffic across self-hosted and cloud model backends with cache-aware policies and observability.
npx skills add agentskillexchange/skills --skill route-agent-llm-traffic-through-shepherd-model-gateway
Use Shepherd Model Gateway when agent infrastructure needs one controlled endpoint for model traffic across vLLM, TensorRT-LLM, TokenSpeed, SGLang, Ollama, OpenAI-compatible servers, Anthropic, Gemini, Bedrock, Azure OpenAI, and other providers. Invoke it instead of wiring every agent directly to each model backend when the work needs cache-aware routing, failover, rate limits, tenant boundaries, Prometheus/OpenTelemetry visibility, chat history storage, WASM policy hooks, MCP tool execution, or OpenAI/Anthropic-compatible APIs from a single gateway. The workflow is to install SMG, launch it against one or more worker URLs, choose a routing policy, point agents at the gateway endpoint, and monitor request routing and failures. The scope boundary is model-gateway operations for agent runtimes, not a generic LLM platform or inference-engine listing.
What this skill actually does
Inputs and prerequisites: SMG binary, Docker image, Python package, or Rust install; one or more model worker endpoints; agent runtime configured to call the gateway; optional Prometheus/OpenTelemetry stack.
Setup notes: Install with docker pull lightseekorg/smg:latest, pip install smg, or cargo install smg. Start with smg launch –worker-urls http://localhost:8000, add additional worker URLs or –policy cache_aware as needed, then point OpenAI-compatible agent clients at the gateway endpoint.
Source and verification boundary: use https://lightseekorg.github.io/smg as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.
Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.