"error", "exception", "500", "failing"?
WORKFLOW 1: Debug Errors
Full-stack observability for AI agents — debug errors from logs, trace latency across services, manage alerts, respond to incidents with runbooks, and protect SLO error budgets via Datadog, Grafana, or New Relic.
What this skill teaches
An MCP server tells an agent which actions are available. A skill adds the judgment around those actions: how to recognize the job, which sequence to follow, what to avoid, and how to decide that the result is complete.
Orchestrate full-stack observability — query logs, search traces, monitor metrics, manage alerts, handle incidents, track SLOs, and execute runbooks. Use when debugging errors, investigating latency, checking service health, managing alerts, responding to incidents, reviewing SLO burn rate, or finding runbooks.
Architecture
The skill stays readable and portable because it contains instructions rather than service credentials or business data. ADK-Rust supplies it to the agent, the agent chooses from its reviewed tool boundary, and the connected MCP server performs the authenticated operation.
The agent receives a goal expressed in ordinary language.
Matches intent, supplies the decision guide, and narrows the tool boundary.
Plans the workflow and streams each meaningful step through the runtime.
Executes authenticated operations against the system that owns the capability.
The skill's completion rules shape the evidence returned to the user.
Decision guide
These routes come directly from the skill instructions. They help the model recognize intent and select a focused tool or workflow instead of improvising across the entire capability surface.
"error", "exception", "500", "failing"?
WORKFLOW 1: Debug Errors
"slow", "latency", "timeout", "p99"?
WORKFLOW 2: Trace Latency
"health", "CPU", "memory", "disk"?
WORKFLOW 3: System Health
"alert", "firing", "paging"?
WORKFLOW 4: Alert Management
"incident", "outage", "down"?
WORKFLOW 5: Incident Response
"SLO", "error budget", "reliability"?
WORKFLOW 6: SLO Tracking
"dashboard", "overview"?
WORKFLOW 7: Dashboards
Unclear?
get_system_health first for overall picture
Proven workflows
A workflow joins several tool calls into a task the user actually recognizes. The skill explains the sequence and the intended result while ADK-Rust streams the agent's progress through the shared runtime.
Logs → traces → root cause
Find the slow span
CPU/memory/disk overview
Triage + runbook + acknowledge
Declare + investigate + resolve
Budget remaining + forecast
Tool boundary
This allowlist is declared by the skill. It keeps the agent focused on the actions needed for this job while mcp-observability retains responsibility for authentication, validation, and the connected system.
Working rules
Operating boundaries
Install and connect
Install the repository where your ADK-Rust skill loader can discover it, connect mcp-observability, and confirm the declared tools are available before asking the agent to use the workflow.
git clone https://github.com/zavora-ai/skill-observability-monitoring.git \
~/.skills/skills/observability-monitoringlet skills = SkillLoader::from_dir("~/.skills/skills").await?;
let skill = skills.load("observability-monitoring").await?;
let agent = LlmAgentBuilder::new("agent")
.instruction(skill.instructions())
.tools(skill.allowed_tools(toolset)?)
.build()?;The repository's compatibility statement: Requires mcp-observability server connected (Datadog, Grafana Cloud, New Relic, or Custom API).
Official documentation
The repository remains authoritative for its exact instructions, examples, helper scripts, assets, MCP requirements, license, and later updates.
Keep exploring
Source record
Repository metadata for this skill entry.