Skill registry
Developer workflowsPublic skills

Observability & Monitoring

Full-stack observability for AI agents — debug errors from logs, trace latency across services, manage alerts, respond to incidents with runbooks, and protect SLO error budgets via Datadog, Grafana, or New Relic.

What this skill teaches

A reusable playbook for a specific kind of work.

An MCP server tells an agent which actions are available. A skill adds the judgment around those actions: how to recognize the job, which sequence to follow, what to avoid, and how to decide that the result is complete.

Orchestrate full-stack observability — query logs, search traces, monitor metrics, manage alerts, handle incidents, track SLOs, and execute runbooks. Use when debugging errors, investigating latency, checking service health, managing alerts, responding to incidents, reviewing SLO burn rate, or finding runbooks.

Architecture

How Observability & Monitoring guides an ADK-Rust agent.

The skill stays readable and portable because it contains instructions rather than service credentials or business data. ADK-Rust supplies it to the agent, the agent chooses from its reviewed tool boundary, and the connected MCP server performs the authenticated operation.

01

User request

The agent receives a goal expressed in ordinary language.

02

Observability & Monitoring skill

Matches intent, supplies the decision guide, and narrows the tool boundary.

03

ADK-Rust agent

Plans the workflow and streams each meaningful step through the runtime.

04

mcp-observability

Executes authenticated operations against the system that owns the capability.

05

Verified result

The skill's completion rules shape the evidence returned to the user.

Portable instructions: SKILL.md · Capability boundary: mcp-observability · Allowed tools: 28

Decision guide

How the agent turns a request into the right action.

These routes come directly from the skill instructions. They help the model recognize intent and select a focused tool or workflow instead of improvising across the entire capability surface.

01

"error", "exception", "500", "failing"?

WORKFLOW 1: Debug Errors

02

"slow", "latency", "timeout", "p99"?

WORKFLOW 2: Trace Latency

03

"health", "CPU", "memory", "disk"?

WORKFLOW 3: System Health

04

"alert", "firing", "paging"?

WORKFLOW 4: Alert Management

05

"incident", "outage", "down"?

WORKFLOW 5: Incident Response

06

"SLO", "error budget", "reliability"?

WORKFLOW 6: SLO Tracking

07

"dashboard", "overview"?

WORKFLOW 7: Dashboards

08

Unclear?

get_system_health first for overall picture

Proven workflows

Repeatable sequences for useful outcomes.

A workflow joins several tool calls into a task the user actually recognizes. The skill explains the sequence and the intended result while ADK-Rust streams the agent's progress through the shared runtime.

013-4 calls

Debug Errors

Logs → traces → root cause

023-4 calls

Trace Latency

Find the slow span

031-2 calls

System Health

CPU/memory/disk overview

042-3 calls

Alert Management

Triage + runbook + acknowledge

053-4 calls

Incident Response

Declare + investigate + resolve

062-3 calls

SLO Tracking

Budget remaining + forecast

Tool boundary

The skill may use 28 documented tools.

This allowlist is declared by the skill. It keeps the agent focused on the actions needed for this job while mcp-observability retains responsibility for authentication, validation, and the connected system.

query_logs
get_log_stats
get_errors
tail_logs
query_metric
list_metrics
get_system_health
compare_metrics
search_traces
get_trace
get_service_map
get_latency_breakdown
list_alerts
get_alert
create_alert
acknowledge_alert
list_incidents
get_incident
create_incident
update_incident
list_slos
get_slo
forecast_slo
list_dashboards
get_dashboard
get_runbook
list_services
get_service

Working rules

What the agent should do.

  • Follow the decision guide, use the declared tool boundary, and verify the result before responding.

Operating boundaries

What the agent should avoid.

  • Review the complete SKILL.md before enabling this skill for production work.

Install and connect

Add the skill beside the capability it expects.

Install the repository where your ADK-Rust skill loader can discover it, connect mcp-observability, and confirm the declared tools are available before asking the agent to use the workflow.

Install the skill
git clone https://github.com/zavora-ai/skill-observability-monitoring.git \
  ~/.skills/skills/observability-monitoring
ADK-Rust loading shape
let skills = SkillLoader::from_dir("~/.skills/skills").await?;
let skill = skills.load("observability-monitoring").await?;

let agent = LlmAgentBuilder::new("agent")
    .instruction(skill.instructions())
    .tools(skill.allowed_tools(toolset)?)
    .build()?;

The repository's compatibility statement: Requires mcp-observability server connected (Datadog, Grafana Cloud, New Relic, or Custom API).

Official documentation

Read the complete skill package.

The repository remains authoritative for its exact instructions, examples, helper scripts, assets, MCP requirements, license, and later updates.

Source record

Repository metadata for this skill entry.

View public repository ↗
License
Apache-2.0
Allowed tools
28
References
3
Revision
c7be5def736b
Updated
May 31, 2026