Agentes de Voz en Tiempo Real

Los agentes en tiempo real permiten interacciones de voz con asistentes de IA mediante transmisión de audio bidireccional. La crate adk-realtime proporciona una interfaz unificada para construir agentes habilitados para voz que funcionan con la API de Realtime de OpenAI y la API de Gemini Live de Google.

Visión General

Los agentes en tiempo real difieren de los LlmAgents basados en texto en varias formas clave:

CaracterísticaLlmAgentRealtimeAgent
EntradaTextoAudio/Texto
SalidaTextoAudio/Texto
ConexiónSolicitudes HTTPWebSocket
LatenciaSolicitud/respuestaStreaming en tiempo real
VADN/ADetección de voz del lado del servidor

Arquitectura

              ┌─────────────────────────────────────────┐
              │              Agent Trait                │
              │  (name, description, run, sub_agents)   │
              └────────────────┬────────────────────────┘
                               │
       ┌───────────────────────┼───────────────────────┐
       │                       │                       │
┌──────▼──────┐      ┌─────────▼─────────┐   ┌─────────▼─────────┐
│  LlmAgent   │      │  RealtimeAgent    │   │  SequentialAgent  │
│ (text-based)│      │  (voice-based)    │   │   (workflow)      │
└─────────────┘      └───────────────────┘   └───────────────────┘

RealtimeAgent implementa el mismo trait Agent que LlmAgent, compartiendo:

  • Instrucciones (estáticas y dinámicas)
  • Registro y ejecución de herramientas
  • Callbacks (before_agent, after_agent, before_tool, after_tool)
  • Transferencias de sub-agentes

Inicio Rápido

Instalación

Agregue a su Cargo.toml:

[dependencies]
adk-realtime = { version = "2.0.0", features = ["openai"] }

# For Vertex AI Live (Google Cloud with ADC auth)
# adk-realtime = { version = "2.0.0", features = ["vertex-live"] }

# For LiveKit WebRTC bridge
# adk-realtime = { version = "2.0.0", features = ["livekit"] }

# For all transports (except WebRTC which needs cmake)
# adk-realtime = { version = "2.0.0", features = ["full"] }

Uso Básico

use adk_realtime::{
    RealtimeAgent, RealtimeModel, RealtimeConfig, ServerEvent,
    openai::OpenAIRealtimeModel,
};
use std::sync::Arc;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("OPENAI_API_KEY")?;

    // Create the realtime model
    let model: Arc<dyn RealtimeModel> = Arc::new(
        OpenAIRealtimeModel::new(&api_key, "gpt-realtime")
    );

    // Build the realtime agent
    let agent = RealtimeAgent::builder("voice_assistant")
        .model(model.clone())
        .instruction("You are a helpful voice assistant. Be concise.")
        .voice("alloy")
        .server_vad()  // Enable voice activity detection
        .build()?;

    // Or use the low-level session API directly
    let config = RealtimeConfig::default()
        .with_instruction("You are a helpful assistant.")
        .with_voice("alloy")
        .with_modalities(vec!["text".to_string(), "audio".to_string()]);

    let session = model.connect(config).await?;

    // Send text and get response
    session.send_text("Hello!").await?;
    session.create_response().await?;

    // Process events
    while let Some(event) = session.next_event().await {
        match event? {
            ServerEvent::TextDelta { delta, .. } => print!("{}", delta),
            ServerEvent::AudioDelta { delta, .. } => {
                // Play audio (delta is base64-encoded PCM)
            }
            ServerEvent::ResponseDone { .. } => break,
            _ => {}
        }
    }

    Ok(())
}

Proveedores Compatibles

ProveedorModeloTransporteFeature FlagFormato de Audio
OpenAIgpt-realtimeWebSocketopenaiPCM16 24kHz
OpenAIgpt-realtimeWebRTCopenai-webrtcOpus
Googlegemini-live-2.5-flash-native-audioWebSocketgeminiPCM16 16kHz/24kHz
GoogleGemini vía Vertex AIWebSocket + OAuth2vertex-livePCM16 16kHz/24kHz
LiveKitCualquiera (puente a Gemini/OpenAI)WebRTClivekitPCM16

Nota: gpt-realtime es el último modelo en tiempo real de OpenAI con calidad de voz, emoción y capacidades de llamada a funciones mejoradas.

Opciones de Transporte

ADK-Realtime admite múltiples capas de transporte:

  • WebSocket (predeterminado): Conexión directa a OpenAI o Gemini. Sencillo, de baja latencia, funciona en todas partes.
  • Vertex AI Live: Se conecta a Gemini a través de Google Cloud con autenticación OAuth2 (Credenciales predeterminadas de la aplicación). Úselo cuando necesite autenticación empresarial e integración con GCP.
  • LiveKit WebRTC: Puente WebRTC de grado de producción. Dirige el audio a través de un servidor LiveKit para escenarios escalables y multi-participantes.
  • OpenAI WebRTC: Conexión WebRTC directa a OpenAI con códec Opus y canales de datos. Requiere cmake para construir la biblioteca Opus C.

RealtimeAgent Builder

El RealtimeAgentBuilder proporciona una API fluida para configurar agentes:

let agent = RealtimeAgent::builder("assistant")
    // Required
    .model(model)

    // Instructions (same as LlmAgent)
    .instruction("You are helpful.")
    .instruction_provider(|ctx| format!("User: {}", ctx.user_name()))

    // Voice settings
    .voice("alloy")  // Options: alloy, coral, sage, shimmer, etc.

    // Voice Activity Detection
    .server_vad()  // Use defaults
    .vad(VadConfig {
        mode: VadMode::ServerVad,
        threshold: Some(0.5),
        prefix_padding_ms: Some(300),
        silence_duration_ms: Some(500),
        interrupt_response: Some(true),
        eagerness: None,
    })

    // Tools (same as LlmAgent)
    .tool(Arc::new(weather_tool))
    .tool(Arc::new(search_tool))

    // Sub-agents for handoffs
    .sub_agent(booking_agent)
    .sub_agent(support_agent)

    // Callbacks (same as LlmAgent)
    .before_agent_callback(|ctx| async { Ok(()) })
    .after_agent_callback(|ctx, event| async { Ok(()) })
    .before_tool_callback(|ctx, tool, args| async { Ok(None) })
    .after_tool_callback(|ctx, tool, result| async { Ok(result) })

    // Realtime-specific callbacks
    .on_audio(|audio_chunk| { /* play audio */ })
    .on_transcript(|text| { /* show transcript */ })

    .build()?;

Detección de Actividad de Voz (VAD)

VAD permite un flujo de conversación natural al detectar cuándo el usuario comienza y deja de hablar.

let agent = RealtimeAgent::builder("assistant")
    .model(model)
    .server_vad()  // Uses sensible defaults
    .build()?;

Configuración Personalizada de VAD

use adk_realtime::{VadConfig, VadMode};

let vad = VadConfig {
    mode: VadMode::ServerVad,
    threshold: Some(0.5),           // Speech detection sensitivity (0.0-1.0)
    prefix_padding_ms: Some(300),   // Audio to include before speech
    silence_duration_ms: Some(500), // Silence before ending turn
    interrupt_response: Some(true), // Allow interrupting assistant
    eagerness: None,                // For SemanticVad mode
};

let agent = RealtimeAgent::builder("assistant")
    .model(model)
    .vad(vad)
    .build()?;

VAD Semántico (Gemini)

Para los modelos Gemini, puede usar VAD semántico que considera el significado:

let vad = VadConfig {
    mode: VadMode::SemanticVad,
    eagerness: Some("high".to_string()),  // low, medium, high
    ..Default::default()
};

Llamada a Herramientas

Los agentes en tiempo real admiten la llamada a herramientas durante las conversaciones de voz:

use adk_realtime::{config::ToolDefinition, ToolResponse};
use serde_json::json;

// Define tools
let tools = vec![
    ToolDefinition {
        name: "get_weather".to_string(),
        description: Some("Get weather for a location".to_string()),
        parameters: Some(json!({
            "type": "object",
            "properties": {
                "location": { "type": "string" }
            },
            "required": ["location"]
        })),
    },
];

let config = RealtimeConfig::default()
    .with_tools(tools)
    .with_instruction("Use tools to help the user.");

let session = model.connect(config).await?;

// Handle tool calls in the event loop
while let Some(event) = session.next_event().await {
    match event? {
        ServerEvent::FunctionCallDone { call_id, name, arguments, .. } => {
            // Execute the tool
            let result = execute_tool(&name, &arguments);

            // Send the response
            let response = ToolResponse::new(&call_id, result);
            session.send_tool_response(response).await?;
        }
        _ => {}
    }
}

Transferencias entre Múltiples Agentes

Transfiera conversaciones entre agentes especializados:

// Create sub-agents
let booking_agent = Arc::new(RealtimeAgent::builder("booking_agent")
    .model(model.clone())
    .instruction("Help with reservations.")
    .build()?);

let support_agent = Arc::new(RealtimeAgent::builder("support_agent")
    .model(model.clone())
    .instruction("Help with technical issues.")
    .build()?);

// Create main agent with sub-agents
let receptionist = RealtimeAgent::builder("receptionist")
    .model(model)
    .instruction(
        "Route customers: bookings → booking_agent, issues → support_agent. \
         Use transfer_to_agent tool to hand off."
    )
    .sub_agent(booking_agent)
    .sub_agent(support_agent)
    .build()?;

Cuando el modelo llama a transfer_to_agent, el RealtimeRunner gestiona la transferencia automáticamente.

Formatos de Audio

FormatoFrecuencia de MuestreoBitsCanalesCaso de Uso
PCM1624000 Hz16MonoOpenAI (predeterminado)
PCM1616000 Hz16MonoEntrada de Gemini
G711 u-law8000 Hz8MonoTelefonía
G711 A-law8000 Hz8MonoTelefonía
use adk_realtime::{AudioFormat, AudioChunk};

// Create audio format
let format = AudioFormat::pcm16_24khz();

// Work with audio chunks
let chunk = AudioChunk::new(audio_bytes, format);
let base64 = chunk.to_base64();
let decoded = AudioChunk::from_base64(&base64, format)?;

Tipos de Eventos

Eventos del Servidor

EventoDescripción
SessionCreatedConexión establecida
AudioDeltaFragmento de audio (PCM base64)
TextDeltaFragmento de respuesta de texto
TranscriptDeltaTranscripción de audio de entrada
FunctionCallDoneSolicitud de llamada a Tool
ResponseDoneRespuesta completada
SpeechStartedVAD detectó inicio de habla
SpeechStoppedVAD detectó fin de habla
ErrorOcurrió un error

Eventos del Cliente

EventoDescripción
AudioInputEnviar fragmento de audio
AudioCommitConfirmar búfer de audio
ItemCreateEnviar respuesta de texto o de Tool
CreateResponseSolicitar una respuesta
CancelResponseCancelar respuesta actual
SessionUpdateActualizar configuración

Vertex AI Live (Google Cloud)

Conéctese a Gemini Live a través de Vertex AI con autenticación empresarial (ADC, cuentas de servicio, WIF):

use adk_realtime::gemini::{GeminiLiveBackend, GeminiRealtimeModel};
use adk_realtime::{RealtimeConfig, RealtimeModel};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let project_id = std::env::var("GOOGLE_CLOUD_PROJECT")?;
    let region = std::env::var("GOOGLE_CLOUD_REGION")
        .unwrap_or_else(|_| "us-central1".to_string());

    // Use Application Default Credentials
    let credentials = google_cloud_auth::credentials::Builder::default()
        .build()
        .await?;

    let backend = GeminiLiveBackend::Vertex { credentials, region, project_id };
    let model = GeminiRealtimeModel::new(backend, "models/gemini-live-2.5-flash-native-audio");

    let config = RealtimeConfig::default()
        .with_instruction("You are a helpful voice assistant.");

    let session = model.connect(config).await?;
    session.send_text("Hello from Vertex AI!").await?;
    session.create_response().await?;

    // Process events...
    Ok(())
}

También hay un constructor de conveniencia para ADC:

let model = GeminiRealtimeModel::vertex_adc(
    "us-central1",
    "my-project-id",
    "models/gemini-live-2.5-flash-native-audio",
).await?;

Vertex AI Live con llamada a Tool

El ejemplo vertex_live_tools demuestra la llamada de función a través de una sesión de Vertex AI Live:

use adk_realtime::config::ToolDefinition;
use adk_realtime::events::ToolResponse;
use serde_json::json;

// Declare tools
let tools = vec![
    ToolDefinition {
        name: "get_weather".to_string(),
        description: Some("Get current weather for a city".to_string()),
        parameters: Some(json!({
            "type": "object",
            "properties": {
                "city": { "type": "string" }
            },
            "required": ["city"]
        })),
    },
];

let config = RealtimeConfig::default()
    .with_tools(tools)
    .with_instruction("Use tools to answer questions about weather.");

let session = model.connect(config).await?;

// Handle FunctionCallDone events and send ToolResponse back
while let Some(event) = session.next_event().await {
    match event? {
        ServerEvent::FunctionCallDone { call_id, name, arguments, .. } => {
            let result = match name.as_str() {
                "get_weather" => json!({"temperature": "22°C", "condition": "sunny"}),
                _ => json!({"error": "unknown tool"}),
            };
            session.send_tool_response(ToolResponse::new(&call_id, result)).await?;
        }
        ServerEvent::TextDelta { delta, .. } => print!("{delta}"),
        ServerEvent::ResponseDone { .. } => break,
        _ => {}
    }
}

Banderas de Características

CaracterísticaDependenciasCaso de Uso
vertex-livegemini + google-cloud-authVertex AI Live con autenticación ADC/cuenta de servicio
livekitlivekit + livekit-apiPuente LiveKit WebRTC
openai-webrtcopenai + str0m + audiopusOpenAI WebRTC con Opus (requiere cmake)
fullopenai + gemini + vertex-live + livekitTodos los transportes excepto WebRTC
full-webrtcfull + openai-webrtcTodo (requiere cmake)

Puente LiveKit WebRTC

Para aplicaciones de voz en producción, el puente LiveKit enruta el audio a través de un servidor LiveKit para escenarios escalables y multi-participante.

LiveKitConfig

Configure de forma segura las credenciales de LiveKit. Las claves API y los secretos se almacenan usando secrecy::SecretString y se redactan en la salida de depuración:

use adk_realtime::livekit::{LiveKitConfig, LiveKitRoomBuilder};

let config = LiveKitConfig::new(
    "wss://your-server.livekit.cloud",
    std::env::var("LIVEKIT_API_KEY")?,
    std::env::var("LIVEKIT_API_SECRET")?,
)?;

LiveKitConfig::new() valida el formato de la URL y rechaza las credenciales vacías en el momento de la construcción.

LiveKitRoomBuilder

Un constructor de tipado de estados para conectar a salas LiveKit. El campo identity es requerido en tiempo de compilación — connect() solo está disponible después de que se establece:

let bundle = LiveKitRoomBuilder::new(config)
    .identity("my-agent")           // required — enables connect()
    .name("Voice Agent")            // optional display name
    .room_name("session-room-123")  // optional — auto-generated if omitted
    .auto_subscribe(true)           // subscribe to remote tracks
    .with_audio(24_000, 1)          // publish a local audio track (sample rate, channels)
    .connect()
    .await?;

// The bundle contains everything you need
let room = bundle.room;
let mut events = bundle.events;
let audio_source = bundle.audio_source;  // for publishing audio
let audio_track = bundle.audio_track;

Conectando Audio

Utilice las utilidades del puente para conectar audio de LiveKit a un RealtimeRunner:

use adk_realtime::livekit::{LiveKitEventHandler, bridge_input};

// Wrap your event handler to publish model audio to LiveKit
let lk_handler = LiveKitEventHandler::new(inner_handler, audio_source, 24000, 1);

// Bridge participant audio from LiveKit into the RealtimeRunner
tokio::spawn(bridge_input(remote_track, runner));

Ejemplos

Ejecute los ejemplos incluidos:

# OpenAI Realtime (WebSocket)
cargo run -p adk-realtime --example openai_session_update --features openai

# Vertex AI Live (requires gcloud auth application-default login)
cargo run -p adk-realtime --example vertex_live_voice --features vertex-live
cargo run -p adk-realtime --example vertex_live_tools --features vertex-live

# LiveKit Bridge (requires LiveKit server)
cargo run -p adk-realtime --example livekit_bridge --features livekit,openai
cargo run -p adk-realtime --example livekit_gemini_bridge --features livekit,gemini

# Debug utilities
cargo run -p adk-realtime --example debug_gemini --features gemini
cargo run -p adk-realtime --example debug_livekit_auth --features livekit

# OpenAI WebRTC (requires cmake)
cargo run -p adk-realtime --example openai_webrtc --features openai-webrtc

Mejores Prácticas

  1. Use VAD del Servidor: Permita que el servidor maneje la detección de voz para una menor latencia
  2. Manejar interrupciones: Habilite interrupt_response para conversaciones naturales
  3. Mantenga las instrucciones concisas: Las respuestas de voz deben ser breves
  4. Pruebe con texto primero: Depure la lógica de su Agent con texto antes de añadir audio
  5. Maneje los errores con elegancia: Los problemas de red son comunes con las conexiones WebSocket

Comparación con OpenAI Agents SDK

La implementación en tiempo real de ADK-Rust sigue el patrón de OpenAI Agents SDK:

CaracterísticaOpenAI SDKADK-Rust
Clase base de AgentAgentAgent trait
Agent en tiempo realRealtimeAgentRealtimeAgent
ToolsDefiniciones de funcionesTool trait + ToolDefinition
Transferenciastransfer_to_agentsub_agents + Tool auto-generada
RetrollamadasGanchosbefore_* / after_* retrollamadas

Anterior: ← Graph Agents | Siguiente: Proveedores de Modelos →

Agentes de Voz en Tiempo Real - Documentación ADK-Rust | ADK-Rust