リアルタイム音声エージェント

リアルタイムエージェントは、双方向オーディオストリーミングを使用して、AIアシスタントとの音声ベースのインタラクションを可能にします。adk-realtime crateは、OpenAIのRealtime APIおよびGoogleのGemini Live APIと連携して動作する音声対応エージェントを構築するための統一されたインターフェースを提供します。

概要

リアルタイムエージェントは、テキストベースのLlmAgentsとはいくつかの重要な点で異なります。

特徴LlmAgentRealtimeAgent
入力テキスト音声/テキスト
出力テキスト音声/テキスト
接続HTTP リクエストWebSocket
レイテンシリクエスト/レスポンスリアルタイムストリーミング
VAD該当なしサーバーサイド音声検出

アーキテクチャ

              ┌─────────────────────────────────────────┐
              │              Agent Trait                │
              │  (name, description, run, sub_agents)   │
              └────────────────┬────────────────────────┘
                               │
       ┌───────────────────────┼───────────────────────┐
       │                       │                       │
┌──────▼──────┐      ┌─────────▼─────────┐   ┌─────────▼─────────┐
│  LlmAgent   │      │  RealtimeAgent    │   │  SequentialAgent  │
│ (text-based)│      │  (voice-based)    │   │   (workflow)      │
└─────────────┘      └───────────────────┘   └───────────────────┘

RealtimeAgent は同じ Agent トレイトを LlmAgent として実装しており、以下を共有します。

  • 指示 (静的および動的)
  • ツールの登録と実行
  • コールバック (before_agent, after_agent, before_tool, after_tool)
  • サブエージェントへの引き渡し

クイックスタート

インストール

あなたの Cargo.toml に追加してください:

[dependencies]
adk-realtime = { version = "2.0.0", features = ["openai"] }

# For Vertex AI Live (Google Cloud with ADC auth)
# adk-realtime = { version = "2.0.0", features = ["vertex-live"] }

# For LiveKit WebRTC bridge
# adk-realtime = { version = "2.0.0", features = ["livekit"] }

# For all transports (except WebRTC which needs cmake)
# adk-realtime = { version = "2.0.0", features = ["full"] }

基本的な使い方

use adk_realtime::{
    RealtimeAgent, RealtimeModel, RealtimeConfig, ServerEvent,
    openai::OpenAIRealtimeModel,
};
use std::sync::Arc;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("OPENAI_API_KEY")?;

    // Create the realtime model
    let model: Arc<dyn RealtimeModel> = Arc::new(
        OpenAIRealtimeModel::new(&api_key, "gpt-realtime")
    );

    // Build the realtime agent
    let agent = RealtimeAgent::builder("voice_assistant")
        .model(model.clone())
        .instruction("You are a helpful voice assistant. Be concise.")
        .voice("alloy")
        .server_vad()  // Enable voice activity detection
        .build()?;

    // Or use the low-level session API directly
    let config = RealtimeConfig::default()
        .with_instruction("You are a helpful assistant.")
        .with_voice("alloy")
        .with_modalities(vec!["text".to_string(), "audio".to_string()]);

    let session = model.connect(config).await?;

    // Send text and get response
    session.send_text("Hello!").await?;
    session.create_response().await?;

    // Process events
    while let Some(event) = session.next_event().await {
        match event? {
            ServerEvent::TextDelta { delta, .. } => print!("{}", delta),
            ServerEvent::AudioDelta { delta, .. } => {
                // Play audio (delta is base64-encoded PCM)
            }
            ServerEvent::ResponseDone { .. } => break,
            _ => {}
        }
    }

    Ok(())
}

サポートされているプロバイダー

プロバイダーモデルトランスポート機能フラグオーディオ形式
OpenAIgpt-realtimeWebSocketopenaiPCM16 24kHz
OpenAIgpt-realtimeWebRTCopenai-webrtcOpus
Googlegemini-live-2.5-flash-native-audioWebSocketgeminiPCM16 16kHz/24kHz
GoogleGemini Vertex AI 経由WebSocket + OAuth2vertex-livePCM16 16kHz/24kHz
LiveKit任意 (Gemini/OpenAI へのブリッジ)WebRTClivekitPCM16

注釈: gpt-realtime は、音声品質、感情、および関数呼び出し機能が向上した OpenAI の最新のリアルタイムモデルです。

転送オプション

ADK-Realtime は複数の転送レイヤーをサポートしています。

  • WebSocket (デフォルト): OpenAI または Gemini への直接接続。シンプルで低遅延、どこでも動作します。
  • Vertex AI Live: OAuth2 認証 (Application Default Credentials) を使用して Google Cloud 経由で Gemini に接続します。エンタープライズ認証と GCP 統合が必要な場合に使用します。
  • LiveKit WebRTC: プロダクショングレードの WebRTC ブリッジ。スケーラブルな多人数参加型シナリオのために、音声を LiveKit サーバー経由でルーティングします。
  • OpenAI WebRTC: Opus コーデックとデータチャネルを備えた OpenAI への直接 WebRTC 接続。Opus C ライブラリをビルドするには cmake が必要です。

RealtimeAgent ビルダー

RealtimeAgentBuilder は、エージェントを構成するための流暢な API を提供します。

let agent = RealtimeAgent::builder("assistant")
    // Required
    .model(model)

    // Instructions (same as LlmAgent)
    .instruction("You are helpful.")
    .instruction_provider(|ctx| format!("User: {}", ctx.user_name()))

    // Voice settings
    .voice("alloy")  // Options: alloy, coral, sage, shimmer, etc.

    // Voice Activity Detection
    .server_vad()  // Use defaults
    .vad(VadConfig {
        mode: VadMode::ServerVad,
        threshold: Some(0.5),
        prefix_padding_ms: Some(300),
        silence_duration_ms: Some(500),
        interrupt_response: Some(true),
        eagerness: None,
    })

    // Tools (same as LlmAgent)
    .tool(Arc::new(weather_tool))
    .tool(Arc::new(search_tool))

    // Sub-agents for handoffs
    .sub_agent(booking_agent)
    .sub_agent(support_agent)

    // Callbacks (same as LlmAgent)
    .before_agent_callback(|ctx| async { Ok(()) })
    .after_agent_callback(|ctx, event| async { Ok(()) })
    .before_tool_callback(|ctx, tool, args| async { Ok(None) })
    .after_tool_callback(|ctx, tool, result| async { Ok(result) })

    // Realtime-specific callbacks
    .on_audio(|audio_chunk| { /* play audio */ })
    .on_transcript(|text| { /* show transcript */ })

    .build()?;

音声活動検出 (VAD)

VAD は、ユーザーが話し始めたり話し終えたりするのを検出することで、自然な会話の流れを可能にします。

let agent = RealtimeAgent::builder("assistant")
    .model(model)
    .server_vad()  // Uses sensible defaults
    .build()?;

カスタム VAD 設定

use adk_realtime::{VadConfig, VadMode};

let vad = VadConfig {
    mode: VadMode::ServerVad,
    threshold: Some(0.5),           // Speech detection sensitivity (0.0-1.0)
    prefix_padding_ms: Some(300),   // Audio to include before speech
    silence_duration_ms: Some(500), // Silence before ending turn
    interrupt_response: Some(true), // Allow interrupting assistant
    eagerness: None,                // For SemanticVad mode
};

let agent = RealtimeAgent::builder("assistant")
    .model(model)
    .vad(vad)
    .build()?;

セマンティック VAD (Gemini)

Gemini モデルの場合、意味を考慮するセマンティック VAD を使用できます。

let vad = VadConfig {
    mode: VadMode::SemanticVad,
    eagerness: Some("high".to_string()),  // low, medium, high
    ..Default::default()
};

ツール呼び出し

リアルタイムエージェントは、音声会話中にツール呼び出しをサポートします。

use adk_realtime::{config::ToolDefinition, ToolResponse};
use serde_json::json;

// Define tools
let tools = vec![
    ToolDefinition {
        name: "get_weather".to_string(),
        description: Some("Get weather for a location".to_string()),
        parameters: Some(json!({
            "type": "object",
            "properties": {
                "location": { "type": "string" }
            },
            "required": ["location"]
        })),
    },
];

let config = RealtimeConfig::default()
    .with_tools(tools)
    .with_instruction("Use tools to help the user.");

let session = model.connect(config).await?;

// Handle tool calls in the event loop
while let Some(event) = session.next_event().await {
    match event? {
        ServerEvent::FunctionCallDone { call_id, name, arguments, .. } => {
            // Execute the tool
            let result = execute_tool(&name, &arguments);

            // Send the response
            let response = ToolResponse::new(&call_id, result);
            session.send_tool_response(response).await?;
        }
        _ => {}
    }
}

マルチエージェントハンドオフ

特殊なエージェント間で会話を転送します。

// Create sub-agents
let booking_agent = Arc::new(RealtimeAgent::builder("booking_agent")
    .model(model.clone())
    .instruction("Help with reservations.")
    .build()?);

let support_agent = Arc::new(RealtimeAgent::builder("support_agent")
    .model(model.clone())
    .instruction("Help with technical issues.")
    .build()?);

// Create main agent with sub-agents
let receptionist = RealtimeAgent::builder("receptionist")
    .model(model)
    .instruction(
        "Route customers: bookings → booking_agent, issues → support_agent. \
         Use transfer_to_agent tool to hand off."
    )
    .sub_agent(booking_agent)
    .sub_agent(support_agent)
    .build()?;

モデルが transfer_to_agent を呼び出すと、RealtimeRunner がハンドオフを自動的に処理します。

オーディオ形式

フォーマットサンプルレートビットチャンネルユースケース
PCM1624000 Hz16モノラルOpenAI (デフォルト)
PCM1616000 Hz16モノラルGemini 入力
G711 u-law8000 Hz8モノラルテレフォニー
G711 A-law8000 Hz8モノラルテレフォニー
use adk_realtime::{AudioFormat, AudioChunk};

// Create audio format
let format = AudioFormat::pcm16_24khz();

// Work with audio chunks
let chunk = AudioChunk::new(audio_bytes, format);
let base64 = chunk.to_base64();
let decoded = AudioChunk::from_base64(&base64, format)?;

イベントの種類

サーバーイベント

イベント説明
SessionCreated接続確立
AudioDeltaオーディオチャンク (base64 PCM)
TextDeltaテキスト応答チャンク
TranscriptDelta入力オーディオトランスクリプト
FunctionCallDoneツール呼び出しリクエスト
ResponseDone応答完了
SpeechStartedVADが音声開始を検出
SpeechStoppedVADが音声終了を検出
Errorエラーが発生しました

クライアントイベント

イベント説明
AudioInputオーディオチャンクを送信
AudioCommitオーディオバッファをコミット
ItemCreateテキストまたはツール応答を送信
CreateResponse応答をリクエスト
CancelResponse現在の応答をキャンセル
SessionUpdate設定を更新

Vertex AI ライブ (Google Cloud)

エンタープライプ認証 (ADC、サービスアカウント、WIF) を使用して、Vertex AI 経由で Gemini Live に接続します:

use adk_realtime::gemini::{GeminiLiveBackend, GeminiRealtimeModel};
use adk_realtime::{RealtimeConfig, RealtimeModel};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let project_id = std::env::var("GOOGLE_CLOUD_PROJECT")?;
    let region = std::env::var("GOOGLE_CLOUD_REGION")
        .unwrap_or_else(|_| "us-central1".to_string());

    // Use Application Default Credentials
    let credentials = google_cloud_auth::credentials::Builder::default()
        .build()
        .await?;

    let backend = GeminiLiveBackend::Vertex { credentials, region, project_id };
    let model = GeminiRealtimeModel::new(backend, "models/gemini-live-2.5-flash-native-audio");

    let config = RealtimeConfig::default()
        .with_instruction("You are a helpful voice assistant.");

    let session = model.connect(config).await?;
    session.send_text("Hello from Vertex AI!").await?;
    session.create_response().await?;

    // Process events...
    Ok(())
}

ADC 用の便利なコンストラクタもあります:

let model = GeminiRealtimeModel::vertex_adc(
    "us-central1",
    "my-project-id",
    "models/gemini-live-2.5-flash-native-audio",
).await?;

Vertex AI ライブとツール呼び出し

vertex_live_tools の例は、Vertex AI ライブセッションでの関数呼び出しを示しています:

use adk_realtime::config::ToolDefinition;
use adk_realtime::events::ToolResponse;
use serde_json::json;

// Declare tools
let tools = vec![
    ToolDefinition {
        name: "get_weather".to_string(),
        description: Some("Get current weather for a city".to_string()),
        parameters: Some(json!({
            "type": "object",
            "properties": {
                "city": { "type": "string" }
            },
            "required": ["city"]
        })),
    },
];

let config = RealtimeConfig::default()
    .with_tools(tools)
    .with_instruction("Use tools to answer questions about weather.");

let session = model.connect(config).await?;

// Handle FunctionCallDone events and send ToolResponse back
while let Some(event) = session.next_event().await {
    match event? {
        ServerEvent::FunctionCallDone { call_id, name, arguments, .. } => {
            let result = match name.as_str() {
                "get_weather" => json!({"temperature": "22°C", "condition": "sunny"}),
                _ => json!({"error": "unknown tool"}),
            };
            session.send_tool_response(ToolResponse::new(&call_id, result)).await?;
        }
        ServerEvent::TextDelta { delta, .. } => print!("{delta}"),
        ServerEvent::ResponseDone { .. } => break,
        _ => {}
    }
}

機能フラグ

機能依存関係ユースケース
vertex-livegemini + google-cloud-authVertex AI Live と ADC/サービスアカウント認証
livekitlivekit + livekit-apiLiveKit WebRTC ブリッジ
openai-webrtcopenai + str0m + audiopusOpenAI WebRTC と Opus (cmake が必要)
fullopenai + gemini + vertex-live + livekitWebRTC を除くすべてのトランスポート
full-webrtcfull + openai-webrtcすべて (cmake が必要)

LiveKit WebRTC Bridge

運用環境の音声アプリケーション向けに、LiveKit bridge はスケーラブルな多人数参加型シナリオのためにLiveKit サーバーを介して音声をルーティングします。

LiveKitConfig

LiveKit 認証情報を安全に設定します。API のキーとシークレットはsecrecy::SecretString を使用して保存され、デバッグ出力では編集されます。

use adk_realtime::livekit::{LiveKitConfig, LiveKitRoomBuilder};

let config = LiveKitConfig::new(
    "wss://your-server.livekit.cloud",
    std::env::var("LIVEKIT_API_KEY")?,
    std::env::var("LIVEKIT_API_SECRET")?,
)?;

LiveKitConfig::new() はURL 形式を検証し、構築時に空の認証情報を拒否します。

LiveKitRoomBuilder

LiveKit ルームに接続するための型状態ビルダー。identity フィールドはコンパイル時に必須であり、connect() は設定後にのみ利用可能です。

let bundle = LiveKitRoomBuilder::new(config)
    .identity("my-agent")           // required — enables connect()
    .name("Voice Agent")            // optional display name
    .room_name("session-room-123")  // optional — auto-generated if omitted
    .auto_subscribe(true)           // subscribe to remote tracks
    .with_audio(24_000, 1)          // publish a local audio track (sample rate, channels)
    .connect()
    .await?;

// The bundle contains everything you need
let room = bundle.room;
let mut events = bundle.events;
let audio_source = bundle.audio_source;  // for publishing audio
let audio_track = bundle.audio_track;

音声のブリッジング

bridge ユーティリティを使用して、LiveKit の音声をRealtimeRunner に接続します。

use adk_realtime::livekit::{LiveKitEventHandler, bridge_input};

// Wrap your event handler to publish model audio to LiveKit
let lk_handler = LiveKitEventHandler::new(inner_handler, audio_source, 24000, 1);

// Bridge participant audio from LiveKit into the RealtimeRunner
tokio::spawn(bridge_input(remote_track, runner));

含まれている例を実行します。

# OpenAI Realtime (WebSocket)
cargo run -p adk-realtime --example openai_session_update --features openai

# Vertex AI Live (requires gcloud auth application-default login)
cargo run -p adk-realtime --example vertex_live_voice --features vertex-live
cargo run -p adk-realtime --example vertex_live_tools --features vertex-live

# LiveKit Bridge (requires LiveKit server)
cargo run -p adk-realtime --example livekit_bridge --features livekit,openai
cargo run -p adk-realtime --example livekit_gemini_bridge --features livekit,gemini

# Debug utilities
cargo run -p adk-realtime --example debug_gemini --features gemini
cargo run -p adk-realtime --example debug_livekit_auth --features livekit

# OpenAI WebRTC (requires cmake)
cargo run -p adk-realtime --example openai_webrtc --features openai-webrtc

ベストプラクティス

  1. サーバーVADを使用する: 低レイテンシのためにサーバーに音声検出を処理させる
  2. 割り込みを処理する: 自然な会話のためにinterrupt_response を有効にする
  3. 指示を簡潔にする: 音声応答は簡潔であるべきです
  4. まずテキストでテストする: 音声を追加する前にテキストでエージェントロジックをデバッグする
  5. エラーを適切に処理する: WebSocket 接続ではネットワークの問題がよく発生します

OpenAI Agents SDK との比較

ADK-Rust のリアルタイム実装は、OpenAI Agents SDK パターンに従います。

機能OpenAI SDKADK-Rust
Agent 基底クラスAgentAgent trait
リアルタイム AgentRealtimeAgentRealtimeAgent
ツール関数定義Tool trait + ToolDefinition
ハンドオフtransfer_to_agentsub_agents + 自動生成された tool
コールバックフックbefore_* / after_* callbacks

前へ: ← Graph Agents | 次へ: モデルプロバイダー →