Skip to content

Release Notes: v0.7.0

Release Date: 2026-07-05

Summary

Adds Voice Agent realtime provider implementing corereal.Provider interface for native voice-to-voice conversations with ~100-300ms latency using Deepgram's Voice Agent API.

Highlights

  • Voice Agent Realtime Provider: Native voice-to-voice conversations without STT → LLM → TTS pipeline overhead
  • ~100-300ms Latency: Direct audio streaming to Deepgram Voice Agent API
  • Function Calling Support: Tool use during voice conversations
  • CLI Testing Tool: deepgram-voice-test for debugging and testing

Added

Voice Agent Realtime Provider

New omnivoice/realtime package implementing corereal.Provider:

import (
    "github.com/plexusone/omni-deepgram/omnivoice/realtime"
    corereal "github.com/plexusone/omnivoice-core/realtime"
)

// Create provider with API key
provider, err := realtime.New(
    realtime.WithAPIKey("your-api-key"),
    realtime.WithGreeting("Hello! How can I help you?"),
    realtime.WithInputSampleRate(48000),  // Match your audio source
)

// Start bidirectional audio stream
audioCh, transcriptCh, err := provider.ProcessAudioStream(ctx, audioIn, corereal.ProcessConfig{
    Instructions: "You are a helpful voice assistant.",
    Voice:        "aura-2-thalia-en",
})

// Receive audio and transcripts
for {
    select {
    case chunk := <-audioCh:
        // Play agent audio response
    case transcript := <-transcriptCh:
        // Log conversation
    }
}

Key Components:

File Purpose
provider.go Main Provider with ProcessAudioStream implementation
chan_handler.go AgentMessageChan for SDK event handling
convert.go Config mapping to Deepgram SettingsOptions
factory.go Gateway factory for provider creation
options.go Functional options pattern for configuration

Configuration Options

realtime.New(
    realtime.WithAPIKey(apiKey),
    realtime.WithModel("aura-2-thalia-en"),      // Voice model
    realtime.WithLanguage("en-US"),               // Language
    realtime.WithGreeting("Hello!"),              // Agent greeting
    realtime.WithInputSampleRate(48000),          // Input sample rate
    realtime.WithOutputSampleRate(24000),         // Output sample rate
    realtime.WithExperimental(true),              // Enable latency metrics
    realtime.WithThinkProvider(map[string]any{    // Custom LLM
        "type": "anthropic",
        "model": "claude-sonnet-4-20250514",
    }),
)

Function Calling

audioCh, transcriptCh, err := provider.ProcessAudioStream(ctx, audioIn, corereal.ProcessConfig{
    Instructions: "You can check the weather.",
    Functions: []corereal.FunctionDeclaration{
        {
            Name:        "get_weather",
            Description: "Get weather for a location",
            Parameters:  json.RawMessage(`{"type":"object","properties":{"location":{"type":"string"}}}`),
        },
    },
    OnFunctionCall: func(id, name, args string) (string, error) {
        // Handle function call
        return `{"temperature": 72}`, nil
    },
})

CLI Testing Tool

New deepgram-voice-test CLI for debugging:

export DEEPGRAM_API_KEY="your-api-key"
go run ./cmd/deepgram-voice-test \
    --greeting "Hello! How can I help?" \
    --voice aura-2-thalia-en \
    --duration 60s

Fixed

  • Use SDK types (interfacesv1.Functions) for function declarations instead of map[string]any for type safety (f4c813a)

Tests

  • Unit tests for configuration and options
  • Integration tests for Voice Agent provider (requires DEEPGRAM_API_KEY)

Dependencies

Dependency Version
omnivoice-core v0.15.0
deepgram-go-sdk v3.7.0
spf13/cobra v1.10.2

Installation

go get github.com/plexusone/omni-deepgram@v0.7.0

Migration Guide

From v0.6.1

No breaking changes. The new realtime provider is additive.

To use the realtime provider with existing applications:

  1. Import the realtime package
  2. Create a provider with realtime.New()
  3. Call ProcessAudioStream() with your audio input channel

Audio Sample Rates

The realtime provider defaults to:

  • Input: 48kHz (matches LiveKit and most WebRTC sources)
  • Output: 24kHz (Deepgram TTS output)

If your audio source uses a different sample rate, configure it explicitly:

realtime.WithInputSampleRate(16000)  // For 16kHz audio