Release Notes: v0.7.0¶
Release Date: 2026-07-05
Summary¶
Adds Voice Agent realtime provider implementing corereal.Provider interface for native voice-to-voice conversations with ~100-300ms latency using Deepgram's Voice Agent API.
Highlights¶
- Voice Agent Realtime Provider: Native voice-to-voice conversations without STT → LLM → TTS pipeline overhead
- ~100-300ms Latency: Direct audio streaming to Deepgram Voice Agent API
- Function Calling Support: Tool use during voice conversations
- CLI Testing Tool:
deepgram-voice-testfor debugging and testing
Added¶
Voice Agent Realtime Provider¶
New omnivoice/realtime package implementing corereal.Provider:
import (
"github.com/plexusone/omni-deepgram/omnivoice/realtime"
corereal "github.com/plexusone/omnivoice-core/realtime"
)
// Create provider with API key
provider, err := realtime.New(
realtime.WithAPIKey("your-api-key"),
realtime.WithGreeting("Hello! How can I help you?"),
realtime.WithInputSampleRate(48000), // Match your audio source
)
// Start bidirectional audio stream
audioCh, transcriptCh, err := provider.ProcessAudioStream(ctx, audioIn, corereal.ProcessConfig{
Instructions: "You are a helpful voice assistant.",
Voice: "aura-2-thalia-en",
})
// Receive audio and transcripts
for {
select {
case chunk := <-audioCh:
// Play agent audio response
case transcript := <-transcriptCh:
// Log conversation
}
}
Key Components:
| File | Purpose |
|---|---|
provider.go |
Main Provider with ProcessAudioStream implementation |
chan_handler.go |
AgentMessageChan for SDK event handling |
convert.go |
Config mapping to Deepgram SettingsOptions |
factory.go |
Gateway factory for provider creation |
options.go |
Functional options pattern for configuration |
Configuration Options¶
realtime.New(
realtime.WithAPIKey(apiKey),
realtime.WithModel("aura-2-thalia-en"), // Voice model
realtime.WithLanguage("en-US"), // Language
realtime.WithGreeting("Hello!"), // Agent greeting
realtime.WithInputSampleRate(48000), // Input sample rate
realtime.WithOutputSampleRate(24000), // Output sample rate
realtime.WithExperimental(true), // Enable latency metrics
realtime.WithThinkProvider(map[string]any{ // Custom LLM
"type": "anthropic",
"model": "claude-sonnet-4-20250514",
}),
)
Function Calling¶
audioCh, transcriptCh, err := provider.ProcessAudioStream(ctx, audioIn, corereal.ProcessConfig{
Instructions: "You can check the weather.",
Functions: []corereal.FunctionDeclaration{
{
Name: "get_weather",
Description: "Get weather for a location",
Parameters: json.RawMessage(`{"type":"object","properties":{"location":{"type":"string"}}}`),
},
},
OnFunctionCall: func(id, name, args string) (string, error) {
// Handle function call
return `{"temperature": 72}`, nil
},
})
CLI Testing Tool¶
New deepgram-voice-test CLI for debugging:
export DEEPGRAM_API_KEY="your-api-key"
go run ./cmd/deepgram-voice-test \
--greeting "Hello! How can I help?" \
--voice aura-2-thalia-en \
--duration 60s
Fixed¶
- Use SDK types (
interfacesv1.Functions) for function declarations instead ofmap[string]anyfor type safety (f4c813a)
Tests¶
- Unit tests for configuration and options
- Integration tests for Voice Agent provider (requires
DEEPGRAM_API_KEY)
Dependencies¶
| Dependency | Version |
|---|---|
| omnivoice-core | v0.15.0 |
| deepgram-go-sdk | v3.7.0 |
| spf13/cobra | v1.10.2 |
Installation¶
Migration Guide¶
From v0.6.1¶
No breaking changes. The new realtime provider is additive.
To use the realtime provider with existing applications:
- Import the realtime package
- Create a provider with
realtime.New() - Call
ProcessAudioStream()with your audio input channel
Audio Sample Rates¶
The realtime provider defaults to:
- Input: 48kHz (matches LiveKit and most WebRTC sources)
- Output: 24kHz (Deepgram TTS output)
If your audio source uses a different sample rate, configure it explicitly: