VoiceEngine

Shared voice infrastructure for Elixir — speech-to-text / text-to-speech behaviours, pluggable adapters, and a session-scoped VoiceBridge orchestrator that wires STT → turn → TTS over Phoenix channels.

VoiceEngine owns no business logic. Your host application supplies the conversation turn_fn, a Phoenix Endpoint/PubSub, and a voice profile. The library handles capture, transcription, synthesis, and broadcasting.

Adapters

AdapterSTTTTSNotes
VoiceEngine.ElevenLabsCloud. Requires an ElevenLabs API key.
VoiceEngine.SparkSelf-hosted. Targets a Whisper STT + Orpheus TTS box you run yourself.

Adapters implement the VoiceEngine.STT and VoiceEngine.TTS behaviours, so you can drop in your own.

Installation

def deps do
[
{:voice_engine, "~> 0.1"}
]
end

VoiceEngine starts its own supervision tree (a Registry + a DynamicSupervisor for bridges) when your app boots — no extra wiring required.

Configuration

Each host app configures voice profiles under its own app env. Pass the OTP app name when reading config:

# config/runtime.exs
config :my_app, VoiceEngine.Config,
current_profile: :eleven_labs,
profiles: %{
eleven_labs: [
stt: VoiceEngine.ElevenLabs.STT,
tts: VoiceEngine.ElevenLabs.TTS,
common: [api_key: System.get_env("ELEVENLABS_API_KEY")],
stt_opts: [model: "scribe_v2", language: "en", content_type: "audio/webm"],
tts_opts: [
voice_id: "your-elevenlabs-voice-id",
model: "eleven_multilingual_v2",
format: :mp3,
sample_rate: 44_100
]
],
spark: [
stt: VoiceEngine.Spark.STT,
tts: VoiceEngine.Spark.TTS,
# Self-hosted services — point these at your own boxes:
stt_opts: [stt_base_url: "http://localhost:8100"],
tts_opts: [tts_base_url: "http://localhost:8101", voice: "tara"]
]
}

Read the active profile:

VoiceEngine.Config.current(:my_app)
VoiceEngine.Config.profile(:my_app, :spark)

The Spark adapter has no built-in base URLstt_base_url / tts_base_url are required. Point them at the Whisper/Orpheus services you host yourself.

The Bridge

VoiceEngine.Bridge is a per-session GenServer that runs the full loop: inbound audio → STT → your turn_fn → TTS → outbound audio, broadcasting status and audio chunks over a Phoenix channel.

VoiceEngine.BridgeSupervisor.ensure_bridge(
session_id: session.id,
config: VoiceEngine.Config.current(:my_app),
turn_fn: &MyApp.Conversation.turn(session.id, &1),
endpoint: MyAppWeb.Endpoint,
pubsub: MyApp.PubSub
)

Drive it from a LiveView or channel:

VoiceEngine.Bridge.voice_start(session_id)
VoiceEngine.Bridge.audio_chunk(session_id, chunk)
VoiceEngine.Bridge.voice_stop(session_id)

Incoming audio is persisted to disk before transcription, so a recording is never lost if STT fails or times out — it can be retried.

License

MIT