mob_speech
Speech-to-text for apps built with Mob, as a plugin.
Speech-to-text only. Text-to-speech stays in mob core (Mob.Speech).
Engines are pluggable. The default :platform engine uses Android's
SpeechRecognizer and iOS's SFSpeechRecognizer with an AVAudioEngine
input tap. MobSpeech.Engine.Fake plays a script, so tests and agents can
drive dictation without a microphone. Other packages add engines by
implementing MobSpeech.Engine.
Installation
# mix.exs
{:mob_speech, "~> 0.1"}
# mob.exs
config :mob, :plugins, [:mob_speech]
Apps generated by mix mob.new already trust the first-party signing key in
config :mob, :trusted_plugins. To add the entry by hand:
config :mob, :trusted_plugins, %{
mob_speech: "ed25519:nc56w+1Kx0gIt/4EkHxnMZCKHMzp4+S5kS/HoSzEZkg="
}
Once you activate the plugin, the build merges in:
- Android:
android.permission.RECORD_AUDIO. - iOS:
NSSpeechRecognitionUsageDescription, which you can override in your ownInfo.plist, and theSpeechandAVFoundationframeworks.NSMicrophoneUsageDescriptionmust already be in yourInfo.plist. Apps generated bymix mob.newhave it; core owns the microphone key, so this plugin doesn't declare it. - Permissions: the plugin registers the
:speechcapability forMob.Permissions.request(socket, :speech). On Android that'sRECORD_AUDIO. On iOS it's speech-recognition authorisation plus the microphone record permission, and the result is:grantedonly when both are granted.
On Android 11 and later, if MobSpeech.available?() returns false even
though a recogniser is installed, add the package-visibility query to
AndroidManifest.xml (inside <manifest>):
<queries>
<intent><action android:name="android.speech.RecognitionService" /></intent>
</queries>
Usage
# Ask once (the platform engine needs [:speech]).
socket = Enum.reduce(MobSpeech.permissions(), socket, &Mob.Permissions.request(&2, &1))
# Hold-to-talk: listen on press, stop on release.
socket = MobSpeech.listen(socket, language: "en-US")
socket = MobSpeech.stop(socket) # deliver the final
socket = MobSpeech.cancel(socket) # abort, no final
def handle_info({:speech, :state, state}, socket), do: ... # :listening | :processing | :idle
def handle_info({:speech, :partial, text}, socket), do: ...
def handle_info({:speech, :final, text}, socket), do: ...
def handle_info({:speech, :error, reason}, socket), do: ...
| Function | Returns | Notes |
|---|---|---|
listen(socket, opts \\ []) |
socket | language: (BCP-47, default device locale), prefer_offline: (default false), partial_results: (default true), stop_timeout_ms: (default 2000 or the engine's), engine: (:platform or a module), to: (pid, default self()). Other options go to the engine. Listening again cancels the previous session. |
stop(socket) |
socket | :processing, then a final (or an error), then idle. No-op once the recognition has ended. |
cancel(socket) |
socket | idle, no final. |
available?(engine \\ :platform) |
boolean | false on a host build without the NIF. On iOS it stays false until speech recognition is authorised. |
permissions(engine \\ :platform) |
[atom] |
capabilities to request first. |
A press shorter than ~300 ms has nothing to recognise. Android's recogniser
usually answers it with :client or :no_speech, so hold-to-talk apps should
cancel very short presses themselves. On Android, the platform engine also
takes silence_ms: (default 10 000), so a pause while the button is held
doesn't end the recognition.
These guarantees hold for every engine:
- After a final or an error you get exactly one
{:speech, :state, :idle}, and nothing else for thatlisten. :processingis sent when you callstop/1. It is never sent when the recogniser endpoints by itself, because during hold-to-talk the button is still held.- The recogniser may finish by itself before you call
stop/1. In that case final and idle arrive first, and the laterstop/1does nothing. - If the final is empty, or the recogniser reports
:no_speechafter partials, you get{:speech, :final, last_partial}. - After
stop/1, if the recogniser doesn't answer withinstop_timeout_ms, it is cancelled and you get the last partial (or:no_speech). The Google recogniser can take 10–20 s to report afterstopListening. - If a permission is missing,
listen/2still returns the socket. The screen then gets{:speech, :error, :permission}followed by idle.
Error reasons: :no_speech, :language, :permission, :service_permission,
:network, :audio, :busy, :client, :server, :unavailable,
:too_many_requests, or {:unknown, code}. What each means:
:language: no model or language pack for the locale.:service_permission: on Android, the recognition service (the Google app) has no microphone access itself, even though your app does.
Testing without a microphone
MobSpeech.listen(socket,
engine: MobSpeech.Engine.Fake,
script: [:listening, {:partial, "hello"}, {:wait, 100}, {:partial, "hello world"}],
on_stop: [{:final, ""}] # empty final → you get "hello world"
)
Development
mix setup # deps + git hooks (format, credo --strict, compile; tests when mix.exs changes)
mix test
License
MIT