Skip to content

Repository files navigation

HA Voice Add-ons

A small Home Assistant add-on repository for speaker verification and Speech-to-Text over an OpenAI-compatible API (Groq, OpenAI, LocalAI, …).

Repository: https://github.com/kdinya/ha-voice-addons

Add-on Wyoming port Purpose
Voice Match 10350 Verifies that it's really you talking, strips background noise / other speakers, then proxies the cleaned audio to an upstream Wyoming STT add-on. Fully self-contained.
Wyoming OpenAI STT 10300 Speech-to-Text (Ukrainian / Russian by default) via any OpenAI-compatible transcription API.

Installation

Tested with Home Assistant OS, Core 2026.8.x, Supervisor 2026.07.x, on amd64 and aarch64.

  1. Settings → Add-ons → Add-on store → ⋮ → Repositories
  2. Add:
    https://github.com/kdinya/ha-voice-addons
    
  3. Refresh the store, install the add-on(s) you need, and start them.
  4. Add a Wyoming Protocol integration pointed at the relevant port.

Typical pipeline

Home Assistant microphone
    → Voice Match (10350)            ← speaker verification
        → Wyoming OpenAI STT (10300) ← transcription (Groq / OpenAI / …)
            → Assist / automations

You can also use either add-on on its own — Voice Match works with any Wyoming STT service as its upstream, and Wyoming OpenAI STT works as a standalone Wyoming STT provider without Voice Match in front of it.


Voice Match — quick overview

  • Fully self-contained — no dependency on a third-party base image.
  • Hot-reload: after enrolling a new voice, no add-on restart is needed.
  • Runs on CPU only (amd64 / aarch64) — no GPU required.
  • Web UI (Ingress panel) for recording samples, checking their quality, enrollment, and managing voiceprints.
  • First start takes 1–3 minutes while the ECAPA-TDNN model downloads into /data; subsequent starts are fast (cached).

Setup:

  1. In Configuration, set upstream_uri ("Wyoming STT address"), e.g. tcp://homeassistant:10300 — or use Scan in the web UI.
  2. Open the Voice Match side-panel:
    • pick or create a speaker;
    • record from the microphone or upload files;
    • check sample quality;
    • Enroll → the voiceprint is active immediately.
  3. Add a Wyoming Protocol integration on port 10350 and select Voice Match as the STT engine in your Assist pipeline.
Option Description Default
upstream_uri Where verified audio is forwarded tcp://homeassistant:10300
verify_threshold Voice similarity threshold (0–1); higher = stricter 0.35
extraction_threshold Strips other voices/background before STT 0.30
stt_languages Comma-separated languages advertised to HA uk,ru
require_speaker_match Reject audio from unenrolled speakers true
log_level DEBUG / INFO / WARNING / ERROR INFO

Full reference: voice-match/DOCS.md. Security notes: SECURITY.md.

Enrolling a voice

  1. Pick an existing speaker or type a new name (lowercase, a-z0-9_-).
  2. Record 3–5 samples (3–10 seconds of clean speech each).
  3. Check all samples → remove any marked weak/bad.
  4. Enroll → the voiceprint is created and active immediately (no restart required).

Wyoming OpenAI STT — quick overview

  • Speech-to-Text via any OpenAI-compatible transcription API (Groq, OpenAI, LocalAI, …).
  • Ships tuned defaults for Ukrainian / Russian auto-detection, but you can point it at any model/language your provider supports.
Option Description Example
base_url OpenAI-compatible API endpoint https://api.groq.com/openai/v1
api_key Provider API key gsk_...
model Model name at the provider whisper-large-v3-turbo
languages Space-separated languages; first = preferred on ambiguity uk ru
language_mode auto (detect) or fixed (force first language) auto
log_level DEBUG / INFO / WARNING / ERROR INFO

Full reference: wyoming-openai-stt/DOCS.md.


Releases

Each add-on is versioned independently via its own config.yaml, but this repo publishes one combined GitHub Release per release event — push a single tag, and .github/workflows/release.yml builds one release body containing the changelog section from both add-ons:

git tag 3.0.1
git push origin 3.0.1

Push exactly one tag, matching the version you want as the release title (normally the voice-match version, since it's the primary add-on). The workflow pulls the matching ## <version> section from voice-match/CHANGELOG.md; for wyoming-openai-stt/CHANGELOG.md it uses the same version if present, otherwise falls back to that file's latest ## section — so the release always shows both add-ons' current state.

Do not push a second, add-on-prefixed tag (e.g. wyoming-openai-stt-2.0.0) to "also" release the other add-on, and don't run gh release create by hand — either one produces a duplicate, empty release, since the workflow only understands a single bare version tag. See RELEASING.md for the full process and how to clean up if a duplicate release slips through.

CI (.github/workflows/ci.yml) runs on every push/PR: syntax checks, a regression test for the languages code-injection fix, add-on smoke tests, a check that CHANGELOG.md has an entry for the current config.yaml version, and a Dockerfile lint + build for both add-ons.

License

See LICENSE.

About

Home Assistant Voice Add-ons — Voice Match + OpenAI STT

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages