Playback uses piper-tts to generate narration audio locally — no cloud dependency, no API key, no network call during synthesis.
Set one or more voices in meta.yaml:
voices:
- northern_english_male
- southern_english_femaleThe pipeline generates a full output set per voice. Omit the field to use the
defaultVoices from playback.config.ts.
| Voice | Quality | Sample rate | Notes |
|---|---|---|---|
alan |
medium | 22 050 Hz | |
alba |
medium | 22 050 Hz | |
northern_english_male |
medium | 22 050 Hz | |
southern_english_female |
low | 22 050 Hz | No medium model exists |
aru_09 |
medium | 22 050 Hz | Multi-speaker model; see below |
southern_english_female only has a low quality model. The pipeline handles
this automatically — no config change needed.
All five voices have tuned lengthScale, noiseScale, and noiseW entries in
VOICE_CONFIG in src/runner/piper.ts and are fully supported by the pipeline.
aru_09 uses the Liverpool ARU Speech Corpus model (en_GB-aru-medium), which
contains 12 speakers. The speaker field in voices.example.yaml selects
speaker 4 (female, Received Pronunciation). The model file downloads once
regardless of how many speaker entries reference it. To add other ARU speakers,
copy the aru_09 block in your XDG catalogue, change the key name, and set a
different speaker ID (from the speaker_id_map in en_GB-aru-medium.onnx.json).
| Limitation | Notes |
|---|---|
southern_english_female is low quality only |
No medium model exists for this voice. The audio quality is noticeably lower than the medium models. |
| Prosody varies across segments | Each narration field is a separate piper call. VITS samples fresh noise each time, so adjacent segments can sound like different reads. See Audio variance below. |
| No seed control | The ONNX runtime does not expose a random seed. Reproducible synthesis requires batching all text into one call. |
| CPU-only synthesis | piper uses ONNX CPU inference. A 7-segment episode takes several seconds per segment on Apple silicon. The pipeline runs segments sequentially to avoid CPU thrashing. |
| No mid-episode voice switching | All segments in an episode use the same voice. |
src/substitutions.ts maps literal strings to phonetic spellings before
synthesis:
| Input | Replacement |
|---|---|
GOV.UK |
guv yew-kay |
govuk |
guv yew-kay |
Add entries here for any acronym, brand name, or technical term that piper mispronounces. Put longer or more-specific entries first.
The sections below cover the synthesis pipeline internals. They are for
contributors working on src/runner/piper.ts, src/runner/ffmpeg.ts, or the
audio mix.
Each .onnx.json config bakes in default inference parameters. All four voices
share identical defaults:
| Parameter | Default | Effect |
|---|---|---|
noise_scale |
0.667 |
Prosody variation (pitch, emphasis). Higher = more expressive but less consistent. |
noise_w |
0.8 |
Duration variation (phoneme widths). Higher = more speed variation between calls. |
length_scale |
1.0 |
Speaking rate multiplier. >1 = slower, <1 = faster. |
| (sample rate) | 22050 |
Fixed by the model — the inference call cannot change it. |
The piper CLI can override all three:
--noise-scale Generator noise (default from model config)
--noise-w-scale Phoneme width noise (default from model config)
--length-scale Phoneme length (default from model config)
buildAudioFilterComplex() in src/runner/ffmpeg.ts delay-normalises each
.wav and mixes them:
[1:a]loudnorm=I=-16:TP=-1.5:LRA=11[norm0];[norm0]adelay=0|0[a0];
[2:a]loudnorm=I=-16:TP=-1.5:LRA=11[norm1];[norm1]adelay=7950|7950[a1];
…
[a0][a1]…amix=inputs=N:duration=longest:normalize=0[aout]
Key choices:
loudnormnormalises each segment to -16 LUFS before mixing so no single clip dominatesadelaypositions each clip at the correct start time in millisecondsamix normalize=0keeps each stream at constant volume. The default (normalize=1) divides the sum by active input count, which causes volume to rise as earlier segments end — the "shouting" effect
Narration sounds like several different people saying one sentence each. Pitch and speaking rate shift noticeably between segments despite all audio coming from the same model.
The pipeline calls piper once per narration field. VITS samples from a noise
distribution on every inference call. With noise_scale = 0.667 and
noise_w = 0.8, each independent call draws fresh noise samples — so each
segment sounds like a distinct read rather than a continuous delivery.
call 1 → "First, clone the repository…" → noise sample A
call 2 → "Git downloads the repository…" → noise sample B
call 3 → "Let's see what's inside." → noise sample C
No shared noise state exists across calls, and the ONNX runtime provides no supported way to fix a seed.
Each synthesis call receives a single isolated sentence. The eSpeak-NG
phonemiser applies sentence-boundary prosody independently to every call —
short sentences get an aggressive falling tone, adjacent sentences each have
their own prosodic arc. This occurs even with noise_scale 0 and noise_w 0.
loudnorm=I=-16:TP=-1.5:LRA=11 applied per segment in ffmpeg buffers and
analyses each clip independently. A short segment and a long segment get
different gain curves, which can introduce audible differences between adjacent
segments even when the raw audio is consistent.
Pass lower noise parameters to piper. Narrows variance without eliminating it:
// src/runner/piper.ts — synthesise()
'--noise-scale', '0.33', // was: 0.667
'--noise-w-scale', '0.4', // was: 0.8These values are half the defaults. Benchmark by ear against a 7-segment tape.
These could also appear in PlaybackConfig for per-episode tuning.
Concatenate all narration segments into a single piper call, then split the
resulting audio using silence detection (ffmpeg -af silencedetect) or a
marker tone. Removes prosody discontinuity almost entirely. Requires changes to
src/runner/piper.ts and src/extractor/tts.ts.
Adding trailing silence reduces the hard-cut feel between segments without fixing prosody:
--sentence-silence 0.15The northern_english_male model uses the en-gb-x-rp eSpeak voice. The
alba and alan models may use different eSpeak phonemisers and could have
better cross-segment continuity. Worth benchmarking once the pipeline supports them.