Blog/LLMOps & evals
Local text-to-speech at scale: narrating 95 articles with open models
Three open TTS models on one laptop: how 95 articles became 153 minutes of narration, from SQLite rows through chunked synthesis and AAC at 192 kbit/s to a player that follows you down the page.
Balázs Csorba··11 min read
- Text-to-speech
- Audio
- Kokoro
- FFmpeg
- Vue

Key takeaways
- Local TTS inverts the economics of narration: 95 articles and 153 minutes of audio cost about 35 minutes of generation and zero API fees — the marginal cost of one more article is electricity.
- Model choice at this scale is a speed-quality Pareto, not a benchmark fight: Piper was fastest, Bark never finished, and Kokoro v1.0 (82M, ONNX) won on prosody at about four times real time.
- Chunk at sentence boundaries up to 400 characters with 80 ms of silence between pieces; smarter break points buy almost nothing a listener would notice.
- WAV is a working format, not a shipping format: AAC at 192 kbit/s with faststart cut 406 MB to 181 MB and makes duration and seeking work before the file is fully downloaded.
- In a prerendered player, media events beat assumptions: loadedmetadata can fire before hydration, so duration must be re-read on durationchange, canplay and mount — subscribe to state, not notifications.
Every article on this site now has a narration: 95 posts, 153 minutes of audio, generated on one laptop without a single API call. This post is the full account — how three open text-to-speech models were compared, how the pipeline turns SQLite rows into sentence-sized synthesis chunks, why the shipped format is AAC at 192 kbit/s, and how the player is built so the audio follows you down the page.
The numbers first, because they frame every decision below: 95 articles (45 blog posts and 50 tool reviews), 153 minutes of finished narration, about 35 minutes of generation time, 406 MB of intermediate WAV reduced to 181 MB of shipped AAC, and zero euros in API fees. Everything ran on a single Apple Silicon laptop with 16 GB of memory.
Why local text-to-speech
Hosted text-to-speech is excellent and improving every quarter — but it is metered. At this volume the bill scales with every word, every re-render after an edit, every experiment with voice or speed. A local pipeline inverts the economics: the marginal cost of one more article is a few cents of electricity and two minutes of waiting. It is also reproducible — the same text, the same model version and the same voice produce the same file next year, which matters when an edited article must be re-narrated without the voice drifting.
There is a second reason: the text never leaves the machine. The draft of an unpublished article is read by the same process that will publish it, on the same machine, and when a paragraph changes, rerunning one slug regenerates one file — not ninety-five.
Three models, one laptop
Piper was the starting point because it is boring in the best way: a small ONNX model (the LibriTTS high voice is about 137 MB), espeak-ng phonemization, 22.05 kHz output, generation at roughly five times real time. The result is clean and consistent — a competent newsreader — but it is also even. Long articles sound slightly flat.
The first obstacle was packaging rather than quality: the prebuilt macOS binary ships x86_64 only, and on Apple Silicon it refuses to link against the arm64 espeak-ng library with an architecture mismatch. The Python package sidesteps the binary entirely and was running in minutes. A useful reminder that "it does not run" is often a toolchain problem, not a model problem.
Bark was the most tempting of the three: expressive, capable of laughs and sighs, with published samples that make any other model sound like a metronome. It was also the one that never produced a usable sentence. The current PyTorch changed the default of torch.load to weights_only, so checkpoint unpickling failed until patched; once loading worked, inference on CPU was slow enough that 95 articles would have taken most of a day. Expressiveness was not worth that bill — Bark stays a research toy for this workload.
Kokoro v1.0 through the kokoro-onnx bindings was the find: 82 million parameters, about 325 MB of ONNX, 24 kHz output, 54 voices, and generation at about four times real time with noticeably better prosody than Piper — sentences land their emphasis, and numbers and abbreviations are handled without the halting quality typical of smaller models. It is what every article on this site sounds like now.
| Model | Weights | Output | Speed on this machine | Verdict |
|---|---|---|---|---|
| Piper, LibriTTS high | about 137 MB, ONNX | 22.05 kHz WAV | about 5x real time | Fast and consistent, slightly flat |
| Bark by Suno | about 2 GB, PyTorch | 24 kHz WAV | never finished in budget | Expressive, impractical here |
| Kokoro v1.0 | 82M, about 325 MB, ONNX | 24 kHz WAV | about 4x real time | Natural prosody; the shipped voice |
The generation pipeline
The content database is the single source of truth, so the pipeline reads from it directly: title, description, key takeaways and FAQ on one side, the article block list on the other. Nothing is scraped from the rendered page. If a paragraph exists in the post, it is narrated — and if it is a table, a callout or a code block, it is flattened into something a listener can follow: table rows become comma-separated lines, callout titles are spoken like headings, code is read as written.
Two constraints shape the middle of the pipeline. Models truncate long input, so articles must be split; and prosody must not be cut in half, so splits must land on sentence boundaries. The chunker is deliberately plain: accumulate sentences up to 400 characters, flush at the boundary, join the pieces with 80 milliseconds of silence. Smarter break points — paragraph ends, headings as intonation resets — bought almost nothing; the sentence is the unit listeners actually notice.
The chunker
def chunks(text, limit=400):
"""Accumulate sentences up to limit characters; never split mid-sentence."""
out, buf = [], ""
for sentence in text.split(". "):
cand = (buf + " " + sentence).strip()
if len(cand) > limit and buf:
out.append(buf + ".")
buf = sentence
else:
buf = cand
if buf:
out.append(buf)
return out
pieces = []
for part in chunks(article_text):
audio, _ = model.create(part, voice="af_bella")
pieces.append(audio)
pieces.append(np.zeros(int(24000 * 0.08))) # 80 ms between chunks
sf.write(tmp_wav, np.concatenate(pieces), 24000)One voice reads all 95 articles. It is af_bella, chosen after rendering the same paragraph in several of the 54 voices. Consistency is the point: the archive should sound like one publication, not a lottery. The narration reads the English text of each article — including from the German and Hungarian pages — and is labelled as the English narration of the piece, not as a translation of it.
WAV is a working format, not a shipping format
The synthesis step writes 24 kHz, 16-bit PCM: correct, seekable, uncompressed — and about 406 MB for 153 minutes of audio. On disk that is harmless; over the wire it is a mistake, and on a site that prerenders everything it would be the single largest class of asset by an order of magnitude.
The shipped format is AAC in an M4A container at 192 kbit/s, mono. That bitrate is generous for speech — 96 to 128 kbit/s already sounds transparent at 24 kHz — but 192 leaves headroom and costs about a megabyte per article. The container matters as much as the codec: faststart moves the moov atom to the front of the file, so a browser knows the duration and can seek before the download finishes. Without it, players sit at 0:00 and refuse to scrub until the last byte arrives.
ffmpeg -i article.wav -c:a aac -b:a 192k -ac 1 -movflags +faststart article.m4a406 MB became 181 MB — a 55 percent reduction, every file under 1.5 MB, delivered with ordinary range requests and cache headers. The WAV files never reach the build output: they are deleted the moment the encoder succeeds.
The player
Audio that arrives late is audio nobody hears, so the player loads metadata only: the header is fetched, the duration is learned, and no actual samples are downloaded until the reader presses play. Nothing autoplays — a page that talks at you is a page people close.
Each article gets an inline bar under the table of contents: play and pause, a seekable progress line with current and total time, and a speed toggle cycling 1x, 1.25x, 1.5x and 2x. The seek control is a native range input styled to the site, so keyboard navigation and screen-reader semantics are inherited rather than reimplemented.
The inline bar is useless once you scroll past it, so it hands over to a sticky tab pinned to the right edge of the viewport. An IntersectionObserver watches the bar: the moment it leaves the screen, the tab appears; scroll back up, and it steps aside. The tab fills from the bottom as playback advances — progress is legible at a glance — and it honours prefers-reduced-motion like the rest of the interface.
What broke on the way
- Packaging beat models twice. Piper shipped an x86_64-only binary for an arm64 machine, and Bark’s checkpoint loading failed after PyTorch flipped the weights_only default. Neither failure had anything to do with speech.
- Version drift is real. The kokoro-onnx package passed the speed parameter as int32 while the shipped model expects float — a one-line patch, found by reading a two-line stack trace instead of guessing.
- Verify the audio, not just the file. ffprobe on every output caught duration problems before they reached a browser; a file that exists is not a file that plays.
- Old formats die hard. The first pass left WAVs in the build output, caught only because a request for one returned the wrong byte count.
The numbers
| Metric | Value |
|---|---|
| Articles narrated | 95 (45 blog posts, 50 tool reviews) |
| Finished audio | 153 minutes |
| Generation wall time | about 35 minutes, 4.5x real time |
| Peak memory | under 2 GB |
| Intermediate WAV | 406 MB |
| Shipped AAC | 181 MB, about 1 MB per article |
| API cost | 0 EUR |
Read together, the numbers say that the cost of this pipeline is patience, not money. A full re-run — after a voice change or an edit that touches every article — is well under an hour on a machine that stays usable throughout. That is the argument for local text-to-speech at this scale: not that it beats a frontier hosted voice on expressiveness, but that it makes narration a default rather than a budget line.
What I would do differently
Start with Kokoro. The shootout was not wasted — comparing models is what makes the choice defensible — but the shipped pipeline would have been identical with Kokoro as the only candidate. Second: encode straight from synthesis to AAC and never persist WAVs; the intermediate format added a cleanup step and 406 MB of files that never needed to exist. Third: treat the player’s media events as state to subscribe to, not notifications to catch — and the duration bug never happens.
What remains is scope. The narration is English-only for now — one voice, one language, 95 articles — and per-locale voices are the obvious next step if the German and Hungarian audiences ask for them. Until then the English narration doubles as the pronunciation guide for every product name on the site, which is its own quiet utility.
Sources
Frequently asked questions
Why not use a hosted text-to-speech API?
Because at 95 articles the metered cost starts to matter with every edit and re-render, and because reproducibility matters more: the same text and model version produce the same file next year. Hosted voices still win on peak expressiveness; they do not win on the economics of a whole archive.
Which model sounds the best?
Kokoro v1.0, and it is what ships: 82 million parameters, natural emphasis, clean handling of numbers and product names. Piper with the LibriTTS voice is a close second and noticeably faster. Bark is the most expressive of the three but was too slow to be usable at this volume.
How long does the full run take?
About 35 minutes for all 95 articles on one Apple Silicon laptop — roughly 4.5 times real time — with peak memory under 2 GB. A single article regenerates in about 20 seconds.
Why AAC (M4A) instead of MP3 or Opus?
At 192 kbit/s the codec differences are inaudible for speech, so the decision is about packaging: AAC plays natively everywhere including Safari, faststart exposes the duration before the download completes, and Opus would be smaller but has uneven browser support outside WebM.
Does the narration exist in German and Hungarian?
Not yet. The narration reads the English original of each article; the player interface itself is fully localized in English, German and Hungarian.