No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Morten Olsen 60a9641d9f
AMD GPU
2026-09-09 00:19:07 +02:00
corpus/audio init 2026-09-08 16:27:36 +02:00
results AMD GPU 2026-09-09 00:19:07 +02:00
scripts AMD GPU 2026-09-09 00:19:07 +02:00
src/zerolatency AMD GPU 2026-09-09 00:19:07 +02:00
.gitignore init 2026-09-08 16:27:36 +02:00
.python-version init 2026-09-08 16:27:36 +02:00
DECISION.md update 2026-09-08 17:46:44 +02:00
FUTURE.md stuff 2026-09-08 18:52:02 +02:00
LEARNINGS.md AMD GPU 2026-09-09 00:19:07 +02:00
mise.toml init 2026-09-08 16:27:36 +02:00
PRESENTATION.md AMD GPU 2026-09-09 00:19:07 +02:00
pyproject.toml AMD GPU 2026-09-09 00:19:07 +02:00
README.md update 2026-09-08 17:46:44 +02:00
Taskfile.yml AMD GPU 2026-09-09 00:19:07 +02:00
ticket.md init 2026-09-08 16:27:36 +02:00
USER-TEST.md update 2026-09-08 17:46:44 +02:00
uv.lock AMD GPU 2026-09-09 00:19:07 +02:00

zero-latency — offline voice assistant POC

Proof of concept for a fully offline voice assistant optimized for near-zero perceived response latency by doing LLM + TTS work while the user is still speaking. See ticket.md for the brief, LEARNINGS.md for findings, DECISION.md for the incremental-KV decision gate, USER-TEST.md for the live test protocol, and results/*/report.md for benchmark outcomes.

Current headline (M4, fully offline): ~0.47 s median speak-end → speak-start vs ~1.63 s for the traditional VAD→STT→LLM→TTS pipeline — a 3.5× speedup, at 100% speculation hit rate and less total LLM compute than the baseline.

Architecture

mic / corpus replay (16 kHz, 40 ms chunks)
  ├─ Silero VAD ─────────────────────────────┐
  ├─ Nemotron streaming ASR (sherpa-onnx) ───┤ partials + stable/unstable split
  └─ smart-turn v3 (semantic) ───────────────┤
                                ▼
              adaptive endpointing (required silence shrinks
              as semantic completion confidence rises)
                                ▼
        speculative LLM (llama-server, Qwen3-4B Q4) ── abortable
                                ▼
        clause-chunked Piper TTS into a playback buffer
                                ▼
        turn committed → play buffer | user continued → discard

Benchmark modes:

  • A — baseline: VAD endpoint → final STT → LLM → TTS → play
  • B — conventional speculation (NVIDIA/pipecat style)
  • C — B + incremental KV prefill: every stable transcript extension is pushed into llama-server's slot cache (cache_prompt + --cache-reuse) as a n_predict: 0 warm request, so decode starts against a hot cache

Setup (macOS, Apple Silicon)

Requires: mise or uv + Python 3.12, llama-server (brew install llama.cpp), task.

task setup     # deps + ~3.5 GB of models
task corpus    # build the 55-utterance benchmark corpus (uses macOS `say`)
task bench     # full A/B/C benchmark (~45 min)
task report RUN=results/<dir>
task live      # talk to it

Layout

  • src/zerolatency/ — pipeline (asr/, turn/, llm/, tts/, pipeline/), benchmark harness (bench/), instrumentation (events.py), live mode (live.py)
  • scripts/spikes/ — standalone component micro-benchmarks with printed measurements
  • corpus/ — benchmark utterances (spec in bench/corpus.py, audio + ground truth built)
  • results/ — traces (JSONL, one event stream per run), metrics, reports