No description
- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| corpus/audio | ||
| results | ||
| scripts | ||
| src/zerolatency | ||
| .gitignore | ||
| .python-version | ||
| DECISION.md | ||
| FUTURE.md | ||
| LEARNINGS.md | ||
| mise.toml | ||
| PRESENTATION.md | ||
| pyproject.toml | ||
| README.md | ||
| Taskfile.yml | ||
| ticket.md | ||
| USER-TEST.md | ||
| uv.lock | ||
zero-latency — offline voice assistant POC
Proof of concept for a fully offline voice assistant optimized for near-zero
perceived response latency by doing LLM + TTS work while the user is still
speaking. See ticket.md for the brief, LEARNINGS.md for findings,
DECISION.md for the incremental-KV decision gate, USER-TEST.md for the
live test protocol, and results/*/report.md for benchmark outcomes.
Current headline (M4, fully offline): ~0.47 s median speak-end → speak-start vs ~1.63 s for the traditional VAD→STT→LLM→TTS pipeline — a 3.5× speedup, at 100% speculation hit rate and less total LLM compute than the baseline.
Architecture
mic / corpus replay (16 kHz, 40 ms chunks)
├─ Silero VAD ─────────────────────────────┐
├─ Nemotron streaming ASR (sherpa-onnx) ───┤ partials + stable/unstable split
└─ smart-turn v3 (semantic) ───────────────┤
▼
adaptive endpointing (required silence shrinks
as semantic completion confidence rises)
▼
speculative LLM (llama-server, Qwen3-4B Q4) ── abortable
▼
clause-chunked Piper TTS into a playback buffer
▼
turn committed → play buffer | user continued → discard
Benchmark modes:
- A — baseline: VAD endpoint → final STT → LLM → TTS → play
- B — conventional speculation (NVIDIA/pipecat style)
- C — B + incremental KV prefill: every stable transcript extension is
pushed into llama-server's slot cache (
cache_prompt+--cache-reuse) as an_predict: 0warm request, so decode starts against a hot cache
Setup (macOS, Apple Silicon)
Requires: mise or uv + Python 3.12, llama-server
(brew install llama.cpp), task.
task setup # deps + ~3.5 GB of models
task corpus # build the 55-utterance benchmark corpus (uses macOS `say`)
task bench # full A/B/C benchmark (~45 min)
task report RUN=results/<dir>
task live # talk to it
Layout
src/zerolatency/— pipeline (asr/,turn/,llm/,tts/,pipeline/), benchmark harness (bench/), instrumentation (events.py), live mode (live.py)scripts/spikes/— standalone component micro-benchmarks with printed measurementscorpus/— benchmark utterances (spec inbench/corpus.py, audio + ground truth built)results/— traces (JSONL, one event stream per run), metrics, reports