- Python 69.1%
- HTML 27.4%
- JavaScript 3.5%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| apps/viewer | ||
| assets | ||
| docs | ||
| qwen | ||
| services | ||
| .gitignore | ||
| me.wav | ||
| mise.toml | ||
| persona.md | ||
| README.md | ||
| Taskfile.yml | ||
Local 3D Conversational Avatar
A self-hosted 3D avatar that listens through a microphone, answers with a local LLM, speaks with local text-to-speech, and animates a GLB face in the browser. No hosted inference service or project-specific infrastructure is required.
How it works
microphone -> streaming ASR -> local Qwen3 LLM -> Piper or Qwen3-TTS
| | |
+--------------- barge-in -------------+ v
Audio2Face-3D
|
synchronized audio +
ARKit blendshape frames
|
browser viewer
The default path uses Piper because it starts quickly and synthesizes speech in real time on a laptop. Qwen3-TTS is optional: it clones the checked-in reference voice locally, but has substantially higher latency on Apple Silicon.
Requirements
The supported quickstart targets an Apple Silicon Mac. You need:
- Homebrew
- mise for Task and uv
- a microphone, headphones, and a Chromium-based browser
- several gigabytes of free disk space for the local models
Use headphones before starting a conversation. Speaker output feeds back into the microphone because the pipeline has no acoustic echo cancellation. The barge-in detector then mistakes the avatar's voice for yours and repeatedly interrupts playback.
Quickstart
Install the toolchain and download the models used by the Piper path:
brew install mise
mise trust
mise install
task setup
task setup installs llama.cpp, synchronizes both Python environments, and
downloads the ASR, turn-detection, LLM, Piper, and Audio2Face models. Downloads
are cached, and Task skips models that are already present.
Put on your headphones, then start the avatar in three terminals:
# Terminal 1: web server and face service
task up
# Terminal 2: microphone conversation using local Piper TTS
task talk
# Terminal 3: open the viewer
task open
Click Enable audio in the viewer. Speak after the avatar has loaded; press Ctrl-C in the first two terminals to stop the session.
If task is not on your shell path after mise install, prefix commands with
mise exec --, for example mise exec -- task setup.
Optional Qwen voice cloning
Install audio.cpp and the 2.7 GB Qwen3-TTS Q8 model once:
task setup-qwen
Then keep task up running and replace task talk with:
task talk-qwen
This command starts the local Qwen TTS server, waits for it to become healthy,
runs the conversation, and stops the server on exit. See
qwen/README.md for the standalone server command and model
details.
Make the avatar your own
Replace the 3D avatar
Go to KeenTools, create an avatar with their simple
avatar creator, and export it as GLB. Replace assets/avatar.glb with the
exported file, keeping the same filename. The included viewer loads the new
avatar automatically the next time it starts.
Replace the Qwen voice reference
Piper uses its downloaded voice model and does not read the reference clip.
Qwen clones assets/voice/reference.wav, which is
a redistributable LibriSpeech sample included as a generic default.
To use a different voice:
- Record a clean mono WAV containing 3–30 seconds of one consenting speaker.
- Replace
assets/voice/reference.wav. - Set
default_voice_preset.reference_textinqwen/qwen-server.jsonto the exact words audible in the recording. - Restart
task talk-qwenso the reference cache is rebuilt.
Do not paraphrase the transcript or include inaudible words. See
assets/voice/README.md for the default clip's source,
CC BY 4.0 attribution, and recording guidance.
Change the character
Edit persona.md to change the conversational character. Keep
responses short and retain the final-position emotion-tag instructions if you
want the LLM to drive expressions such as [smiling] and [thinking].
Useful commands
| Command | Purpose |
|---|---|
task setup |
Install the default local stack and download its models |
task setup-qwen |
Install optional local Qwen3-TTS support |
task up |
Run the static viewer server and face service |
task talk |
Run the complete local conversation with Piper |
task talk-qwen |
Run the complete local conversation with Qwen voice cloning |
task talk-mic |
Drive the face directly from a microphone; no ASR, LLM, or TTS |
task speak-test |
Animate the checked-in reference clip without the voice pipeline |
task open |
Open the live viewer |
task open-calibration |
Open the morph-target calibration viewer |
task audio-devices |
List microphone and output device names |
task test-face |
Run the face-service protocol tests |
task --list |
List every available task |
Pass extra voice-pipeline arguments after --. For example:
task talk -- --in-dev "USB Microphone"
Call and OBS audio
OBS Virtual Camera carries video only. The optional call workflow uses two
BlackHole buses so the avatar does not hear its own response. Run
task audio-setup once, then task audio-routing for the exact OBS and call
application settings. Use task talk-call instead of task talk when call
audio arrives through BlackHole. Press H or double-click the live viewer to
hide its controls; use http://localhost:8765/apps/viewer/live.html?clean=1 for
an OBS browser source.
Repository layout
| Path | Purpose |
|---|---|
assets/avatar.glb |
Replaceable GLB avatar |
assets/voice/ |
Replaceable Qwen voice reference and its attribution |
persona.md |
Replaceable local LLM system prompt |
apps/viewer/ |
Three.js live and calibration viewers |
services/voice/ |
Streaming ASR, turn detection, local LLM, Piper/Qwen TTS, and face bridge |
services/face/ |
Audio2Face-3D engine and synchronized WebSocket service |
qwen/ |
Local audio.cpp/Qwen3-TTS configuration and launcher |
Downloaded models
Model weights are fetched during setup and are not committed:
- ASR:
csukuangfj/sherpa-onnx-nemotron-speech-streaming-en-0.6b-160ms-int8-2026-04-25 - Turn detection:
pipecat-ai/smart-turn-v3andcsukuangfj/vad - LLM:
bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF - Piper:
csukuangfj/vits-piper-en_US-joe-medium; its model card identifies the source dataset as CC0 - Face animation:
nvidia/Audio2Face-3D-v2.3-Mark - Optional cloned TTS:
audio-cpp/audio.cpp-gguf
Review each upstream model's terms before redistribution or deployment. The conversation pipeline currently supports English input; Qwen3-TTS latency may be too high for natural live conversation on some machines.