VocoLoco: OmniVoice for the browser
k2-fsa/OmniVoice converted for text-to-speech that runs entirely in the browser with VocoLoco: weights for VocoLoco's own WebGPU engine (engine/), and ONNX exports for ONNX Runtime Web, the fallback for browsers without the engine's WebGPU features.
Try it live: https://magkino.github.io/vocoloco_tts/
Files
WebGPU engine (engine/)
Each model is a JSON manifest (tensor names, shapes, offsets) and one binary file with the weights.
| File | Size | Description |
|---|---|---|
engine/omnivoice-f16.json + .bin |
1.1 GB | Main TTS model, f16 weights |
engine/decoder-f32.json + .bin |
85 MB | Audio decoder (tokens to 24 kHz waveform), f32 |
engine/encoder-f32.json + .bin |
622 MB | Audio encoder for voice cloning, f32 |
engine/whisper-i8.json + .bin |
936 MB | Whisper large-v3-turbo for automatic transcription of voice references, int8 weights (one f16 scale per 32 rows) |
ONNX (ONNX Runtime Web)
| File | Size | Description |
|---|---|---|
omnivoice-main-split.onnx + omnivoice-main.onnx_data_00–_04 |
2.3 GB | Main TTS model (FP32, sharded; shards listed in omnivoice-main-manifest.json) |
omnivoice-decoder.onnx |
83 MB | Audio decoder (tokens to waveform) |
omnivoice-encoder-fixed.onnx |
624 MB | Audio encoder for voice cloning |
Shared
| File | Size | Description |
|---|---|---|
omnivoice-config.json |
<1 KB | Model config |
tokenizer.json, tokenizer_config.json |
11 MB | Qwen2 BPE text tokenizer |
Usage
These files are made for VocoLoco, a fully client-side TTS application (live). No server required.
Architecture
- Backbone: Qwen3-0.6B (28 transformer layers)
- Audio codec: HiggsAudioV2 (8 codebooks, 24 kHz output)
- Generation: iterative masked diffusion (8-32 steps)
- Voice cloning: zero-shot via reference audio encoding
- Voice design: text-based control (gender, pitch, accent)
License
The files carry the licences of the models they are converted from:
| Files | Converted from | Licence |
|---|---|---|
engine/omnivoice-f16.*, omnivoice-main-split.onnx + data shards, omnivoice-config.json |
OmniVoice main model by Xiaomi Corp. (k2-fsa) | CC BY-NC (non-commercial) |
engine/decoder-f32.*, engine/encoder-f32.*, omnivoice-decoder.onnx, omnivoice-encoder-fixed.onnx |
Higgs Audio v2 tokenizer by Boson AI, as shipped in OmniVoice's audio_tokenizer/ |
Boson Higgs Audio 2 Community License (LICENSE-higgs-audio-2), based on the Meta Llama 3 Community License (LICENSE-llama-3); see NOTICE |
engine/whisper-i8.* |
Whisper large-v3-turbo by OpenAI | MIT (LICENSE-whisper) |
tokenizer.json, tokenizer_config.json |
Qwen3 tokenizer (Qwen2 BPE) by Alibaba Cloud, as used by OmniVoice | Apache 2.0 (LICENSE-qwen) |
Built with Higgs Materials licensed from Boson AI USA, Inc., Copyright Boson AI USA, Inc., All Rights Reserved and Meta Llama 3 licensed under the Meta Llama 3 Community License, Copyright Meta Platforms, Inc., All Rights Reserved.
Use of the audio tokenizer files must follow the Llama 3 Acceptable Use Policy; products or services with more than 100,000 annual active users need an expanded licence from Boson AI.
Licence history (OmniVoice main model): The ONNX exports were uploaded on 2026-04-09 (commit), when the OmniVoice model card listed the weights as Apache 2.0 (card as of 2026-04-05,
license: apache-2.0). On 2026-07-03 the OmniVoice authors relicensed the pre-trained model as CC BY-NC, citing their training data (model card change). The weight files themselves were not changed. This repository follows the current upstream licence; the engine files were converted after the change.
Attribution
Based on OmniVoice by Xiaomi Corp (k2-fsa), the Higgs Audio v2 tokenizer by Boson AI and Whisper by OpenAI.