VocoLoco: OmniVoice for the browser

k2-fsa/OmniVoice converted for text-to-speech that runs entirely in the browser with VocoLoco: weights for VocoLoco's own WebGPU engine (engine/), and ONNX exports for ONNX Runtime Web, the fallback for browsers without the engine's WebGPU features.

Try it live: https://magkino.github.io/vocoloco_tts/

Files

WebGPU engine (engine/)

Each model is a JSON manifest (tensor names, shapes, offsets) and one binary file with the weights.

File Size Description
engine/omnivoice-f16.json + .bin 1.1 GB Main TTS model, f16 weights
engine/decoder-f32.json + .bin 85 MB Audio decoder (tokens to 24 kHz waveform), f32
engine/encoder-f32.json + .bin 622 MB Audio encoder for voice cloning, f32
engine/whisper-i8.json + .bin 936 MB Whisper large-v3-turbo for automatic transcription of voice references, int8 weights (one f16 scale per 32 rows)

ONNX (ONNX Runtime Web)

File Size Description
omnivoice-main-split.onnx + omnivoice-main.onnx_data_00–_04 2.3 GB Main TTS model (FP32, sharded; shards listed in omnivoice-main-manifest.json)
omnivoice-decoder.onnx 83 MB Audio decoder (tokens to waveform)
omnivoice-encoder-fixed.onnx 624 MB Audio encoder for voice cloning

Shared

File Size Description
omnivoice-config.json <1 KB Model config
tokenizer.json, tokenizer_config.json 11 MB Qwen2 BPE text tokenizer

Usage

These files are made for VocoLoco, a fully client-side TTS application (live). No server required.

Architecture

  • Backbone: Qwen3-0.6B (28 transformer layers)
  • Audio codec: HiggsAudioV2 (8 codebooks, 24 kHz output)
  • Generation: iterative masked diffusion (8-32 steps)
  • Voice cloning: zero-shot via reference audio encoding
  • Voice design: text-based control (gender, pitch, accent)

License

The files carry the licences of the models they are converted from:

Files Converted from Licence
engine/omnivoice-f16.*, omnivoice-main-split.onnx + data shards, omnivoice-config.json OmniVoice main model by Xiaomi Corp. (k2-fsa) CC BY-NC (non-commercial)
engine/decoder-f32.*, engine/encoder-f32.*, omnivoice-decoder.onnx, omnivoice-encoder-fixed.onnx Higgs Audio v2 tokenizer by Boson AI, as shipped in OmniVoice's audio_tokenizer/ Boson Higgs Audio 2 Community License (LICENSE-higgs-audio-2), based on the Meta Llama 3 Community License (LICENSE-llama-3); see NOTICE
engine/whisper-i8.* Whisper large-v3-turbo by OpenAI MIT (LICENSE-whisper)
tokenizer.json, tokenizer_config.json Qwen3 tokenizer (Qwen2 BPE) by Alibaba Cloud, as used by OmniVoice Apache 2.0 (LICENSE-qwen)

Built with Higgs Materials licensed from Boson AI USA, Inc., Copyright Boson AI USA, Inc., All Rights Reserved and Meta Llama 3 licensed under the Meta Llama 3 Community License, Copyright Meta Platforms, Inc., All Rights Reserved.

Use of the audio tokenizer files must follow the Llama 3 Acceptable Use Policy; products or services with more than 100,000 annual active users need an expanded licence from Boson AI.

Licence history (OmniVoice main model): The ONNX exports were uploaded on 2026-04-09 (commit), when the OmniVoice model card listed the weights as Apache 2.0 (card as of 2026-04-05, license: apache-2.0). On 2026-07-03 the OmniVoice authors relicensed the pre-trained model as CC BY-NC, citing their training data (model card change). The weight files themselves were not changed. This repository follows the current upstream licence; the engine files were converted after the change.

Attribution

Based on OmniVoice by Xiaomi Corp (k2-fsa), the Higgs Audio v2 tokenizer by Boson AI and Whisper by OpenAI.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gigsu/vocoloco-onnx

Finetuned
Qwen/Qwen3-0.6B
Finetuned
k2-fsa/OmniVoice
Quantized
(40)
this model