Qwen2.5-Math-1.5B-Instruct for ExecuTorch
ExecuTorch exports of Qwen/Qwen2.5-Math-1.5B-Instruct (revision aafeb0fc6f22) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.
Files
Windows in this repo: XNNPACK (CPU) at 2k to 32k; Vulkan (GPU) at 2k to 32k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).
| Backend | Target | File | Window | Size | Smoke test |
|---|---|---|---|---|---|
| XNNPACK (CPU) | any arm64 | xnnpack/Qwen2.5-Math-1.5B-Instruct-8da4w-gptq-2k.pte |
2,048 tokens | 1.11 GB | passed ("Paris") |
| XNNPACK (CPU) | any arm64 | xnnpack/Qwen2.5-Math-1.5B-Instruct-8da4w-gptq-4k.pte |
4,096 tokens | 1.11 GB | passed ("Paris") |
| XNNPACK (CPU) | any arm64 | xnnpack/Qwen2.5-Math-1.5B-Instruct-8da4w-gptq-8k.pte |
8,192 tokens | 1.12 GB | passed ("Paris") |
| XNNPACK (CPU) | any arm64 | xnnpack/Qwen2.5-Math-1.5B-Instruct-8da4w-gptq-16k.pte |
16,384 tokens | 1.14 GB | passed ("Paris") |
| XNNPACK (CPU) | any arm64 | xnnpack/Qwen2.5-Math-1.5B-Instruct-8da4w-gptq-32k.pte |
32,768 tokens | 1.17 GB | passed ("Paris") |
| Vulkan (GPU) | any arm64 | vulkan/Qwen2.5-Math-1.5B-Instruct-vulkan-8da4w-2k.pte |
2,048 tokens | 1.40 GB | structure checked (no host NPU runtime) |
| Vulkan (GPU) | any arm64 | vulkan/Qwen2.5-Math-1.5B-Instruct-vulkan-8da4w-4k.pte |
4,096 tokens | 1.41 GB | structure checked (no host NPU runtime) |
| Vulkan (GPU) | any arm64 | vulkan/Qwen2.5-Math-1.5B-Instruct-vulkan-8da4w-8k.pte |
8,192 tokens | 1.42 GB | structure checked (no host NPU runtime) |
| Vulkan (GPU) | any arm64 | vulkan/Qwen2.5-Math-1.5B-Instruct-vulkan-8da4w-16k.pte |
16,384 tokens | 1.44 GB | structure checked (no host NPU runtime) |
| Vulkan (GPU) | any arm64 | vulkan/Qwen2.5-Math-1.5B-Instruct-vulkan-8da4w-32k.pte |
32,768 tokens | 1.49 GB | structure checked (no host NPU runtime) |
Tokenizer: tokenizer.json, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.
Memory
- XNNPACK (CPU) at 2,048 tokens: the KV cache costs 57,344 bytes per token (fp32), 117,440,512 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 4,096 tokens: the KV cache costs 57,344 bytes per token (fp32), 234,881,024 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 8,192 tokens: the KV cache costs 57,344 bytes per token (fp32), 469,762,048 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 16,384 tokens: the KV cache costs 57,344 bytes per token (fp32), 939,524,096 bytes for the whole window, allocated in full when the model loads.
- XNNPACK (CPU) at 32,768 tokens: the KV cache costs 57,344 bytes per token (fp32), 1,879,048,192 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 2,048 tokens: the KV cache costs 57,344 bytes per token (fp32), 117,440,512 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 4,096 tokens: the KV cache costs 57,344 bytes per token (fp32), 234,881,024 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 8,192 tokens: the KV cache costs 57,344 bytes per token (fp32), 469,762,048 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 16,384 tokens: the KV cache costs 57,344 bytes per token (fp32), 939,524,096 bytes for the whole window, allocated in full when the model loads.
- Vulkan (GPU) at 32,768 tokens: the KV cache costs 57,344 bytes per token (fp32), 1,879,048,192 bytes for the whole window, allocated in full when the model loads.
How it was made
- XNNPACK (CPU) any arm64: ExecuTorch 1.4.0
export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1. - Vulkan (GPU) any arm64: ExecuTorch 1.4.0
export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, the Vulkan delegate, prefill chunk 2048, fp32 KV cache. Built by run 1.
License
A quantized derivative of Qwen/Qwen2.5-Math-1.5B-Instruct, distributed under the same terms (apache-2.0).
The upstream license files are included unchanged: LICENSE.
- Downloads last month
- 113
Model tree for experimentalmachines/Qwen2.5-Math-1.5B-Instruct-ExecuTorch
Base model
Qwen/Qwen2.5-1.5B