Instructions to use litert-community/Phi-4-mini-reasoning with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/Phi-4-mini-reasoning with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/Phi-4-mini-reasoning \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/Phi-4-mini-reasoning with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Phi-4-mini-reasoning β LiteRT-LM (blockwise int4)
LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).
Measured on device (edge-compat): Galaxy S26 Β· LiteRT-LM prebuilt-adac974b Β· GPU Β· decode 13.1 tok/s Β· prefill 321 tok/s Β· TTFT 3.26 s Β· all 1706 ops delegated (2026-10-05); Mac Studio M4 Max Β· LiteRT-LM 0.17.1 Β· GPU Β· decode 114.9 tok/s Β· prefill 1254 tok/s Β· TTFT 213 ms Β· all 1706 ops delegated (2026-10-05); Galaxy S26 Β· LiteRT-LM prebuilt-adac974b Β· CPU Β· decode 7.0 tok/s Β· prefill 52 tok/s Β· TTFT 20.82 s (2026-10-05). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/phi4-mini-reasoning/CARD.md
Phi-4-mini-reasoning β LiteRT-LM (blockwise int4)
microsoft/Phi-4-mini-reasoning converted to
the LiteRT-LM (.litertlm) format for on-device inference with Google's
LiteRT-LM runtime (the engine behind the official
litert-community/* models).
Phi-4-mini-reasoning is a dense 3.8B math/reasoning model from Microsoft (implemented as
Phi3ForCausalLM, 32 layers) β it solves problems with a <think>β¦</think> chain-of-thought, then
the answer.
| File | model.litertlm β int4 block 32 (~2.6 GB) |
| Quantization | int4 weights (symmetric) + OCTAV optimal-clipping; embeddings INT8 (externalized section) |
| Compute | integer |
| Context (KV cache) | 4096 |
| Base model | microsoft/Phi-4-mini-reasoning |
2026-10-06: GPU graph rewrite. Only the prefill/decode graph of model.litertlm changed. Its two DYNAMIC_UPDATE_SLICE KV-cache writes per layer are now one STABLEHLO_COMPOSITE odml.cache_update. Its BATCH_MATMUL(adjY) attention products are now odml.runtime_bmm. A signature input param_tensor INT32[1,1,1,7] was added; the runtime fills start/end. A graph exported with litert-torch's --apply_gpu_composites has this same shape, and the previous file was exported without that flag. The attention mask is applied as an ADD on a [bk, g, T, C] view of the logits, with the FLOAT32 mask broadcast; the mask input stays FLOAT32. This file has no 16-token prefill signature: a version with one lost final answers on the Mac GPU with fp16 activations (8 questions 7/8 instead of 8/8, GSM8K 26/30 instead of 28/30). Every other bundle section and the graph's weight region are byte-identical to the previous file (checked per section and per buffer). The new file is 2,784,187,376 bytes, sha256 d507262ca3d01d93d59c573a82e6f2eeb3aec21d9a67c31fb4c0f93d17225b8f. The previous file was 2,783,974,384 bytes, sha256 0e6b9daeeb45dc4b9d5ce65d09985a6ad8e861418f00fa658f74ea200976225f. With fp32 activations on the Mac GPU, the new file's answers are byte-identical to the previous file's on 8 of 8 questions. With the default fp16 activations, the new file also answered one long prompt correctly. On the Mac CPU, the new file's decode and prefill are Γ1.000 and Γ1.010 of the previous file's. The speed rows below ran both files in the same window, with the Mac protocol under Performance. The Raspberry Pi 5, Pixel 8a and iPhone 17 Pro rows further down were measured on the previous file.
| Mac Studio M4 Max, GPU, fp16 activations | previous file | new file |
|---|---|---|
| 8 questions, correct final answers | 8/8 | 8/8 |
| GSM8K, 30 questions (max tokens 2048, greedy) | 28/30 | 28/30 |
| Prefill, 256-token prompt | 1158 tok/s | 1254 tok/s |
| Decode, 256 tokens | 80.6 tok/s | 114.9 tok/s |
| TTFT, 16-token prompt | 0.130 s | 0.113 s |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
β οΈ It's a reasoning model β give it room to think
This model emits a <think>β¦</think> chain-of-thought, then a \boxed{} answer. Run it with
max_tokens β₯ 2048 β at a short limit it gets cut off before the answer. (All quality numbers below
were measured at 2048.)
Performance
Measured on the new file (see the 2026-10-06 note).
Apple M4 Max (macOS)
Mac Studio M4 Max, litert-lm 0.17.1 CLI, benchmark --cache no, 256 prompt and 256 decode tokens. GPU = WebGPU on Metal with fp16 activations, median of 3 processes. CPU = XNNPACK, median of 2 processes. The TTFT (16-token prompt) column comes from separate runs with a 16-token prompt and 32 decode tokens.
| Backend | Prefill (256) | Decode | TTFT | TTFT (16-token prompt) | Init |
|---|---|---|---|---|---|
| GPU (WebGPU on Metal) | 1254 tok/s | 114.9 tok/s | 0.21 s | 0.113 s | 3.2 s |
| CPU (XNNPACK) | 116 tok/s | 19.7 tok/s | 2.70 s | 1.596 s | 2.4 s |
Galaxy S26 (Android)
Samsung Galaxy S26 (SM-S942Q, Android 16), litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18. GPU = OpenCL delegate with the bundle's default fp16 activations, cold init (--disable_cache=true), 3 iterations in one process. CPU = XNNPACK with 4 threads and the XNNPACK weight cache, 2 iterations in one process. Each range runs from the lowest to the highest iteration in one process; the first GPU iteration follows the cold init. Peak is the process high-water mark (VmHWM).
| Backend | Prompt / decode tokens | Prefill | Decode | TTFT | Init | Peak (VmHWM) |
|---|---|---|---|---|---|---|
| GPU (OpenCL) | 1024 / 256 | 158.9β321.6 tok/s | 12.77β19.14 tok/s | 3.24β6.52 s | 15.36 s | 1.57 GB |
| CPU (XNNPACK, 4 threads) | 1024 / 256 | 41.3β61.9 tok/s | 6.83β7.25 tok/s | 16.68β24.95 s | 2.41 s | 4.28 GB |
Accuracy note
Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought, max_tokens 2048, identical prompt and answer-extraction for every row).
| Configuration | GSM8K |
|---|---|
| bf16 (reference) | 89.0% |
| LiteRT int4 β block 32 | 81.0% (β8 pt) |
int4 (block 32) is at parity (β8 pt). Why block 32 (not block 128)? This is a precision-sensitive math model: the coarser block-128 int4 dropped to 74% (β15 pt) and degenerated on some prompts, while block 32 holds at 81%. So only the block-32 build is published.
Galaxy S26 β GPU and CPU
The new file runs on the Galaxy S26 GPU backend, and 5 of 5 gate prompts were answered correctly. The OpenCL GPU delegate (LITERT_CL) takes every node of both transformer signatures, in one partition each:
| signature | nodes on LITERT_CL |
|---|---|
prefill_128 |
1706 / 1706 |
decode |
1581 / 1581 |
XNNPACK takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128. The Mac GPU (WebGPU) delegates the same node counts. Measured on a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18; GPU and CPU speed on this phone is in the Performance tables above.
GPU wiring, including the Gallery import toggle: GPU guide.
Pixel 8a β CPU and GPU (previous file)
Measured on the previous file (sha256 0e6b9daeeb45dc4b9d5ce65d09985a6ad8e861418f00fa658f74ea200976225f) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical and the new file's Mac CPU decode and prefill are Γ1.000 and Γ1.010 of the previous file's, so the CPU row is expected to hold, while the GPU row is not comparable because the graph changed.
On a Google Pixel 8a (Tensor G3, 8 GB, Android 16), measured 2026-09-05 with litert_lm_advanced_main built from the litert-lm v0.16.0 tree (2026-08-17 build), --benchmark --sampler_backend=cpu, a real 499-token prompt read from a file (--input_prompt_file; 511 tokens after the chat template) and --max_output_tokens=256. Every listed run is shown and the median reported; init time is excluded.
| backend | prefill tok/s (511 tokens) | decode tok/s (256 tokens) | time to first token | runs (prefill ; decode) |
|---|---|---|---|---|
| CPU (XNNPACK) | 19.4 | 3.77 | 29.70 s | 13.12 / 25.66 ; 3.51 / 4.03 |
| GPU (LiteRT GPU / OpenCL) | 76.6 | 6.83 | 6.81 s | 75.89 / 77.33 ; 6.75 / 6.91 |
On this handset the GPU backend is faster at prefill and faster at decode than the CPU backend for this bundle (6.8 vs 3.8 tok/s decode). One further Pixel 8a CPU run ended without a benchmark report (the adb session dropped mid-run) and is not listed.
Usage
# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
--model_path model.litertlm \
--backend gpu \
--input_prompt "A bat and a ball cost \$1.10. The bat costs \$1.00 more than the ball. How much is the ball?"
The .litertlm bundle carries the tokenizer and prompt template (Phi format β <|user|>β¦<|end|><|assistant|>),
so no separate tokenizer files are needed.
Run on Android
Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β no computer or
adbneeded. The manual steps below are only required on older builds or for sideloading a local file.
The official Google AI Edge Gallery app runs
.litertlm models on-device:
- Install a recent Gallery (package
com.google.ai.edge.gallery, 1.0.15+ supports.litertlm). - Download
model.litertlmand push it:adb push model.litertlm /sdcard/Download/ - In the app tap +, pick the file, choose the GPU backend, and raise the max-tokens setting (β₯2048).
Run on desktop (LiteRT-LM CLI)
The same .litertlm bundle runs on macOS / Linux / Windows with the official
LiteRT-LM CLI β including as a
local OpenAI-compatible API server:
pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/Phi-4-mini-reasoning model.litertlm phi-4-mini-reasoning
litert-lm run phi-4-mini-reasoning # interactive chat in the terminal
litert-lm serve
Run on iPhone
Verified on iPhone 17 Pro (LiteRT-LM Swift runtime): loads and generates correct answers (previous file). This is a ~2.6 GB bundle (Phi's 200K-token vocab makes a large externalized embedder), so it sits near the iOS memory ceiling β if you hit "embedding lookup model is not initialized" (a low-memory symptom), reboot the phone to free RAM and reload.
Conversion
Converted with the official litert-torch
converter. Phi-4-mini uses the Phi3ForCausalLM arch with LongRoPE + a (nominal) sliding window;
two export-time adjustments are needed for current litert-torch:
- LongRoPE: replace
Phi3RotaryEmbedding.forwardwith a static version (the@dynamic_rope_updateseq-len branch is data-dependent under torch.export; for cache β€ original_max=4096 the short factor is always correct). - Sliding window: set
config.sliding_window=None(it is 262144 β« context, i.e. full-causal) so the standard causal mask path is used.
Recipe: blockwise-32 int4 + OCTAV, embeddings INT8, KV cache 4096, externalize_embedder=True.
Graph rewrite (2026-10-06). The current file was made from the previous one (Hub revision ea890746c4f4) with gpu_graph_retrofit.py from tools/gpu_graph/ in hf-to-litertlm. It makes the graph changes listed in the 2026-10-06 note; every original weight buffer keeps its bytes. The 16-token prefill copy (prefill_bucket_clone.py) is not applied to this file.
python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
previous.litertlm is the previous file, retrofit.litertlm the result published here as model.litertlm; LITERT_LM_CLI names the litert-lm CLI the script uses to unpack and pack the bundle.
A rebuild with this command gives every bundle section byte-identical to this file (sha256 per section). Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.
2026-08-28 β start_token fix (weights unchanged)
The LiteRT-LM engine prepends the metadata start_token to every prompt, but this model's reference prompt has no leading BOS at all β the bundle was feeding an extra <|endoftext|> the model was never trained on. The start token has been removed.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file β only the file's own sha256 differs. What changed is the input: the token stream the model reads for a given conversation can differ from the previous file's, and it now matches this model's own reference chat stream. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file in either direction. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-28, re-download.
2026-08-29 β default system prompt restored (weights unchanged)
The upstream chat template emits a default system turn whenever the caller sends no system message β for this model: Your name is Phi, an AI math expert developed by Microsoft.. The converter's template probe renders the template with a system message already present, so that block was never seen and never reached the bundle: with no system message the model was running without the default system turn it was tuned with. The chat template in model.litertlm now emits the block exactly once when no system message is given. In model.litertlm, the block is not emitted when you pass a system message; the upstream template also wraps a caller's system message in its own preamble, and this file passes it through unchanged, exactly as it did before. The restored block adds 15 prefill tokens to a conversation that sends no system message, so time-to-first-token grows by that much; per-token speed is unchanged.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file per token β only the file's own sha256 differs. What changed is the input: with no system message, the prompt now renders byte-identical to the upstream chat template's output, verified on the LiteRT-LM runtime. A system message you pass yourself renders as before. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-29, re-download.
2026-08-31 β thought channel declared (metadata only, weights unchanged)
model.litertlm now declares the reasoning channel in its metadata (LlmMetadata.channels: channel name thought, markers <think>β¦</think> exactly as this model emits them). Without the declaration the runtime has no way to tell the reasoning apart from the answer: the raw thinking streamed inline into the visible text, and a thinking_token_budget was silently ignored (the API returns OK and only logs a warning). With the channel declared, LiteRT-LM returns the reasoning separated in channels["thought"] and the thinking budget takes effect.
Metadata-only change: every section of the bundle except LlmMetadata is byte-identical to the previous file (verified per section, tokenizer included), so the weights, the graph, the tokenizer and the chat template are unchanged and the speed and accuracy numbers on this card still describe this file β only the file's own sha256 differs. Verified on the LiteRT-LM runtime (litert-lm-api 0.16.1): the visible answer stays clean, the reasoning lands in channels["thought"], and on one bundle of this batch thinking_token_budget=16 was confirmed to truncate the reasoning at exactly 16 tokens where it was a no-op before. If you downloaded before 2026-08-31, re-download to get the channel-aware file.
Raspberry Pi 5 (CPU) (previous file)
Measured on the previous file (sha256 0e6b9daeeb45dc4b9d5ce65d09985a6ad8e861418f00fa658f74ea200976225f) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical and the new file's Mac CPU decode and prefill are Γ1.000 and Γ1.010 of the previous file's, so the throughput columns are expected to hold.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (minβmax in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
model.litertlm |
14.1 (13.9β14.2) | 1.6 (1.6β1.6) | 21.1 s | 4.2 GB |
License
MIT, inherited from the base model microsoft/Phi-4-mini-reasoning.
- Downloads last month
- 915
Model tree for litert-community/Phi-4-mini-reasoning
Base model
microsoft/Phi-4-mini-reasoning