Instructions to use org2ai/Wald-26B-A4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-26B-A4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-26B-A4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("org2ai/Wald-26B-A4B") model = AutoModelForMultimodalLM.from_pretrained("org2ai/Wald-26B-A4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use org2ai/Wald-26B-A4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-26B-A4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-26B-A4B
- SGLang
How to use org2ai/Wald-26B-A4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-26B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use org2ai/Wald-26B-A4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-26B-A4B
Wald-26B-A4B
An open decision model built on Gemma 4 26B-A4B, a mixture of experts with 26B parameters and about 4B active per token. You send a state and typed options, and it returns a probability for every option. It is served through a Jev-compatible POST /v1/systemone API on your own GPU.
Why Wald-26B-A4B
- Probabilities, not text: you get a probability for every option and have nothing to parse.
- Decision Index 0.3: 62.30, complete public suite, one pass (our run with the official 0.3 kit). Jev 1.13 scores 57.96. Auto scores ≈ 65.9 (estimate from a 14 % sample; Auto exceeds the 1,000 ms admission latency, so one pass is the default).
- JevBench-XL TEST: 75.3 % (Jev 1.13: 67.85 %).
- JevBench public set: 219 / 231 (macro 95.3 %). Jev 1.13 gets 200 / 231.
- Optional thinking: an Auto mode can use Gemma 4's native thinking when the top probability is below 0.7; the default is one pass.
- Fast: median 53 ms in one pass, measured serially with one request in flight on a warm server with CUDA graphs. This is processing time on one RTX PRO 6000, over 200 Decision Index 0.3 requests.
- Open: Apache-2.0, BF16 weights (52 GB), served on one 96 GB GPU.
Scores
| Benchmark | Wald-26B-A4B | Jev 1.13 (hosted) |
|---|---|---|
| Decision Index 0.3, public suite | 62.30 (one pass, complete suite); Auto ≈ 65.9 (estimate*) | 57.96 (board public column) |
| JevBench public set (231) | 219 (macro 95.3 %) | 200 (macro 88.0 %) |
| JevBench-XL TEST (8,177) | 75.26 % | 67.85 % |
| newdomainTEST (3,125 held-out decisions) | 72.74 | — |
62.30 is our run of the complete public suite (140,620 requests, all answered) with the official 0.3 kit, in one pass, the default setting.
* Auto has not been run on the complete suite: it exceeds the 1,000 ms admission latency, so one pass is the default. The estimate comes from a fixed 14 % sample of 20,241 requests, read and paired item by item against a checkpoint whose complete-suite score we know (54.18), then projected to the full suite.
All scores are self-run and are not leaderboard results.
Quick start
hf download org2ai/Wald-26B-A4B --local-dir ./Wald-26B-A4B && cd Wald-26B-A4B
./run.sh "$PWD" # POST /v1/systemone on :8000, one pass (one 96 GB NVIDIA GPU, uv)
EFFORT=auto ./run.sh "$PWD" # optional: Auto 0.7 (thinks when unsure; slower)
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "The customer wants to return a damaged kettle.",
"questions": {"route": {"type": "choice", "instructions": "Choose the support queue.",
"criteria": {"returns": "Returns and refunds", "delivery": "Delivery tracking", "other": "Other"}}}}'
Requests and responses use the Jev-compatible systemone format (choice, noul and score questions). run.sh installs vLLM 0.30.0 and transformers 5.17.0 in a local uv environment. It starts vLLM on loopback with the evaluated arguments (serving.json), then the reader in reference/. At start-up it checks every reader file against reference/source-pins.json.
How it works
- One pass (default): a single forward pass reads a probability for every option from the model's option-letter logits, using Gemma 4's chat template. No temperature table is applied (T = 1).
- Auto 0.7 (optional): if the top probability is below 0.7, the model thinks once in Gemma 4's native thought channel (at most 512 tokens), then reads the options again.
- More than 26 options: a one-pass knockout over all options.
- Effort is set at server start. Unlike Wald-4B, it is not chosen per request. To offer both modes, run two readers on one vLLM.
- Text only. The vision tower is in the weights, unchanged from the base model, but the server runs vLLM with
--language-model-only. We have not evaluated image decisions.
Details
Training. This is a single full-parameter pass of decision training from Gemma-4-26B-A4B-it, with the router, embeddings and vision tower frozen. The checkpoint (
05901-c33) is the snapshot at two thirds of that pass, after 94.3M training tokens; the full pass covers about 142M. Thoughts are the base model's own: no thought text was trained.Data. About 17 % of training tokens come from the train splits of public benchmarks that the Decision Index also draws on, reformatted as decisions. The rest is broad decision data:
- task-specific decision sets, about 32 %;
- general decision tasks (rules and policies, classification, entailment and fact checking, tables and documents, judging), about 25 %;
- hard decisions written and/or labelled by large frontier models, including robustness cases with perturbed states, about 15 %;
- web and GUI actions, about 11 %;
- math and step-by-step reasoning, about 2 %.
Public datasets are used as train splits only. A contamination scan against the full Decision Index 0.3 public suite found 0 strict hits.
Limits. All scores are self-run. Confidence is not correctness, so set thresholds on your own data. Answers after a thought are sampled and can vary slightly; one pass is deterministic. Yes/no (
noul) probabilities inside (0.19, 0.81) are moved to the nearer edge by the decisive floor (see Serving), so the returned P(yes) of those answers is 0.81 or 0.19, not the value as read.Licence. Apache-2.0 for the weights and the code (NOTICE). The base model, Gemma 4, is also Apache-2.0. The smaller 4B model in the same family is org2ai/Wald-4B.
Serving
These are the evaluated runtime settings, which run.sh reproduces:
- vLLM 0.30.0 and transformers 5.17.0, BF16, on one NVIDIA RTX PRO 6000 (96 GB);
- environment:
VLLM_USE_FLASHINFER_SAMPLER=0,VLLM_USE_DEEP_GEMM=0,VLLM_MOE_USE_DEEP_GEMM=0; - vLLM arguments:
--max-model-len 131072 --gpu-memory-utilization 0.90 --max-num-seqs 128 --max-logprobs 64 --seed 0 --language-model-only(CUDA graphs on).
The reader is python -m eval.systemone_vllm (with PYTHONPATH=reference), run with these arguments:
- always:
--template chat --prompt-format plain --wide knockout --budget 512 --max-model-len 131072 --noul-decisive-floor 0.81; - one pass: add
--gate 0; - Auto 0.7: add
--gate 0.7 --native-think gemma --tokenizer <weights dir>.
Noul answers use a decisive floor of 0.81, as Wald-4B v2.1: a P(yes) strictly between 0.19 and 0.81 moves to 0.81 or 0.19 on its own side (exactly 0.5 goes to no), and the value before the floor is returned as noul_calibrated. Choice and score answers are unchanged. serving.json sets noul_decisive_floor; NOUL_DECISIVE_FLOOR=0 ./run.sh turns it off.
Thoughts are sampled at temperature 0.6, top-p 0.95 and top-k 20, with a seed derived from the request and the question id. A smaller GPU needs a lower --gpu-memory-utilization or --max-model-len, which departs from the evaluated setup.
- Downloads last month
- -
Model tree for org2ai/Wald-26B-A4B
Space using org2ai/Wald-26B-A4B 1
Collection including org2ai/Wald-26B-A4B
Evaluation results
- Macro accuracy (self-run) on JevBench public set (231 items)self-reported95.300