Instructions to use selorahomes/Selora-AI-LLM-14B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use selorahomes/Selora-AI-LLM-14B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S # Run inference directly in the terminal: llama cli -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S # Run inference directly in the terminal: llama cli -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S # Run inference directly in the terminal: ./llama-cli -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Use Docker
docker model run hf.co/selorahomes/Selora-AI-LLM-14B:Q5_K_S
- LM Studio
- Jan
- vLLM
How to use selorahomes/Selora-AI-LLM-14B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "selorahomes/Selora-AI-LLM-14B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "selorahomes/Selora-AI-LLM-14B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/selorahomes/Selora-AI-LLM-14B:Q5_K_S
- Ollama
How to use selorahomes/Selora-AI-LLM-14B with Ollama:
ollama run hf.co/selorahomes/Selora-AI-LLM-14B:Q5_K_S
- Unsloth Desktop
- Pi
How to use selorahomes/Selora-AI-LLM-14B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "selorahomes/Selora-AI-LLM-14B:Q5_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use selorahomes/Selora-AI-LLM-14B with Docker Model Runner:
docker model run hf.co/selorahomes/Selora-AI-LLM-14B:Q5_K_S
- Lemonade
How to use selorahomes/Selora-AI-LLM-14B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull selorahomes/Selora-AI-LLM-14B:Q5_K_S
Run and chat with the model
lemonade run user.Selora-AI-LLM-14B-Q5_K_S
List all available models
lemonade list
- Hermes Agent
How to use selorahomes/Selora-AI-LLM-14B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default selorahomes/Selora-AI-LLM-14B:Q5_K_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use selorahomes/Selora-AI-LLM-14B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf selorahomes/Selora-AI-LLM-14B:Q5_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "selorahomes/Selora-AI-LLM-14B:Q5_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Selora Homes: selorahomes.com Selora AI Home Assistant Integration: github.com/SeloraHomes/ha-selora-ai
Selora AI LLM 14B
Selora AI LLM is an instruction-tuned language model for Home Assistant: five small LoRA specialists on one shared base, hot-swapped per request by the Selora AI integration. This is the larger line, on Qwen3 14B, for deployments with about 16 GB of memory. The 1.7B line runs on far less.
- command โ control devices ("turn off the kitchen lights")
- automation โ author home automations and blueprints
- answer โ answer questions about the home's live state
- clarification โ ask a follow-up when a request is ambiguous
- utilities โ docs-grounded maintenance, troubleshooting, and setup help
Both lines are trained on the same corpus, against the same system prompts, and emit the same compact JSON slim envelopes, so everything the 1.7B card documents about the specialists, their output shapes, and setting the model up in Home Assistant applies here unchanged. This card covers what differs: the files, the memory, and the measured scores.
A single fused model for runtimes that cannot swap adapters is published separately as selorahomes/Selora-AI-LLM-14B-ollama.
Base: Qwen3 14B ยท Format: GGUF Q5_K_S base + 5 per-specialist LoRA adapters (F16) ยท License: Apache-2.0
Results
Measured on the open allenporter/home-assistant-datasets suite โ the harness behind the Home Assistant LLM leaderboard โ at temperature 0, over the full suite, with llama.cpp on CUDA at the shipping quant.
Two modes are reported, because they measure different things:
- model only โ the integration's deterministic overrides are switched off, so every turn is decided by the model's own output. This is the number a retrain moves.
- net + model โ those overrides in the loop. Close to the shipped path, not
identical: approval and the safe-command allowlist are off for the run, and
questionsreads high on calendar cases, whose fixtures supply event data a real Home Assistant does not.
| dataset | cases | model only | net + model |
|---|---|---|---|
| assist | 460 | 78.3% [74.3, 81.8] | 88.0% [84.8, 90.7] |
| assist-mini | 196 | 93.9% [89.6, 96.5] | 100.0% [98.1, 100.0] |
| questions | 370 | 47.6% [42.6, 52.7] | 54.1% [49.0, 59.1] |
| automations | 20 | 100.0% [83.9, 100.0] | 100.0% [83.9, 100.0] |
95% Wilson intervals. automations is four blueprint tasks, five samples each,
scored separately because the harness collects those replies without checking
them: a reply passes if Home Assistant accepts it as an automation blueprint
with every requested input declared and wired. The automation is not run.
The fused single-model build scores lower, most of all with the overrides off; its card has the numbers.
Files
| file | role | size |
|---|---|---|
qwen3_14b_base.Q5_K_S.gguf |
shared base, Qwen3 14B at Q5_K_S | 10.3 GB (9.6 GiB) |
selora-command.f16.gguf |
request โ Home Assistant service calls | 51 MB |
selora-automation.f16.gguf |
automations, including blueprint YAML | 96 MB |
selora-answer.f16.gguf |
answers about the home's live state | 39 MB |
selora-clarification.f16.gguf |
asks when a request is ambiguous | 26 MB |
selora-utilities.f16.gguf |
documentation lookup with citations | 51 MB |
Modelfile.{intent} |
Ollama recipe per specialist (base + LoRA + system prompt) | โ |
prompts/{intent}_system_prompt.txt |
the system prompt each adapter was trained against | โ |
manifest.json |
version, per-file size and sha256, training recipe | โ |
LICENSE / NOTICE |
Apache 2.0 text and the Qwen3 attribution | โ |
The five specialists share one base, so the whole set costs about 10.5 GB of weights rather than five separate 14B models.
Requirements
| memory | ~16 GB for the model, KV cache and Home Assistant together |
| serving | llama.cpp with LoRA support (llama-server), or Ollama 0.30+ per specialist |
| context | 4096 tokens โ the budget the corpus and manifest.json (runtime.ctx_size) are built to |
Why Q5_K_S
The quant is chosen to fit a 16 GB deployment with the adapters, the KV cache and Home Assistant running alongside. Q4_K_M is smaller, but measurably hurt documentation-lookup accuracy.
Running it
llama-server (Home Assistant integration runtime)
Download the base and all five adapters into one directory, then:
llama-server \
--model qwen3_14b_base.Q5_K_S.gguf \
--lora-init-without-apply \
--lora selora-command.f16.gguf,selora-automation.f16.gguf,selora-answer.f16.gguf,selora-clarification.f16.gguf,selora-utilities.f16.gguf \
--ctx-size 4096
One --lora with comma-separated paths, not one flag per adapter: llama.cpp
keeps only the last value when --lora is repeated, so the per-flag form loads
a single adapter and silently answers every intent with it.
POST to /lora-adapters to switch the active adapter before each
/v1/chat/completions call. Resolve slot ids by filename from the
/lora-adapters listing rather than assuming an order.
Ollama
From the same directory, one named model per specialist:
ollama create selora-14b-command -f Modelfile.command
ollama create selora-14b-automation -f Modelfile.automation
ollama create selora-14b-answer -f Modelfile.answer
ollama create selora-14b-clarification -f Modelfile.clarification
ollama create selora-14b-utilities -f Modelfile.utilities
Ollama does not hot-swap LoRAs in-process, so this path suits direct chat and
scripting; for the Home Assistant integration use llama-server. For a single
model instead of five, use the
fused build.
Prompt format
Training and inference must match byte-for-byte. ChatML, with a /no_think
prefix on the user turn so Qwen3 returns bare JSON:
<|im_start|>system
{system prompt}<|im_end|>
<|im_start|>user
/no_think {user message}<|im_end|>
<|im_start|>assistant
<think>
</think>
The empty <think></think> block is part of the prompt, not the output, and
training includes it (masked). The chat template embedded in these GGUFs always
emits it, so llama.cpp and Ollama supply it whatever the client asks for. The
user message carries USER REQUEST,
EXISTING AUTOMATIONS, IMPORTANT, and AVAILABLE ENTITIES, in that order;
utilities adds RELEVANT DOCS. The
1.7B card
has the details.
Generation parameters
{
"temperature": 0.0,
"repeat_penalty": 1.0,
"repeat_last_n": 256,
"max_tokens": 384,
"stop": ["<|im_end|>", "<|endoftext|>"]
}
Bump max_tokens to 1536 for automation requests.
Training
Base: Qwen3 14B โ the post-trained
checkpoint, not Qwen3-14B-Base. Fine-tuned with
PEFT on CUDA as QLoRA (4-bit base,
paged_adamw_8bit, learning rate 5e-5, effective batch 8, sequence length
4096). Each specialist is its own LoRA, trained on a synthetic Home Assistant
corpus generated from curated home specifications and a service matrix โ no
scraped forum or user data โ with loss on the assistant turn only.
manifest.json records the exact recipe of the release you are on: steps,
epochs and example count per specialist, held-out eval loss, and the software
stack each adapter was trained with.
Versioning
main is the latest release. Each release is also tagged v<x.y.z>, and
manifest.json records the version, per-file checksums and the training recipe
for whatever revision you are on.
License
Apache 2.0, inherited from Qwen3-14B.
See LICENSE and NOTICE.
- Downloads last month
- 430
5-bit
16-bit