Instructions to use schnapper79/lumikabra-123B_v0.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use schnapper79/lumikabra-123B_v0.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="schnapper79/lumikabra-123B_v0.2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("schnapper79/lumikabra-123B_v0.2") model = AutoModelForCausalLM.from_pretrained("schnapper79/lumikabra-123B_v0.2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use schnapper79/lumikabra-123B_v0.2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "schnapper79/lumikabra-123B_v0.2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schnapper79/lumikabra-123B_v0.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/schnapper79/lumikabra-123B_v0.2
- SGLang
How to use schnapper79/lumikabra-123B_v0.2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "schnapper79/lumikabra-123B_v0.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schnapper79/lumikabra-123B_v0.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "schnapper79/lumikabra-123B_v0.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schnapper79/lumikabra-123B_v0.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use schnapper79/lumikabra-123B_v0.2 with Docker Model Runner:
docker model run hf.co/schnapper79/lumikabra-123B_v0.2
lumikabra-123B v0.2
This is lumikabra. It's based on Mistral-Large-Instruct-2407 , merged with Magnum-v2-123B, Luminum-v0.1-123B and Tess-3-Mistral-Large-2-123B.
I shamelessly took this idea from FluffyKaeloky. Like him, i always had my troubles with each of the current large mistral based models. Either it gets repetitive, shows too many GPTisms, is too horny or too unhorny. RP and storytelling is always a matter of taste, and i found myself swiping too often for new answers or even fixing them when I missed a little spice or cleverness.
Luminum was a great improvement, mixing a lot of desired traits, but I still missed some spice, another sauce. So i took Luminum, added magnum again and also Tess for knowledge and structure.
This is a second version with another mixture of the same sauce. It is different than v0.1, not worse not better, just a little different. Again, I believe it is just a matter of taste, which answers one prefers and like the most.
Quants
Merge Details
Merge Method
This model was merged using mergekit with the della_linear merge method using mistralai_Mistral-Large-Instruct-2407 as a base.
Configuration
The following YAML configuration was used to produce this model:
models:
- model: anthracite-org_magnum-v2-123b
parameters:
weight: 0.24
density: 0.5
- model: FluffyKaeloky_Luminum-v0.1-123B
parameters:
weight: 0.34
density: 0.8
- model: migtissera_Tess-3-Mistral-Large-2-123B
parameters:
weight: 0.24
density: 0.9
merge_method: della_linear
base_model: mistralai_Mistral-Large-Instruct-2407
parameters:
epsilon: 0.05
lambda: 1
int8_mask: true
dtype: bfloat16
- Downloads last month
- 15