AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1
Viewer • Updated • 99.9k • 3.78k • 153
How to use Varadrajan/LFM2.5-1.2B-CyberSec-Fenrir with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-1.2B-Instruct")
model = PeftModel.from_pretrained(base_model, "Varadrajan/LFM2.5-1.2B-CyberSec-Fenrir")A LoRA adapter for LiquidAI/LFM2.5-1.2B-Instruct fine-tuned on the Fenrir v2.1 cybersecurity dataset (89,000 examples).
Transforms the base model into a cybersecurity knowledge assistant that:
Prompt: "Which MITRE ATT&CK technique covers credential dumping from LSASS memory?"
| Base Model | Fine-Tuned |
|---|---|
| ❌ Hallucinated T1059.001 ("Credential Dumping via LSASS Memory" — this ID doesn't exist) | ✅ Correct T1003.001 (OS Credential Dumping: LSASS Memory), mentions Mimikatz, Kerberos tickets |
Prompt: "EDR alert: powershell.exe spawned by winword.exe with encoded command. What's happening?"
| Base Model | Fine-Tuned |
|---|---|
| Generic advice: "check process details, review user activity" | Precise: T1059.001 (PowerShell) + T1027 (Obfuscated Files), process genealogy analysis, NIST CSF DE.AE-2 mapping |
Prompt: "What is CWE-79?"
| Base Model | Fine-Tuned |
|---|---|
| ❌ Wrong: "Use of Unnecessary Privileges" | ✅ Identifies input validation weakness, mentions XSS, maps to ATT&CK T1059/T1203, NIST CSF PR.DS-2 |
| Parameter | Value |
|---|---|
| Base model | LiquidAI/LFM2.5-1.2B-Instruct |
| Dataset | Fenrir v2.1 (89,000 decontaminated rows) |
| Method | LoRA SFT (rank 16, alpha 16) |
| Target modules | q_proj, k_proj, v_proj, out_proj, in_proj, w1, w2, w3 |
| Optimizer | AdamW 8-bit, lr=2e-4, linear schedule |
| Epochs | 1 |
| Effective batch size | 8 (micro=2 × grad_accum=4) |
| Max sequence length | 2,048 tokens |
| Train loss | 1.088 |
| Eval loss | 0.999 |
| Training time | 8h 47m on RTX 4060 Laptop (8GB) |
| Peak VRAM | 3.66 GB |
| Adapter size | ~22 MB |
Training followed Liquid AI's official LoRA recipe with train_on_responses_only masking.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "LiquidAI/LFM2.5-1.2B-Instruct"
adapter_id = "Varadrajan/LFM2.5-1.2B-CyberSec-Fenrir"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
base = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="cuda")
model = PeftModel.from_pretrained(base, adapter_id)
messages = [{"role": "user", "content": "What is CWE-79 and how do you prevent it?"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to("cuda")
output = model.generate(inputs, max_new_tokens=512, do_sample=False, repetition_penalty=1.05)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Base model
LiquidAI/LFM2.5-1.2B-Base