πŸ‡ WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B-AWQ (4-bit INT4)

This repository provides an optimized 4-bit AWQ (Activation-aware Weight Quantization) release of WhiteRabbitNeo/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B.

By utilizing AWQ with GEMM kernels, this model reduces the original 15.2 GB FP16 footprint down to 5.58 GB total on disk, enabling deployment on cloud (Tesla T4, L4) and consumer GPUs (RTX 3060/4060) while maintaining core instruction-following and security-analysis capabilities.


⚑ Quantization & Model Specifications

  • Method: 4-bit AWQ (INT4)
  • Kernel: GEMM (Optimized for NVIDIA Tensor Cores)
  • Group Size (q_group_size): 128
  • Zero Point: True
  • Modules Kept in FP16: ["lm_head"] (unquantized to preserve precision across 152,064 vocabulary logits)
  • Calibration Dataset: pileval (general text)
  • Calibration Samples: 16 samples (max_calib_seq_len=256, apply_clip=False)
  • Files on Disk: 5.58 GB total (Shard 1: 4.48 GB, Shard 2: 1.09 GB)
  • VRAM Requirements:
    • Static Weights: ~5.5 GB
    • Active Execution (8K context): ~6.0 GB
    • Active Execution (Full 32K context): ~7.3–8.0 GB (GQA with 4 KV heads = ~56 KB/token KV cache, plus activation buffers)
    • Leaves plenty of headroom on 16 GB GPUs (like Colab Tesla T4) for long-context generation.

πŸ’‘ Calibration & Quality Note: This build was quantized with an agile calibration pass on general English text (pileval, 16 samples Γ— 256 tokens) to establish base quantization scales. Because pileval does not contain domain-specific exploit or software vulnerability corpora, users deploying for mission-critical security evaluations or esoteric language syntax are encouraged to validate domain performance on their specific evaluation suites.


πŸš€ Quick Start & Tested Environments

1. High-Throughput Serving with vLLM (Recommended)

Requires vllm>=0.6.0:

vllm serve Sharunkrish/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B-AWQ --max-model-len 8192

(vLLM automatically detects AWQ parameters from config.json and selects optimized GEMM/Marlin kernels).

2. Hugging Face Transformers (gptqmodel Backend)

gptqmodel provides modern ExLlamaV2/GEMM AWQ execution compatible with newer transformers:

pip install -U "transformers>=4.45.0" accelerate gptqmodel
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "Sharunkrish/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B-AWQ"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)

prompt = "Analyze this function for memory leaks and privilege escalation vectors:"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Note on autoawq: Standalone autoawq (autoawq>=0.2.6) is pinned to specific transformers 4.x versions and may fail on newer releases. If using recent transformers packages, gptqmodel or vLLM is recommended.


πŸ”¬ Compatibility & Configuration Details

  • RoPE Configuration: config.json includes both modern rope_parameters and legacy top-level "rope_theta": 1000000.0 / "torch_dtype": "bfloat16" to ensure full compatibility across both transformers 4.x and 5.x without silent frequency scaling degradation.
  • Architecture: Qwen2ForCausalLM (28 layers, 3584 hidden size, 18944 intermediate size, 28 attention heads / 4 KV heads, 32,768 max positions).
  • Chat Template: Ships with standard chat_template.jinja compatible with Qwen2 chat formatting.
Downloads last month
56
Safetensors
Model size
8B params
Tensor type
I32
Β·
BF16
Β·
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Sharunkrish/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B-AWQ