Instructions to use AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor") model = AutoModelForCausalLM.from_pretrained("AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor
- SGLang
How to use AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor with Docker Model Runner:
docker model run hf.co/AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor
Model Summary
Qwen3-4B-Instruct-2507-Insight-Extractor is a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507. It converts a raw customer-support ticket into a single structured JSON object across 9 fields, for automated ticket classification and downstream entity-extraction workflows.
| Base model | Qwen/Qwen3-4B-Instruct-2507 |
| Fine-tuning method | LoRA Adapters via Unsloth + TRL SFTTrainer |
| Task | Structured JSON extraction from support-ticket text |
| Training data | AbhishekG711/processed_support_tickets (train split) |
| Target repo | AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor |
Base Model Architecture (Qwen3-4B-Instruct-2507)
| Type | Dense, causal decoder-only Transformer |
| Parameters | 4.0B total (~3.6B non-embedding) |
| Layers | 36 |
| Attention | Grouped-Query Attention — 32 query heads / 8 KV heads |
| Native context length | 262,144 tokens |
| Reasoning modes | Non-thinking only. Unlike the 0.6B and 1.7B checkpoints, the "-Instruct-2507" line was released specifically without a thinking mode — it never emits <think>...</think> blocks. |
| Base model license | Apache 2.0 |
Intended Use
Given a raw support ticket, produce one JSON object for automated triage, routing, sentiment/urgency dashboards, or downstream alerting — not a conversational assistant.
Fine-Tuning Training Details
| Hyperparameter | Value |
|---|---|
| Quantization | None (load_in_4bit=False) |
LoRA rank (r) |
32 |
| LoRA alpha | 64 (= 2 × r) |
| LoRA dropout | 0.0 |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| LoRA bias | none |
| Max sequence length | 4096 (intentional, confirmed — see note above) |
| Per-device train batch size | 4 |
| Target global batch size | 32 |
| Gradient accumulation steps | 8 (= 32 // 4) |
| Epochs | 1 |
| Learning rate | 2e-4 |
| LR scheduler | linear |
| Warmup steps | 10 |
| Weight decay | 0.01 |
| Optimizer | adamw_8bit |
| Precision | bf16 if supported, else fp16 (auto-detected) |
| Completion-only loss | True |
| Sequence packing | False |
| Eval strategy | every 25 steps |
| Save strategy | every 50 steps, save_total_limit=3 |
| Load best model at end | True (metric_for_best_model="eval_loss", greater_is_better=False) |
| Gradient checkpointing | "unsloth" mode |
| Seed | 3407 |
| Experiment tracker | report_to="none" |
Evaluation
| Metric | Value |
|---|---|
Validation loss (eval_loss, best checkpoint) |
0.222669 |
How to Prompt This Model
Same system/user structure as the rest of the Insight-Extractor family:
System prompt:
# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.
User message template:
Analyze the following customer support email and extract structured insights.
<raw ticket text>
Expected output: one JSON object with exactly these 9 keys, in this order: is_actionable, summary, sentiment, category, intent, aspect, urgency, reported_cause, entities.
Inference example
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
SYSTEM_PROMPT = """
# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.
"""
USER_INSTRUCTION = "Analyze the following customer support email and extract structured insights.\n\n"
ticket_text = "My order #4471 arrived damaged and support hasn't replied in 3 days."
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": USER_INSTRUCTION + ticket_text},
]
# See the architecture note above: this base model is non-thinking-only, so do not assume
# enable_thinking behaves the same way here as it does for the 0.6B/1.7B checkpoints.
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
- Downloads last month
- 502
Model tree for AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor
Base model
Qwen/Qwen3-4B-Instruct-2507