neuralbroker commited on
Commit
b93d2d5
·
verified ·
1 Parent(s): 2414aad

Update MODEL_CARD.md (v2.1 production)

Browse files
Files changed (1) hide show
  1. MODEL_CARD.md +185 -172
MODEL_CARD.md CHANGED
@@ -1,185 +1,180 @@
1
  ---
2
  language:
3
- - en
 
4
  library_name: llama-cpp-python
5
  pipeline_tag: text-generation
6
  tags:
7
- - code-generation
8
- - coding-assistant
9
- - gguf
10
- - llama.cpp
11
- - qwen2.5
12
- - python
13
- - javascript
14
- - fine-tuned
 
 
15
  base_model:
16
- - Qwen/Qwen2.5-1.5B-Instruct
 
17
  ---
18
 
19
  # BlitzKode
20
 
21
- **BlitzKode** is a fine-tuned AI coding assistant built by **Sajad** using the Qwen2.5-1.5B base model. It's packaged as a GGUF format model for fast local inference with llama.cpp.
 
 
 
22
 
23
- > Created by [Abdulla Sajad](https://github.com/neuralbroker)
24
- > Project: [neuralbroker/blitzkode](https://github.com/neuralbroker/blitzkode)
 
 
25
 
26
  ---
27
 
28
- ## Model Summary
29
 
30
- | Property | Value |
31
- |----------|-------|
32
- | **Model Name** | BlitzKode |
33
- | **Version** | 2.0 |
34
- | **Base Model** | Qwen/Qwen2.5-1.5B-Instruct |
35
- | **Model Format** | GGUF (F16, ~3GB) |
36
- | **Primary Runtime** | llama.cpp / llama-cpp-python |
37
- | **Artifact** | `blitzkode.gguf` |
38
- | **Context Window** | 2048 tokens |
39
- | **Creator** | Sajad |
40
- | **License** | MIT (also see Qwen2.5 upstream license) |
41
 
42
  ---
43
 
44
  ## Architecture
45
 
46
- - **Model Type**: Transformer-based LLM (1.5B parameters)
47
- - **Architecture**: Qwen2
48
- - **Quantization**: GGUF F16 (~3GB)
49
- - **Vocabulary**: 151,936 tokens
50
- - **Inference**: CPU/GPU with llama.cpp (configurable via BLITZKODE_GPU_LAYERS)
 
 
 
 
 
51
 
52
  ---
53
 
54
  ## Training Pipeline
55
 
56
- BlitzKode was fine-tuned through a 4-stage pipeline:
57
 
58
- ### 1. SFT (Supervised Fine-Tuning)
59
- Applies LoRA fine-tuning to coding-style prompts and responses using PEFT library.
 
 
60
 
61
- ### 2. Reward-based SFT continuation
62
- Applies additional SFT with heuristic reward functions for code correctness, formatting, and reasoning. Note: this stage uses standard SFT training, not a full GRPO implementation.
63
 
64
- ### 3. DPO (Direct Preference Optimization)
65
- Trains on handcrafted preference pairs to improve clarity and answer quality.
 
 
66
 
67
- ### 4. Merge & Export
68
- Merges LoRA adapters into base model and converts to GGUF format.
69
 
70
- ### Training Frameworks
71
- - HuggingFace Transformers
72
- - PEFT (LoRA)
73
- - TRL (DPO/GRPO)
74
- - llama.cpp (inference/export)
75
 
76
- ---
 
77
 
78
- ## Training Data
 
 
 
79
 
80
- Custom curated coding datasets covering:
81
- - Algorithm implementation
82
- - Data structures
83
- - Code explanations
84
- - Programming concepts
85
- - Bug fixing scenarios
86
 
87
- ---
 
 
88
 
89
- ## Features
90
-
91
- - **Multi-language Code Generation** - Python, JavaScript, Java, C++, TypeScript, SQL
92
- - **Code Explanation** - Clear comments and documentation
93
- - **Bug Fixing** - Debug and fix code issues
94
- - **Algorithm Assistance** - Data structures and algorithms
95
- - **Offline Operation** - Runs locally without internet
96
- - **Fast Inference** - Optimized CPU inference
97
- - **Modern UI** - Professional dark interface
98
 
99
  ---
100
 
101
- ## Intended Use
102
 
103
- ### Best For
104
- - Local offline coding assistance
105
- - Algorithm and data structure help
106
- - Code generation and explanation
107
- - Educational programming support
108
- - Code review and debugging
109
 
110
- ### Out of Scope
111
- - Production code without expert review
112
- - Security-critical applications
113
- - Multi-modal tasks (images not supported)
114
- - Long-context repository analysis
 
115
 
116
- ---
 
117
 
118
- ## API & Usage
119
 
120
- ### Running the Server
121
 
122
- ```bash
123
- # Install dependencies
124
- pip install llama-cpp-python fastapi uvicorn pydantic
 
 
 
 
 
125
 
126
- # Start server
127
- python server.py
128
 
129
- # Open browser
130
- # http://localhost:7860
131
- ```
132
 
133
- ### API Endpoints
134
 
135
- | Endpoint | Method | Description |
136
- |----------|--------|-------------|
137
- | `/` | GET | Web UI |
138
- | `/health` | GET | Health check |
139
- | `/info` | GET | API info |
140
- | `/generate` | POST | Generate response |
141
- | `/generate/stream` | POST | Stream tokens |
142
 
143
- ### API Example
 
144
 
145
- ```bash
146
- # Generate code
147
- curl -X POST http://localhost:7860/generate \
148
- -H "Content-Type: application/json" \
149
- -d '{"prompt": "Write hello world in python"}'
150
  ```
151
 
152
- ### Python Usage
153
 
154
  ```python
155
- from llama_cpp import Llama
 
156
 
157
- llm = Llama(
158
- model_path="blitzkode.gguf",
159
- n_ctx=2048,
160
- n_threads=8,
161
- )
162
 
163
- prompt = """<|im_start|>system
164
- You are BlitzKode, a coding assistant.<|im_end|>
165
- <|im_start|>user
166
- Write a hello world in Python<|im_end|>
167
- <|im_start|>assistant
168
- """
169
-
170
- result = llm(prompt, max_tokens=256)
171
- print(result["choices"][0]["text"])
172
  ```
173
 
174
- ---
175
-
176
- ## Prompt Format
177
 
178
- Uses ChatML-style template:
179
 
180
  ```
181
  <|im_start|>system
182
- You are BlitzKode, an AI coding assistant created by Sajad. You are an expert in Python, JavaScript, Java, C++, and other programming languages. Write clean, efficient, and well-documented code. Keep responses concise and practical.<|im_end|>
 
 
183
  <|im_start|>user
184
  {your prompt}<|im_end|>
185
  <|im_start|>assistant
@@ -187,89 +182,107 @@ You are BlitzKode, an AI coding assistant created by Sajad. You are an expert in
187
 
188
  ---
189
 
190
- ## Configuration
191
 
192
- The server supports environment variables:
 
 
 
 
 
193
 
194
- | Variable | Default | Description |
195
- |----------|---------|-------------|
196
- | `BLITZKODE_MODEL_PATH` | `blitzkode.gguf` | Model file path |
197
- | `BLITZKODE_FRONTEND_PATH` | `frontend/index.html` | UI path |
198
- | `BLITZKODE_HOST` | `0.0.0.0` | Server host |
199
- | `BLITZKODE_PORT` | `7860` | Server port |
200
- | `BLITZKODE_THREADS` | CPU count | CPU threads |
201
- | `BLITZKODE_N_CTX` | `2048` | Context window |
202
- | `BLITZKODE_BATCH` | `128` | Batch size |
203
- | `BLITZKODE_MAX_PROMPT_LENGTH` | `4000` | Max prompt chars |
204
 
205
  ---
206
 
207
  ## Limitations
208
 
209
- - **Text-only input** - No image/vision support
210
- - **2048 token context** - CPU-friendly but limited
211
- - **Verify outputs** - Always review generated code before use
212
- - **Small model** - May occasionally produce incorrect code
 
 
213
 
214
  ---
215
 
216
- ## Project Structure
217
 
218
- ```
219
- BlitzKode/
220
- ├── server.py # FastAPI backend (v1.6)
221
- ├── blitzkode.gguf # Quantized model (~3GB)
222
- ├── frontend/
223
- │ └── index.html # Web UI
224
- ├── tests/
225
- │ └── test_server.py # HTTP tests
226
- ├── scripts/
227
- │ ├── train_sft.py # SFT training
228
- │ ├── train_grpo.py # GRPO training
229
- │ ├── train_dpo.py # DPO training
230
- │ ├── export_gguf.py # Model export
231
- │ └── test_inference.py # Inference test
232
- ├── checkpoints/ # LoRA checkpoints
233
- ├── datasets/ # Training data
234
- ├── MODEL_CARD.md # This file
235
- └── README.md # Project docs
236
- ```
237
 
238
  ---
239
 
240
- ## Version History
241
 
242
- | Version | Date | Changes |
243
- |---------|------|---------|
244
- | 1.6 | Current | CPU optimization, faster inference |
245
- | 1.5 | Earlier | Added streaming support |
246
- | 1.0 | Initial | Base model release |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
247
 
248
  ---
249
 
250
  ## License
251
 
252
- MIT License - See README.md for details.
253
 
254
- Also comply with upstream Qwen base model license when redistributing.
 
 
 
 
 
 
 
 
255
 
256
  ---
257
 
258
  ## Contact
259
 
260
- - **GitHub**: https://github.com/neuralbroker/blitzkode
261
- - **Portfolio**: https://neuralbroker.vercel.app
262
- - Issues and contributions welcome!
 
263
 
264
  ---
265
 
266
  ## Citation
267
 
268
  ```bibtex
269
- @software{blitzkode2026,
270
- author = {Sajad},
271
- title = {BlitzKode - AI Coding Assistant},
272
- year = {2026},
273
- url = {https://github.com/neuralbroker/blitzkode}
274
  }
275
  ```
 
1
  ---
2
  language:
3
+ - en
4
+ license: mit
5
  library_name: llama-cpp-python
6
  pipeline_tag: text-generation
7
  tags:
8
+ - code-generation
9
+ - coding-assistant
10
+ - gguf
11
+ - llama.cpp
12
+ - qwen2.5
13
+ - python
14
+ - javascript
15
+ - fine-tuned
16
+ - lora
17
+ - peft
18
  base_model:
19
+ - Qwen/Qwen2.5-1.5B-Instruct
20
+ - Qwen/Qwen2.5-0.5B-Instruct
21
  ---
22
 
23
  # BlitzKode
24
 
25
+ **BlitzKode** is a local AI coding assistant fine-tuned from the Qwen2.5 family. It
26
+ ships as a **GGUF model** (1.5B, F16, ~3 GB) for fast offline inference with
27
+ llama.cpp, and as a **LoRA adapter** (0.5B, ~100 MB) for PEFT-based research and
28
+ further fine-tuning.
29
 
30
+ > **Creator:** [Sajad (neuralbroker)](https://github.com/neuralbroker)
31
+ > **GitHub:** <https://github.com/neuralbroker/blitzkode>
32
+ > **GGUF model:** [`neuralbroker/blitzkode`](https://huggingface.co/neuralbroker/blitzkode)
33
+ > **LoRA adapter:** [`neuralbroker/blitzkode-lora-0.5b`](https://huggingface.co/neuralbroker/blitzkode-lora-0.5b)
34
 
35
  ---
36
 
37
+ ## Model Variants
38
 
39
+ | Variant | Version | Base Model | Format | Size | Runtime |
40
+ |---|---|---|---|---|---|
41
+ | **GGUF** (production) | 2.0 | `Qwen/Qwen2.5-1.5B-Instruct` | GGUF F16 | ~3 GB | llama.cpp / llama-cpp-python |
42
+ | **LoRA adapter** (research) | 2.1 | `Qwen/Qwen2.5-0.5B-Instruct` | PEFT safetensors | ~100 MB | PEFT + Transformers |
 
 
 
 
 
 
 
43
 
44
  ---
45
 
46
  ## Architecture
47
 
48
+ | Property | GGUF (1.5B) | LoRA Adapter (0.5B) |
49
+ |---|---|---|
50
+ | **Model type** | Transformer (Qwen2) | Transformer (Qwen2) + LoRA |
51
+ | **Parameters** | 1.5 B | 0.5 B + adapter weights |
52
+ | **Quantization** | GGUF F16 | bfloat16 / float16 |
53
+ | **LoRA rank (r)** | — | 16 |
54
+ | **LoRA alpha** | — | 32 |
55
+ | **LoRA target modules** | — | q, k, v, o, gate, up, down projections |
56
+ | **Context window** | 2 048 tokens | 2 048 tokens |
57
+ | **Vocabulary** | 151 936 | 151 936 |
58
 
59
  ---
60
 
61
  ## Training Pipeline
62
 
63
+ BlitzKode was produced by a **4-stage fine-tuning pipeline**:
64
 
65
+ ### Stage 1 — SFT (Supervised Fine-Tuning)
66
+ LoRA fine-tuning (`r=32`, base: Qwen2.5-1.5B-Instruct) on 71 curated algorithmic
67
+ coding problems covering arrays, strings, trees, dynamic programming, graphs,
68
+ sorting, hash tables, binary search, and more.
69
 
70
+ - **Adapter checkpoint:** `checkpoints/sft-1.5b-v1/`
71
+ - **Library:** PEFT + HuggingFace Transformers
72
 
73
+ ### Stage 2 — Reward-SFT
74
+ Continued SFT with heuristic reward functions to reinforce code correctness,
75
+ formatting quality, and concise explanation style. This is a standard SFT
76
+ training loop using scalar reward signals, **not** full GRPO.
77
 
78
+ - **Adapter checkpoint:** `checkpoints/grpo-v1/` *(label is historical)*
79
+ - **Library:** TRL / Transformers
80
 
81
+ ### Stage 3 — DPO (Direct Preference Optimization)
82
+ Preference optimization on handcrafted chosen/rejected pairs to improve answer
83
+ clarity, reduce verbosity, and penalize hallucinated APIs or filenames.
 
 
84
 
85
+ - **Adapter checkpoint:** `checkpoints/dpo-v1/`
86
+ - **Library:** TRL
87
 
88
+ ### Stage 4 — Continued LoRA SFT (Published Adapter)
89
+ Final LoRA fine-tuning (`r=16`, base: **Qwen2.5-0.5B-Instruct**) on 99 samples
90
+ drawn from the 199-sample full dataset. Training ran for 50 steps; final loss
91
+ reached **~0.48**.
92
 
93
+ - **Adapter checkpoint:** `checkpoints/available-lora-0.5b-full/final` ✅ *(publicly available)*
94
+ - **Library:** PEFT + Transformers
 
 
 
 
95
 
96
+ ### Stage 5 — Merge & Export (GGUF)
97
+ LoRA adapters from Stage 1–3 were merged into the 1.5B base model using
98
+ `merge_and_unload()`, then converted to GGUF F16 format with llama.cpp.
99
 
100
+ - **Script:** `scripts/export_gguf.py`
101
+ - **Artifact:** `blitzkode.gguf` (~3 GB, git-ignored)
 
 
 
 
 
 
 
102
 
103
  ---
104
 
105
+ ## Training Data
106
 
107
+ **Total: 199 samples across 3 subsets**
 
 
 
 
 
108
 
109
+ | Subset | Count | Source | License | Purpose |
110
+ |---|---|---|---|---|
111
+ | Curated algorithmic problems | 71 | Custom (local) | MIT | Core coding skills: arrays, strings, trees, DP, graphs, sorting, searching |
112
+ | MetaMathQA samples | 100 | [`meta-math/MetaMathQA`](https://huggingface.co/datasets/meta-math/MetaMathQA) | CC BY 4.0 | Math reasoning transfer to improve step-by-step problem solving |
113
+ | Python/JavaScript patterns | 28 | Custom (local) | MIT | Practical patterns: decorators, context managers, data classes, async, CLI tools |
114
+ | **Total** | **199** | | | |
115
 
116
+ See [`datasets/MANIFEST.md`](datasets/MANIFEST.md) for full dataset provenance,
117
+ preprocessing notes, and per-sample license details.
118
 
119
+ ---
120
 
121
+ ## Features
122
 
123
+ - **Multi-language code generation** — Python, JavaScript, Java, C++, TypeScript, SQL
124
+ - **Code explanation** — clear inline comments and documentation
125
+ - **Bug fixing** — debug and fix common code issues
126
+ - **Algorithm assistance** — data structures and algorithms (LeetCode-style)
127
+ - **Offline operation** — fully local, no internet required at inference time
128
+ - **Fast CPU inference** — GGUF F16 runs on commodity CPUs
129
+ - **Modern web UI** — React/Vite chat interface with SSE streaming
130
+ - **REST API** — FastAPI backend with streaming and optional web-search augmentation
131
 
132
+ ---
 
133
 
134
+ ## Usage
 
 
135
 
136
+ ### Production: GGUF with llama.cpp
137
 
138
+ ```bash
139
+ # Clone and install
140
+ git clone https://github.com/neuralbroker/blitzkode
141
+ cd blitzkode
142
+ pip install -r requirements.txt
 
 
143
 
144
+ # Build the frontend
145
+ cd frontend && npm install && npm run build && cd ..
146
 
147
+ # Start the server (place blitzkode.gguf in repo root first)
148
+ python server.py
149
+ # Open http://localhost:7860
 
 
150
  ```
151
 
152
+ ### Research: LoRA Adapter with PEFT
153
 
154
  ```python
155
+ from peft import PeftModel
156
+ from transformers import AutoModelForCausalLM, AutoTokenizer
157
 
158
+ base_model_id = "Qwen/Qwen2.5-0.5B-Instruct"
159
+ adapter_repo = "neuralbroker/blitzkode-lora-0.5b"
 
 
 
160
 
161
+ tokenizer = AutoTokenizer.from_pretrained(base_model_id, trust_remote_code=True)
162
+ model = AutoModelForCausalLM.from_pretrained(
163
+ base_model_id, torch_dtype="auto", device_map="auto", trust_remote_code=True
164
+ )
165
+ model = PeftModel.from_pretrained(model, adapter_repo)
166
+ model.eval()
 
 
 
167
  ```
168
 
169
+ ### Prompt Format (ChatML)
 
 
170
 
171
+ All variants use the Qwen ChatML template:
172
 
173
  ```
174
  <|im_start|>system
175
+ You are BlitzKode, an AI coding assistant created by Sajad. You are an expert
176
+ in Python, JavaScript, Java, C++, and other languages. Write clean, efficient,
177
+ and well-documented code. Keep responses concise and practical.<|im_end|>
178
  <|im_start|>user
179
  {your prompt}<|im_end|>
180
  <|im_start|>assistant
 
182
 
183
  ---
184
 
185
+ ## Intended Use
186
 
187
+ ### Best For
188
+ - Local offline coding assistance
189
+ - Algorithm and data structure problem solving
190
+ - Code generation and explanation
191
+ - Educational programming support
192
+ - Code review, refactoring, and debugging
193
 
194
+ ### Out of Scope
195
+ - Production code without thorough expert review
196
+ - Security-critical or cryptographic applications
197
+ - Multi-modal tasks (images not supported)
198
+ - Long-context repository analysis (> 2 048 tokens)
 
 
 
 
 
199
 
200
  ---
201
 
202
  ## Limitations
203
 
204
+ - **Text-only input** — no image or file-upload support
205
+ - **2 048-token context** — CPU-friendly but limits long conversation history
206
+ - **Verify all outputs** — always review and test generated code
207
+ - **Small model** — 0.5B–1.5B scale; may produce incorrect code on complex tasks
208
+ - **No real-time data** — knowledge cutoff follows the Qwen2.5 base model
209
+ - **Math reasoning** — MetaMathQA transfer helps basic reasoning; not a math specialist
210
 
211
  ---
212
 
213
+ ## Environment Variables (Inference Server)
214
 
215
+ | Variable | Default | Description |
216
+ |---|---|---|
217
+ | `BLITZKODE_GPU_LAYERS` | `0` | Number of layers to offload to GPU |
218
+ | `BLITZKODE_THREADS` | system | CPU inference thread count |
219
+ | `BLITZKODE_N_CTX` | `2048` | Context window size |
220
+ | `BLITZKODE_BATCH` | `512` | llama.cpp batch size |
221
+ | `BLITZKODE_PRELOAD_MODEL` | `false` | Load model at startup vs first request |
 
 
 
 
 
 
 
 
 
 
 
 
222
 
223
  ---
224
 
225
+ ## Project Structure
226
 
227
+ ```text
228
+ BlitzKode/
229
+ server.py # FastAPI backend (inference + search)
230
+ blitzkode.gguf # GGUF model artifact (~3 GB, git-ignored)
231
+ frontend/ # React/Vite web UI
232
+ scripts/
233
+ train_sft.py # Stage 1: SFT training
234
+ train_reward_sft.py # Stage 2: Reward-SFT
235
+ train_dpo.py # Stage 3: DPO
236
+ train_available.py # Stage 4: LoRA fine-tune (0.5B)
237
+ export_gguf.py # Merge & convert to GGUF
238
+ push_to_hub.py # Push adapter to HuggingFace Hub
239
+ build_full_dataset.py # Dataset builder (algorithmic + HF datasets)
240
+ datasets/
241
+ MANIFEST.md # Dataset provenance and license info
242
+ checkpoints/
243
+ available-lora-0.5b-full/ # Published LoRA adapter (0.5B)
244
+ tests/
245
+ test_server.py # HTTP integration tests
246
+ docs/
247
+ PROJECT_OVERVIEW.md # Architecture and design notes
248
+ README.md # Full project documentation
249
+ MODEL_CARD.md # This file
250
+ ```
251
 
252
  ---
253
 
254
  ## License
255
 
256
+ **MIT** — see [LICENSE](https://github.com/neuralbroker/blitzkode/blob/main/LICENSE).
257
 
258
+ You must also comply with the upstream Qwen2.5 license when redistributing any
259
+ fine-tuned weights derived from it.
260
+
261
+ - [Qwen2.5-0.5B-Instruct license](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct)
262
+ - [Qwen2.5-1.5B-Instruct license](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
263
+
264
+ Training data subsets carry their own licenses:
265
+ - MetaMathQA: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)
266
+ - Custom/local samples: MIT
267
 
268
  ---
269
 
270
  ## Contact
271
 
272
+ - **GitHub Issues:** <https://github.com/neuralbroker/blitzkode/issues>
273
+ - **Portfolio:** <https://neuralbroker.vercel.app>
274
+
275
+ Contributions and feedback are welcome!
276
 
277
  ---
278
 
279
  ## Citation
280
 
281
  ```bibtex
282
+ @software{blitzkode2025,
283
+ author = {Sajad},
284
+ title = {BlitzKode: A Local AI Coding Assistant},
285
+ year = {2025},
286
+ url = {https://github.com/neuralbroker/blitzkode}
287
  }
288
  ```