Post
622
π JackOD-9B-Coder β a 9B merge built to FINISH agentic coding tasks. Four-way omnimerge_v2 over Qwen3.5-9B: Jack = Qwopus3.5-9B-Coder (0.30), O = Ornith-1.5-9B (0.15), D = DeltaCoder (0.55). MTP head kept.
π Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 β merge / base / DeltaCoder / Qwopus / Ornith:
β‘ LiveCodeBench v6 (55 hard) β 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
β HumanEval β 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
β HumanEval+ β 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
π€ MultiPL-E β 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
π IFEval β 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
π― LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
π And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 Β· base 18/55 Β· Qwopus 8/55 Β· Ornith 2/55 Β· JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
π€ tool-eval-bench hardmode, 5 seeds: JackOD 144.4 Β±4.7, second behind Ornith 145.6 Β±4.2, above base 142.0 β all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst β it stops too early. The merge does both.
π§ Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 β the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
π§ͺ Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
π¦ 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
π danielcherubini/Qwen3.5-DeltaCoder-9B Β· ornith-ai/Ornith-1.5-9B
π ManniX-ITA/JackOD-9B-Coder
π ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
π https://ollama.com/mannix/JackOD-9B-Coder
π Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 β merge / base / DeltaCoder / Qwopus / Ornith:
β‘ LiveCodeBench v6 (55 hard) β 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
β HumanEval β 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
β HumanEval+ β 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
π€ MultiPL-E β 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
π IFEval β 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
π― LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
π And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 Β· base 18/55 Β· Qwopus 8/55 Β· Ornith 2/55 Β· JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
π€ tool-eval-bench hardmode, 5 seeds: JackOD 144.4 Β±4.7, second behind Ornith 145.6 Β±4.2, above base 142.0 β all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst β it stops too early. The merge does both.
π§ Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 β the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
π§ͺ Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
π¦ 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
π danielcherubini/Qwen3.5-DeltaCoder-9B Β· ornith-ai/Ornith-1.5-9B
π ManniX-ITA/JackOD-9B-Coder
π ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
π https://ollama.com/mannix/JackOD-9B-Coder