Files
imladris/hw/dixie
kyleandClaude Opus 5.5 72cae29d5c dixie: unified KV per model; Qwen thinking sampling for the 9B
kv-unified lets one request use a model's whole 16k (split slots capped each at
8k and every Honcho dialectic prompt of 11-17k tokens failed). 9B gets temp 0.6 /
top-p 0.95 / top-k 20 / min-p 0: in a 20-batch replay of Honcho's deriver the
0.8 default dropped 7 batches (empty or wrong-shape output); with these, 3,
and 1 with a format instruction on rift. VRAM unchanged at 11.5 GB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-22 12:56:16 -07:00
..