Our read of the spike report, the numbers that actually matter, and the questions we need answered before we switch the digest over from Jev.
A family reads a daily digest of school email. Each item gets sorted into a topic; a fixed rules table then decides show / hide. Only the classifier changes. The rules don't, so the show-or-hide outcome matters more than the topic name itself.
Goal: run the classifier on our own machine — CPU, no GPU pressure (the GPU is shared with the resident text model), no paid fallback. It picks one of 56 topics, plus a weak child and action answer.
These are the numbers from the report. Treat them as reported, not yet reproduced; we'll recompute them from the pushed code + raw eval.
Every score above is measured against Jev's answer, not against truth. Jev re-answers the same item identically only 97.1% of the time, and matched a human label just 11 of 20 on a disagreement set. So 87.6% currently means "agrees with the teacher," not "is right." That's not a criticism of the result — it's the limit of what can be claimed right now.
We're not just "copying Jev." The point of two label tiers is that Jev is a black box we don't fully trust, so we keep volume and trust separate:
Train on silver, verify and clean with gold. Gold is what turns "same-as-Jev" into an honest accuracy estimate, and the only way to find where Jev systematically errs — so we don't teach the student its mistakes. Without gold, a student could score 99% "same as Jev" and still be wrong.
1. Is the silver used raw, or gold-corrected before training? If raw, the student still learns Jev's errors and 87.6% is honestly "just agrees with Jev." If gold cleaned the silver, the system beats a pure copy — and that difference is the whole value of the two-tier setup. Which is it?
2. How big is the gold really? 31 across 56 topics is far too small to estimate accuracy. Is the fresh, unopened 100-item set human gold, and is it held out from all tuning (not just the test split)?
It is not a plain linear probe on frozen vectors — it's LoRA on the encoder plus a prototype head. Correct us anywhere this is off:
Qwen/Qwen3-Embedding-0.6B → item vector (~1024-d)The shortlist idea from "embedding + reranker" survived into the cascade. The trained head won on accuracy-per-watt. Click to expand.
| Model / method | Same topic | Speed / memory | Outcome |
|---|---|---|---|
| GLiNER2.5-Decide 340M (shipped) | 24.1% | ~0.5s GPU | drop — over-picks one topic |
| GLiNER2.5-Decide 1B | 23.8% | — | drop |
| GLiClass-instruct-large | 1.2% | — | drop — scores saturate |
| BGE-small embeddings, zero-shot | 29.4% | CPU | drop |
| Embedding + reranker | 56–59% | — | drop — but shortlist reused |
| TF-IDF (word-count) | 65–67% | CPU | drop |
| Embedding + simple head | 65–67% | CPU | drop |
| BGE-small, trained | < Qwen3 | CPU | drop |
| Qwen3.5-4B over all 56 topics | 57.5–75.3% | ~7 GiB GPU | drop alone — child only 45.9% |
| GLiNER-340M fine-tuned, 10-topic | 81.9% | 3.1 GiB / 1.3s CPU | not chosen — costs GPU |
| Cascade (shortlist → Qwen3.5-4B) | 87.6% | 0.41s · 9.3 GiB | helper too much GPU |
| Qwen3-Embedding-0.6B + trained head | 87.6% | 0.20s CPU · 2.2 GiB · 0 GPU | chosen |
| Clef-flash 9B (Cloudflare) | 84.8% | 1.47s · 8.3 GiB | drop — slow, drops items |
| Intern-Decision-4B | 79.0% | 0.8s · 3.7–9 GiB | parked — not better |
| Jev-Style 0.8B | 88%* | 0.55s GPU / 7.6s CPU | drop |
| Kev-0.8B | 86%* | 0.53s GPU / 10s CPU | drop |
| Kev-4B | — | 46s CPU | drop — too slow |
| Open-Jev-DeBERTa | — | — | can't run — 512 vs 843–1923 tokens |
* scored on a different set with a show/hide measure; the chosen head scored 93% on that same measure. Test set = 105 items, so a 2–3 point gap is 2–3 items — inside the noise.
| Metric | Head | Cascade | GLiNER-FT | Clef | Intern | Jev↔Jev |
|---|---|---|---|---|---|---|
| Same topic, val/test | 75.3/87.6 | 88.2/87.6 | 84.7/81.9 | 77.6/84.8 | 80.0/79.0 | 96.5/97.1 |
| Same show/hide, val/test | 83.5/90.5 | 92.9/94.3 | 91.8/90.5 | 85.9/92.4 | 84.7/89.5 | 97.7/98.1 |
| Same final result (of 190) | 176 | 185 | 181 | 176 | 174 | 189 |
| Important items dropped ↓ | 7 | 3 | 5 | 11 | 8 | 0 |
| Time / item ↓ | 0.20s | 0.18–0.41s | 0.09s | 1.47s | 0.80s | — |
| GPU memory ↓ | 0 | 9.3 | 3.1 | 8.3 | 9.05 | — |
GPU has 16 GiB; the resident text model uses 6.7–11.4 GiB of it, so anything ≥9 GiB doesn't fit alongside it. The head fits precisely because it's 0.
The report is Claude-drafted, so these are mostly "confirm against the real code / data" items.
Drop these into the model repo — each folder's README has the exact schema. Priority top to bottom.
src/train_head.py + config.yaml | the actual training code + hyperparameters top |
eval/per_item_test.jsonl | per-item raw output → we recompute 87.6 / 0.572 / 90.5 without re-running |
src/evaluate.py + rules_table.json | metric definitions + the fixed show/hide rules |
checkpoints/lora + checkpoints/head | adapter + 56 prototypes (seed 0), with shapes |
data/*.jsonl + taxonomy + 31 labels | train/val/test, descriptions, human gold → private data repo, not here |
requirements.txt + how-to-run | torch/peft versions, base-model path, seed/config flags |
Verification checklist — what we'll confirm in the code. Tick as we go (saved in your browser only).
While waiting for src/, we asked: is the spike reproducible from the report + checklist with zero code? Answer: yes — Kodep/jev-topic-head-example implements the whole thing: LoRA r16/α32/dropout .05 on attention+FFN, 56 prototypes seeded from topic descriptions, scale init 20, CE + smoothing .05, AdamW 3e-4 / 2e-3, warmup 10% → decay 5%, 1/√count sampling with hard pairs ×2, early stop on val macro-F1, 3 seeds.
| What it proves | the spec is complete enough to rebuild: pipeline trains, early-stops (~epoch 4), ships a 57,345-param head (56×1024+1 exactly), and reproduces the report's "errors hide in low confidence" pattern |
| What it does NOT prove | quality — it trained on synthetic data with the real split's shape (279/85/105, 56 topics, 15 unseen); the 0.956 synthetic macro-F1 measures label cleanliness, not brains |
| Ready-made | seed-0 weights (39 MB adapter + 230 KB head via LFS — base model auto-downloads, we ship only our delta) + predict.py; hf download and try in two commands |
| Intended use | a diff target: when your train_head.py lands, the deltas between our guesses (last-token pooling? bf16 = weights or autocast?) become concrete questions instead of unknowns |
The ask is unchanged from block 07: src/ + config.yaml first (they answer the silver/gold question directly), then data access — and we rerun this pipeline on real labels to face 92/105 and 0.572 honestly.