Review & open questions

Replacing Jev with a local topic classifier

Our read of the spike report, the numbers that actually matter, and the questions we need answered before we switch the digest over from Jev.

Draft for discussion 3 Oct 2026 Claude-drafted → verify in code Model repo →

01 The job, in one place

A family reads a daily digest of school email. Each item gets sorted into a topic; a fixed rules table then decides show / hide. Only the classifier changes. The rules don't, so the show-or-hide outcome matters more than the topic name itself.

Emailschool & team messages
›
Readerbig model splits into items
›
ClassifierJev today → our model
›
Rules tablefixed show/hide, no AI
›
Digestwhat the family sees

Goal: run the classifier on our own machine — CPU, no GPU pressure (the GPU is shared with the resident text model), no paid fallback. It picks one of 56 topics, plus a weak child and action answer.

02 The result — and what it measures

These are the numbers from the report. Treat them as reported, not yet reproduced; we'll recompute them from the pushed code + raw eval.

87.6%
Same topic as Jev
top-1, test 92/105
0.572
Macro-F1
each topic weighted equally
90.5%
Same show/hide as Jev
test 95/105
0 / 0.2s
GPU / item
2.2 GiB RAM, CPU only
⚠ Same as Jev ≠ correct

Every score above is measured against Jev's answer, not against truth. Jev re-answers the same item identically only 97.1% of the time, and matched a human label just 11 of 20 on a disagreement set. So 87.6% currently means "agrees with the teacher," not "is right." That's not a criticism of the result — it's the limit of what can be claimed right now.

03 Why silver + gold (the part we want to make sure of)

We're not just "copying Jev." The point of two label tiers is that Jev is a black box we don't fully trust, so we keep volume and trust separate:

Silver · volume

Jev's predictions

Cheap, abundant, usually right but noisy.
~279 training items
+
Gold · trust

Human labels

Expensive, scarce, human-verified. The only real yardstick.
31 items so far

Train on silver, verify and clean with gold. Gold is what turns "same-as-Jev" into an honest accuracy estimate, and the only way to find where Jev systematically errs — so we don't teach the student its mistakes. Without gold, a student could score 99% "same as Jev" and still be wrong.

The crux — two things we need to confirm

1. Is the silver used raw, or gold-corrected before training? If raw, the student still learns Jev's errors and 87.6% is honestly "just agrees with Jev." If gold cleaned the silver, the system beats a pure copy — and that difference is the whole value of the two-tier setup. Which is it?

2. How big is the gold really? 31 across 56 topics is far too small to estimate accuracy. Is the fresh, unopened 100-item set human gold, and is it held out from all tuning (not just the test split)?

04 The method, as we understand it

It is not a plain linear probe on frozen vectors — it's LoRA on the encoder plus a prototype head. Correct us anywhere this is off:

Base
Qwen/Qwen3-Embedding-0.6B → item vector (~1024-d)
Change
LoRA on attention + feed-forward · rank 16, α 32, dropout 0.05 · rest frozen
Head
56 prototype vectors · score = cosine(item, proto) × learned scale (init 20)
Seed head
Prototypes start from topic-description embeddings (14 topics have no training item)
Loss
softmax cross-entropy, label smoothing 0.05 · single-label
Optim
AdamW: LoRA 3e-4 (wd 0.01), head 2e-3 (no wd) · warmup 10% → decay 5%
Train
batch 16 · ≤12 epochs · early-stop on val macro-F1 · bf16
Balance
sample ∝ 1/√count · hard pairs ×2 · no augmentation / synthetic
Seeds
3 seeds; macro-F1 0.49–0.59 across them; shipped = seed 0
Also
child + action answers — weak (matches Jev 51–65)

05 Everything that was tried

The shortlist idea from "embedding + reranker" survived into the cascade. The trained head won on accuracy-per-watt. Click to expand.

▶ Full model comparison (18 rows)
Model / methodSame topicSpeed / memoryOutcome
GLiNER2.5-Decide 340M (shipped)24.1%~0.5s GPUdrop — over-picks one topic
GLiNER2.5-Decide 1B23.8%—drop
GLiClass-instruct-large1.2%—drop — scores saturate
BGE-small embeddings, zero-shot29.4%CPUdrop
Embedding + reranker56–59%—drop — but shortlist reused
TF-IDF (word-count)65–67%CPUdrop
Embedding + simple head65–67%CPUdrop
BGE-small, trained< Qwen3CPUdrop
Qwen3.5-4B over all 56 topics57.5–75.3%~7 GiB GPUdrop alone — child only 45.9%
GLiNER-340M fine-tuned, 10-topic81.9%3.1 GiB / 1.3s CPUnot chosen — costs GPU
Cascade (shortlist → Qwen3.5-4B)87.6%0.41s · 9.3 GiBhelper too much GPU
Qwen3-Embedding-0.6B + trained head87.6%0.20s CPU · 2.2 GiB · 0 GPUchosen
Clef-flash 9B (Cloudflare)84.8%1.47s · 8.3 GiBdrop — slow, drops items
Intern-Decision-4B79.0%0.8s · 3.7–9 GiBparked — not better
Jev-Style 0.8B88%*0.55s GPU / 7.6s CPUdrop
Kev-0.8B86%*0.53s GPU / 10s CPUdrop
Kev-4B—46s CPUdrop — too slow
Open-Jev-DeBERTa——can't run — 512 vs 843–1923 tokens

* scored on a different set with a show/hide measure; the chosen head scored 93% on that same measure. Test set = 105 items, so a 2–3 point gap is 2–3 items — inside the noise.

▶ Finalists, side by side (190 items)
MetricHeadCascadeGLiNER-FTClefInternJev↔Jev
Same topic, val/test75.3/87.688.2/87.684.7/81.977.6/84.880.0/79.096.5/97.1
Same show/hide, val/test83.5/90.592.9/94.391.8/90.585.9/92.484.7/89.597.7/98.1
Same final result (of 190)176185181176174189
Important items dropped ↓7351180
Time / item ↓0.20s0.18–0.41s0.09s1.47s0.80s—
GPU memory ↓09.33.18.39.05—

GPU has 16 GiB; the resident text model uses 6.7–11.4 GiB of it, so anything ≥9 GiB doesn't fit alongside it. The head fits precisely because it's 0.

06 Questions for Andrei

The report is Claude-drafted, so these are mostly "confirm against the real code / data" items.

  1. Silver raw, or gold-corrected?
    Was any Jev label changed by your human gold before training, or is silver used as-is and gold only for reporting? This decides whether the model is a copy of Jev or better.
  2. How big is gold, really?
    31 human labels across 56 topics can't establish accuracy. Is the unopened 100-item set the real gold, and fully held out from tuning?
  3. What is the "94.3% outcome" for the cascade?
    It sits next to 87.6% test accuracy but isn't defined. Different metric? A confident-subset score? We want the exact definition.
  4. The test split was used for decisions.
    The report calls it dev data now, but the headline 87.6% comes from it. Can we quote the val number, or a true held-out number, instead?
  5. The sports-class skew.
    ~half the items are sports and almost always right. What's the non-sports agreement number (you cite ~78%) — and does the headline hold with sports down-weighted?
  6. Child + action heads.
    They match Jev only 51–65. Are they in scope for the switch, or topic-only for now?

07 What we need to reproduce it

Drop these into the model repo — each folder's README has the exact schema. Priority top to bottom.

src/train_head.py + config.yamlthe actual training code + hyperparameters top
eval/per_item_test.jsonlper-item raw output → we recompute 87.6 / 0.572 / 90.5 without re-running
src/evaluate.py + rules_table.jsonmetric definitions + the fixed show/hide rules
checkpoints/lora + checkpoints/headadapter + 56 prototypes (seed 0), with shapes
data/*.jsonl + taxonomy + 31 labelstrain/val/test, descriptions, human gold → private data repo, not here
requirements.txt + how-to-runtorch/peft versions, base-model path, seed/config flags

Verification checklist — what we'll confirm in the code. Tick as we go (saved in your browser only).

0 of 9 confirmed

08 The synthetic twin — we rebuilt the recipe from your docs alone

While waiting for src/, we asked: is the spike reproducible from the report + checklist with zero code? Answer: yes — Kodep/jev-topic-head-example implements the whole thing: LoRA r16/α32/dropout .05 on attention+FFN, 56 prototypes seeded from topic descriptions, scale init 20, CE + smoothing .05, AdamW 3e-4 / 2e-3, warmup 10% → decay 5%, 1/√count sampling with hard pairs ×2, early stop on val macro-F1, 3 seeds.

What it provesthe spec is complete enough to rebuild: pipeline trains, early-stops (~epoch 4), ships a 57,345-param head (56×1024+1 exactly), and reproduces the report's "errors hide in low confidence" pattern
What it does NOT provequality — it trained on synthetic data with the real split's shape (279/85/105, 56 topics, 15 unseen); the 0.956 synthetic macro-F1 measures label cleanliness, not brains
Ready-madeseed-0 weights (39 MB adapter + 230 KB head via LFS — base model auto-downloads, we ship only our delta) + predict.py; hf download and try in two commands
Intended usea diff target: when your train_head.py lands, the deltas between our guesses (last-token pooling? bf16 = weights or autocast?) become concrete questions instead of unknowns

The ask is unchanged from block 07: src/ + config.yaml first (they answer the silver/gold question directly), then data access — and we rerun this pipeline on real labels to face 92/105 and 0.572 honestly.