Pavel Nakonechnyy

Skill-Selection with JEV: the same quality, 12–17x less waiting

Published by Pavel Nakonechnyy on in IT Management.

Key Takeaways

  • The challenge: KATE’s skill-selection step ran an LLM call for every complex request, adding 35-50s of latency per message even when the request needed no skill at all.

  • The approach: Replace the generative decision with a local, deterministic ranker, and prove the swap with a benchmark.

  • The outcome: The local ranker keeps held-out quality (F1 0.714, exact 6/10, precision 0.833) at roughly 12-17x the speed of the 12B baseline. No remote model beats it on held-out exact accuracy.

  • The lesson: Validate the rule you ship. A held-out set you never tune against is the only honest quality number.

Skill Injection: What This Benchmark Serves

What skill injection does

KATE is an AI assistant that helps working professionals across Business Analysis, Project Management, Product Management, and Marketing. She has a set of skills — reusable operating instructions (research, writing, summarize, memoryrecall, networking). Just like in popular harnesses like OpenCode and Claude, each skill lives as a plain markdown file in workspace/skills/ whose YAML header carries a short name and description, and whose body carries the full instructions.

Before KATE executes a request, she decides which skills the request needs. She then injects the full instructions of the selected skills into the conversation as system messages, so the agent executes with the exact procedure for the task. This is just-in-time injection: only the instructions the current task needs enter the context window. Small local models cannot hold every skill’s instructions at once without losing focus, so selective injection is what makes them usable as autonomous workers.

The AS IS design before JEV

Originally the selector was a structured-output LLM call. The system built a prompt that listed every skill’s name and description, asked the model to return the appropriate skill names as a strict JSON array, parsed the reply, and injected the matching instructions.

Production exposed three weaknesses:

  • Cost on no-skill tasks. The model ran once per eligible message, even when the request needed no skill. Each call took 9–19 s on the original provider and 35–50 s on the local 12B generator. A large share of those calls produced an empty result: latency and tokens spent on nothing.

  • No reliable way to say “none”. Given a prompt with no relevant skill, the model still tended to name a plausible one (“always pick something” syndrome). The injected instructions then bloated the context of a request that did not need them, causing the context cost the injection step exists to prevent.

  • Stochastic output. A generative model returned a different verdict on each run, making the behavior hard to reproduce or audit.

The TO BE design we arrived at

Working assumption (drives the top-2 selection rule): no prompt needs more than two skills active at once.

With the arrival of JEV, a technique that made classification tasks much faster and cheaper, the idea came to use a small LLM model for the decision-making. JEV does not generate natural-language text. It ranks a fixed list of options and returns each option’s probability, so its output is deterministic and cheap.

GPU-less hosting is much cheaper, so we decided to try self-hosting using JEV-CPU (https://huggingface.co/Meanblock/JEV-CPU), a fork of SemIf. Early results have shown very poor quality on 0.6B and 2B models, and we arrived at Qwen3.5-4B.

The selector builds one ranking criterion: each skill is an option, plus one synthetic no_skill option. JEV scores all options from the user prompt. The rules:

  1. If no_skill ranks first, the whole selection is empty.

  2. Otherwise, take the top-2 skills whose probability is at least 0.15.

  3. If JEV errors, or returns an empty selection that is not an explicit abstention, the gemma-4 structured-output model runs as a safety fallback. This guarantees the injection step never silently drops all context on a misfiring classifier.

The injected result is identical in form to the old design: the full instruction text of each selected skill becomes a system message. Depending on the skills selected, the harness can run some other tasks and inject information into context. What changed is the decision itself: one cheap, deterministic probability pass (~3 s on CPU) instead of a 30–50s stochastic generation, plus a genuine “no skill” verdict.

This benchmark measures whether that swap preserved selection quality. The held-out split is the honest answer; the rest of this document is the evidence.

Purpose

These are the options measured to compare three primary classifiers plus variants and remote baselines for proactive skill selection:

  • JEV-CPU per-skill yay/nay — one 2-option criterion per skill (yay/nay), shared decision state = user prompt. Historical behavior on Qwen3-0.6B: saturates (every skill >= 0.99 yay) and is unusable. On Qwen3.5-4B the framing unblocks (best threshold F1 0.727) but stays a control. Criterion wording is deliberately cautious (“Default to ‘nay’ unless the request clearly and unambiguously needs this exact skill”) to fight the classifier’s over-selection bias.

  • JEV-CPU multi-option ranking (jevrank) — ONE criterion whose options are all skills; the softmax ranks the skills by probability. Selection = skills whose probability clears a guard (> minProb), capped at top-N (harness rule: JevSkillSelectTopN = 2, JevSkillSelectMinProb = 0.15). It became the shipped path.

  • Unsloth gemma-4 12B — the previous default: a structured-output model call that returns the selected skill names from a JSON schema.

The benchmark answers: which classifier selects the right skills per prompt, at what speed?

Dataset

Dataset was composed of 8 synthetic skills (memoryrecall, networking, research, routes, writing, image, summarize, translate), each with a description, and 34 labeled prompts, derived from KATE’s 25k conversations in real usage data. The set includes single- and multi-skill, incl. a prompt in Russian, no-skill-expected prompts, skill-vocabulary decoys (the word “wifi”/”photo”/”translator” appears but no skill action is requested), and a single-skill routes prompt with a memoryrecall lure. Each prompt’s ground truth is the subset of skills that a correct selector should choose.

First 24 prompts are the TRAIN set, the set used for all rule-tuning history, and 10 semantically-new prompts became HELD-OUT (never used in tuning).

Metrics

  • Precision / Recall / F1 over all skill x prompt instances.

  • Exact-prompt accuracy: strict subset equality between selected and expected per prompt.

  • Speed — client wall-clock seconds per prompt (whole classify call).

Results

The final benchmark run compares three JEV-based selectors, local AS IS gemma selector, and four remote model selectors (Groq gpt-oss-20b / gpt-oss-120b, Opencode big-pickle, Kilo nemotron-3-ultra).

Quality — ALL 34 prompts (macro P/R/F1 over TRAIN+HELD-OUT 272 skill instances)

Classifier Precision Recall F1 Exact acc. exact TP FP FN
JEV (argmax yay) 0.923 0.414 0.571 0.559 19/34 12 1 17
jevrank (top-2, minProb 0.15) + no_skill 0.808 0.724 0.764 0.647 22/34 21 5 8
jevrank-any (no top-N cap, minProb 0.15) 0.778 0.724 0.750 0.618 21/34 21 6 8
gemma-4 0.743 0.897 0.813 0.647 22/34 26 9 3
gpt-oss-20b (Groq) 0.667 0.897 0.765 0.559 19/34 26 13 3
gpt-oss-120b (Groq) 0.676 0.862 0.758 0.529 18/34 25 12 4
big-pickle (Opencode) 0.619 0.897 0.732 0.588 20/34 26 16 3
nemotron-3-ultra (Kilo) 0.697 0.793 0.742 0.618 21/34 23 10 6

Quality — TRAIN split (used to tune the rule)

Classifier Precision Recall F1 Exact acc. exact TP FP FN
JEV (argmax yay) 0.889 0.381 0.533 0.542 13/24 8 1 13
JEV (best threshold 0.3) 0.929 0.619 0.743 0.667 16/24 13 1 8
jevrank (top-2, minProb 0.15) + no_skill 0.800 0.762 0.780 0.667 16/24 16 4 5
jevrank (no top-N cap, thr 0.1, train-tuned) 0.773 0.810 0.791 0.667 16/24 17 5 4
gemma-4 0.750 0.857 0.800 0.625 15/24 18 6 3
gpt-oss-20b (Groq) 0.667 0.857 0.750 0.542 13/24 18 9 3
gpt-oss-120b (Groq) 0.680 0.810 0.739 0.500 12/24 17 8 4
big-pickle (Opencode) 0.655 0.905 0.760 0.625 15/24 19 10 2
nemotron-3-ultra (Kilo) 0.708 0.810 0.756 0.667 16/24 17 7 4

Quality — HELD-OUT split (the honest generalization number)

Classifier Precision Recall F1 Exact acc. exact TP FP FN
JEV (argmax yay) 1.000 0.500 0.667 0.600 6/10 4 0 4
JEV (best thr 0.3, fixed from train) 1.000 0.625 0.769 0.700 7/10 5 0 3
jevrank (top-2, minProb 0.15) + no_skill 0.833 0.625 0.714 0.600 6/10 5 1 3
jevrank (no top-N cap, thr 0.1, fixed from train) 0.833 0.625 0.714 0.600 6/10 5 1 3
gemma-4 0.727 1.000 0.842 0.700 7/10 8 3 0
gpt-oss-20b (Groq) 0.667 1.000 0.800 0.600 6/10 8 4 0
gpt-oss-120b (Groq) 0.667 1.000 0.800 0.600 6/10 8 4 0
big-pickle (Opencode) 0.538 0.875 0.667 0.500 5/10 7 6 1
nemotron-3-ultra (Kilo) 0.667 0.750 0.706 0.500 5/10 6 3 2

Speed (same 34-case run)

Classifier avg s/prompt (client) notes
gpt-oss-20b (Groq) 0.722 structured output, 0 failures
gpt-oss-120b (Groq) 0.798 structured output, 0 failures
jevrank (1 criterion + no_skill) 2.990 server total 2.983 s, forward 2.978 s/decision, avg 346 tokens/prompt
big-pickle (Opencode) 3.074 structured output, 0 failures
JEV (yay/nay, 8 criteria) 12.621 server total 12.616 s, forward 12.603 s/decision, avg 1442 tokens/prompt
nemotron-3-ultra (Kilo) 24.793 structured output, 0 failures
gemma-4 34.90 / 46.91 / 49.76 local 12B structured output, 0 failures, no queueing

Findings

  • Held-out evaluation led to a mild, expected drop for jevrank (train 0.780 → held-out 0.714 F1).

  • No remote model beats jevrank on held-out exact accuracy or precision.

  • jevrank’s held-out errors are 3 false negatives — all multi-skill drops where the second skill fell under the guard; plus 1 false positive (FP) on the C29 translator decoy.

  • jevrank’’’s held-out precision (0.833) beats gemma-4 (0.727), yet loses on F1.

  • jevrank is 4x faster than the 8-criterion yay/nay framing (one 346-token softmax pass vs 8 passes over 1441 tokens) and ~17x faster than gemma-4.

  • top-2 + 0.15 minProb + no_skill harness rule produces the best scores among the jevrank variants on the full 34-prompt set.

  • The no_skill option fixes the weak spot: on every no-skill prompt the ranker now abstains cleanly instead of committing to a plausible skill.

Evidence

All the artefacts, including the benchmarks code can be found at GitHub.

  • Per-case CSV (one row per classifier x prompt, appended across runs): skillSelectionBenchmark.csv (the probs column records the per-skill probability map for the jev/jevrank rows).

  • The 0.6B-era rows: skillSelectionBenchmark.qwen3-06b.csv

  • Some of the intermediary runs: skillSelectionBenchmark.qwen3-5b-c01-c14.csv, skillSelectionBenchmark.qwen3-5b-c01-c24.csv, skillSelectionBenchmark.qwen3-5b-c01-c24-noskill.csv

Outcomes

KATE is a self-hosted AI assistant that helps professionals across Business Analysis, Project Management, and Marketing. Before using its tools, she must decide which skills a request needs. The original decision ran a structured-output LLM call. It added latency and tokens to every message. It could not abstain. The business problem was simple: every chat message paid the cost of a full generative model, most of the time for an empty result.

The assignment became to move that selection to a local, deterministic, cheap server and measure whether quality survives.

The move that followed sounds like a downgrade until you look at the mechanism: we took the generative model off the decision and gave the job to a small ranker running a 4-billion-parameter model on CPU. A good selection names the right skills, names none when none apply, and never inflates the context.

The swap became an experiment with a clear definition of done.

We devised our own synthetic tests to measure quality against. Every change to the rule went into a decision log with a date and a rationale to ensure traceability. We evaluated multiple dimensions to make the final decision, balancing between accuracy, quality (over-selection vs under-selection), speed, and compute cost.

The decision was to ship the jevrank with gemma-4 as fallback, not because it outscored gemma on every metric (it did not), but because 0.833 held-out precision at ~3s was the better product trade against gemma’s higher recall at 12–17x the cost.

None of this required a machine-learning engineer. It required the working habits of analysis: define the acceptable result before building, split the evidence so the evaluation stays honest, and measure the artifact that ships. The stake here was one decision step in a personal assistant. The lesson is the same at enterprise scale.

Conclusion

The JEV skill-selection case is deliberately narrow: one decision step inside an assistant. It shows that data-driven decision making in Product Management and Business Analysis is not only for large roadmap bets or enterprise-scale launches; it is a discipline that applies to small technical design decisions too.

The team treated the selector as a product decision, not an implementation detail. We defined what “good” meant before building, measured quality, and compared the shipped ranker against generative and remote baselines. The outcome was not just a faster classifier; it was a defensible decision: held-out quality preserved at roughly 12–17x the speed.

The broader lesson for PM/BA is that leverage often sits less in the model or the technology than in the framing, boundaries, and evidence around the decision. The decision is the product. Even a small swap, from a generative call to a local ranker, deserves a definition of done, honest evaluation on the artifact that ships, and traceability from question to result. That is how teams move from guessing to deciding: by using the right data to decide what actually reaches users.

2