Model: Jeff by firelex (code MIT), weights Jeff-Qwen3.5-0.8B (Apache 2.0), run locally with MLX on a MacBook Pro with an M4 Pro chip, on September 29, 2026. One forward pass reads the message and the options and gives a probability for each option.
Data: the Banking77 test set by PolyAI (CC BY 4.0), 3,080 real online banking questions in 77 intents. We grouped 74 intents into 7 support teams and dropped 3 that no single team owns (supported cards and currencies, fiat currency support, country support), leaving 2,960 messages. The mapping is in the repo.
"How sure" is the highest of Jeff's probabilities for that message. Jeff also returns a rescaled confidence field; we use the raw top probability because it is what the threshold acts on.
Accuracy counts a message as right when Jeff's top team matches our team for its Banking77 intent. Banking77 has label quirks and our grouping has judgment calls, so some confident "mistakes" are arguable: several messages about a card being declined while shopping are filed under declined_transfer, and "Is there a fee for topping up" counts as Fees, not Top ups. We kept every label as is.
Cost uses list prices per million tokens: Claude Sonnet 5.5 $2 input and $10 output, Claude Opus 5.5 $4 and $20. Jeff runs on your own machine, so its cost is counted as zero. The big model accuracy on escalated messages is your input, not measured here.
Timing is wall time per HTTP request to the local Jeff server, one message at a time, including tokenizing the message and all 7 team descriptions. The Jeff authors report 28 ms on an M4 Max for their own benchmark prompts.