🔍 Read the full analysis: 24 Ways To Use Jev: A Practical Decision-Model Playbook on ThorstenMeyerAI.com
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer has published a practical playbook mapping 24 use cases for Jev, a decision-model service that answers typed questions with calibrated probabilities. Three uses run live in his publishing operation covering roughly 90,000 decisions, 12 are strong fits, 7 need measurement first, and 2 are poor fits.
Thorsten Meyer published a playbook on September 29, 2026 mapping 24 concrete use cases for Jev, a decision-model service that returns calibrated, typed answers instead of prose — and says 15 of them are ready to build or already running, including three live in his own publishing operation that have processed roughly 90,000 decisions so far.
According to the playbook, Jev does not write, summarise or extract. Users send a state (text or JSON) and a set of typed questions, and receive calibrated answers that code can branch on directly, with no prose to parse. A single call carrying the state and all questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens, according to the author’s figures. Three answer types are supported: noul (a probability of yes, for gates and filters), choice (one option with probability and confidence, for routing and classification), and score (a position on ordered levels, for quality and priority).
The author’s central claim concerns confidence. In his measurement on a 31-topic classification task, Jev agreed with a frontier LLM 97 to 99% of the time when its confidence was 0.8 or higher, but only 42% of the time below 0.5 confidence. The recommended pattern across nearly every use case is therefore the same: act on clear cases, route the gray zone to a person or a smarter system.
The 24 uses are tagged by fit: 3 live in the author’s publishing operation, 12 strong fits meeting all four conditions of his fit test, 7 tagged “measure first” because a visibly failing heuristic is unproven, and 2 poor fits. The three live uses are a relevance gate (about 10,000 story-site pairings judged in three days, with only 22% clearly on-topic), a language check (78,889 articles scanned for $2.01, finding 1,576 non-English pieces and fixing 1,553), and a classifier fallback that showed 89% agreement with a frontier LLM overall.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
The Economics of Cheap Checks
The playbook frames its argument in economic terms: when a single check costs fractions of a cent, applying it to every item rather than a sample becomes affordable. The author cites his one-night scan of 78,889 articles for $2.01 as an example — a task the playbook says would cost substantially more with a frontier LLM. For teams running publishing, commerce or moderation pipelines, the playbook proposes a design pattern: use a cheap decision model for the clear majority of cases and reserve human or LLM attention for the low-confidence band.
The fit test also serves as a constraint on deployment. The author identifies two poor fits, including a same-event dedupe use case where his own canary test found zero duplicates to fix. The playbook presents this as evidence that a cheap, narrow question still requires a measured, visibly failing heuristic before deployment.
The Four-Condition Fit Test
The playbook requires all four conditions before any deployment: high volume (thousands of small calls, not a handful of big ones), a narrow question (no multi-step reasoning), cheap errors (a wrong answer costs little, or unsure cases route to something smarter), and a visibly failing heuristic — measured, not assumed. If an existing keyword rule works, the advice is to keep it.
The recommended rollout is staged: replay 300 to 500 real past decisions, compare results overall and per confidence band, and read 20 disagreements manually to judge who was right. Deployment is only advised where the high-confidence band reaches 95% accuracy, starting behind an off-by-default flag, canary-tested on 5 to 10 units, then rolled out. Each of the 24 use cases is documented with the question to ask, the question type, and the rule acting on the answer, such as “route when confidence is 0.8 or higher, otherwise send to a person.”
The mapped uses span publishing, commerce, software, business operations and the home. Publishing examples include a thin-source detector (prompted by the author’s finding that 88% of news items he processes start from a bare headline), a disclosure check that catches paraphrases a regex misses, headline quality scoring, and comment moderation with auto-approve and auto-hide thresholds at 0.9 confidence.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, ThorstenMeyerAI.com
Open Questions in the Playbook
Seven of the 24 use cases are tagged “measure first” because the fourth condition — a visibly failing heuristic — remains unproven. These include the thin-source detector, product-roundup fit checks and headline quality scoring. The playbook does not yet include error-rate measurements for these cases.
All performance figures come from the author’s own measurements and his own publishing operation, not from independent evaluation. The 97–99% agreement figure was measured against a single frontier LLM on one 31-topic task, and the playbook does not address how the results generalise to other domains, languages or question types. The article also appears to be the first part of a longer series: the commerce section is cut off mid-list, suggesting additional use cases and results may follow.
From Playbook to Production
The playbook prescribes a staged adoption path for readers: pick a strong-fit use case, replay 300 to 500 past decisions, verify the high-confidence band reaches 95%, then canary on a small slice before rollout. For the author’s own operation, the playbook indicates that measurements on the seven “measure first” cases are pending — including the thin-source detector, which the author says was prompted by his finding that 88% of processed news items begin as bare headlines. The apparent continuation of the commerce and remaining sections suggests follow-up instalments with more use cases and measured results.
Key Questions
What is Jev, according to the playbook?
A decision-model service that takes a state (text or JSON) plus typed questions and returns calibrated answers — probabilities, choices or scores with confidence — that software can branch on without parsing prose. One call takes about 0.3–0.9 seconds and costs about $0.04 per million input tokens, per the author’s figures.
How many of the 24 use cases are actually running?
Three are live in the author’s publishing operation, covering roughly 90,000 decisions so far. Twelve more are tagged as strong fits that meet all four fit-test conditions, seven need a measurement first, and two are poor fits.
What are the four conditions of the fit test?
High volume (thousands of small calls), a narrow question (no multi-step reasoning), cheap errors (wrong answers cost little or route onward), and a heuristic that fails visibly — measured rather than assumed.
In his 31-topic classification measurement, Jev agreed with a frontier LLM 97–99% of the time at confidence 0.8 or higher, but only 42% below 0.5. These are the author’s own measurements, not an independent benchmark.
Which use case failed, and why?
Same-event dedupe was tagged a poor fit. Despite being cheap and narrow, the author’s canary test found zero duplicates, meaning condition four — a visibly failing heuristic — was not met, so there was no problem for Jev to solve.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
