🔍 Read the full analysis: A Field Guide To Jev: 24 Uses For AI Decision Models on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Thorsten Meyer published a field guide on Sept. 29 mapping 24 uses for Jev, an AI decision model, and says three are running in his publishing operation. The guide classifies 12 other uses as strong fits, seven as requiring measurement and two as poor fits.
Meyer describes Jev as a system that receives text or JSON and typed questions, then returns answers that software can use to make decisions. Its answer types include a yes-or-no probability, a choice among options, or a score on ordered levels. He says it does not write or summarize content. A call takes about 0.3 to 0.9 seconds, and he reports a cost of about $0.04 per million input tokens.
The guide’s three live examples are a relevance check matching stories to sites, a language check, and a topic classifier used as a fallback when a primary language model errs. Meyer says the language check scanned 78,889 articles for $2.01, found 1,576 non-English items and fixed 1,553. He also reports judging about 10,000 story-and-site pairings in three days, with 22% clearly on topic, and 89% agreement with a frontier LLM for the classifier overall.
Meyer says Jev agreed with that model on the 31-topic classification 97% to 99% of the time when its confidence was at least 0.8, compared with 42% below 0.5. These are figures from his own measurement, not independently verified results. His proposed operating pattern is to act on high-confidence answers and route uncertain cases elsewhere; the application’s code determines what action follows.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Why It Matters
The guide describes criteria Meyer uses to assess whether a classification or screening task may suit an AI model: high volume, a narrow question, low-cost errors or escalation for uncertain answers, and evidence that an existing heuristic fails. He notes that a low per-call price alone does not establish whether a tool improves a process.
In the relevance gate, Meyer says clearly off-topic pairings can be dropped while uncertain cases continue through the existing publishing path. For disclosure checks, he proposes routing potential misses to human review rather than automatically publishing. The guide does not provide independent evaluations of all proposed use cases.
Background
Meyer says users should first replay 300 to 500 past decisions in shadow mode, compare results overall and by confidence band, then review 20 disagreements to determine which answer was right. He recommends wiring in a use only where the high-confidence band reaches 95%, placing the feature behind a flag that is off by default, and starting with a 5% to 10% canary.
The guide separates ideas by readiness. Its publishing examples include a thin-source detector, product relevance checks, disclosure detection, headline quality and comment moderation. Meyer labels some of these “measure first” because the existing system’s error rate has not been established. He calls same-event deduplication a poor fit for now, saying his canary found no duplicates to address. The supplied source excerpt ends partway through the guide’s commerce section, so it does not provide the full list of all 24 uses.
“Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly.”
— Thorsten Meyer, guide author
What’s Next
Meyer recommends testing a proposed use against historical decisions in shadow mode, reviewing disagreements and checking accuracy within confidence bands. If results meet his stated threshold, he recommends a feature flag and a small canary before broader use. The guide does not announce a separate product launch or give a date for further results. It says measure-first ideas require evaluation before a decision about deployment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
