🔍 Read the full analysis: A 24-Use Playbook For Jev And AI Decision-Making on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A Sept. 29 article maps 24 potential uses for Jev, a tool that returns typed answers to narrow decision questions rather than generating prose. Its author says three uses are live in a publishing operation, 12 meet a four-part fit test, seven need measurement and two are poor fits.
Thorsten Meyer published a 24-use playbook for Jev on Sept. 29, describing three applications already running in his publishing operation and rating 12 additional uses as strong fits. The guide matters to teams considering AI-assisted routine decisions because it sets out a test for when Jev may help, how to route uncertain answers, and when existing rules should remain in place.
Meyer describes Jev as a system that receives text or JSON plus typed questions, then returns answers software can act on. The available answer types include a yes-or-no probability, a choice among options with probabilities and confidence, and a score on ordered levels. The article says Jev does not write, summarize or extract information; the application code decides what to do with its answers.
According to Meyer, a call containing the state and all questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. In his measurement of a 31-topic classification task, Jev agreed with a frontier large language model 97% to 99% of the time when confidence was at least 0.8, and 42% of the time below 0.5. Those figures are the author’s reported results, not an independent evaluation.
The three live examples are a story-to-site relevance gate, an English-language check and a fallback topic classifier. Meyer reports that a scan of 78,889 articles cost $2.01; the language check flagged 1,576 non-English items, of which 1,553 were fixed. The classifier reportedly agreed with a frontier model 89% of the time overall. The article also lists 12 strong fits, seven cases that need measurement first and two poor fits across 24 mapped applications.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test for Routine AI Decisions
The playbook’s central argument is that a model call becomes useful when it can cheaply handle many small decisions, while software or staff retain control over what happens next. Meyer’s four conditions are high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that the current heuristic fails visibly. The last condition is intended to prevent adding AI where a simple rule already works.
That distinction affects both cost and risk. A high-confidence answer can trigger an automated action under a defined rule; a low-confidence answer can go to a person or a more capable system. Meyer recommends first replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements. His proposed threshold is to wire in a use only where the high-confidence band reaches 95%, then use a separate feature flag, a 5% to 10% canary and a staged rollout.
The reported scan offers a concrete example of the potential economics: a large batch of article checks cost $2.01, according to Meyer. But the article’s own categories also make clear that a low per-call cost is not enough. A task still needs a real, measured problem and a safe rule for handling mistakes.
From Publishing Checks to Other Uses
The playbook begins with publishing, where Meyer says 88% of the news items he processes start from a bare headline. He proposes a thin-source detector to decide whether material contains enough verifiable facts to support a factual report. That case is tagged “measure first” because the guide has not established that the existing heuristic fails enough to justify deployment. Other publishing examples include checking disclosures, moderating comments, assessing headline quality and matching products to roundup topics.
The article says a same-event deduplication idea is a poor fit for now: a canary found zero duplicates, so there was no demonstrated failure to address. By contrast, disclosure checks and comment moderation are presented as strong fits. Meyer says a missed disclosure can create compliance risk and recommends routing such cases to human review rather than publishing automatically.
The supplied article excerpt starts to introduce commerce and customer operations but ends before detailing those applications. It therefore supports the overall 24-use count and category breakdown, but does not provide the remaining use-case descriptions or enough detail to assess them individually.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, author of the playbook
Evidence Still Comes From One Operation
The results in the article are reported by Meyer; the supplied material does not include an independent audit, detailed evaluation data or the full set of 24 use cases. It does not identify the frontier model used for comparison, explain the sampling and labeling method behind the agreement figures, or give confidence-band sample sizes. Those omissions make it difficult to judge how well the reported accuracy would transfer to other organizations or tasks.
The article also leaves open whether the seven measure-first cases will meet the fit test after testing, and whether the recommended 95% high-confidence threshold is suitable for decisions with different costs of error. Meyer describes a staged rollout, but the excerpt does not report outcomes from canaries beyond the zero-duplicate finding.
Measure Before Wider Deployment
Meyer’s proposed next step for a prospective use is to replay 300 to 500 real past decisions, compare Jev’s answers across confidence bands and examine 20 disagreements. If a high-confidence band reaches the article’s 95% threshold, the deployment plan calls for a separate flag, initially off, followed by a canary covering 5% to 10% of units and a gradual rollout.
For the seven measure-first applications, the next milestone is evidence that the existing heuristic fails often enough to justify a trial. The article excerpt does not specify a publication schedule for further results or provide details on the remaining commerce, software, operations and home examples.
Key Questions
What is Jev, according to the playbook?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns structured answers such as probabilities, category choices or scores. The application code determines what action follows.
How many uses are live or considered strong fits?
The article maps 24 uses: three are live in Meyer’s publishing operation, 12 are rated strong fits, seven need measurement first, and two are poor fits.
What evidence does the article give for the live applications?
Meyer reports that a scan of 78,889 articles cost $2.01 and that an English-language check found 1,576 non-English items, with 1,553 fixed. These are results reported by the author, not independently verified in the supplied material.
What should happen when Jev is uncertain?
The playbook recommends routing uncertain answers to a person or a more capable system. It says software should act on clear cases according to a defined rule rather than treating every answer as equally reliable.
What does the excerpt leave out?
Although it gives the total of 24 use cases and the fit breakdown, the supplied text cuts off as the commerce and customer operations section begins. It does not show the rest of those examples or provide independent validation of the reported performance figures.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
