Blog / Practical guide
When AI should judge—and code should count
Use models to interpret language, then use deterministic operations for counts and shares. Know when a group-level question needs a different kind of evidence.
Suppose you have a collection of game reviews. You want to know which mention performance problems, which praise the story, and what share of the collection recommends the game.
There are two jobs here. Interpreting a review requires judgment. Counting the resulting labels requires arithmetic. Giving each job to the right operation makes the result easier to inspect.
Turn interpretation into structured answers
A review might say “the story kept me playing even when the frame rate didn’t.” A model can help classify the topics and assess the sentiment expressed toward each one. Keep the original review, the question and the returned answer together.
Once each review has a label, a deterministic operation can calculate how many received that label and what share of the included records they represent. It does not need to reinterpret the text to do the count.
The count can be exact relative to the stored answers while the answers themselves contain mistakes. Always keep those two ideas separate.
Make the denominator visible
“Thirty percent mention performance” is incomplete without a definition of the collection. Does it include every review, only English-language reviews, or only reviews that passed a relevance filter? Were failed or missing answers excluded?
State the numerator and denominator together. Keep a separate count of records that could not be classified. If labels can overlap, their shares may add up to more than 100%; that is expected for multiple-topic questions, but should be clear to the reader.
Counting is not always the whole answer
Some questions concern the meaning of a group rather than the frequency of its labels. To understand a discussion, a model may need to read the related records together.
Our internal October 5 experiments illustrate both cases:
| Task | Counting record-level predictions | Reading the group |
|---|---|---|
| Medication-use classification | 89% agreement | 85% agreement |
| Tennis match-result inference | 59% agreement | 98% agreement |
These are reported results for specific tasks, not a general ranking of methods. Tennis turns can be individually ambiguous while a conference contains a direct statement of the result. Counting those ambiguous predictions does not recover the missing context. See the grouping guide for the prediction units and limitations.
Watch for arithmetic disguised as interpretation
In our historical NBA experiment, identifying the scoring team in individual events reached 99.9%. Yet asking for game closeness reached only 26%, below a 38% majority-label baseline. All games fit in context, so missing context did not explain that result.
A sensible workflow separates extracting the necessary events from calculating a total. The calculation must also model the rules correctly: identifying a team does not tell you how many points an event was worth. Validate the extraction and the arithmetic separately against known games.
Jevreduce’s current tally operations handle counts, shares and score averages over answers. They are not a promise of arbitrary numerical summation or a complete sports scoring engine.
Build a small, inspectable loop
Start with a question and a few records whose answers you can check. Review the model judgments. Count the stored answers with a clearly defined denominator. Open a surprising result and inspect the underlying records.
If the problem is an unclear label, refine the question. If it is an incorrect filter or denominator, fix the data operation. A single aggregate percentage cannot tell you which part went wrong.
Explore the review example or join the waitlist with a collection you want to analyze. The product is coming soon; the review scenario above is illustrative, and the reported experiments are not guarantees for another dataset.