Blog / Experiments
What we learned from 714,480 rows
Seven datasets taught us to separate processing scale from answer quality, choose the right unit of analysis, and check the questions as carefully as the model.
A large run can finish successfully and still answer the wrong question. That is the most useful lesson from our recent Jevreduce experiments.
Our October 5, 2026 internal test report covers seven datasets and 714,480 rows, with no failed rows reported. The collection spans reviews, public text and sports records. Most datasets contained 100,000 rows; software reviews contained 114,155 and medication reviews 100,325.
Those numbers tell us that the processing completed in these runs. To understand whether the answers were useful, we had to look at each task separately.
What we measured
We asked structured questions about individual records and groups of related records, then compared outputs with dataset labels. The tasks included identifying sentiment, inferring a label from related text, and interpreting historical sports events with the final outcome hidden.
This article summarizes our team’s reported results. It is not an independent benchmark or a controlled comparison against other tools. Different tasks use different prediction units, labels and scoring rules. Averaging their percentages into one product accuracy score would hide the most useful information.
Context changes the task
Individual tennis question-and-answer turns produced 54% agreement with match-result labels. Reading turns together as press conferences produced 98% across the scored conference groups. There were 3,055 groups, of which 3,054 whole conferences were scored.
A short answer about preparation may reveal little about whether a player won. A conference can contain a direct discussion of the result. The grouped task has different evidence and a different unit of prediction; it is not simply the same test with a better percentage.
The pattern appeared in public political text too: individual tweets produced 81% agreement with party labels, while grouped politician accounts produced 218 correct answers out of 219, or 99.5%. Again, an account contains more context than a tweet. Read the grouping guide before treating either number as a forecast for another dataset.
Counting and interpreting are different operations
For the medication task, counting the classifications of individual reviews produced 89% agreement, compared with 85% for asking the model about a group. In the tennis task, counting turn-level classifications reached 59%, compared with 98% for reading the conferences together.
Neither operation wins everywhere. Use counting when the question concerns the distribution of record-level labels. Use group context when related records jointly contain the evidence needed for an answer. Our guide to model judgments and counts walks through this distinction.
Extracting a fact is not the same as calculating a result
The NBA experiment made the limitation especially clear. Scoring-team extraction from individual events reached 99.9%. Game-winner inference reached 74%, and closeness classification reached 26%, below the 38% majority-label baseline.
All NBA games fit in context in that run. We cannot blame the errors on truncated groups. Recognizing the team that scored does not establish that a model can reliably reconstruct a scoreboard. These were retrospective tasks, not predictions about future games.
Make the test fit the question
For a new collection, define a useful unit first: an article, a review, an interview, or a project. Write down the allowed answers. Check a sample with known labels, including ambiguous examples and the most common answer as a simple baseline.
Then inspect disagreements. A confusing label set, a missing record or a question with the wrong scope can look like a model failure. Conversely, a successful import says nothing about whether the labels mean what you intended.
We are using these lessons to build a workflow where you can inspect inputs, review answers and try another question without losing the earlier run. Explore an example or join the waitlist with a collection you would like to understand.
Source and scope: Jevreduce internal scale-test report dated October 5, 2026. Results are reported team measurements, rounded as in the report. Raw run artifacts are not published here; reproducibility and transfer to other datasets have not been established. No claim about million-row performance, customer pricing or comparative speed is made from these tests.