Blog / Experiments
Why grouping records changes AI answers
One sentence, a whole interview, or an entire account: choosing what counts as a record can matter as much as choosing the question.
Before asking a model a better question, check whether you have given it the right piece of evidence.
An interview answer, an entire interview and a collection of interviews are three different inputs. Each can support different conclusions. Grouping related records is useful when the evidence for an answer is spread across them—but it also changes what you are measuring.
A tennis example
In our October 5, 2026 internal tests, predicting the match-result label from individual tennis question-and-answer turns produced 54% agreement. The corresponding task over grouped press conferences produced 98%. The report lists 3,055 conference groups, with 3,054 whole conferences scored.
A single turn might discuss training or playing conditions without saying who won. Other turns in the same conference can supply that missing context. Reading them together can make the label much easier to identify.
It would be misleading to describe the difference as a controlled 44-point improvement on the same test. A turn-level prediction and a conference-level prediction have different denominators and different available information. Both the input and the prediction unit changed.
A majority vote answers another question
We also counted the predictions from individual tennis turns. That approach produced 59% agreement, compared with 98% for the grouped reading task. The reported majority-label baseline was 71%.
Counting says which answer appears most often in the turn-level outputs. Group reading can use one decisive statement even when most of the other turns contain little evidence. Repeated weak evidence is not necessarily stronger than one explicit statement.
The opposite pattern appeared in the medication-review task: counting individual classifications reached 89%, while the group question reached 85%. The useful operation depends on the question. If you want a distribution of review labels, a count may be exactly right.
Choose groups that match the question
Start with a reliable grouping key. For an interview, that might be the interview ID. For project updates, it might be the project ID and a defined date range. For reviews, it could be a product ID.
A similar title or display name is not always the same entity. Combining two unrelated projects can create a confident answer from the wrong evidence. Inspect uncertain matches, and keep records that could not be matched visible.
Be equally careful about time. Combining comments before and after a change may hide the distinction you want to study. A question about “current concerns” needs a defined window rather than an unlimited history.
More context still has limits
Large groups may not fit in a model request. If you select only some records, document what was included and why. A sample of a long conversation is different from the whole conversation, even if both share the same group ID.
Nor does fitting everything guarantee success. In our historical NBA task, every game fit in context, yet closeness classification reached 26%, below the 38% majority baseline. Providing the full evidence did not make the model’s interpretation reliable.
Test at the level you will use
If your output is one answer per interview, review and score interviews. If it is one answer per article, review articles. Track omitted or incomplete groups, and compare against a simple baseline using the same unit.
That makes the result easier to trust and much easier to debug. Explore the sports and community examples for ways to apply this thinking to your own collection.
These percentages summarize reported internal experiments dated October 5, 2026. They are task-specific label-agreement results, not independently reproduced benchmarks. Sports tasks used historical records with final outcomes hidden; they do not measure future forecasting performance.