Better labels make AI classification easier to check

Two people can read the same survey response and choose different categories without either being careless. The labels may overlap, depend on unstated context, or ask for distinctions the text never makes.

AI classification has the same problem. Before changing models, check whether the task has a clear set of answers.

A label is part of the question

Consider a community survey with the categories “events,” “activities” and “meetups.” Without definitions, a comment about a weekend gathering could fit all three.

You could merge the categories, or define the distinction you actually need: organized public events, recurring clubs, and informal gatherings. Give examples of each and decide how to handle a response that covers more than one.

Also decide whether a record needs one label or several. Asking for a single “main topic” is different from asking which topics are mentioned. The same response can be correctly labeled differently under those two questions.

What our baseline tests revealed

An earlier internal baseline covering 40,000 rows reported 95.6% agreement on Amazon sentiment, but 35.9% on a consumer-complaint issue task with 84 answer options.

The report identified ambiguity among issue labels and the presence of options belonging to other products. A plausible next experiment is to identify the product first, then restrict the issue choices to that product. The report proposed this approach; it did not establish an improved result from it.

The lesson is to test the label design. A smaller, relevant choice set may make the task clearer, but that is a hypothesis until checked against held-out examples. These task-specific percentages are not a comparison of products or a universal model accuracy score.

Do not quietly change the answer space

In one binary sentiment task, the planner introduced “Mixed” and “Neutral” alongside the expected answers. That produced 109 extra-category responses counted as wrong.

Those categories might be useful in another project. Here they changed the task being evaluated. If the reference labels allow only positive or negative, adding a third answer needs an explicit decision about how the evaluation should treat it.

For your own collection, define an “unclear” option if it serves a real purpose. Keep it distinct from missing data or a failed request. Then decide what happens to those records when you calculate a share.

Review disagreements, not just the headline number

Look at examples from each category. Include the boundaries: short replies, mixed opinions, unfamiliar terms and records with too little context. Compare the model’s answer with your judgment and the evidence in the source.

When you disagree, write down why. Was the question underspecified? Did two labels overlap? Did the model miss a clear statement? Different causes need different changes.

Keep some reviewed examples aside when refining the question. Improving a rubric on the same few examples does not establish that it works on new records. Compare with a simple baseline, such as choosing the most common label, so a high score on an imbalanced collection does not mislead you.

Write the rubric you wish every reviewer had

A useful rubric includes the question, allowed labels, short definitions, boundary examples, a rule for ambiguity and the unit of analysis. Preserve it alongside the results, so a later run can be compared with the earlier one.

Start with a small collection and make the decisions visible. Explore the community feedback example to see a possible workflow, or join the Jevreduce waitlist to share what you would like to classify.

The baseline numbers above come from the Jevreduce team’s internal accuracy-and-cost test report reviewed in September 2026. They have not been independently reproduced for this article; label changes described as proposed remain unmeasured.