Why F1 Score Fails Ordinal Tasks (and What to Use Instead)
F1 treats a near-miss and a catastrophe as the same mistake. For ordered labels — essay scores, star ratings, severity grades — that’s a blind spot. Here’s the math that proves it, and the metrics that don’t lie.
Imagine you’re grading essays on a scale of 1 to 5. Two AI graders look at an essay that deserves a 5.
Grader A gives it a 4. Slightly harsh, but reasonable — it saw most of what makes the essay good.
Grader B gives it a 1. It completely misread the essay.
Now the uncomfortable question: if you evaluate both graders with F1 score, can the metric tell them apart?
No. F1 scores them identically. And that is the whole problem with using F1 on ordinal tasks.
Ordinal is not nominal
Most classification tutorials quietly assume your labels are nominal — categories with no natural order. Cat, dog, bird. Mistaking a cat for a dog is exactly as wrong as mistaking it for a bird. For nominal labels, F1 is a fine citizen.
Ordinal labels are different. They carry an order, and the order carries meaning:
-
Essay scores: 1 < 2 < 3 < 4 < 5
-
Star ratings, pain scales, cancer stages, credit ratings
Here, how wrong you are matters. Predicting 4 when the truth is 5 is a near-miss. Predicting 1 is a catastrophe. A metric that can’t distinguish the two is grading with a blindfold on.
The experiment: same F1, different models
Let me make this concrete. I simulated 100 essays with true scores from 1 to 5 (20 essays per score). Two models. Both get exactly 60% of essays spot-on.
-
Model A — whenever it’s wrong, it’s off by exactly one level. A 5 becomes a 4. A near-miss, every single time.
-
Model B — whenever it’s wrong, it’s off by two levels. A 5 becomes a 3. Consistently worse judgment.
Now the scores:
| Macro-F1 | MAE | QWK | |
|---|---|---|---|
| Model A (off by one) | 0.605 | 0.40 | 0.894 |
| Model B (off by two) | 0.607 | 0.80 | 0.565 |
Read that again. F1 calls it a tie — 0.605 vs 0.607, a difference so small it’s noise. Meanwhile MAE says Model B’s errors are twice as large, and Quadratic Weighted Kappa rates Model A as excellent (0.894) and Model B as mediocre (0.565).
Why F1 is blind (the 30-second version)
F1 is built from counts: true positives, false positives, false negatives — per class. A prediction is either a hit or a miss. There is no “close”. The formula has nowhere to put the distance between the predicted label and the true label:
Precision and recall are defined over sets of correct vs. incorrect decisions. “Off by one” and “off by four” both land in the same bucket: incorrect. Macro-averaging across classes doesn’t fix it either — averaging blind scores just gives you a blind average. The blindness is per-decision, not per-aggregation.
(Accuracy has the exact same disease, by the way. It’s just F1 with better PR.)
What to use instead
MAE (Mean Absolute Error). The simplest honest metric for ordinal tasks: the average distance between prediction and truth, measured in label units. “Off by 0.4 levels on average” actually means something. No tuning, no parameters, and it respects order by construction.
QWK (Quadratic Weighted Kappa). The field standard wherever ordinal grading matters — essay scoring, medical imaging, any rubric-based evaluation. Two things make it better than plain MAE:
-
Quadratic weights — a miss by two levels is penalized four times as much as a miss by one. In most real rubrics, far misses genuinely are disproportionately worse.
-
Chance correction — it measures agreement beyond what you’d get by guessing the label distribution, so it can’t be gamed by always predicting the majority class.
The penalty weight for confusing class with class (out of classes):
This is the metric I used in my own IndoASAG work on automated short-answer grading for Indonesian — late-fusion stacking of IndoBERT with morphological features reached a test QWK of 0.890. I didn’t pick QWK because it’s fashionable. I picked it because when a student’s answer deserves a 4, giving it a 1 is a genuinely worse failure than giving it a 3 — and the metric should know that.
And always: look at the confusion matrix. No single number survives contact with a weird error pattern. A quick glance at the matrix tells you where the model is confused, which is usually more actionable than any scalar.
The honest nuance
F1 isn’t “wrong”. It answers a specific question — did you hit the exact class? — and answers it well. If your task genuinely only cares about exact hits, F1 is fine.
The sin is never F1 itself. It’s the mismatch between the question your task asks (“how close are the predictions?”) and the question your metric answers (“how exact are they?”). Most ordinal tasks ask the first question. So most ordinal tasks shouldn’t be graded with F1.
The decision rule
-
Labels have no order (cat / dog / bird) → F1, accuracy, the usual nominal toolkit.
-
Labels have an order (scores, ratings, stages) → MAE for a quick honest read, QWK when you need the standard.
-
Either way → open the confusion matrix before you believe any number.
Your metric is a lens. F1 is a fine lens — just don’t use it to judge distance.
Discussion & Comments