The Nutrition Dex

Dietary Assessment

Floor and Ceiling Effects

Also known as: range restriction, saturation effect

The compression that occurs when an instrument or scoring scheme runs out of room at either end of its range, hiding real differences between the best or worst performers.

By James Oliver · Editor & Publisher ·

Key takeaways

  • A ceiling effect makes several genuinely different systems look identical because the scale cannot separate them.
  • A floor effect does the same at the bottom, which is why very poor performers often cluster.
  • Scoring schemes capped at 10 or 100 produce ceilings well before the underlying capability does.
  • In accuracy testing, an easy meal set creates an artificial ceiling: everything scores well and the test discriminates nothing.
  • The fix is a harder or wider test, not finer gradations on the existing scale.

Floor and ceiling effects occur when a measurement instrument or a scoring scheme runs out of range, so genuinely different subjects return indistinguishable values.

They are a property of the test rather than of what is being tested, which makes them easy to mistake for a real finding: "these four apps are basically equivalent" and "our test could not tell them apart" produce the same table.

The two forms

  • Ceiling effect. The task is easy enough that most subjects approach the maximum. Differences that exist are compressed into the last few points of the scale, where measurement noise dominates.
  • Floor effect. The task is hard enough that most subjects approach the minimum. The same compression, at the other end.

How it appears in accuracy testing

The most common instance is an easy meal set. Flat plated food and packaged items are handled well by every competent system, so a test built mostly from them produces a cluster of good scores and discriminates almost nothing.

Note how this interacts with composition bias: composition bias moves the level of a result, and a ceiling effect destroys its resolution. A single easy test set can do both at once, producing high scores that also fail to rank.

Scoring schemes create their own ceilings

A 0–10 or 0–100 score has a hard upper bound that the underlying capability does not. Once several products sit at 9.4–9.7, the remaining differences are being expressed in a range narrower than the scheme's own noise, and small changes in weighting reorder the table without anything real changing.

This is one reason ranked scores are less informative than the measurements underneath them, and why a published ranking should always expose the figures it was computed from.

What fixes it

A harder or wider test — not finer gradations on the existing scale.

Adding decimal places to a compressed score creates the appearance of resolution without adding any. The corrective is to include cases that actually separate the subjects: deep vessels, layered dishes, low light, small portions. If everything still clusters after that, the finding that the systems are equivalent is real.

Frequently asked

What is a ceiling effect in app testing?

It is when the test is easy enough that most systems approach the maximum score, so real differences between them are compressed into the last few points where measurement noise dominates. The usual cause is a meal set made mostly of flat plated food and packaged items, which every competent system handles well. The resulting table looks like a finding that the apps are equivalent, when what it actually shows is that the test could not tell them apart.

How is a ceiling effect different from composition bias?

Composition bias moves the level of a result — an easier menu produces better numbers for everyone. A ceiling effect destroys the resolution of the result — it removes the ability to rank. An easy meal set typically causes both at once, which is why it produces high scores that are also unable to separate the products being compared.

How do you fix a ceiling effect?

By making the test harder or wider, not by adding decimal places to the score. Extra precision on a compressed scale creates the appearance of resolution without adding any. Include the cases that actually separate systems — deep vessels, layered dishes, low light, small portions — and if everything still clusters afterwards, then the equivalence is a real finding rather than an artefact.

References

  1. Terwee CB, Bot SDM, de Boer MR, et al.. "Quality criteria were proposed for measurement properties of health status questionnaires". Journal of Clinical Epidemiology , 2007 .
  2. Streiner DL, Norman GR, Cairney J. "Health Measurement Scales: A Practical Guide to Their Development and Use". Oxford University Press , 2015 .

Related terms