Star ratings are read more often than the reviews they accompany, and they are the part most likely to be quoted. The compression involved removes exactly the information a reader needs to act on.

Identical ratings describe different things

A middling score can mean a competent work with no ambition, or an ambitious one that fails in interesting ways.

Those are entirely different recommendations, and a reader who would enjoy one would be poorly served by the other.

The rating cannot separate them, because it reports a position on a single scale rather than the shape of the judgement.

The scale is not shared between critics

Reviewers use the range differently. Some reserve the top mark for exceptional work, while others award it to anything that achieves what it set out to do.

Publications rarely publish their conventions, so the same figure carries different meanings across outlets and sometimes within one.

Comparing ratings from different sources therefore compares scales as much as it compares works.

Genre expectations are baked in silently

Critics generally assess a work against what it is attempting, so a strong entry in a modest genre and a flawed attempt at something ambitious can receive the same mark.

The review explains that framing; the rating does not carry it, and once separated the number implies a comparison the critic never made.

This is the most common source of confusion when ratings are quoted in isolation from the text.

The rating is often not the critic's

At many publications the score is applied by an editor, and some critics decline to assign one at all.

Where a rating has to fit a house style, it may be adjusted for consistency with other coverage rather than reflecting the individual verdict.

Readers reasonably assume the number and the text come from the same judgement, and that assumption does not always hold.

The useful information is in the reasons

A review that explains what a work does well and where it fails allows a reader to predict their own response, which a figure cannot.

Finding two or three critics whose tastes diverge from your own in known directions is more informative than any aggregate, because the disagreement is legible.

Ratings remain useful as a filter for what to read about, which is a narrower job than the one they are generally asked to do.