A score you cannot interrogate is just an opinion with a decimal point.
Every product on this site carries two scores. This page explains how both are produced, what each criterion reads, what counts as evidence and what does not — and where the method reaches its limits.
One number forces two different questions into the same answer: is this product any good, and what do people who own it say. Those can disagree — and when they disagree, that is usually the most useful thing on the page.
Reads specifications, certifications and documentation.
A product can carry one and not the other. That is three honest statements rather than one invented one.
What stays constant is the structure. What the eight criteria are depends entirely on what is being compared — flow rate matters for a water filter and means nothing for a mattress.
Relative weight only. The exact weights are part of the calculation and are not published — see What we publish, and what we don't.
Equal weighting produced a result we thought was wrong. A certified system that removes PFAS ranked below a cheaper uncertified one, because filter life and warranty together outweighed contaminant performance and certification two to one. Certification was worth an eighth of the score. It is now worth about a fifth.
Ease of Installation puts more than half the category into two adjacent values. Filter Life is scored from a published interval band, which puts most reverse osmosis products on the same value. A criterion that returns nearly the same answer for every product is not measuring anything, and pretending otherwise by weighting it heavily would be worse than admitting it.
End of the worked example. Everything after this point is the method itself and holds whatever is being compared.
There are two different things you could want from a page like this, and only one of them is yours.
You need it to read a score. Without it, a 7 is a number with no meaning.
That is the work. It took months to build, correct and re-correct. It is the product.
Everything you need to judge whether a score is defensible is on this page. Everything you would need to clone the scoring engine is not. If a score does not square with what is described here, that is a discrepancy worth reporting.
Everything above describes how a usability score is produced. This is what came out, across all 301 products currently carrying one.
10 is the best system it is possible to buy, not the best on this page. The anchor tables are written against the market, not against the catalogue — a 10 on Flow Rate is the largest residential demand with headroom, and a 10 on Certification Coverage is a listing covering substantially every claim the product makes, verifiable in a certifier's own database. Very few products are that, and none of them is every one of those things at once.
The alternative is to anchor the top of each scale to whatever the best product in the catalogue happens to do, which makes every score a rank in disguise and moves them all whenever anything is added. It also guarantees a 10 exists, which is the flattery this whole page is written against.
A product page shows two rings and no third number. The tier badge beside them, and the line saying where the product places in its category, are both computed from a figure that is not on the page: the plain average of the two rings, weighted equally.
Worked from the middle of this catalogue: a usability score of 5.29 beside a social score of 6.00 carries a 5.65, and it is that 5.65 the badge and the field position read. The two halves are kept apart everywhere a reader can see them, because when they disagree the disagreement is the most useful thing on the page — but one number is needed to place a product against its peers, and this is it. That 5.65 is an illustration built from the two medians, and it is deliberately not the same object as the figure in the next paragraph: the middle of the blends and the blend of the middles are different quantities, and on this catalogue they land a few hundredths apart.
Across the 301 products carrying both halves, that blended figure runs a median of 5.70, a middle half between 5.31 and 6.12, and a best of 7.56. It is higher and tighter than the usability distribution above, because the social half runs higher than the usability half across most of the catalogue — which is itself a finding, and the reason the two are never averaged into a single displayed score.
The badge itself is a position rather than a verdict — Top, Upper, Mid, Lower — and it is cut inside a product's own category rather than across the whole site, because a 7 in Whole House and a 7 under a sink were earned against different tables. It used to be a fixed cut applied across everything, with the bottom band labelled as a judgement. That was replaced on 10 September 2026, and no stored score changed when it was.
Every figure in this section is read from the live catalogue when the page is built, not written into it.
This is the section that excludes the most, and it is where most of the work goes. Each of these is a tell we check for by hand, on every pool.
Publishing a method means publishing its limits.
Several of the examples on this page are our own errors, caught and fixed. That is deliberate.
When an anchor table changes, every score taken against the old one is invalid and gets recomputed. We do not silently edit a table and leave old scores standing under it.
A specification, a certification claim, a score you think the method does not support. Corrections are made on the evidence, regardless of who is asking.
contact@stuffvsthings.com Position is not for sale. See our affiliate disclosure for the commercial side.
The social score
Unlike the criteria above, this half does not change by category. The five places people look before buying something are the same whether it is a water filter or a mattress.
Every product shows all five slots. An empty pool stays visible, with the reason underneath it where we know it — because an omitted row and a row nobody checked look identical to a reader. Two populated pools is the minimum for an overall: one pool is an anecdote wearing an average.
The cascade
For each pool we work down three levels and stop at the first that yields real signal.
Every summary opens by declaring which level it reached. Falling back is normal, not a concession — but a pool describing different hardware is weaker evidence, so brand-level findings are scored down. A pool is blank only when nothing exists at any level anywhere, and that has to be earned by searching.