Methodology

How we score

A score you cannot interrogate is just an opinion with a decimal point.

Every product on this site carries two scores. This page explains how both are produced, what each criterion reads, what counts as evidence and what does not — and where the method reaches its limits.

Report a score you disagree with We would rather hear it than not.
301 products currently carrying a usability score
8 criteria, each scored against a fixed table
8.10 the highest score anything on this site holds
0 products we have tested ourselves. We say so below.
01

Two scores, kept separate

One number forces two different questions into the same answer: is this product any good, and what do people who own it say. Those can disagree — and when they disagree, that is usually the most useful thing on the page.

5.3
The usability score
Our assessment, against eight fixed tables

Reads specifications, certifications and documentation.

It does not read reviews.
6.0
The social score
What five independent pools say

Aggregates what independent sources say, across the five places people already look.

It does not read specifications.

A product can carry one and not the other. That is three honest statements rather than one invented one.

02

Eight criteria, chosen per category

What stays constant is the structure. What the eight criteria are depends entirely on what is being compared — flow rate matters for a water filter and means nothing for a mattress.

Worked example
Water Filtration — the first category on this site, and so far the only one. Everything from here to the end of this section describes its eight criteria and nothing else. A second category gets its own eight, its own anchor tables and its own weighting, written up here the same way; none of it carries over.
Criterion What it reads Weight
Contaminant Reduction The count of distinct contaminant classes the brand documents, plus certification as evidence of published performance. Never a count of filter media.
Heaviest
Certification Coverage The share of the product's published claims that a listed certification actually covers. Never a count of standards held.
Heavy
Flow Rate The published GPM figure, against household demand.
Above standard
Filter Life The shortest routine service interval — what the owner actually has to do, and how often.
Weighted down — see below
Build Quality Housing and tank material, valve and fittings, construction quality.
Standard
Warranty & Support Published duration of the manufacturer's promise. Conditions go in the text, not the number.
Standard
Ease of Installation Connection type, bypass and unions, whether a plumber is needed, quality of documentation.
Weighted down — see below
Value — computed, not assigned Deviation from the price its specification predicts, against genuine peers. Never an absolute price.
Above standard

Relative weight only. The exact weights are part of the calculation and are not published — see What we publish, and what we don't.

Why certification carries so much

Equal weighting produced a result we thought was wrong. A certified system that removes PFAS ranked below a cheaper uncertified one, because filter life and warranty together outweighed contaminant performance and certification two to one. Certification was worth an eighth of the score. It is now worth about a fifth.

Why two criteria are weighted down

Ease of Installation puts more than half the category into two adjacent values. Filter Life is scored from a published interval band, which puts most reverse osmosis products on the same value. A criterion that returns nearly the same answer for every product is not measuring anything, and pretending otherwise by weighting it heavily would be worse than admitting it.

End of the worked example. Everything after this point is the method itself and holds whatever is being compared.

03

What we publish, and what we don't

There are two different things you could want from a page like this, and only one of them is yours.

The scale is published

You need it to read a score. Without it, a 7 is a number with no meaning.

What each criterion measures Which inputs it reads — and which it deliberately ignores Where the top and bottom of each scale sit What counts as evidence, and what gets thrown out Where the method reaches its limits
The calculation is not

That is the work. It took months to build, correct and re-correct. It is the product.

The exact cut points in each table The weight each criterion carries The models behind the computed criteria The internal procedures that produce and audit a score

Everything you need to judge whether a score is defensible is on this page. Everything you would need to clone the scoring engine is not. If a score does not square with what is described here, that is a discrepancy worth reporting.

04

What the scores actually look like

Everything above describes how a usability score is produced. This is what came out, across all 301 products currently carrying one.

Median5.29
Middle half — the darker bars4.81 – 6.05
Highest on the site8.10
Scoring 9 or above0

Why nothing is near 10

10 is the best system it is possible to buy, not the best on this page. The anchor tables are written against the market, not against the catalogue — a 10 on Flow Rate is the largest residential demand with headroom, and a 10 on Certification Coverage is a listing covering substantially every claim the product makes, verifiable in a certifier's own database. Very few products are that, and none of them is every one of those things at once.

The alternative is to anchor the top of each scale to whatever the best product in the catalogue happens to do, which makes every score a rank in disguise and moves them all whenever anything is added. It also guarantees a 10 exists, which is the flattery this whole page is written against.

The number you never see

A product page shows two rings and no third number. The tier badge beside them, and the line saying where the product places in its category, are both computed from a figure that is not on the page: the plain average of the two rings, weighted equally.

Worked from the middle of this catalogue: a usability score of 5.29 beside a social score of 6.00 carries a 5.65, and it is that 5.65 the badge and the field position read. The two halves are kept apart everywhere a reader can see them, because when they disagree the disagreement is the most useful thing on the page — but one number is needed to place a product against its peers, and this is it. That 5.65 is an illustration built from the two medians, and it is deliberately not the same object as the figure in the next paragraph: the middle of the blends and the blend of the middles are different quantities, and on this catalogue they land a few hundredths apart.

Across the 301 products carrying both halves, that blended figure runs a median of 5.70, a middle half between 5.31 and 6.12, and a best of 7.56. It is higher and tighter than the usability distribution above, because the social half runs higher than the usability half across most of the catalogue — which is itself a finding, and the reason the two are never averaged into a single displayed score.

The badge itself is a position rather than a verdict — Top, Upper, Mid, Lower — and it is cut inside a product's own category rather than across the whole site, because a 7 in Whole House and a 7 under a sink were earned against different tables. It used to be a fixed cut applied across everything, with the bottom band labelled as a judgement. That was replaced on 10 September 2026, and no stored score changed when it was.

Every figure in this section is read from the live catalogue when the page is built, not written into it.

05

The social score

Unlike the criteria above, this half does not change by category. The five places people look before buying something are the same whether it is a water filter or a mattress.

Reddit
6.2
Web coverage
5.5
YouTube
7.0
Amazon
Not sold on Amazon — no returns window, no third-party price
Retailer
5.3

Every product shows all five slots. An empty pool stays visible, with the reason underneath it where we know it — because an omitted row and a row nobody checked look identical to a reader. Two populated pools is the minimum for an overall: one pool is an anecdote wearing an average.

The cascade

For each pool we work down three levels and stop at the first that yields real signal.

Level 1 · strongest Product Coverage naming this exact model.
Level 2 Product line The brand's family sharing a chassis and media.
Level 3 · scored down Brand Sentiment about the company and its products generally.

Every summary opens by declaring which level it reached. Falling back is normal, not a concession — but a pool describing different hardware is weaker evidence, so brand-level findings are scored down. A pool is blank only when nothing exists at any level anywhere, and that has to be earned by searching.

06

What does not count as an independent source

This is the section that excludes the most, and it is where most of the work goes. Each of these is a tell we check for by hand, on every pool.

01 A manufacturer's own review widget is marketing However precise the number looks, and whichever third-party platform serves it. What matters is who controls collection.
02 Syndicated retailer reviews are the same widget in a retailer's badge “Customer review from {Brand}” on a big-box listing is the manufacturer's file republished — often across the retailer's whole catalogue for that brand.
03 A mixed pool's star average is not usable One system reads 4.6 stars alongside a 60% recommend rate. Those do not describe the same population — and the gap is the finding.
04 A brand-store aggregate is not a product rating One figure across everything a brand sells, usually dominated by cartridges. Legitimate brand-level evidence, scored down for scope. Never presented as a product figure.
05 A seller rating is not a product rating “Sold and shipped by {Brand}, 4.5 from 145 seller reviews” sits where a product average would. It measures despatch.
06 A five-star-only count is not a rating “3,428 five-star reviews” is a numerator with the denominator withheld.
07 Identical star-and-count pairs mean one review file Two products with genuinely independent populations do not land on the same count. Across two unrelated retailers, the tell is stronger still.
08 A replacement cartridge's rating is evidence about the cartridge Real evidence about what owners live with, sometimes the largest pool available — but not about the housing, the fittings or the valve.
09 Identical text across unrelated communities is excluded as a class The tell is verbatim sentence repetition where there is no shared membership. Any unusually fluent recommendation gets one sentence searched first.
10 A reviewer excluded once is excluded everywhere Re-deciding brand by brand means the question gets re-asked whenever a pool looks thin — which is exactly when the answer is least trustworthy.
11 A competitor's review of a brand is not neutral Several comparison pieces in circulation are marketing by rivals — including at least one written by a company whose own products appear on this site.
12 Read the distribution, not the average One profile reads 4.0 from 54 reviews — which resolves to 50% five-star against 39% one-star. That is two populations, not one average.
07

What this method cannot do

Publishing a method means publishing its limits.

Scores are judgement applied consistently, not measurement We do not test products. Nobody scoring a catalogue this size does. What the anchors buy is consistency — two products scored months apart are comparable because both were scored against the same fixed table.
A criterion scored from a band cannot separate close values
Water Filtration
Water filtration's Filter Life reads a published interval band, so it cannot separate a brand publishing 6 months from one publishing 12. That is why it carries a low weight — the honest response to a criterion that discriminates poorly is to weight it down, not to pretend it discriminates.
Total cost of ownership is not scored
Water Filtration
A $40-a-year cartridge in front of a ten-year tank is a better proposition than three changes a year, and the filter life criterion cannot see the difference. Where it matters, it is in the review text.
Warranty conditions are not scored
Water Filtration
Every lifetime warranty in this category carries some. Reading them into a number requires a legal judgement per product that we are not qualified to make.
Nothing here is advice about your water
Water Filtration
The right system depends on a test of your own supply, on your plumbing, and sometimes on local code. A score is a comparison between products, not a recommendation for your house.
08 · Corrections

A method with no recorded failures has not been examined

Several of the examples on this page are our own errors, caught and fixed. That is deliberate.

When an anchor table changes, every score taken against the old one is invalid and gets recomputed. We do not silently edit a table and leave old scores standing under it.

Found something wrong?

A specification, a certification claim, a score you think the method does not support. Corrections are made on the evidence, regardless of who is asking.

contact@stuffvsthings.com Position is not for sale. See our affiliate disclosure for the commercial side.
Stuff VS Things
Logo
Compare items
  • Total (0)
Compare
0
Shopping cart