How we compile these figures
This page describes what happens to a drive’s listing between being scraped and appearing on a comparison. It includes the parts that go wrong.
We do not review drives
Nothing here is a review in the usual sense. We do not buy, test or handle the drives. This page describes how we handle records — what a manufacturer publishes about a drive — not how we handle the drive.
Building the corpus
Listings are gathered across the searches that define the category, then filtered to the products that actually belong. Search results are full of things that are not drives: enclosures, adapters, cables, docking stations and rack hardware, many of which carry a capacity figure in their own title and would otherwise be read as drives.
Every drive is assigned to a class from its own title and specification table, never from the search that surfaced it. That distinction is not academic — searches for one kind of drive return the other kind in quantity, and a class built from search terms is wrong in both directions.
What this category’s listings get wrong
- Amazon files the SATA bus rate as a read speed. “6 Gb/s” arrives as 600 MB/s, and once as 6,000 when the bits-and-bytes slip runs the other way. No spinning drive sustains either figure, so anything at or above 300 MB/s is treated as the interface rather than the platter and the field reads not stated.
- Capacity searches return the same drives repeatedly. A class built from search terms would have held nine near-identical groups of the same products, so classes here are assigned from each drive’s own record instead.
- Capacities are stated in binary and decimal interchangeably. Bands here follow the binary figures the drives actually report, so a 4 TB drive is placed by its 4,096 GB rather than a rounded 4,000.
Baselines
A baseline is the median of a specification within one class and one capacity band, with values beyond 1.5× the interquartile range excluded so a single outlier cannot move it. Baselines are published only where the group holds at least 8 drives; below that a median describes the sample rather than the market, and the page says no baseline can be published rather than printing a number that looks authoritative and is not.
A specification stated by fewer than half a group’s drives is withheld from that group’s baseline for the same reason.
Prices
We publish price tiers, never exact amounts: a price captured on a scrape date is wrong by the time you read it. The one derived figure we do print is price per terabyte, because it is the only number that makes drives of different capacities comparable and no listing states it — and it is shown against a class median rather than as a claim about what anything costs today.
Where this fails
- A published figure can be optimistic, or measured under conditions the listing does not give. We can check a figure against its own class; we cannot check it against a drive.
- Automated cleaning catches the defects it has rules for, and every rule here exists because something got through first. The next class of defect is one we have not seen yet.
- Coverage is uneven. Some fields are stated by nearly every listing and some by a third, and a field stated by few drives supports a much weaker conclusion.