[M1] Methodology

How this index is built, classified and corrected.

Every quantitative claim on Datacenter.computer traces back to a database query over sourced records. This page documents exactly what is counted, how it is classified, what is modelled and what is unknown.

01

What “indexed” means

The index contains 227 individually sourced data-centre facilities. Each record was compiled from public operator documentation, cloud-region documentation, peering registries, regulatory or planning filings, or public announcements.

Indexed facilities are individually sourced physical data-centre records: each has a real facility name, operator, city, country, traceable primary source, verification status, confidence classification and last-verified date. Candidate records awaiting source verification are held in an internal queue and are never counted, published or indexed. A separate modelled metro-level market-coverage research layer is never described as a facility.

Sourced facilities
227 records with a public basis
Accelerator-linked
0 facilities with a publicly reported accelerator deployment
Operators
11 distinct operators
Countries
27 countries
Cities / metros
87 locations
Recorded capacity
0 MW across 0 records with a capacity figure
Modelled research layer
182 metro-level market signals — analytical context only, never counted as facilities
02

Global records vs enriched profiles

A sourced facility record is a specific site, campus or cloud region with a named operator and a public basis. It receives a permanent ID and a public page.

A modelled market signal is a metro-level density index derived from published metro facility totals. It contains no facility name, no operator, no capacity, no PUE and no accelerator data. Market signals are excluded from facility counts, public facility pages, rankings, the API and every XML sitemap. A synthetic facility generator previously used for market coverage has been permanently retired.

03

Data-source hierarchy

Records are classified by the strongest available source, in this order:

Operator Confirmed
Published by the operator itself (site page, spec sheet, press release).
Government Confirmed
Planning permission, grid allocation, or regulatory filing.
Third-Party Confirmed
Registry or independent dataset (e.g. peering registries, exchange listings).
Publicly Reported
Credible press or industry reporting without operator confirmation.
Modelled
Derived figure. Never presented as a facility.
Unverified
Flagged for review; excluded from ranking claims.
04

Capacity classification

Capacity figures are not comparable unless their basis is stated, so every MW value carries a classification: IT Load, Utility Allocation, Campus Capacity, Building Capacity, Estimated or Unknown. Where the basis is not published, the figure is classified as Estimated and displayed as such. Campus-scale figures are never presented as single-building IT load.

05

Efficiency and accelerator claims

PUE values are shown as operator reported or estimated. Datacenter.computer does not meter facilities, does not audit power and does not run independent verification, so no figure on this site is described as measured, audited or independently verified. Accelerator information reflects publicly reported deployments; where a deployment is not publicly supported, the field reads “Not publicly disclosed”.

06

Confidence scoring

Each record carries one of five levels — Confirmed, High, Medium, Estimated, Unverified — set by source strength, field completeness and recency of the last check against the source. Records without a stored source URL default to Medium and are flagged for review in the internal data-quality report.

07

Validation and duplicate handling

An automated validation pass runs over the whole index and checks for:

duplicate IDs and slugs; the same facility under slightly different names; operator/facility name confusion; city–country mismatches; coordinates outside the stated country; impossible coordinates; cloud regions presented as single buildings; planned campuses marked operational; campus capacity presented as building IT load; unsupported accelerator and PUE claims; suspiciously repeated capacity or PUE values; and records missing a source or a verification date.

Uncertain records are flagged for human review, never deleted automatically. Records with error-severity issues are withheld from public pages and from the sitemaps until resolved.

08

Update frequency

The index is reviewed continuously and re-published with each dataset release. A record not checked against its source within 180 days is flagged stale. Sitewide statistics are computed at build/request time from the live dataset — no public number on this site is hard-coded.

09

Ranking methodology

Rankings order records by the stated field (recorded capacity, reported PUE, accelerator class) and only include records whose relevant field has a public basis. Records with the field unknown are excluded from that ranking rather than assigned a placeholder value. Rankings are editorial comparisons of published data, not endorsements or certifications.

10

Corrections policy

Operators and readers can dispute any field. Send the facility URL, the disputed field, the corrected value and a supporting public source via the contact page. Every submission enters a moderation queue; nothing is published without review. Corrections that change a quantitative field also update the record’s source, confidence level and last-verified date.

11

Known limitations

Coordinates are metro/campus-level approximations, not surveyed addresses. Cloud regions can span multiple buildings and are indexed as one record. Capacity figures mix IT load and utility allocation across operators because operators publish them inconsistently. Accelerator counts are rarely disclosed and are omitted rather than estimated. Not every record currently stores a machine-readable source URL; those records are flagged in the internal data-quality report and are being backfilled.