- Data
- Methodology
Methodology
How MarketCode values property and builds its indices
The valuation, the five indices, the market cube, the rankings and the auction discount are MarketCode models built on public and licensed records. This page says how each is built, what its confidence measures mean and where it is weak, in the same terms the tools use when they answer. Last revised 6 September 2026.
- UK addresses
- 34.7M
- sales since 1995
- 28M
- energy certificates
- 30M
- current-value estimates
- 16.8M
Everything is keyed on the UPRN
Every address in Great Britain carries a Unique Property Reference Number (UPRN), assigned by local authorities and published through Ordnance Survey AddressBase. MarketCode joins every property-level record to it: sales, energy certificates, council tax bands, listings, titles and building footprints.
A question about "12 Lordship Lane" is answered by resolving the address to its UPRN, then reading what has been matched to it.
Matching is conservative. A sale that cannot be placed on a UPRN with confidence stays in the sold-price record for area statistics and is attached to no property.
A council tax band in the VOA list that is not linked to the UPRN is reported as not found, and the response says the gap is ours. Every field in the full record carries a source and a confidence score, so a reader can trust a cell, not just a row.
Where it is weak
- Flats in subdivided buildings are the weakest join: some datasets record the building, not the unit, and a per-unit count over them over-reports.
- Ordnance Survey covers Great Britain; HM Land Registry covers England and Wales. A Scottish address resolves but has no sale history here.
The valuation: an estimate, a range and a trust score
The valuation is a MarketCode model, not a licensed third-party AVM. It is trained on the registered sales of England and Wales, joined to floor areas from energy certificates and to the physical form of each building, and it predicts a current value for every property with a floor area on record.
The current release is a gradient-boosted model over spatial cells, re-based every month to the published local index so the estimate moves with the market between retrains.
Three things come back with the number. A range: the model's own uncertainty for that property, wider where the evidence is thin. A back-series: the value carried backwards by the local price-per-square-metre index, ending at the headline figure by construction, so it shows the path rather than a second opinion.
And a trust score, which combines the index sample behind the area, where the floor area came from (a certificate, measured geometry or a fill) and the model's confidence. Two identical numbers can carry different weights.
Comparables come from the same building and street first, with missing attributes backfilled from sibling records so the set is usable rather than sparse.
Where the property has sold, its own last sale is the strongest anchor, and the model is measured against exactly that: how far the estimate sits from the price it actually made, adjusted to today by the index.
Where it is weak
- No floor area, no valuation. A property with no certificate and no measured geometry returns 404 rather than a guess; coverage of floor area is the main lever on coverage of value.
- It is an automated estimate, not a RICS Red Book valuation. It has not inspected the property, does not know its condition and cannot see a lease. Use it to screen, to challenge and to size; not as the valuation a lender relies on alone.
- Blocks with few sales are valued from their neighbours. The trust score says so; the number does not look any different.
The house price index: repeat sales, by district and authority
The sold-price index is a repeat-sales index, built from pairs of sales of the same property. It measures price change without being pulled around by which properties happened to sell that month.
It is published monthly from 1995 for every postcode district and local authority with enough pairs, base period = 100. The forecast tail is flagged, with a lower and upper band on each forecast point.
Each series carries its own provenance. Where an area has enough repeat sales, the method is recorded as area-level. Where it does not, the national index is published against the area and the method says so: a turn in that series is national, not local.
The last settled period is the last month the registry has filled. Registration lags completion by weeks to months, so the most recent points are provisional and will rise as late registrations arrive.
The index is the same series that shapes the valuation back-series, so a value carried to today with it and a value read from the model agree by construction.
Where it is weak
- A repeat-sales index cannot see a property's first sale, and it treats a refurbished property as the same property. Heavy improvement between sales biases it upward slightly.
- Quarterly series are published only where a monthly one would be too noisy; asking for quarterly elsewhere returns an empty series with a note, not a resampled one.
Tools: market_index_series
The asking-price and asking-rent indices: hedonic, from listings
Asking prices lead sold prices by months and include what never sells, so they are modelled separately.
Both asking indices are hedonic: a regression prices the attributes of each listing (type, bedrooms, floor area where stated, location), so a month with more large houses on the market does not read as a price rise.
They are published for postcode districts, local authorities and Great Britain, 2022M01 = 100, with 80% bands and a posterior standard error on every period's move, smoothed toward the parent geography where a district is thin.
Each area is assigned a publication frequency: monthly where the standard error of the monthly move is a percentage point or less, quarterly otherwise. The response says which. A monthly move quoted for a quarterly area is noise.
The rent index has a referee. Every build is compared with the ONS Price Index of Private Rents by local authority. A month where the two disagree by more than a few points a year is not published, so a missing month is a refusal, not an outage.
Where it is weak
- These are asking figures. The level of an asking index is never a price or a rent; the listing statistics carry the medians.
- Listings that are later reduced lose their listing date in the portal data, so the count of new listings is a never-reduced cohort and undercounts by the reduction share.
- The daily stock series began in September 2026 and cannot be reconstructed backwards: the portals record no exit dates.
Tools: asking_price_index_series, asking_rent_index_series, listing_market_series, listing_stock_series
The commercial indices: repeat sales for capital, rating lists for rent
Commercial capital values use the same repeat-sales method as the residential index, quarterly from 1995. Three families exist: national all-sector; national by sector where there are enough pairs (office, retail, hotel, holiday let); and local authority all-sector for nineteen authorities.
There is deliberately no local-by-sector series: a local sector cell is too thin to carry an index. The two published series are shown side by side instead.
Commercial rents come from a different source. The VOA assesses an open-market rent for every non-domestic property at each rating list's valuation date.
Linking the same properties across the 2010, 2017, 2023 and 2026 lists gives a rent index in four steps for every sector and local authority, including the industrial stock that trades in portfolios and never appears in Price Paid.
Where it is weak
- The rent index moves in four steps dated at the valuation dates. Interpolating between them as if it were monthly invents movement.
- Capital and rent are not interchangeable: capital = rent / yield, and yields moved.
Tools: commercial_index_series, commercial_rent_index_series
The market cube: prices by cohort, with the evidence graded
Area prices, price per square metre and turnover are pre-computed for every postcode district and local authority, by asset class, bedroom band and month.
Every cell carries a decision saying how much evidence sits behind it: direct (thirty or more sales), modelled (five to twenty-nine), suppressed (one to four), imputed (a prior borrowed from the parent geography because the cell had no sales) and missing.
Only a small minority of cells are direct. Most are thin or borrowed, which is why the decision is returned with every number.
Turnover is sales as a share of the stock that could have sold: the liquidity measure. It is exact at every grain because it is recomputed from counts.
Rankings across areas aggregate a trailing window with the most recent months dropped, because HM Land Registry registers sales late and the newest month would rank registration lag rather than liquidity. The window is returned with every ranking.
Counts of sales come from the registry directly. The cube's own count is a UPRN-matched sample that runs short of the registry, so the platform answers "how many sales" from the registry and "what price" from the cube, and says which.
Where it is weak
- An imputed price looks exactly like a measured one and repeats across cohorts in the same district. Read the decision before quoting a price.
- Rental value, gross yield and days on market are present in the schema but not populated in the cube; they are named as empty in every response.
Tools: market_facts, market_ranking, market_volume_series, market_volume_ranking
The auction discount: hammer price against the valuation of the same property
Auction outcomes are collected from five auction houses and matched to addresses and UPRNs. Where a lot has an achieved price and the property has a valuation, the ratio between them is the auction discount.
Aggregated by district and period, it is a measure no listings portal can produce: a portal knows a property went to auction, not what it made.
The ratio is only a discount where the sale and the valuation are close in time. Every valuation shares one snapshot date, so an old lot carries years of market movement as well as any discount.
The response therefore breaks the ratio out by year and flags whether the sample is sufficient. Every auction tool says to read those before the headline.
Where it is weak
- About a quarter of lots carry an achieved price and a fifth a net yield, because most auctioneers publish a guide and never the result. Coverage is a constraint, not a caveat.
- Guide and achieved averages are computed over different populations; subtracting one from the other gives no meaningful number.
Tools: auction_discount, auction_comps, auction_stats, distressed_assets
What is read, never modelled
Some attributes are declared facts and are treated as such. A council tax band is read from the VOA list or reported as not found, never predicted. An energy rating is the lodged certificate. A registered proprietor is the register's entry.
A planning designation is a live check against planning.data.gov.uk, and a failed check is reported as unknown rather than absent. Where a field is filled by a model instead (a floor area from geometry, a value, an index), its source says so.
Where it is weak
- Absence is a real answer. A property with no certificate, no listing or no sale is common and each tool says so in its own words rather than returning an error.
Tools: council_tax_band, epc_certificates, planning_designations, ownership_by_title, transactions_by_uprn
Questions about the methodology
Is the MarketCode valuation a RICS Red Book valuation?
+
No. It is an automated estimate with a range and a trust score, built without inspecting the property. It is designed to screen, to challenge a figure and to size a portfolio; a regulated purpose still needs a surveyor, and the tools say so.
How accurate is the valuation?
+
Against the properties' own registered sales, adjusted to today by the index, as a median absolute percentage error across the stock with a floor area. The figure moves with each model release and is published in the changelog, not fixed here.
Why does a thin area still get an index?
+
Because a gap gets filled silently by whoever reads the page. Instead the national index is published against the area, the method field says so, and the page says the turns are national.
Why are there five indices and not one?
+
They measure different things from different sources: sold prices from the registry, asking prices and rents from listings, commercial capital from commercial repeat sales, commercial rent from the rating lists. Each has its own base period and cadence, and the tools name which you are reading.
Where do the counts on the home page come from?
+
From the warehouse as of September 2026, rounded. The gateway does not publish row counts, so they are updated by hand when the pipelines move and dated here.
The sources themselves, with licences and refresh cadence, are on the sources page; the terms in this page are defined in the glossary.
Challenge the number on a property you know.
Bring an address you have valued. We will run it live and show the comparables, the index and the trust score behind ours.