Methodology
How these prices are produced
Every figure on this site comes from a published procurement record and can be traced back to it. This page sets out exactly how a raw contract award becomes a benchmarked unit price, and — just as importantly — which comparisons the underlying data does and does not support.
1. Where the data comes from
All primary data is published by government procurement authorities in the Open Contracting Data Standard (OCDS) format. We ingest each publisher through its own adapter, so a publisher’s local quirks never leak into the shared logic, and we keep the original record verbatim alongside everything we derive from it.
Every publisher is licence-checked before onboarding. Sources whose licence conflicts with how this site operates are not ingested, even where the data would be valuable.
2. The three tiers, and why they are never mixed
A number is only as good as its provenance, so every figure carries a badge saying where it came from. The three tiers are stored separately — they are not one table with a flag — and no page merges them into a single average.
- Published/Derived
Taken directly from the structured procurement record, or divided out of it arithmetically. No estimation, no judgment. This is the default and the only tier that feeds a benchmark unaided.
- AI-Assisted — Estimated
Used only where a record publishes no usable price, or where a record was flagged as a statistical outlier and needed explaining. Every such figure quotes the source text it was read from, carries a confidence score, and is reviewed by a person before it counts. It is never generated for a record that already has a clean structured price.
- Market Reference
Current vendor pricing, gathered on request when procurement coverage for an item is too thin to benchmark. Each result cites the page it came from and is approved by a person before it is cached. These are not procurement prices and are shown separately from them.
3. How a unit price is derived
Two methods, in strict order of preference. Each price on the site shows which one produced it.
- Published. The publisher stated a unit price on the line item. We use it unchanged.
- Derived. No unit price was published, so we divide the award value by the quantity. We only do this when the award covers exactly one line item — splitting a multi-item award across its lines would mean inventing an allocation.
A published figure is always preferred, even when a derived one is also available. A derived figure carries different risk: a wrong quantity silently produces a plausible-looking wrong price, which is why the method is labelled on every row rather than hidden.
If neither method works — no price, no quantity, no currency, or no resolvable item classification — the record produces no price at all. It is kept as raw source data and may become a candidate for an AI-assisted estimate, but it never appears as a published/derived figure.
4. Which records enter a benchmark
Having a price is not enough to count toward a benchmark. A record also has to pass these filters:
- Competitively awarded. Direct awards and single-source contracts are excluded from the statistics, because they are not evidence of a market price. They are still shown in the source-record table, marked as non-competitive, so you can see them.
- Same item classification. Benchmarks are computed per classification code, never across a whole category.
- Same currency. We do not convert currencies. A classification with awards in three currencies produces three benchmarks. Converting a historical award at today’s rate produces a number that is defensible to nobody, and converting at the award-date rate requires a rate source whose own terms allow it — see section 7.
- Same tax basis. See section 6.
- Agreeing on unit. If the records under one classification disagree about the unit of measure — some priced per item, others per box — the whole group is refused rather than averaged. “4.50 per pair” and “450 per box of 100” must not land in one median.
- Actually a unit price. Some publishers record an entire purchase as a single line, with the shopping list written into the description. The amount is then the total for the lot, not the price of one of anything. Two shapes of this are excluded from the medians:
- No unit of measure and a quantity of one. Nothing in the record establishes what one unit is. One such record reads “bottled purified water (3,500 bottles)” with a quantity of 1.
- A quantity of one where the “unit” names a category rather than a measure — service, unit, project, file, building. Here the publisher did fill the field, and the field still does not say how much of anything was bought. One record published as “1 unit” for 508,403 soles turned out, in the buyer’s own purchase order, to be 193,164 kilograms of filter media across five lines priced between 1.39 and 3.73 per kilogram. Read as a unit price it is wrong by a factor of hundreds of thousands.
- Enough records to be a statistic. Below a minimum sample size (currently five records) no benchmark is computed at all. A median and an interquartile range calculated from one observation are that observation and zero — arithmetic that looks like a statistic and is not one. The item page says how many eligible records exist and why no figure is shown.
This has a consequence worth stating plainly rather than leaving you to infer it from thin pages: for publishers that record purchases as lots rather than as priced units, most records are visible on the site and absent from the statistics. That is a deliberate choice about what a benchmark is allowed to claim, and the underlying figures remain published, linked and countable.
5. The statistics
For each eligible group we publish the median, the interquartile range (the 25th to 75th percentile band), and the minimum and maximum observed.
The median rather than the mean, because procurement data is heavily skewed by a small number of very large or very small awards, and a mean would follow them. The IQR rather than a standard deviation, for the same reason — it describes the middle of the distribution without assuming the distribution is normal.
Outliers are flagged using Tukey fences: a record more than 1.5× the IQR below the first quartile or above the third is marked. Flagging is not deletion — a flagged record stays in the data and stays visible. It marks a figure as worth explaining, and explaining it may conclude the price was entirely legitimate.
Every benchmark states the number of records behind it, and no benchmark is published below the minimum sample size described in section 4 — where there are too few records we show the count and say so, rather than computing a median and labelling it lightly. Benchmarks are versioned and timestamped, so a figure cited last month remains reproducible even after new data arrives.
6. What a price includes — tax, scope, delivery
Two unit prices are only comparable if they measure the same thing. Three attributes decide that, and each is shown beside every figure.
Tax. Whether a figure is gross or net of consumption tax (VAT, IVA, ISV). OCDS has no field for this, so we determine it per publisher from that country’s procurement rules and record the citation. Public procurement regimes generally require bidders to quote tax-inclusive prices, so government publishers are treated as tax-inclusive. International organisations are not: UN bodies and similar are frequently tax-exempt, and assuming otherwise would inflate their prices by the local tax rate in a way that would look like a real price difference. Benchmarks are computed within one tax basis and never across two. Tax rates themselves are recorded separately by jurisdiction and commodity, because most regimes zero-rate or reduce-rate whole categories — but we never convert between gross and net. The figures are shown as published, labelled.
Scope. Whether the procurement was open to foreign bidders. This is not a standard OCDS field either; we read it from the publisher’s own name for the procurement method, plus a per-publisher mapping for methods that are domestic by law but do not say so. International procurements are rare — well under 1% in current sources — so they are labelled and filterable rather than split into their own benchmark, and each benchmark reports how many of its records are international.
Delivery terms. Whether the price includes carriage, insurance and duties — what Incoterms describe. OCDS does not carry this, no published extension adds it, and neither do our current publishers, so for procurement records this is almost always shown as unknown. We state that rather than assume. It matters most when comparing a procurement price against a vendor quotation, where the vendor price is typically quoted excluding tax and shipping; where the two are measured differently, the site says so instead of presenting them side by side as equivalent.
What “unknown” means here. It means we have not established the answer, not that the attribute does not apply. We would rather show you a stated gap than a confident guess — a guess is invisible once it is in a table, and you cannot correct for what you cannot see.
7. Currency, and why prices are not converted
Every figure on this site is shown in the currency the publisher used. Nothing is converted today, and where a classification has awards in three currencies it produces three benchmarks rather than one.
That is a deliberate refusal, not a missing feature. A procurement benchmark compares prices across years. Converting a past award at today’s rate restates it in today’s money and quietly turns a currency movement into what looks like a price difference. Converting at the rate on the award date is the defensible method — but it needs a published historical rate series, and using one means relying on whoever publishes it.
The machinery for doing this properly is built and is switched off. It works like this, and this is what you would be reading if a figure on this site were ever shown in dollars:
- The rate used is the one in force on the award date — specifically, the most recently published rate on or before that date. Rate series are not published daily and the gaps are irregular, so “the rate that was in force” is the only definition that gives the same answer every time it is asked.
- A record with no award date is never converted. There is no honest rate for an undated award, and picking one would hide that.
- A rate is fixed to a figure the first time that figure is converted, and is never re-resolved afterwards. If the rate publisher later corrects a historical rate, the corrected rate is stored as a new entry, the old one is kept and marked superseded, and every figure already published keeps the rate it was computed with. A number cited last month does not change because a rate table was edited this month. Restating a published figure is a deliberate act and is recorded as one.
- A converted figure is always secondary and always labelled. The publisher’s own figure, in the publisher’s own currency, stays the primary number on the page. The converted figure sits beneath it, marked “Converted — derived”, and states the rate and the effective date it used.
- Converted figures are never pooled across currencies into one benchmark. Where a converted benchmark is shown it is the same sample as the original, measured differently — and it states how many of its records could be converted, because a dollar median drawn from a third of the sample is not the same claim as the median beside it.
- Outliers are identified in the original currency only. No record becomes an outlier, or stops being one, because an exchange rate moved.
The reason nothing is converted today is the rate source itself. The series this project intended to use is published by an institution that states it is for that institution’s own internal use and that the general public should not use or quote it as a historic rate. We take that at its word. Publishing figures derived from it, and citing it while doing so, would contradict the sentence we would be citing. Until a rate source is in place whose publisher’s own terms allow this use, this site shows prices exactly as they were published.
8. Known limits
These are the things most likely to make a number misleading. They are listed here so you can check.
- Units are not normalised. We refuse to benchmark a group whose records disagree on unit, but within a single unit code we take the publisher’s word for what that unit means. A publisher that records “unit” for both a single item and a pack will produce a wide, real-looking spread.
- No currency conversion. Deliberate, but it means coverage looks thinner than it is when a classification spans several currencies. Section 7 sets out the method that would be used and why it is not switched on.
- Classification quality varies. Items are grouped by the publisher’s own classification code where present. Where it was missing, a code may have been assigned by matching against the standard codelist — those assignments are reviewed, and cite what they matched.
- Award-stage prices only. We record what was awarded. Contract amendments that change what was ultimately paid are not currently modelled.
- Coverage is partial by design. Publishers are onboarded deliberately, one at a time. Absence of a price for an item means we have no data, not that none exists.
- Item names are machine-translated. Publishers write in their own language. Where we hold an English translation we show it, produced by a language model and not reviewed by a person. The publisher’s own words are shown beneath it and are what we store, cite and match on — no translation is ever used to derive a price, to group records into a benchmark, or to assign a classification. Search matches the translation and the original, so either language finds the item. An item whose name still appears in Spanish has not been translated yet.
9. Checking our work
Every published/derived price links back to its publisher, and every link says what it will show you — because not every publisher publishes a page for every record, and a link that does not say so is worse than no link.
There are three kinds, and the difference matters:
- “Award summary (publisher’s page)” or “Signed contract (PDF)” — the publisher’s own page or document for that specific award. Open it and you can check the figure yourself. This is what every Paraguayan record and 45% of Honduran records carry. Where an award summary is shown, the OCDS record for the same process is offered underneath it, so you can read either the page or the data.
- “Publisher’s OCDS record (data)” — the machine-readable record for that process. Less pleasant to read, equally checkable.
- “Publisher’s system — no record page published” — the publisher’s system, and nothing more specific, because that publisher does not publish a page for this record. It is marked Does not show this record on the record itself. This is the remaining 55% of Honduran records: Honduras’s ONCAE publishes a signed-contract PDF for some contracts and a system homepage for the rest, and we will not dress the second up as the first.
We would rather say this plainly than imply a level of traceability we do not have. Until August 2026 this page claimed every figure linked to a checkable record; for the Honduran data that was not true, and the links themselves pointed at an address that never existed — we had inferred it rather than read it from the publisher’s data. Every source link is now taken from what the publisher itself publishes.
That correction had to be made twice. In August 2026 the Paraguayan “award summary” links were found to lead nowhere: we had been building each address out of an identifier in the data, and the publisher files that page under a different one. We removed the constructed link and showed the plain data record instead — less pleasant, but it worked. The readable page is back now, and the difference is the whole point: its address is copied from the publisher’s own record of the award, not assembled by us. Where a publisher does not publish such a page, none is shown.
Publisher documents are served over plain http:// in some cases; we link to the https:// form of the same address. If a document does not load, the same path over http:// is what the publisher published.
Each price also records which version of our derivation logic produced it, and each benchmark records the parameters it was computed under. If a number looks wrong, those two facts are usually enough to establish why.
10. What an OCID is, and why two rows can share one
An OCID — an Open Contracting ID — is the identifier a publisher assigns to one contracting process: a single tender, followed from the notice through to the award and the contract. It looks like ocds-03ad3f-482596-1. The middle part identifies the publisher; the rest is that publisher’s own reference for the process.
It is not a product code, and it does not identify an item. That distinction is the one worth holding on to: the code beside an item name on this site (UNSPSC 42132203, say) says what was bought; the OCID says which purchase it was bought in.
So the same OCID can appear on several rows, and that is correct. A tender that awards twenty line items produces twenty priced records, all carrying the one OCID, because there was one process. Rows sharing an OCID are therefore not independent evidence of a market price — they are one buyer’s decision on one day. We say so here because it changes how a sample of records should be read, and because the OCID column is the only place the page shows it.
An OCID is stable and public. Given one, you can find the process in the publisher’s own system without going through us — which is the point. Every source link on a record is the publisher’s address for that OCID, taken from what they publish, never assembled by us (section 9).
11. Sources, attribution and licences
Every figure here is built from data somebody else published and made open. These are the sources, named as they name themselves, with the licence each is published under. Where a licence requires attribution, this section is how that condition is met — not a formality, and not an afterthought.
Dirección Nacional de Contrataciones Públicas (DNCP), Paraguay
Dataset as published · Creative Commons Attribution 4.0 International (CC BY 4.0)
Attribution is a condition of this licence.
Oficina Normativa de Contratación y Adquisiciones del Estado (ONCAE), Honduras
Dataset as published · Creative Commons Attribution 4.0 International (CC BY 4.0)
Attribution is a condition of this licence.
Agencia Reguladora de Compras Estatales (ARCE), Uruguay
Dataset as published · Licencia de Datos Abiertos de Uruguay v0.1
Attribution is a condition of this licence.
The licence requires a note of origin naming the provider, the licence and the dataset, and stating that changes were made. Unit prices, benchmarks and translations on this site are derived; the publisher's own records are unmodified.
Organismo Especializado para las Contrataciones Públicas Eficientes (OECE), Perú
Dataset as published · Creative Commons Attribution 4.0 International (CC BY 4.0)
Attribution is a condition of this licence.
Dirección de Compras y Contratación Pública (ChileCompra), Chile
Dataset as published · CC0 1.0 Universal (public-domain dedication)
This licence requires no attribution — the source is named here because a reader should know where a figure came from.
Dirección General de Contrataciones Públicas (DGCP), República Dominicana
Dataset as published · Open Data Commons Open Database License (ODbL) 1.0
Attribution is a condition of this licence. It also imposes share-alike on derived databases.
ODbL is a share-alike licence: a database adapted from it, and works produced from that adapted database, must themselves be offered under the ODbL. That is why the derived dataset on this site carries the licence below.
The licence of the data on this site. The unit prices, benchmark statistics and any exported dataset produced here are an adapted database in the sense the Open Database License uses. One source above is published under that licence, and it is share-alike, so this derived dataset is offered under the same terms — the Open Data Commons Open Database License (ODbL) 1.0. That is a condition we took on knowingly when the source was accepted, not an inconvenience discovered afterwards, and it is stated here because the licence requires that anyone receiving the data be told.
Changes were made. Nothing on this site is a verbatim republication of a publisher’s record: prices are derived, grouped, benchmarked and in places translated, and every one of those steps is described on this page. The publisher’s original record is kept unmodified alongside everything derived from it, and each figure links back to it.