We compared close to 8.4 million addresses from OpenStreetMap against the Algolytics reference address database, which we run in production for telecommunications, finance and logistics. We set out to test whether free access to data means it is ready for production - across four dimensions: completeness, correctness, timeliness and heterogeneity. In this article we explain why we produced the report, how we prepared it, what conclusions it leads to, and how much it really costs to bring “free” data up to production quality.
Why did we run this analysis?
Companies that need a nationwide address database naturally consider using freely available data - such as OpenStreetMap or public registers. At first glance it seems that, since address data is publicly available, building your own reference base on top of it is a simple way to reduce costs.
But does free access to data mean it is ready for production use? OpenStreetMap works very well for mapping, analytics and location use cases. The requirements change, however, once the data has to serve as a nationwide, production-grade address database in automated processes - in sales forms, CRM systems, logistics, scoring and anti-fraud workflows. That is when what counts is complete fields, alignment with the Polish address model, stable identifiers, up-to-date records and predictable quality across the whole country.
That is why, instead of relying on intuition, we compared OSM address data with the Algolytics reference address database - a base we have maintained in production for over 10 years. The aim was not to prove that OSM has no value; on the contrary, it is a very useful resource. We wanted to check whether it can serve the same function as a production database: a stable, complete and regularly updated reference base for business processes.
How did we run the analysis?
Our starting point was every OSM object - point and polygon - with at least one “addr” address tag filled in: more than 8.6 million records in total (as of June 2026). Before we could assess their quality, we had to bring them into a single, consistent form aligned with the Polish addressing model: locality name, street, building number and postal code.
The biggest challenge was the ambiguous addr:place tag which - contrary to the OSM community’s recommendations - sometimes denotes a street name, sometimes a locality, and sometimes part of one. It affected more than 3 million records. To bring them into order we used standardisation via the Data Quality service (Algolytics). After removing records with no building number, merging the point and polygon representations and deduplicating (more than 211,000 duplicates), 8,392,914 records remained in the set.
Our reference point was the Algolytics reference address database (as of 15 April 2026), containing 8,803,386 currently active addresses. We have been building it for more than 10 years on the basis of the EMUiA and TERYT registers, the PNA postal-code directory and other public sources, with our own procedures for quality control, filling gaps and handling exceptions, refreshing it at fixed quarterly intervals. It is the base on which we standardise more than 600 million addresses a year at over 99% accuracy.
We compared the OSM data with the Algolytics database across four dimensions: completeness (how many addresses overlap with the reference base), correctness (whether the address fields match and whether the coordinate location is accurate), timeliness (when the data was last modified) and heterogeneity (whether quality is uniform across the country).
What are the main conclusions of the analysis?
We consider five observations to be key:
- More than 2 million records need to be completed or corrected. OSM contains more than 410,000 fewer addresses than the Algolytics database. 7% of addresses (~617,000) have no counterpart in OSM at all, and among those that do match, 17.5% (~1.4 million) differ in at least one address field.
- Correctness looks good only until you reach the identifiers and postal codes. Full agreement on the text fields is 82.5%, but once the TERYT identifiers are taken into account it drops to 36.2% - and without stable identifiers it is very hard to distinguish repeating names (54.5% of localities in Poland have a non-unique name). The weakest address field is the postal code: wrong on average in about 13% of addresses (and in nearly 32% in the Lubuskie Voivodeship).
- The data can be out of date, because volunteers update it selectively. Nearly 45% of records have not been modified since 2019. Updates depend on community activity rather than a steady process - they can cover as little as part of a single locality (as in Bolestraszyce, where after an addressing change in February 2026 some of the addresses in OSM still reflected a state that no longer applied).
- Positional accuracy is one of OSM’s strengths. Because the coordinates largely come from the EMUiA register, the median error is about 3 m, and discrepancies over 50 m affect fewer than 1% of records. This shows that the challenge is not the precision of the coordinates themselves, but the completeness, timeliness and correctness of the address fields.
- The most important conclusion: OSM quality is extremely uneven. The data is not distributed evenly across space - at the municipality (gmina) level, full and correct field coverage ranges from 0% to 100%, while at the same time no municipality in Poland fully reproduces every address in the reference database. Good quality in one region is no guarantee of it in another, which is why a national average is not enough to assess the risk.
What does a “free” address database really cost?
Free access to data does not mean there are no costs. For OSM data to work like a production database, it has to be normalised, enriched with TERYT and postal codes, corrected for errors and maintained over time. Nationwide, that comes to more than 600,000 gaps to fill, more than 1.4 million records to correct and about 60,000 that need their location fixed - roughly 2.05 million records in total.
To make it easier for companies to estimate the cost, we broke the required work - from acquiring and normalising the data, through adding TERYT identifiers and PNA postal codes, to detecting and correcting errors and ongoing maintenance - into man-days (MD), assuming a rate of PLN 800 net per MD. In total we estimated the work at about 27 man-days: the initial data preparation alone costs around PLN 22,000, and each quarterly update a further PLN 7,100 or so to maintain quality.
Added up, the total cost of ownership (TCO) over a five-year horizon comes to about PLN 156,900 - made up of the one-off data preparation (PLN 22,000) and the recurring updates (about PLN 134,900). Hence our key conclusion: the decision to use OSM is worth making not on the basis of the price of downloading the data (zero), but on the total cost of bringing it up to production quality and maintaining it over time. “Free data” does not always mean a free solution.
How can reading the report help you?
We treated the report as a decision-making tool. Reading it lets you:
- assess whether OSM data genuinely meets the requirements of your process;
- estimate your own cost (TCO) of preparing and maintaining the data by entering your own rates and scope of work into the calculation;
- understand regional risk - OSM quality differs between voivodeships and municipalities, so a national average can be misleading for a company operating locally;
- make an informed “build vs buy” decision - whether to build and maintain your own database from free sources, or use a ready-made production base.
What does the full report contain?
The report takes the reader from context, through methodology and results, all the way to costs:
- what OSM address data is and how to understand its quality - from the tag model to preparing the dataset for analysis;
- why public registers alone are not enough - how EMUiA works and where the gaps and errors come from, including in postal codes;
- how the Algolytics reference address database is built and where it runs in production;
- a full analysis of OSM quality: timeliness, completeness, correctness, positional accuracy and heterogeneity;
- detailed statistics broken down across the 16 voivodeships, along with quality maps at the municipality level;
- a full cost calculation: the list of tasks, the man-day estimate and the five-year TCO.
Who is the report for?
We address this study to the people responsible for address-data quality and process automation: data architects and engineers, analysts, logistics, e-commerce, CRM and operations teams, as well as IT decision-makers - in sectors such as finance, telecommunications, insurance, logistics, e-commerce and retail. In short: anywhere a company is considering basing its processes on free address data and wants to work out how much it will really cost.
Access to the report
The full report, “The "free" address database from OpenStreetMap - What production quality really costs?”, is available free of charge at this link. Inside you will find the complete methodology, detailed regional statistics, an analysis of timeliness and positional accuracy, and the cost calculation (TCO, 5 years). If you have any questions or would like to discuss the findings, we encourage you to get in touch.













![Report: Comparison of standardization & geocoding services [ranking]](https://algolytics.com/wp-content/uploads/2026/06/pexels-googledeepmind-17485657-4-1024x576.jpg)

