
Reliable data enables automations that make reliable decisions
Data governance: data quality audit
We review the company's databases and applications, measure duplicates, invalid data, and inconsistent formats, and deliver a report with the solutions.
When do the problems start?
When systems do not control data quality
Most quality problems have the same origin: fields with too much freedom, which each person fills in with their own criteria. The same customer appears under three different names, a province is written in four ways, and statuses are entered by hand. When a list is generated, data is grouped by area, or two systems are cross-referenced, the figures do not add up and nobody knows which one is correct.
Data governance means deciding who is responsible for each set of information, how it is recorded, what is validated on entry, and how often it is reviewed. In practice, it consists of replacing that freedom with fixed criteria: closed lists, single formats, and validations.
The audit is the starting point: it measures the current state of the data and proposes the rules and tools to correct it. The rules are agreed with the company; the tools make sure they are followed.
Migrations and development projects usually require data cleansing and unification processes. The audit can also be commissioned on its own: it ends in a report, without correcting data or implementing changes.
When is it worth reviewing data quality?
Having reliable data is the foundation of any digitalization
Before a migration
If data is migrated as it is, duplicates and errors carry over to the new system. A migration requires defining formats and field mappings: it is the best time to correct them.
When integrating a BI tool
A dashboard shows the data as it is. If the same concept is recorded in several ways, the charts split the figures across categories that are actually the same, and the totals do not match those of other reports. Before implementing a business intelligence tool, it is worth reviewing the data it draws on.
When implementing an ERP, a CRM, or a PIM
The new system has its own formats and validations, and the existing data has to comply with them. It is worth knowing beforehand how many records do not comply, and why.
Before automating or starting an AI project
Automations and artificial intelligence inherit the data: if it is duplicated or incomplete, their results will be too, and they will be presented as correct.
When unifying several sources
Spreadsheets, an old database, and a new program with the same information. Before unifying it, it is necessary to establish which version of each value is the correct one.
Even if no project is planned
Having well-organized information makes any future digitalization easier: every new integration, automation, or application starts from that data.
What a data quality audit reviews
Database structure
Tables, fields, data types, and relationships. Free-text fields where there should be a closed list, and fields that should be mandatory but are not.
Stored values
Empty fields, placeholder values such as "000000000" or "no data", and the same concept written in different formats.
Duplicates
Repeated records: both identical ones and those that describe the same thing in other words, which are the hardest to detect.
Invalid data
Email addresses and phone numbers that do not exist, ID numbers with an incorrect check digit, impossible dates, and postal codes that do not match the municipality.
Software and integrations
Which applications write each piece of data and where it comes in: forms, imports, or API integrations. This locates the source of each error, not just the error.
Stored personal data
What personal data is stored, in which systems, and in which fields. It is a technical inventory: the legal assessment is the responsibility of the company and its data protection advisor.
The data audit report
The audit ends in a report. Correcting the data and implementing the changes is a later step, which the company decides on based on the report.
Figures by source and by type of problem
How many duplicates, how many incomplete or invalid records, and how many different formats there are in each database. With figures, so they can be compared later.
The source of each error
Which form, import, or integration causes it. If only the data is corrected and not the source, the problem comes back.
Data sources and their owners
What information exists, where it is stored, and where it comes from, with a proposed owner for each dataset.
Prioritized proposals
What is validated on entry, what is cleaned in the historical data, and which tools are needed, ordered by impact and by effort.
How to record countries, addresses, phone numbers, and other data without errors
Most errors are avoided by validating the data or enforcing its format at the moment it is recorded, when the person entering it still knows what is correct. Later, the case has to be tracked down, and the error is harder to diagnose.
Countries from a selector
The country is chosen from a standard list, not typed. That way, "Spain", "ES", and "España" stop being three different countries.
Provinces and municipalities from a list
They are selected according to the country and the postal code, instead of being typed freely.
Addresses with a street directory
Street, number, floor, postal code, and municipality in separate fields, and the street checked against a street database. In a free-text field, nothing can be grouped by area or validated.
Phone numbers with country code
Always in international format (+34 600 000 000) and validated when entered. It seems a minor detail until two databases have to be cross-referenced or an SMS has to be sent.
Verified email addresses
The format and the existence of the domain are checked. If it is necessary to know whether the mailbox exists, an external verification service is contracted, with the company's authorization.
ID numbers
Spanish ID numbers (DNI, NIE, and CIF) validated with their check digit, without the data leaving the server.
Dates in a single format
They are entered with a date picker and stored in ISO format (YYYY-MM-DD). They sort without adjustments, and the doubt over whether 03/04 is March or April disappears.
Statuses from a closed list
"Pending", "In progress", and "Closed", always chosen from a list and not typed by hand. If each person writes them their own way, no report by status adds up.
First name and surnames in separate fields
First name, first surname, and second surname (Spanish names have two) in separate fields, whenever possible. This makes it possible to sort, detect duplicates, and personalize communications without errors.
Duplicates that are not identical
Exact duplicates: easy to detect
Two records with the same name and the same identifier are found by a simple query. They are not the main problem.
Duplicates in other words: the hard ones
Two entries that describe exactly the same situation, but one uses a term and the other a synonym. To the system they are two different records; to whoever reads them, it is obvious they are one.
Comparison by meaning, with lightweight AI models
A lightweight model converts the texts into a representation of what they mean and measures how similar they are to each other. Those that are very close are flagged as possible duplicates. It runs on a private server, as private AI, without sending the data to external services.
Always reviewed by a person
The result is a list to review, not a deletion. A system that merges records automatically ends up losing information, and losing it is much worse than having it duplicated.
Cleaning the data once is not enough
Sometimes it is not possible to add validation mechanisms, because there is no access to the source code of the programs.
The one-off cleanup
Someone spends weeks reviewing records manually. The result is good, but it is not reviewed again. Meanwhile, the cause of the duplicates and incomplete fields is still active, and the cleanup ends up being repeated, each time with more data.
The scheduled check
A periodic process can be developed to generate a short report with possible duplicates, incomplete records, and potentially invalid values.
Product information, centralized in Pimcore
In projects whose goal is to organize product information, we use Pimcore. This prevents a company from having the record of the same product, with different data, in several systems at once.
A single source for each piece of data
Pimcore becomes the system where each product record is maintained, with its documents and images in the same repository.
The website and tools read from the PIM
The website, a high-performance website built with Astro, is synchronized with the PIM. And custom-built tools can run searches and calculations on the product data, reading it from the PIM or from a synchronized copy.
Why this is data governance
A PIM requires deciding what data each product has and how it is structured, and it records who changes it. It is data governance with a tool that enforces it.
How a data quality audit is carried out
Inventory of sources and software
Databases, applications, spreadsheets, forms, and integrations. There is usually some data source that management did not know about.
Measuring data quality
Duplicates, invalid data, placeholder values, and different formats, with figures. As a guideline, the inventory and measurement take four to six consulting sessions, although it depends on the number of databases and the volume of data.
Report and proposed solutions
What is validated on entry, what is cleaned in the historical data, and which tools are needed. The rules are agreed with the company: few of them, and ones that can be followed. The audit can end here.
Implementing the changes
If the company decides so: validations at the entry points, cleanup of the historical data, integrations, and scheduled reviews that send alerts. From then on, maintenance consists of reviewing a short report every month.
Frequently asked questions about data governance
What size of company needs data governance?
Does data governance require changing systems?
Do you need access to our databases?
What does data governance have to do with the GDPR?
How long does a data quality audit take?
How much does a data quality audit cost?

Do the numbers in your reports add up?
Tell us which systems you use and which data raises doubts. In a free first conversation, we assess the scope of the audit and whether it should be done before a migration or a new project.




