711 239 085 info@gilsys.com
es ca en
Abstract background

Reliable data enables automations that make reliable decisions

Data governance: data quality audit

We review the company's databases and applications, measure duplicates, invalid data, and inconsistent formats, and deliver a report with the solutions.

When do the problems start?

When systems do not control data quality

Most quality problems have the same origin: fields with too much freedom, which each person fills in with their own criteria. The same customer appears under three different names, a province is written in four ways, and statuses are entered by hand. When a list is generated, data is grouped by area, or two systems are cross-referenced, the figures do not add up and nobody knows which one is correct.

Data governance means deciding who is responsible for each set of information, how it is recorded, what is validated on entry, and how often it is reviewed. In practice, it consists of replacing that freedom with fixed criteria: closed lists, single formats, and validations.

The audit is the starting point: it measures the current state of the data and proposes the rules and tools to correct it. The rules are agreed with the company; the tools make sure they are followed.

Migrations and development projects usually require data cleansing and unification processes. The audit can also be commissioned on its own: it ends in a report, without correcting data or implementing changes.

When is it worth reviewing data quality?

Having reliable data is the foundation of any digitalization

Before a migration

If data is migrated as it is, duplicates and errors carry over to the new system. A migration requires defining formats and field mappings: it is the best time to correct them.

When integrating a BI tool

A dashboard shows the data as it is. If the same concept is recorded in several ways, the charts split the figures across categories that are actually the same, and the totals do not match those of other reports. Before implementing a business intelligence tool, it is worth reviewing the data it draws on.

When implementing an ERP, a CRM, or a PIM

The new system has its own formats and validations, and the existing data has to comply with them. It is worth knowing beforehand how many records do not comply, and why.

Before automating or starting an AI project

Automations and artificial intelligence inherit the data: if it is duplicated or incomplete, their results will be too, and they will be presented as correct.

When unifying several sources

Spreadsheets, an old database, and a new program with the same information. Before unifying it, it is necessary to establish which version of each value is the correct one.

Even if no project is planned

Having well-organized information makes any future digitalization easier: every new integration, automation, or application starts from that data.

What a data quality audit reviews

Database structure

Tables, fields, data types, and relationships. Free-text fields where there should be a closed list, and fields that should be mandatory but are not.

Stored values

Empty fields, placeholder values such as "000000000" or "no data", and the same concept written in different formats.

Duplicates

Repeated records: both identical ones and those that describe the same thing in other words, which are the hardest to detect.

Invalid data

Email addresses and phone numbers that do not exist, ID numbers with an incorrect check digit, impossible dates, and postal codes that do not match the municipality.

Software and integrations

Which applications write each piece of data and where it comes in: forms, imports, or API integrations. This locates the source of each error, not just the error.

Stored personal data

What personal data is stored, in which systems, and in which fields. It is a technical inventory: the legal assessment is the responsibility of the company and its data protection advisor.

The data audit report

The audit ends in a report. Correcting the data and implementing the changes is a later step, which the company decides on based on the report.

Figures by source and by type of problem

How many duplicates, how many incomplete or invalid records, and how many different formats there are in each database. With figures, so they can be compared later.

The source of each error

Which form, import, or integration causes it. If only the data is corrected and not the source, the problem comes back.

Data sources and their owners

What information exists, where it is stored, and where it comes from, with a proposed owner for each dataset.

Prioritized proposals

What is validated on entry, what is cleaned in the historical data, and which tools are needed, ordered by impact and by effort.

How to record countries, addresses, phone numbers, and other data without errors

Most errors are avoided by validating the data or enforcing its format at the moment it is recorded, when the person entering it still knows what is correct. Later, the case has to be tracked down, and the error is harder to diagnose.

Countries from a selector

The country is chosen from a standard list, not typed. That way, "Spain", "ES", and "España" stop being three different countries.

Provinces and municipalities from a list

They are selected according to the country and the postal code, instead of being typed freely.

Addresses with a street directory

Street, number, floor, postal code, and municipality in separate fields, and the street checked against a street database. In a free-text field, nothing can be grouped by area or validated.

Phone numbers with country code

Always in international format (+34 600 000 000) and validated when entered. It seems a minor detail until two databases have to be cross-referenced or an SMS has to be sent.

Verified email addresses

The format and the existence of the domain are checked. If it is necessary to know whether the mailbox exists, an external verification service is contracted, with the company's authorization.

ID numbers

Spanish ID numbers (DNI, NIE, and CIF) validated with their check digit, without the data leaving the server.

Dates in a single format

They are entered with a date picker and stored in ISO format (YYYY-MM-DD). They sort without adjustments, and the doubt over whether 03/04 is March or April disappears.

Statuses from a closed list

"Pending", "In progress", and "Closed", always chosen from a list and not typed by hand. If each person writes them their own way, no report by status adds up.

First name and surnames in separate fields

First name, first surname, and second surname (Spanish names have two) in separate fields, whenever possible. This makes it possible to sort, detect duplicates, and personalize communications without errors.

Duplicates that are not identical

Exact duplicates: easy to detect

Two records with the same name and the same identifier are found by a simple query. They are not the main problem.

Duplicates in other words: the hard ones

Two entries that describe exactly the same situation, but one uses a term and the other a synonym. To the system they are two different records; to whoever reads them, it is obvious they are one.

Comparison by meaning, with lightweight AI models

A lightweight model converts the texts into a representation of what they mean and measures how similar they are to each other. Those that are very close are flagged as possible duplicates. It runs on a private server, as private AI, without sending the data to external services.

Always reviewed by a person

The result is a list to review, not a deletion. A system that merges records automatically ends up losing information, and losing it is much worse than having it duplicated.

Cleaning the data once is not enough

Sometimes it is not possible to add validation mechanisms, because there is no access to the source code of the programs.

The one-off cleanup

Someone spends weeks reviewing records manually. The result is good, but it is not reviewed again. Meanwhile, the cause of the duplicates and incomplete fields is still active, and the cleanup ends up being repeated, each time with more data.

The scheduled check

A periodic process can be developed to generate a short report with possible duplicates, incomplete records, and potentially invalid values.

Product information, centralized in Pimcore

In projects whose goal is to organize product information, we use Pimcore. This prevents a company from having the record of the same product, with different data, in several systems at once.

A single source for each piece of data

Pimcore becomes the system where each product record is maintained, with its documents and images in the same repository.

The website and tools read from the PIM

The website, a high-performance website built with Astro, is synchronized with the PIM. And custom-built tools can run searches and calculations on the product data, reading it from the PIM or from a synchronized copy.

Why this is data governance

A PIM requires deciding what data each product has and how it is structured, and it records who changes it. It is data governance with a tool that enforces it.

How a data quality audit is carried out

Inventory of sources and software

Databases, applications, spreadsheets, forms, and integrations. There is usually some data source that management did not know about.

Measuring data quality

Duplicates, invalid data, placeholder values, and different formats, with figures. As a guideline, the inventory and measurement take four to six consulting sessions, although it depends on the number of databases and the volume of data.

Report and proposed solutions

What is validated on entry, what is cleaned in the historical data, and which tools are needed. The rules are agreed with the company: few of them, and ones that can be followed. The audit can end here.

Implementing the changes

If the company decides so: validations at the entry points, cleanup of the historical data, integrations, and scheduled reviews that send alerts. From then on, maintenance consists of reviewing a short report every month.

Frequently asked questions about data governance

What size of company needs data governance?

Any company that makes decisions with its data; what changes is the scope. In a large organization, with many areas and audit obligations, a data committee, a formal catalog, and owners with dedicated time make sense. In a medium-sized one, a few written agreements, one owner per area, and some automatic reviews are enough. And a startup is where it costs least: formats and validations are set before there are years of data to correct.

Does data governance require changing systems?

Not necessarily. The work starts with the current systems (the ERP, the CRM, spreadsheets, and forms), and it is often enough to change what is validated on entry and what is reviewed afterward. When the same information is in several systems and none of them is the reference, a central source is needed: for product information, a PIM such as Pimcore.

Do you need access to our databases?

Yes: to measure data quality, the data has to be queried. Before access, a data protection and confidentiality agreement is signed, usually provided by the company. If email addresses or phone numbers need to be verified with an external service, it is contracted with the company's authorization.

What does data governance have to do with the GDPR?

The audit identifies what personal data is stored, in which systems, and in which fields, and records it in the report. This information is useful for complying with the regulation, but the legal assessment (whether data can be kept, on what legal basis, and for how long) is the responsibility of the company and its data protection advisor.

How long does a data quality audit take?

It is only a guideline and depends on the number of databases to be analyzed and on the volume of data: as a reference, four to six consulting sessions. If changes are implemented afterward, the timeline depends on their scope: a project of one to three months, or longer if a thorough restructuring is needed. Validations on entry make a difference from the first day in new data; old data takes longer, because what is corrected and what is kept has to be decided case by case.

How much does a data quality audit cost?

It depends on the number of databases, their volume, and whether data has to be verified with external services. The report includes the cost of each proposal, to decide what is implemented and when. Automatic reviews, once configured, only require maintenance.
Abstract background

Do the numbers in your reports add up?

Tell us which systems you use and which data raises doubts. In a free first conversation, we assess the scope of the audit and whether it should be done before a migration or a new project.