Data Quality & DQM: dimensions, management and the right tools
Data quality is one of the defining challenges for any organisation — for decision-making, for finances and for day-to-day performance alike. Bad data is expensive: research from MIT Sloan Management Review estimates that neglecting data quality can cost a business somewhere between 15% and 25% of its revenue.
Those losses show up as missed opportunities tied to poor decisions or a damaged reputation, but also as regulatory penalties and as the sheer time spent hunting down, cleaning and correcting flawed records. The risk is operational, financial, legal and strategic all at once.
Turn that around, and high-quality data lets organisations improve operational performance, keep customers satisfied and stay more competitive — because a business that trusts its data can reorient its strategy quickly and with confidence.
What makes data “good”? The dimensions of data quality
Here is the uncomfortable truth: a data point has no intrinsic quality. You can only judge it once you know what you intend to do with it. What is the end goal? How will it be processed? What does the information actually mean? What are the quality expectations, and why? In other words, quality is defined by the intended use of the people who rely on it — IT, business teams, management.
That, in turn, demands both a broad and a granular understanding of the business processes running across the organisation, along with the standards that govern how data is exchanged internally and with third parties.
Regulation draws firm lines, too. Frameworks such as the GDPR constrain how personal data may be processed across its entire life cycle. A record stored or used outside its legal boundaries cannot be considered high quality — however useful it might otherwise be. (Keeping data within those boundaries is exactly what data sovereignty is designed to protect.)
Once those conditions are in place, data quality can be measured against a range of data quality dimensions: accuracy, completeness, conformity, integrity, consistency, availability, timeliness, intelligibility, comparability and more. Service-level criteria matter as well — how understandable, how accessible and how fresh the data is at the point of use.
Build a trustworthy data foundation to enable AI-driven automation and reliable decision-making. Download the whitepaper to learn how to establish Data Governance as a strategic capability — and unlock scalable automation and AI.
Why and how to set up Data Quality Management (DQM)
A data quality initiative is not simply about loading correct records into your systems. It also means getting rid of erroneous, corrupted or duplicated data, and describing your data precisely — through a data dictionary, for example — so it can actually be put to work. If you want the operational view of that end-to-end effort, we walk through it in detail in running a data quality process.
First, how do you even measure the quality of your data?
Step one is to take stock of your entire data estate, then classify it and identify each dataset by how it is used. Only then can you analyse quality across the data life cycle, against the dimensions you have prioritised for your context and against your business rules. Mapping and cataloguing your data are therefore the foundation of the whole exercise — you cannot measure the quality of an estate you have never inventoried. Building that inventory is precisely the goal of the Data Catalog challenge.
Errors can be technical, but far more often they are human and organisational, and they creep in at different stages of the life cycle and at different points in the information system:
- At collection — through mistaken data entry, intentional or not.
- At sharing — when several versions of the same record start to circulate.
- At export — through rules defined poorly upstream, or a compatibility problem.
- At maintenance — through faulty encoding.
The consequences of “poor quality” are records that are inaccurate, out of date, non-compliant — or simply dormant. A record can be perfectly correct and still be low quality if nobody uses it any more and it no longer delivers value.
Quality checks across the whole data life cycle are essential. If a record is poor at the source, it will be poor at the end — that is the “garbage in, garbage out” principle. To make a Data Quality strategy stick, it pays to run workshops with the business and place quality rules directly inside the applications people work in every day.
— DRAFT quote, attribution pending sign-off (SoftProject Data Governance lead)
The Data Quality Management approach
Data Quality Management (DQM) is the discipline of delivering reliable data that meets the business and technical needs of its users. It covers every process, tool, governance method and internal policy put in place to sustain quality across the data life cycle — and, ultimately, to turn good data into useful insight.
A continuous-improvement programme (such as TDQM) can lean on the four phases of the Deming cycle — plan, do, check, act. More concretely, once the initial mapping is done, six steps recur:
- Data profiling — examining table structures, the relationships between tables, the relevance of the data and the validity of formats.
- Data cleansing — spotting non-quality data and correcting it at source (removing duplicates, filling missing values). This is iterative by nature.
- Standardisation — harmonising data into a shared form that enables interoperability and a common understanding for everyone. This is the same groundwork that makes data flow across systems in the interoperability and data flows challenge.
- Deduplication — removing duplicates within a single file and identifying records that appear across several company files, so only one version survives.
- Enrichment — improving the completeness of corrected, validated data according to how it will be used. Also a continuous process.
- Reporting and monitoring — tracking how quality evolves over time through dashboards and KPIs.
To carry a data quality programme, an upfront awareness campaign — sponsored from the top — makes a real difference. Put a proper communication plan in place, with training and clear, engaging materials, so colleagues understand the stakes and adopt good practice.
— DRAFT quote, attribution pending sign-off (SoftProject Data Governance lead)
Where metadata fits into your data quality strategy
Data quality can no longer be considered separately from the quality of the metadata that describes it. Metadata supplies the context, traceability and interpretability you need to use data in advanced scenarios — artificial intelligence, regulatory governance, cross-team collaboration. Without reliable, rich, standardised metadata, data becomes hard to reuse, easy to misread and prone to bias. With it, information stays discoverable, AI models train faster and explain themselves better, and technical and business teams speak the same language. That is exactly why metadata management sits at the heart of the SoftProject data governance layer.
In practice it is hard to leave AI out of the picture — and AI is a demanding consumer of quality data. An “AI-ready” record is a high-quality record paired with high-quality metadata, a point underlined in World Bank research on AI-ready data. Excluding metadata from a Data Quality programme is no longer a defensible option.
Data for AI → AI for Data: the virtuous loop
A modern Data Quality approach is circular. Reliable, well-structured data and metadata make it possible to build high-performing AI models (Data for AI). Once trained, those models can in turn drive continuous improvement of data and metadata quality — automatically detecting errors, flagging duplicates, enriching descriptions (AI for Data). Data feeds the AI; the AI, in return, strengthens the reliability, clarity and value of the data.
A Data Quality programme cannot happen without metadata management. Metadata is what supplies the context that explains a piece of data. Since the explosion of generative AI, we have all come to appreciate just how much context matters when you want good answers.
— DRAFT quote, attribution pending sign-off (SoftProject Innovation).
Which tools and roles improve the quality of your reference data?
The people who safeguard data quality
Several roles have emerged as organisations put more weight on data quality and on mastering their reference data. The Master Data Manager, usually tied to an MDM solution; the Data Steward, who makes data easier for the business to access; the Data Owner, who is accountable for final data quality. Leadership roles such as the CDO (Chief Data Officer) are typically the first sponsors of these transformations.
A cross-disciplinary team — data quality manager, data architect, data scientists, data steward, data protection officer — is essential to carry the work. But do not stop at people: choosing the right tools matters just as much.
Data quality is a cross-cutting subject — it cannot be the concern of the data or IT teams alone. You have to bring the business back in and support them directly inside their own applications, so they can check and validate the consistency of their data. The tooling to do that now exists.
— DRAFT quote, attribution pending sign-off (SoftProject Data Governance lead)
A toolbox spanning several functional scopes
Optimal data quality needs the right tools — but our advice is to step back and look at your needs and the functional scope of each stage first, rather than jumping straight to “tools and acronyms”. The main functional areas of Data Quality look like this:
- Data quality analysis (DQ)
- Data security
- Compliance
- Monitoring and analytics
- Data discovery
- Data catalogue (metadata)
- Data lineage
- Master data and reference-data management (MDM)
- Data transformation and movement (ETL, ELT, ESB)
- Data sharing and exposure (API management)
- Visualisation and reporting (business intelligence)
Bringing those scopes together under one roof — instead of stitching together point tools — is the thinking behind the Choosing your Data Platform challenge.
Our conviction: pair the Data and Process views to serve data quality
At SoftProject, we are convinced that data and process are inseparable. That is why business process management on the SoftProject platform — powered by X4 BPMS and orchestrated end to end with Phoenix — lets you understand your business processes and the data uses attached to them.
To manage your master data — customers, suppliers, products, financials — the platform’s Master Data Management and data governance capabilities, working alongside MyDataCatalogue and dataspot., let you supervise and automate everything around your data: collect it, move it, enrich it, deliver it. You can model your reference data, build your own indicators and guarantee the relevance, uniqueness and traceability of information across its entire life cycle — all on one platform.
Want to pressure-test your own data quality challenges with a SoftProject expert?
Edouard Cante is responsible for the strategic direction and further development of SoftProject’s product portfolio as Chief Product Officer. With a strong understanding of the market and a high level of innovative drive, he advances customer-centric solutions and ensures the company’s long-term competitiveness.