ETSI publishes framework to measure data quality for AI and digital systems
ETSI has published a Technical Report defining 18 standardised metrics for assessing data quality, covering reliability, completeness, bias, privacy and other characteristics relevant to data-driven and AI systems.
The European Telecommunications Standards Institute (ETSI) has published a new Technical Report establishing a common framework for measuring the quality of datasets used in digital and artificial intelligence systems.
ETSI TR 104 180, ‘Data Solutions; Development and identification of Data Quality Metrics’, defines 18 metrics that organisations can use to assess whether data is suitable for a particular purpose. The framework provides formal definitions and mathematical methods for measuring different aspects of data quality.
The metrics cover several areas, including completeness, accuracy, reliability, consistency, precision, integrity, redundancy and uniqueness. They also address practical factors such as data availability, coverage, lineage, traceability and timeliness.
The framework includes specific measures for fairness, such as label quality, measurement bias and representation bias, as well as privacy-related characteristics including anonymity and confidentiality.
ETSI tested the approach using two different types of datasets: industrial Internet of Things sensor data and demographic data. The exercise showed that different applications require different combinations and relative weights of the metrics. For sensor data, factors such as reliability, accuracy, timeliness and integrity were particularly relevant, while demographic data required greater attention to representation, anonymity, confidentiality and potential bias.
ETSI also developed an open-source Data Quality Validation System as part of the proof of concept. The system can calculate a data quality score based on the 18 metrics defined in the report.
The framework is intended to give dataset owners a consistent way to evaluate and communicate data quality. For AI systems, this can help organisations identify problems with the reliability or representativeness of training and operational data before it is used.
By establishing common measurements, the report provides a basis for more consistent and repeatable data quality assessments across different digital and AI applications.
