LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement
Governance & Risk

Datasheet for Datasets

A datasheet for datasets is a structured document — proposed by Gebru et al. (2018) — describing a dataset's motivation, composition, collection process, labeling, preprocessing, recommended uses, distribution, and maintenance, so downstream model developers can make informed choices.

Full definition

Datasheets sit beneath model cards in the documentation stack: a model card describes the trained artifact, while a datasheet describes the data it learned from. Critical fields include: who collected the data, what consents were obtained, demographic distributions, label sources and quality, and known biases. The EU AI Act's Article 10 effectively requires datasheet-equivalent disclosures for high-risk systems.

Why it matters

Most AI failures trace to data, not algorithms. A datasheet exposes the exact failure modes — coverage gaps, consent ambiguity, label noise — that downstream model cards inherit. Buyers should refuse training-data summaries and demand datasheets.

Example

A facial-recognition vendor publishes a datasheet revealing 86% of training images came from one continent; a procurement committee disqualifies the product on coverage grounds before pilot.

Source & further reading

Primary source: Gebru et al. — "Datasheets for Datasets" (Communications of the ACM) (2021).

Citation policy: this entry is part of the AIDOLS AI Implementation Glossary and may be quoted for research, journalism, and education with attribution to aidolsgroup.com/de/glossary/datasheet-for-datasets/.