LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement
Training & Optimization

Data Labeling

Data labeling is the process of attaching ground-truth annotations to raw data — text, images, audio — so a supervised model can learn from it, ranging from yes/no classification to structured extraction to multi-turn preference comparisons.

Full definition

Modalities span single-label classification, bounding boxes, segmentation, named-entity tagging, span extraction, and pairwise preference comparisons (for RLHF). Quality controls include double labeling with adjudication, golden questions, inter-annotator agreement (Cohen's κ ≥ 0.7 is a common bar), and labeler training. Vendors include Scale AI, Surge, Labelbox, Snorkel, Toloka. Costs range from $0.05 per simple classification to $50+ per complex preference comparison from PhD-level labelers.

Why it matters

Labeling is the largest line item in most fine-tuning budgets. Spending $200k on premium labelers often beats spending $2M on cheap ones — model gains compound when the underlying labels are right. CFOs should challenge any AI budget that does not call out labeling cost explicitly.

Example

A bank pays $400k for 18,000 expert-labeled (regulatory-clause, classification) pairs; the resulting fine-tuned model passes audit at 96% vs 71% for a cheaper crowdsourced alternative.

Source & further reading

Primary source: Sambasivan et al. — "Everyone wants to do the model work, not the data work" (Google, CHI) (2021).

Citation policy: this entry is part of the AIDOLS AI Implementation Glossary and may be quoted for research, journalism, and education with attribution to aidolsgroup.com/nl/glossary/data-labeling/.