LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement
Infrastructure

GPU (in AI context)

A GPU (Graphics Processing Unit) is a massively parallel processor that, for AI workloads, executes the matrix multiplications at the heart of neural networks 10-100× faster than a CPU — and the dominant hardware for both training and inference of modern AI models.

Full definition

NVIDIA's H100, H200, and Blackwell B200 are the de-facto frontier training GPUs; AMD MI300X and Intel Gaudi 3 are credible alternatives. Beyond raw FLOPs, the strategic differentiators are memory bandwidth (HBM3/HBM3e), interconnect (NVLink, InfiniBand), and software stack (CUDA, cuDNN, ROCm). NVIDIA's lead is as much a software-ecosystem moat as a silicon advantage.

Why it matters

GPUs are now a strategic supply-chain item, not a commodity. Multi-quarter waitlists for frontier GPUs influence which companies can train large models and at what cost. Cloud GPU pricing, on-prem GPU clusters, and reserved-capacity contracts are now CFO-level decisions.

Example

Meta's Llama 3.1 405B was trained on more than 16,000 NVIDIA H100 GPUs over multiple months — the kind of compute concentration that, until 2022, only the largest hyperscalers commanded.

Source & further reading

Primary source: NVIDIA — "H100 Tensor Core GPU Architecture" white paper (2022).

Citation policy: this entry is part of the AIDOLS AI Implementation Glossary and may be quoted for research, journalism, and education with attribution to aidolsgroup.com/it/glossary/gpu/.