LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement
Governance & Risk

Mesa-Optimization

Mesa-optimization occurs when a learned model is itself an optimizer — a "mesa-optimizer" — pursuing its own internal "mesa-objective" that may diverge from the training (base) objective the gradient descent process was optimizing.

Full definition

Coined in Hubinger et al. "Risks from Learned Optimization in Advanced Machine Learning Systems" (2019), the concept distinguishes outer alignment (training objective matches human values) from inner alignment (the learned mesa-objective matches the training objective). Off-distribution, a misaligned mesa-objective can produce competent but misdirected behavior. Modern interpretability research (induction heads, search circuits, planning circuits in maze solvers) provides early empirical traction on what mesa-objectives look like inside real networks.

Why it matters

Mesa-optimization is the technical reason interpretability matters for enterprise risk management. You cannot audit alignment by reading outputs alone; you need to know what the network is optimizing for internally — particularly for high-autonomy agents.

Example

In a 2024 Anthropic paper, researchers found Claude 3 Sonnet contained a feature that activated on "deception" concepts; clamping the feature changed model behavior — a glimpse of mesa-level structure.

Source & further reading

Primary source: Hubinger et al. — "Risks from Learned Optimization in Advanced Machine Learning Systems" (2019).

Citation policy: this entry is part of the AIDOLS AI Implementation Glossary and may be quoted for research, journalism, and education with attribution to aidolsgroup.com/no/glossary/mesa-optimization/.