← the late compiler
C_000161 · mathematical foundations · advanced

Fisher Information

A measure of how sharply the likelihood identifies a parameter, equivalently the curvature of the log-likelihood at its maximum.

Step 1 of 4

In words

What it is, why it matters, and what it is like.

Why am I learning this?

Fisher information is the key to understanding how confident a model is in its own answers, why some parameters are easy to learn and others are not, and how modern AI systems like ChatGPT decide what to say next. It underpins the mathematics of large language models, the design of retrieval systems that know which documents are relevant, and the optimisation algorithms that train every neural network. Learning this concept will prepare you for Bayesian Methods, Information Geometry, and Natural Gradient Descent — all of which are fundamental to the advanced machinery of AI.

The idea, in plain terms

Imagine you are trying to find the exact height of a hill by taking measurements from a distance. If the hill is a sharp, narrow peak, a single measurement tells you a lot about where the summit is — the information is 'sharp'. If the hill is a gentle, wide mound, a measurement tells you almost nothing — the information is 'flat'. Fisher information is a measure of exactly this sharpness, but instead of a hill, it is about a parameter you are trying to estimate from data. When you collect data and try to guess a parameter (like the average height of a population), the likelihood function tells you how well each possible parameter value fits the data. If the likelihood is sharply peaked at one value, that value is very well supported by the data — the Fisher information is high. If the likelihood is flat and spread out, many values explain the data almost equally well — the Fisher information is low. In short, Fisher information quantifies how much information your data carries about the exact value of a parameter, which is directly related to how precisely you can estimate it.

An analogy

Think of a dartboard with a bullseye. You are trying to find the centre of the bullseye, but you can only throw darts and see where they land. Your throws are your data. If the dartboard is a standard one, your darts cluster tightly around the bullseye — each dart gives you a lot of information about where the centre is. Fisher information is like the tightness of that cluster: the tighter the cluster, the more information each dart (data point) gives you. Now imagine the bullseye is on a huge, flat wall with no markings. Your darts land all over the place, and it is impossible to guess the centre. That is low Fisher information: your data are uninformative. The analogy has a limit: Fisher information is not about the spread of your data around a target, but about how the likelihood function (how well a parameter explains your data) changes as you move the parameter away from its best value. A sharp likelihood means even a small change in the parameter makes the data much less likely, so you can be sure that parameter is the one. A flat likelihood means small changes barely matter — many parameters are equally plausible.

Definition

Fisher information is a measure of how sharply the likelihood function identifies a parameter — equivalently, how much the log-likelihood curves at its maximum.

Where this sits

You have not yet met the mathematics behind this, but you are about to. Fisher information is a concept in probability and statistics, and it rests on calculus (specifically derivatives, which measure how fast something changes). Once you understand it, you can see why Bayesians care: in Bayesian inference (which treats unknown parameters as random variables), the posterior distribution’s sharpness is determined by the likelihood, and Fisher information sets a limit on how small your uncertainty can be. It also connects to Information Geometry, where Fisher information is the metric on a curved space of probability distributions.

Signal from the Frontier

Get the next essay on mind, machine, and meaning

Essays at the intersection of AI, philosophy, and Indian governance. No promotional content.

We'll send a one-click sign-in link to confirm. No password needed.

Views expressed are personal and do not represent the Government of India or the Government of Uttarakhand.