← the late compiler
C_000194 · mathematical foundations · advanced

Information Geometry

Treating families of probability distributions as curved manifolds, with divergences playing the role of distance and Fisher information as the metric.

Step 1 of 4

In words

What it is, why it matters, and what it is like.

Why am I learning this?

This unlocks the rest of Bayesian Methods: understanding how distributions are shaped, how to measure the difference between them efficiently, and how optimisation over distributions actually moves on a curved surface. It is the foundation for variational inference (which you will meet in Bayesian Methods), for the natural gradient used in some training algorithms, and for understanding why some optimisation methods struggle with high-dimensional spaces. It also sharpens your intuition for Fisher Information, which you have already noted is the curvature of the log-likelihood — here you will see it as the metric that defines distances in distribution space. Without this, later concepts like variational inference and Gaussian processes will feel like incantations. With it, they become geometry.

The idea, in plain terms

Imagine you are a cartographer trying to map a mountain range. You have a set of possible locations (each a point on the map) and you want to know how far apart they are. But the 'distance' you care about is not the straight-line distance in flat space — it is the distance you would have to walk along the actual terrain. That terrain is curved: a mile on the flat plain feels shorter than a mile climbing a steep ridge. So the distance between two points depends on the local slope and curvature of the ground, not just their map coordinates. Information geometry is exactly this: the 'terrain' is the space of probability distributions. Each distribution (e.g., a Gaussian with a certain mean and variance) is a point on a curved surface. The distance between two distributions — how different they are — is not measured by the difference in their parameters (mean and variance), but by how much the likelihood changes as you move from one to the other. This measure is called the Fisher information, and it defines the curvature of the surface. So when you optimise (e.g., find the best distribution to explain data), you are not sliding on a flat plane but on a curved hill, and the best direction to step depends on the local geometry. This is why information geometry reframes optimisation as movement on a curved surface: a step that looks small in parameter space might be a huge leap in distribution space, and vice versa. In plain words: information geometry is the geometry of 'how different are these probability distributions?' measured in units of information, not in units of the parameters themselves.

An analogy

Think of a hiker on a mountain, but the map does not show elevation — it shows only longitude and latitude. The hiker wants to reach the summit as quickly as possible. If they take a step of fixed size (say, one meter) in the direction that seems steepest according to the contour lines of the map, they might be going uphill fast on the left side but downhill on the right, because the slope changes rapidly. On a curved mountain, the 'steepest' direction at one point might not be the best after a short step. So the hiker needs to know not just the direction of steepest ascent but also how the slope itself changes as they move. That is the curvature. Information geometry gives the hiker a rule: measure the steepness (the gradient) but also the rate at which the steepness changes (the Fisher information), and use both to decide the step. In optimisation over distributions, this is the natural gradient: instead of stepping in the direction of steepest descent in parameter space (which is flat coordinates), you step in the direction that moves the distribution most quickly in terms of information. The analogy breaks down, though, because in the mountain, 'distance' is physical — a meter is a meter regardless of where you are. In distribution space, the 'distance' is defined by the Fisher information, which depends on the distribution you are at. So the hiker's map is not fixed; it changes with where they stand. That is why information geometry is not ordinary Euclidean geometry: the metric (the rule for measuring distance) is different at every point.

Definition

Information geometry is the study of the space of probability distributions as a curved manifold (a surface that may be bent), where the distance between distributions is measured by a quantity called a divergence, and the local curvature is captured by the Fisher information matrix, which acts as the metric of the space.

Where this sits

You have already noted that Fisher Information is a measure of how sharply the likelihood identifies a parameter, equivalently the curvature of the log-likelihood at its maximum. In information geometry, this same Fisher information becomes the metric that defines distances on the manifold of distributions — so your previous notes on Fisher Information directly feed into this. You also have notes on Bayesian Inference and Decision Theory: the posterior is a distribution, and if you want to compare two posteriors (e.g., before and after observing data), you need a measure of difference — that is where KL divergence (a divergence used in information geometry) comes in, and you will meet it again in variational inference. This concept also touches on your notes on Singular Value Decomposition and Neural Tangent Kernel, because those are about geometry in high-dimensional spaces, but the connection is not central here.

Signal from the Frontier

Get the next essay on mind, machine, and meaning

Essays at the intersection of AI, philosophy, and Indian governance. No promotional content.

We'll send a one-click sign-in link to confirm. No password needed.

Views expressed are personal and do not represent the Government of India or the Government of Uttarakhand.