Matrix gradients, without the pain

A simple trick for deriving gradients without drowning in subscripts

Contents

In this post I will talk about a simple trick to make computing gradients and backward passes mechanical, without drowning in indices and ruining your eyesight.

To start with, let’s remember how we all first learned to compute gradients. Suppose we have a linear layer $Y = XW$, where $X$ is an $N\times D$ batch of inputs and $W$ is a $D\times F$ weight matrix, followed by some scalar loss $L(Y) \in \reals$. Suppose you already know the gradient of the loss with respect to the outputs, $\grad L(Y)$, an $N\times F$ matrix.

Write $\Lcal(W) = L(XW)$ for the loss as a function of the weights, with $X$ held fixed. Throughout the post, $L$ takes a layer’s output, and $\Lcal$ is the loss as a function of whatever we’re differentiating.

What’s the gradient of the loss with respect to $W$, $\grad \Lcal(W)$?

The rote way is to compute the answer one entry at a time. Pick an entry $W_{ab}$ and apply the chain rule over every entry of $Y$, where $\partial L/\partial Y_{ij}$ is entry $ij$ of $\grad L(Y)$:

$$\frac{\partial \Lcal}{\partial W_{ab}} = \sum_{i,j} \frac{\partial L}{\partial Y_{ij}} \frac{\partial Y_{ij}}{\partial W_{ab}}.$$

Next, expand $Y_{ij} = \sum_k X_{ik}W_{kj}$ to find that $\partial Y_{ij}/\partial W_{ab}$ is $X_{ia}$ when $j = b$ and zero otherwise. That collapses the sum to

$$\frac{\partial \Lcal}{\partial W_{ab}} = \sum_i X_{ia} \frac{\partial L}{\partial Y_{ib}}.$$

Now you have to recognize this sum as a matrix product. The index $i$ is summed over, but it sits first on both $X$ and $\grad L(Y)$, so one of them must be transposed. After staring at it, you land on $\grad \Lcal(W) = X\transpose \grad L(Y)$. That’s a lot of work for one of the simplest layers there is. Stack a few layers or add a matrix inverse, and the bookkeeping grows until you’re tempted to give up and shuffle transposes around until the shapes line up.

Here’s how I would compute the gradient (every symbol here gets defined below):

$$ \begin{align} \dd Y &= X\, \dd W \\ \inner{\grad \Lcal(W)}{\dd W} &= \inner{\grad L(Y)}{\dd Y} \\ &= \inner{\grad L(Y)}{X \dd W} \\ &= \inner{X\transpose \grad L(Y)}{\dd W} \end{align} $$

So $\grad \Lcal(W) = X\transpose \grad L(Y)$. There are no indices, and the transpose comes from moving $X$ across the inner product. Magical.

By the end of the post, you will understand the machinery behind this trick. The trick rests on keeping two things apart, the derivative and the gradient:

  1. The derivative at a point is a linear function.
  2. The gradient is the vector that, when paired with an inner product, represents the derivative.

With these two ideas, every gradient comes from the same short recipe, given at the end. You only need to know matrix multiplication and transposes to follow along.

The derivative at a point is a linear function#

Let $f: \reals^D \to \reals^{D'}$ be a function with vector input and vector output. The derivative of $f$ at some point $x\in \reals^D$ is a linear function written $\dd f_x: \reals^D \to \reals^{D'}$. It is the best linear approximation of how $f$ changes near $x$.

What does it mean to be the best linear approximation? Suppose you know $f(x)$ and you want to estimate $f(x + \dd x)$ for some small change $\dd x \in \reals^D$. Here $\dd x$ is a regular ol’ vector like any other. We could have written $\Delta x$, but the “differential” notation will pay off later. Given the small change $\dd x$, $\dd f_x$ outputs the first-order change in $f$:

$$f(x + \dd x) - f(x) = \dd f_x(\dd x) + r(\dd x),$$

where the leftover error $r(\dd x)$ shrinks faster than $\dd x$ does, meaning $\norm{r(\dd x)}/\norm{\dd x} \to 0$ as $\dd x \to 0$. For small changes, the change in $f$ is the linear function $\dd f_x$ applied to $\dd x$, and everything else is negligible.

Aside: a Jacobian is a derivative written as a matrix

Recall that every linear function between $\reals^D$ and $\reals^{D'}$ can be written as a matrix. For the derivative at a point $x$, $\dd f_x$, its matrix representation is the Jacobian $J_f(x)$:

$$\dd f_x(\dd x) = J_f(x) \dd x.$$

Its entries are $(J_f(x))_{ij} = (\dd f_x(e_j))_i$, where $e_j$ is the $j$th standard basis vector. Sometimes this is written as $J_f(x)_{ij} = \partial_j f_i(x)$. In automatic differentiation, multiplying the Jacobian by a vector like this is called a Jacobian-vector product, or forward-mode differentiation.

Derivatives follow the chain rule#

Suppose $f = g\circ h$, where $g$ and $h$ are vector-valued functions with vector-valued inputs and outputs. Define $y = h(x)$. Then the chain rule says:

$$\dd f_x(\dd x) = \dd g_y(\dd h_x(\dd x)).$$

In words: push the small change $\dd x$ through $h$’s derivative to get the change in $h$, then push the change in $h$ through $g$’s derivative to get the overall change in $f$.

To talk about gradients, we first need inner products.

Inner products turn vectors into numbers#

The inner product of two vectors $a, b \in \reals^D$ is $\inner{a}{b} = a\transpose b = \sum_i a_i b_i$. For two matrices of the same shape, the standard inner product does the same thing, multiplying matching entries and summing them:

$$\inner{A}{B} = \trace{A\transpose B} = \sum_{ij} A_{ij}B_{ij}.$$
  • A matrix can hop to the other side of an inner product by transposing: $\inner{a}{Mb} = \inner{M\transpose a}{b}$ for vectors, and $\inner{A}{MB} = \inner{M\transpose A}{B}$ for matrices.
  • A scalar equals its own transpose, and its own trace: any scalar $c$ can be rewritten as $\trace{c}$.
  • The trace is unchanged when you cycle a product: $\trace{ABC} = \trace{CAB} = \trace{BCA}$. Only cycling works; in general $\trace{ABC} \neq \trace{ACB}$.
  • Inner products are linear in each argument, so constants and sums pull out.

These properties will help us “move terms around” when computing gradients.

The gradient represents the derivative#

The gradient is only defined for scalar-valued functions, such as the loss function $L$ we saw at the beginning. So let’s consider a function $f: \reals^D\to \reals$. The derivative of $f$ at $x$, $\dd f_x$, is a linear function from $\reals^D$ to $\reals$.

Definition: the gradient

The gradient of $f$ at $x$, written $\grad f(x)$, is the unique vector in $\reals^D$ such that for every small change $\dd x\in \reals^D$:

$$\dd f_x(\dd x) = \inner{\grad f(x)}{\dd x}.$$

In other words, the inner product of $\dd x$ with the gradient gives the change in $f$. That’s the sense in which the gradient represents the derivative.

Why does such a vector exist, and why is it unique? Plug in the basis vector $\dd x = e_i$. The $i$th entry of $\grad f(x)$ must equal $\dd f_x(e_i)$, which is the partial derivative $\partial f/\partial x_i$. So the gradient is the vector of partial derivatives you already know, and the definition just describes it without indices. For a matrix input $W$, the same argument says entry $ab$ of $\grad f(W)$ is $\partial f/\partial W_{ab}$, so a gradient always has the same shape as the variable it’s taken with respect to.

As $x$ varies, $\grad f$ is a function from $\reals^D$ to $\reals^D$. For a fixed $x$, $\dd f_x$ takes a small change and returns a number, so it acts like the row vector $\grad f(x)\transpose$, while $\grad f(x)$ itself is a column vector.

How to compute gradients without indices#

Characterizing the gradient as a representation of the derivative gives us a simple way to compute gradients by first computing derivatives (which lets us use the chain rule):

  1. Compute the derivative $\dd f_x(\dd x)$.1
  2. Rearrange the result into the form $\inner{\textit{something}}{\dd x}$, and read off the $\textit{something}$. Uniqueness guarantees that the $\textit{something}$ is the gradient.

Example: the gradient of a quadratic form#

Let $f(x) = x\transpose A x$, where $A$ is a $D\times D$ matrix. We want $\grad f(x)$.

Before we start, let’s clear up differential notation, which the examples use. I used to find it confusing, but now that we know what a derivative is, it’s simple.

Take $g(x) = Ax$. Its derivative applied to a small change $\dd x$ is $\dd g_x(\dd x) = A\,\dd x$, because $g$ is linear, so its best linear approximation is itself. Differential notation writes this as $\dd(Ax) = A\,\dd x$, which says “the derivative of the anonymous function $x \mapsto Ax$, applied to $\dd x$, is $A\,\dd x$.” In general, $\dd(\text{expression})$ means the derivative of $x \mapsto \text{expression}$ applied to the small change $\dd x$, so we never have to name the function. The same reading covers $\dd x$ itself: the identity function $x \mapsto x$ is linear, so $\dd(x) = \dd x$.

In the examples below, $\dd Y = X\,\dd W$ is the change in $Y = XW$ caused by a small change $\dd W$, and $\dd(W\inverse)$ is the change in $W\inverse$ caused by $\dd W$.

Using the product rule and differential notation, step 1 is one line:

$$\dd f_x(\dd x) = \dd(x\transpose A x) = (\dd x)\transpose Ax + x\transpose A\, \dd x.$$

For step 2, we rewrite the result as an inner product with $\dd x$. The first term is a scalar, so it equals its own transpose, $(\dd x)\transpose Ax = x\transpose A\transpose \dd x$. So

$$ \begin{align} \dd f_x(\dd x) &= (x\transpose A\transpose + x\transpose A)\,\dd x \\ &= \inner{(A+A\transpose)x}{\dd x}. \end{align} $$

By uniqueness of the gradient, $\grad f(x) = (A +A\transpose)x$.

Backprop the simple way using the chain rule#

I remember being confused why we learned about the chain rule for derivatives, but all of a sudden once gradients started showing up, they seemed to obey something that only resembled the chain rule. For example, in single-variable calculus the chain rule says $(g\circ h)'(x) = g'(h(x))\, h'(x)$: the outer derivative times the inner derivative. Now take $\Lcal(x) = L(Wx)$ for a scalar loss $L$ and a matrix $W$. Its gradient is

$$\grad \Lcal(x) = W\transpose \grad L(Wx).$$

The inner derivative $W$ shows up transposed, and it comes before the outer gradient instead of after. Where did the transpose come from, and why did the order flip?

Turns out, derivatives compose by the chain rule, but gradients don’t compose directly. The inner function is usually vector-valued, so it has no gradient to multiply by. What we can do is push the chain rule for derivatives through the definition of the gradient.

Suppose $f = g\circ h$ is scalar-valued, so $g$ is scalar-valued but $h$ need not be. For any $x\in \reals^D$ and any small change $\dd x\in \reals^D$, define $y=h(x)$ and $\dd y = \dd h_x(\dd x)$. The chain rule reads $\dd f_x(\dd x) = \dd g_y(\dd y)$. Writing both sides with the definition of the gradient gives:

The backprop equation

$$\inner{\grad f(x)}{\dd x} = \inner{\grad g(y)}{\dd y}$$

This one equation is all of backprop. To find $\grad f(x)$ from $\grad g(y)$, express $\dd y$ in terms of $\dd x$ on the right-hand side, then rearrange until it reads $\inner{\textit{something}}{\dd x}$. That something is $\grad f(x)$.

Doing this once in general gives the gradient’s version of the chain rule. Write the derivative of $h$ as its Jacobian, $\dd y = J_h(x)\,\dd x$ (see the Jacobian aside above), and hop it across the inner product:

$$\inner{\grad g(y)}{J_h(x)\,\dd x} = \inner{J_h(x)\transpose\grad g(y)}{\dd x},$$

so

$$\grad f(x) = J_h(x)\transpose \grad g(y).$$

This is the analog of the chain rule for gradients: the outer gradient, multiplied by the transpose of the inner derivative, in the reverse order. It’s also exactly what backprop computes. Each layer takes the gradient flowing back from the layer above, $\grad g(y)$, and multiplies it by its own transposed Jacobian to get the gradient to pass further back. In practice you never build $J_h$, which for a matrix input would be enormous. Rearranging the inner product gets $J_h(x)\transpose\grad g(y)$ directly.

Writing $\grad$ everywhere gets heavy, so from here on I’ll use bars: $\bar y := \grad L(y)$ is the gradient of the loss with respect to an intermediate value $y$, and likewise $\bar x := \grad \Lcal(x)$ and $\bar W := \grad \Lcal(W)$ for whatever we’re differentiating. This bar notation comes from reverse-mode automatic differentiation, where $\bar y$ is called the adjoint of $y$2.

Example: the gradient with respect to a vector#

Take a single input instead of a batch: the layer is $y = Wx$, where $x$ is a column vector (e.g. a data point) and $W$ is a fixed $F\times D$ weight matrix, and $L(y)$ is a scalar loss on the output. Write $\Lcal(x) = L(Wx)$ for the loss as a function of the input. We want $\bar x$, and we know $\bar y$. Since $y = Wx$, $\dd y = W\dd x$. Then

$$ \begin{align} \inner{\bar x}{\dd x} &= \inner{\bar y}{\dd y} \\ &= \bar y\transpose(W \dd x) \\ &= \inner{W\transpose\bar y}{\dd x}. \end{align} $$

By uniqueness of the gradient, the gradient of the loss with respect to the input is $\bar x = W\transpose \bar y$.

This answers the puzzle from the start of this section. The derivative does follow the chain rule in the usual order: $\dd \Lcal_x(\dd x) = \bar y\transpose W\,\dd x$, so as a row vector the derivative is $\bar y\transpose W$, outer then inner. The gradient is that row vector transposed, and transposing a product reverses its order: $(\bar y\transpose W)\transpose = W\transpose \bar y$. That’s where both the transpose and the flipped order come from.

Aside: Jacobian-vector products and vector-Jacobian products

In the example above, $\dd \Lcal_x(\dd x) = \bar y\transpose W\, \dd x$, and $W$ is the Jacobian of the inner function $x\mapsto Wx$. There are two orders in which to compute this product:

  • A Jacobian-vector product computes $W\dd x$ first, then takes the inner product with $\bar y$. This is forward-mode differentiation, and it handles one direction $\dd x$ at a time.
  • A vector-Jacobian product computes $\bar y\transpose W$ first. This result does not depend on $\dd x$ at all, and its transpose is the full gradient $W\transpose\bar y$. This is reverse-mode differentiation, also called backpropagation.

One vector-Jacobian product gives the whole gradient, while forward mode would need one Jacobian-vector product per input dimension. That is why neural networks are trained with backpropagation.

Example: the gradient with respect to a matrix#

The inner-product definition of the gradient also works for scalar-valued functions of matrices. The only change is that we use the matrix inner product from earlier, $\inner{A}{B} = \trace{A\transpose B}$.

Now we can work through the example from the start of the post properly. Recall the linear layer $Y = XW$, where $X$ is an $N\times D$ batch of inputs, $W$ is a $D\times F$ weight matrix, and $L(Y)$ is a scalar loss. Hold $X$ fixed and treat $W$ as the input, so $\Lcal(W) = L(XW)$. We want $\bar W$, and we know $\bar Y$. The layer is linear in $W$, so $\dd Y = X\,\dd W$. Then

$$ \begin{align} \inner{\bar W}{\dd W} &= \inner{\bar Y}{\dd Y} && \text{backprop equation} \\ &= \inner{\bar Y}{X\,\dd W} \\ &= \inner{X\transpose \bar Y}{\dd W} && \text{hop } X \text{ across} \end{align} $$

By uniqueness of the gradient, the gradient of the loss with respect to the weight matrix is $\bar W = X\transpose \bar Y$. Entry $ab$ of $\bar W$ is $\partial \Lcal/\partial W_{ab}$, the same quantity the rote way computed one entry at a time, and its shape, $D\times F$, matches $W$.

As an exercise, try computing $\bar X$, the gradient with respect to the activations3!

Example: the gradient through a matrix inverse#

Now swap the layer for $y = W\inverse x$, where $W$ is a square invertible matrix, and let $\Lcal(W) = L(W\inverse x)$ be the loss as a function of $W$. We want $\bar W$. Following the steps above, $\inner{\bar W}{\dd W} = \inner{\bar y}{\dd y} = \inner{\bar y}{\dd (W\inverse)\, x}$. But what is $\dd (W\inverse)$?

Notice that $WW\inverse = I$. Take the derivative of both sides. The identity matrix is constant, so its derivative is zero, and the product rule gives the first line below. Multiplying on the left by $W\inverse$ gives the second.

$$ \begin{align} (\dd W)\, W\inverse + W\, \dd(W\inverse) &= 0 \\ \dd (W\inverse) &= -W\inverse (\dd W) W\inverse \end{align} $$

Plugging this back in, and using $W\inverse x = y$:

$$ \begin{align} \inner{\bar W}{\dd W} &= \inner{\bar y}{-W\inverse (\dd W)\, y} \\ &= -\trace{\bar y\transpose W\inverse (\dd W)\, y} \\ &= -\trace{y\,\bar y\transpose W\inverse\, \dd W} && \text{cycle the trace} \\ &= \inner{-(W\inverse)\transpose\bar y\, y\transpose}{\dd W} && \text{definition of } \inner{\cdot}{\cdot} \end{align} $$

So $\bar W = -(W\inverse)\transpose\bar y\, y\transpose$. In code, you get $(W\inverse)\transpose\bar y$ by solving one linear system, $W\transpose v = \bar y$, without ever forming the inverse.

There’s no gradient of $W \mapsto W\inverse x$ to multiply by here, since it outputs a vector. That’s why we work with derivatives and convert to a gradient only at the end.

Some useful identities#

These come up often when you apply the recipe below to neural networks.

  • If $f$ is an elementwise function, like $\exp$, then $\dd (f(x)) = \diag{f'(x)} \dd x$.
  • Since $\diag{z}\onevec = z$ and $\onevec\transpose \diag{z} = z\transpose$, you can switch between an elementwise product and a diagonal matrix: $a \odot b = \diag{a}\, b$.

Summary: a recipe for deriving gradients#

Here is the whole method as a recipe. Say a scalar loss depends on some variable $\theta$ (an input, a weight matrix, anything) through a layer $y = h(\theta)$, and you already know $\bar y$, the gradient of the loss with respect to the layer’s output. Write $\Lcal(\theta) = L(h(\theta))$ for the loss as a function of $\theta$.

Recipe: finding $\bar\theta = \grad \Lcal(\theta)$

  1. Check that the function is scalar-valued. Gradients only exist for scalar functions. If yours isn’t, you want a derivative or a Jacobian, not a gradient.
  2. Write down the backprop equation. Start from $\inner{\bar\theta}{\dd \theta} = \inner{\bar y}{\dd y}$, where the right-hand side uses the gradient you already know.
  3. Take the derivative of the layer. Write $\dd y$ in terms of $\dd \theta$, holding every other input fixed. Use linearity, the chain rule, and useful identities.
  4. Move everything except $\dd\theta$ into the left slot. Transpose a matrix to move it across the inner product, rewrite a scalar as its own trace, and cycle the trace, until you have $\inner{\textit{something}}{\dd \theta}$.
  5. Read off the gradient. By uniqueness, the $\textit{something}$ is $\bar\theta$.

Takeaways#

  • The derivative $\dd f_x$ is a linear function that maps a small change $\dd x$ to the change in $f$. For a scalar function, the gradient is the unique vector with $\dd f_x(\dd x) = \inner{\grad f(x)}{\dd x}$.
  • Derivatives compose by the chain rule. Gradients don’t compose directly, because the inner function usually has no gradient. Their version of the chain rule is $\grad f(x) = J_h(x)\transpose \grad g(y)$, and computing it is backprop.
  • Backprop is the chain rule written with inner products: $\inner{\bar x}{\dd x} = \inner{\bar y}{\dd y}$. Rewrite $\dd y$ in terms of $\dd x$ and rearrange.

  1. Remember, before taking any gradients, ask yourself if the function is a scalar function. ↩︎

  2. In PyTorch, $\bar y$ is the grad_output that a custom backward function receives. ↩︎

  3. Answer: $\dd Y = (\dd X)\,W$, so $\bar X = \bar Y W\transpose$. ↩︎

Cite this post
@article{wang2026derivativesandgradients,
  author = {Alex Wang},
  title = {Matrix gradients, without the pain},
  journal = {alexwang.ai},
  year = {2026},
  month = {Oct},
  url = {https://alexwang.ai/posts/derivatives-and-gradients/}
}

More writing