LLM & Probabilistic Approaches

Denoising Score Matching: Why a Conditional Target Learns the Marginal Score

Diffusion models need the marginal score to reverse a noising process. This note explains why a tractable conditional Gaussian target learns that score and in what sense the two losses are equivalent.

Where Denoising Score Matching Enters a Diffusion Model

A diffusion model first defines a forward process that gradually adds Gaussian noise to data. At a noise level t, a common parameterization is

xt =αtx0 +σtε, ε∼N(0,I).

For sufficiently large noise, the distribution of xt becomes close to a simple Gaussian. Generation starts from that noisy distribution and moves in the reverse direction. The reverse transition—expressed as a reverse-time SDE, an ODE, or a discrete denoising update—requires the score of the marginal distribution at the current noise level:

sθ(xt,t) ≈∇xt logqt(xt).

This vector tells the sampler how the noisy state should be corrected at that level. Without it, the forward noising process is easy to run, but its reverse dynamics are unavailable.

The difficulty is that qt(xt) is the result of averaging the forward kernel over the unknown data distribution. Its score cannot normally be evaluated for a training sample. What can be evaluated is the conditional score given the clean datum:

∇xt logq(xt|x0) =αtx0−xtσt2 =−εσt.

Training can sample x0, t, and ε, construct xt, and regress against this known target. Implementations may predict the score, the added noise ε, or a denoised quantity; under Gaussian parameterizations these are related by known rescalings. Denoising score matching is the argument that makes this trainable target legitimate: its population optimum is the marginal score required by the reverse process.

The rest of this note fixes one noise level σ to isolate that argument. A diffusion model applies the same reasoning across many levels and conditions the network on t, usually with a level-dependent loss weight.

Denoising score matching replaces an inaccessible marginal target with an accessible conditional one. The score of the full noisy-data distribution is usually unknown because that distribution is a mixture over the unknown data distribution. The score of a Gaussian corruption kernel, however, is available in closed form for every clean-noisy training pair.

The replacement looks suspicious at first. A conditional score points from one noisy sample toward the particular clean sample that generated it. The marginal score cannot depend on that hidden clean sample. It must describe the geometry of the entire noisy-data distribution. Why should regression against the first target recover the second?

The answer is an exact conditional-expectation identity. The two squared-error objectives are not numerically identical, but they differ only by a term independent of the model parameters. Their minimizers and parameter gradients are therefore the same.

Clean Samples, Noisy Samples, and Two Distributions

Let a clean sample be drawn from the data distribution,

x∼pdata(x).

We corrupt it with isotropic Gaussian noise,

x~=x+σε, ε∼N(0,I).

This defines the conditional corruption kernel

qσ(x~|x) =N(x,σ2I).

This distribution is known. We chose it. For a fixed clean point x, it says where the perturbed point x~ can land.

The marginal noisy-data distribution is different:

qσ(x~) =∫pdata(x) qσ(x~|x)dx.

It is the mixture obtained after corrupting every possible clean sample. Even though the Gaussian kernel is explicit, the mixture generally is not: the data distribution is available through samples, not as a tractable density.

This distinction is the whole problem. The conditional density is easy to evaluate, while the marginal density is the distribution whose score the model must ultimately learn.

The Direct Objective Is Inaccessible

Let sθ(x~) be a neural network that predicts a vector. The desired target is the marginal score,

sθ(x~) ≈∇x~ logqσ(x~).

The score is the gradient of log density, not the density itself. At a noisy point, it indicates the local direction in which log density increases fastest. For a smoothed data distribution, this often points toward a region containing more probability mass. Calling it a “direction back to the data manifold” is useful intuition, but it is only approximate: at an ambiguous point, the score combines several possible clean explanations.

The most direct squared-error objective would be

Jexplicit(θ) =Ex~ [‖ sθ(x~) −∇x~ logqσ(x~) ‖2].

The target in this expression requires the intractable marginal mixture. This objective states the right problem but does not yet give a usable training rule.

The Conditional Gaussian Score Is Available

Denoising score matching instead uses the score of the conditional corruption kernel:

JDSM(θ) =Ex,x~ [‖ sθ(x~) −∇x~ logqσ(x~|x) ‖2].

For Gaussian corruption, the target is explicit:

∇x~ logqσ(x~|x) =x−x~σ2.

Every clean-noisy pair therefore supplies a supervised regression target. For that pair, the vector points from the noisy sample toward its generating clean sample, scaled by inverse noise variance.

This target is noisy in a statistical sense. The same noisy location can be compatible with multiple clean samples, and those clean samples imply different regression targets. The network cannot identify which hidden clean point generated a location from the noisy location alone. Under squared loss, it learns their conditional mean.

The Identity That Connects the Two Scores

The marginal score is exactly that conditional mean:

∇x~ logqσ(x~) =Ex|x~ [∇x~ logqσ(x~|x)].

The derivation is short. Differentiate the marginal density under the integral sign:

∇x~ qσ(x~) =∫pdata(x) ∇x~ qσ(x~|x)dx.

Using ∇q=q∇logq, divide by the marginal density. The ratio inside the integral becomes the posterior density over clean samples,

pdata(x) qσ(x~|x) qσ(x~) =qσ(x|x~).

Substitution gives the conditional-expectation identity above. The usual regularity conditions are doing real work here: differentiation must be allowed to pass through the integral, the relevant expectations must exist, and the marginal density must be positive where its log score is evaluated. Gaussian smoothing makes these conditions mild in many standard settings, but the algebra should not be read as assumption-free.

Exact Equivalence of the Objectives

Define

a(x~) =∇x~logqσ(x~), b(x,x~) =∇x~logqσ(x~|x).

Then a(x~)=E[b|x~]. Conditional bias-variance decomposition gives

JDSM(θ) =Jexplicit(θ) +Ex~ [Ex|x~ [‖b−a‖2]].

The last term is the conditional variance of the vector target, more precisely its expected squared deviation from the conditional mean. It contains no θ. Hence

arg minθJDSM(θ) =arg minθJexplicit(θ), ∇θJDSM =∇θJexplicit.

This is the precise meaning of “the two losses are equivalent.” Their numerical values need not match. The denoising objective includes irreducible target variance, so it is generally larger by a constant. What matches is the optimization problem with respect to θ.

Why the Learned Vector Is a Denoising Direction

For Gaussian corruption, insert the conditional score into the expectation identity:

s*(x~) =Ex|x~ [x−x~σ2].

Rearranging yields

E[x|x~] =x~+σ2 s*(x~).

The score correction moves a noisy point toward the posterior mean of the clean sample. This is stronger and more precise than saying that every score vector points back to one original datum. When several clean explanations are plausible, the learned vector averages them according to the posterior induced by the corruption process.

That averaging is also the reason sample-wise training works. Each training pair provides a target aimed at one clean sample. Across many pairs, squared-error regression estimates the conditional mean of those targets. The conditional mean is the marginal score. Denoising supervision is therefore not a heuristic substitute for score learning; under the stated conditions, it is a tractable regression formulation of the same parameter optimization problem.

The scope of the claim should remain narrow. It does not say that a finite neural network reaches the population optimum, that finite-sample training is unbiased in every implementation, or that optimization finds the global minimizer. It says that at the population-objective level, replacing the marginal score with the conditional corruption score changes the loss only by a parameter-independent constant. That is the core equivalence.

Reference

**Stanford CME296 Diffusion & Large Vision Models Spring 2026 Lecture 2 - Score matching**