Stats#

brainmaze_utils.stat.combine_gauss_distributions(mu1, std1, N1, mu2, std2, N2)#

Recalculates a normal 1-D distribution given two subsets of data.

brainmaze_utils.stat.combine_mvgauss_distributions(mu1, var1, N1, mu2, var2, N2)#

Pooled mean and covariance of two multivariate data subsets from their summaries.

Exact for population (ddof=0) statistics: if mu_i/var_i are the mean and the biased covariance (numpy.cov(..., bias=True)) of subset i, the result equals the mean and biased covariance of the concatenated data:

Sigma = (N1*Sigma1 + N2*Sigma2) / N + N1*N2 / N**2 * (mu2 - mu1)(mu2 - mu1)^T

with N = N1 + N2.

Parameters:
  • mu1 (array_like) – Means, shape (d,) or (1, d).

  • mu2 (array_like) – Means, shape (d,) or (1, d).

  • var1 (array_like) – Covariance matrices, shape (d, d).

  • var2 (array_like) – Covariance matrices, shape (d, d).

  • N1 (int or float) – Number of samples in each subset.

  • N2 (int or float) – Number of samples in each subset.

Returns:

  • mu_combined (numpy.ndarray) – Pooled mean, same shape as mu1.

  • var_combined (numpy.ndarray) – Pooled (d, d) covariance (symmetric).

Notes

Note

Changed after v2.0.0: The between-group term used the element-wise square (mu2 - mu1)**2 instead of the outer product, so off-diagonal (cross-covariance) terms were wrong - e.g. +1.5 instead of the true -1.0 for groups whose means move in opposite directions.

brainmaze_utils.stat.kl_divergence(mu1, std1, mu2, std2)#

Parametric KL-Divergence between 2 normal 1-D distributions.

Normal Distribution

brainmaze_utils.stat.kl_divergence_mv(mu1, var1, mu2, var2)#

Multidimensional parametric KL-Divergence between 2 normal distributions.

KL-Divergence

Trace

brainmaze_utils.stat.kl_divergence_nonparametric(pk, qk, eps=None)#

KL divergence D(P || Q) between discrete distributions (e.g. histograms with identical bins), in nats.

The last axis holds the bins. A 1-D input is one distribution; an N-D input is a stack of distributions (e.g. (n_features, n_bins), one histogram per row), and the result is the sum of the per-row divergences (as in v2.0.0). Each distribution is normalised to sum to 1 along the last axis first, so raw counts may be passed. Bins with p == 0 contribute 0.

Parameters:
  • pk (array_like) – Non-negative, finite weights/counts with the same shape, bins on the last axis. Every distribution (row) must have positive total mass.

  • qk (array_like) – Non-negative, finite weights/counts with the same shape, bins on the last axis. Every distribution (row) must have positive total mass.

  • eps (float, optional) – If None (default) no smoothing is applied and the result is inf whenever some bin has p > 0 and q == 0 (P is not absolutely continuous w.r.t. Q). If given (finite, > 0), eps is added to every bin of both distributions (after normalisation) and they are re-normalised, giving a finite, smoothed estimate.

Returns:

sum(p * log(p / q)) over all bins (and rows), >= 0; inf as described above.

Return type:

float

Raises:

ValueError – If the shapes differ, any value is negative or not finite, a distribution has zero total mass, or eps is not a finite positive number.

Notes

Note

Changed after v2.0.0: Bins with q == 0, p > 0 were silently dropped (returning e.g. 0.0 instead of inf). Inputs were not normalised: for already normalised histograms the result is unchanged (also for 2-D stacks of normalised rows, which is how brainmaze-eeg calls it), but raw counts now give the KL divergence of the normalised histograms instead of sum(c_p * log(c_p / c_q)), which scaled with the number of samples and was not a divergence. Invalid input (negative, NaN/inf, zero mass, different shapes) now raises ValueError.