Stats#
- brainmaze_utils.stat.combine_gauss_distributions(mu1, std1, N1, mu2, std2, N2)#
Recalculates a normal 1-D distribution given two subsets of data.
- brainmaze_utils.stat.combine_mvgauss_distributions(mu1, var1, N1, mu2, var2, N2)#
Pooled mean and covariance of two multivariate data subsets from their summaries.
Exact for population (
ddof=0) statistics: ifmu_i/var_iare the mean and the biased covariance (numpy.cov(..., bias=True)) of subseti, the result equals the mean and biased covariance of the concatenated data:Sigma = (N1*Sigma1 + N2*Sigma2) / N + N1*N2 / N**2 * (mu2 - mu1)(mu2 - mu1)^Twith
N = N1 + N2.- Parameters:
mu1 (array_like) – Means, shape
(d,)or(1, d).mu2 (array_like) – Means, shape
(d,)or(1, d).var1 (array_like) – Covariance matrices, shape
(d, d).var2 (array_like) – Covariance matrices, shape
(d, d).N1 (int or float) – Number of samples in each subset.
N2 (int or float) – Number of samples in each subset.
- Returns:
mu_combined (numpy.ndarray) – Pooled mean, same shape as
mu1.var_combined (numpy.ndarray) – Pooled
(d, d)covariance (symmetric).
Notes
Note
Changed after v2.0.0: The between-group term used the element-wise square
(mu2 - mu1)**2instead of the outer product, so off-diagonal (cross-covariance) terms were wrong - e.g. +1.5 instead of the true -1.0 for groups whose means move in opposite directions.
- brainmaze_utils.stat.kl_divergence(mu1, std1, mu2, std2)#
Parametric KL-Divergence between 2 normal 1-D distributions.
- brainmaze_utils.stat.kl_divergence_mv(mu1, var1, mu2, var2)#
Multidimensional parametric KL-Divergence between 2 normal distributions.
- brainmaze_utils.stat.kl_divergence_nonparametric(pk, qk, eps=None)#
KL divergence
D(P || Q)between discrete distributions (e.g. histograms with identical bins), in nats.The last axis holds the bins. A 1-D input is one distribution; an N-D input is a stack of distributions (e.g.
(n_features, n_bins), one histogram per row), and the result is the sum of the per-row divergences (as in v2.0.0). Each distribution is normalised to sum to 1 along the last axis first, so raw counts may be passed. Bins withp == 0contribute 0.- Parameters:
pk (array_like) – Non-negative, finite weights/counts with the same shape, bins on the last axis. Every distribution (row) must have positive total mass.
qk (array_like) – Non-negative, finite weights/counts with the same shape, bins on the last axis. Every distribution (row) must have positive total mass.
eps (float, optional) – If
None(default) no smoothing is applied and the result isinfwhenever some bin hasp > 0andq == 0(P is not absolutely continuous w.r.t. Q). If given (finite,> 0),epsis added to every bin of both distributions (after normalisation) and they are re-normalised, giving a finite, smoothed estimate.
- Returns:
sum(p * log(p / q))over all bins (and rows),>= 0;infas described above.- Return type:
float
- Raises:
ValueError – If the shapes differ, any value is negative or not finite, a distribution has zero total mass, or
epsis not a finite positive number.
Notes
Note
Changed after v2.0.0: Bins with
q == 0, p > 0were silently dropped (returning e.g. 0.0 instead of inf). Inputs were not normalised: for already normalised histograms the result is unchanged (also for 2-D stacks of normalised rows, which is how brainmaze-eeg calls it), but raw counts now give the KL divergence of the normalised histograms instead ofsum(c_p * log(c_p / c_q)), which scaled with the number of samples and was not a divergence. Invalid input (negative, NaN/inf, zero mass, different shapes) now raisesValueError.