Classifiers#
- class brainmaze_eeg.classifiers.KDEBayesianCausalModel(*args, **kwargs)#
- fit(X, y)#
Fit the transform pipeline and the per-state densities.
- Parameters:
X (np.ndarray, shape (n_epochs, n_features)) – Feature matrix (one row per epoch).
y (np.ndarray of str, shape (n_epochs,)) – Labels; every state needs enough rows for a non-singular density (more rows than retained dimensions).
- class brainmaze_eeg.classifiers.KDEBayesianModel(fbands=[[0.5, 5], [4, 9], [8, 14], [11, 16], [14, 20], [20, 30]], segm_size=30, fs=200, bands_to_erase=[], filter_bands=True, nfft=12000, window_smooth_n=3, window_std=1, cat_bias={'AWAKE': 1, 'N2': 1, 'N3': 1, 'REM': 1}, Selector2=True, n_jobs=10, standardize=True, log_lik_floor='auto', log_lik_quantile=0.001, log_lik_margin=30.0)#
Single-channel sleep classifier: spectral features -> standardisation -> feature selection -> PCA -> z-score -> per-state density model -> equal-prior posterior, smoothed over epochs, with out-of-distribution epochs flagged
'UNKNOWN'.Pipeline fitted by
fit(X, y)(X: raw feature matrix, e.g. fromextract_features_bulk):classes are balanced by duplicating minority-class rows (
balance_classes); the balanced copy is used to fit every transform below, the unbalanced data to fit the densities;if
standardize(default):StandardScaler(mean/std of the balanced training data, frozen). The raw features mix units (Hz, [0, 1] ratios, log10 ratios) and steps 3-4 are not scale-invariant: without this step a feature’s influence depends on its unit (PR #69 review R4: discriminative features x 0.1 -> chance-level accuracy without any error). With it the whole model is invariant to per-feature affine rescaling of the input;RFECVwith a linearSVRregressing the label-encoded class index (alphabetical order) selects features (step=5,>= 4features);PCAModulekeeps the components explaining>= 98 %of variance;ZScoreModule(trainable): mean/std of the balanced training data, stored and re-used unchanged at prediction time (no test-set statistics);optional
SelectFromModel(L1LinearSVC,max_features=4) ifSelector2(ValueErrorif it keeps no feature);one
gaussian_kdeper state on the transformed unbalanced training data, and the out-of-distribution floor (seelog_lik_floor).
Validation status (PR #69): every class is exercised end to end on synthetic data with known structure (feature-level Gaussian classes and synthetic sine + noise “EEG”), against a QDA / Gaussian naive Bayes reference. This checks the wiring and the numerics; it is not a validation of sleep-staging accuracy on real recordings.
- Parameters:
fbands (list of [low, high], Hz) – Bands for the spectral features.
segm_size (float, seconds) – Epoch length (also used for annotation bounds in
predict_signal).fs (float, Hz) – Sampling rate the feature extractor expects (signals are resampled to it in
extract_features_bulk).bands_to_erase (list of [low, high], Hz) – Spectral ranges ignored by the extractor (e.g. stimulation artefacts).
filter_bands – Kept for API compatibility, unused.
filter_order – Kept for API compatibility, unused.
nfft (int) – FFT length of the extractor.
window_smooth_n – Gaussian smoothing window (taps, std in taps) applied to the scores across consecutive epochs.
window_smooth_n=1disables smoothing.window_std – Gaussian smoothing window (taps, std in taps) applied to the scores across consecutive epochs.
window_smooth_n=1disables smoothing.cat_bias (dict state -> float) – Multiplicative per-state bias applied to the smoothed scores.
Selector2 (bool) – Enable step 6.
n_jobs (int) – Parallel jobs for
RFECV(default 10, as hard-coded up to v1.0.0).standardize (bool) – Enable step 2 (default True).
Falsereproduces the v1.0.0 pipeline.log_lik_floor – Out-of-distribution floor, see
_EpochClassifierMixin(default'auto', 0.001, 30 nats).log_lik_floor=Nonedisables flagging.log_lik_quantile – Out-of-distribution floor, see
_EpochClassifierMixin(default'auto', 0.001, 30 nats).log_lik_floor=Nonedisables flagging.log_lik_margin – Out-of-distribution floor, see
_EpochClassifierMixin(default'auto', 0.001, 30 nats).log_lik_floor=Nonedisables flagging.
- extract_features(signal, return_names=False)#
- fit(X, y)#
Fit the transform pipeline and the per-state densities.
- Parameters:
X (np.ndarray, shape (n_epochs, n_features)) – Feature matrix (one row per epoch).
y (np.ndarray of str, shape (n_epochs,)) – Labels; every state needs enough rows for a non-singular density (more rows than retained dimensions).
- transform(X)#
- class brainmaze_eeg.classifiers.KDEBayesianModelNC(fbands=[[0.5, 5], [4, 9], [8, 14], [11, 16], [14, 20], [20, 30]], segm_size=30, fs=200, bands_to_erase=[], filter_bands=True, filter_order=5001, nfft=12000, window_smooth_n=3, window_std=1, cat_bias={'AWAKE': 1, 'N2': 1, 'N3': 1, 'REM': 1}, Selector2=True, n_jobs=10, standardize=True, log_lik_floor='auto', log_lik_quantile=0.001, log_lik_margin=30.0)#
Same as
KDEBayesianModelbut without the z-score step after PCA (NC = no centering/normalisation): features -> [StandardScaler] -> RFECV -> PCA -> [SelectFromModel] -> KDE. Parameters as forKDEBayesianModel; the input standardisation (standardize=True, default) is the scale fix of review R4, not the omitted z-score (standardize=Falsereproduces v1.0.0).__name__was"KDEBayesianModel"up to v1.0.0 (copy-paste); it is now"KDEBayesianModelNC".- extract_features(signal, return_names=False)#
- fit(X, y)#
Fit the transform pipeline and the per-state densities.
- Parameters:
X (np.ndarray, shape (n_epochs, n_features)) – Feature matrix (one row per epoch).
y (np.ndarray of str, shape (n_epochs,)) – Labels; every state needs enough rows for a non-singular density (more rows than retained dimensions).
- transform(X)#
- class brainmaze_eeg.classifiers.MVGaussBayesianCausalModel(*args, **kwargs)#
MVGaussBayesianModeldensities + the Markov-chain filter ofKDEBayesianCausalModel.__name__was"MVGaussBayesianModel"up to v1.0.0 (copy-paste); it is now"MVGaussBayesianCausalModel".
- class brainmaze_eeg.classifiers.MVGaussBayesianModel(fbands=[[0.5, 5], [4, 9], [8, 14], [11, 16], [14, 20], [20, 30]], segm_size=30, fs=200, bands_to_erase=[], filter_bands=True, nfft=12000, window_smooth_n=3, window_std=1, cat_bias={'AWAKE': 1, 'N2': 1, 'N3': 1, 'REM': 1}, Selector2=True, n_jobs=10, standardize=True, log_lik_floor='auto', log_lik_quantile=0.001, log_lik_margin=30.0)#
KDEBayesianModelwith one multivariate Gaussian (sample mean and covariance) per state instead of a KDE.
- class brainmaze_eeg.classifiers.Mapper#
Co-registration of a feature space onto a template distribution.
A map is a per-dimension affine transform
x -> scale(x, s) + t(scaling about the data mean, then translation;tinSHIFT_RANGE,sinSCALE_RANGE) chosen to minimise a cost against the template:fit_genetic: sum over dimensions of the KL divergence between the 200-bin marginal histograms of the template and of the mapped data;fit_genetic_likelihood:1000 - sumof a fitted model’s class likelihoods of the mapped data (model._likelihood);fit_map: random search (self.Nrandom transforms), returns all costs.
Usage:
create_template(x_ref[, y_ref])thenmap(x, name)(fits once pername, cached inself.MAPS).xarrays are (n_samples, n_dims) with the samen_dimsas the template.The KL costs come from
brainmaze_utils.stat.kl_divergence_nonparametricon the (n_dims, 200) histogram array: the sum of the per-dimension divergences. Every bin is floored at 1e-9 inget_probabilities, so theinfthat brainmaze_utils >= 3.0.0 returns forq == 0, p > 0cannot occur. (An intermediate state of brainmaze_utils PR #28 averaged over dimensions instead, which would have rescaled the storedMAPS[name]['cost']but not the selected transform; the released behaviour keeps the sum, so costs are comparable across versions.)- create_template(x, y=None, name='Template')#
- fit_genetic(x, y=None, popsize=15, seed=None, **de_kwargs)#
Differential-evolution search for the transform minimising the KL cost to the template (
_cost).- Parameters:
x (np.ndarray, shape (n_samples, n_dims))
y (ignored)
popsize (int) –
differential_evolutionpopulation multiplier.seed (int or None) –
differential_evolutionseed (None = not reproducible, as before).**de_kwargs – Passed to
scipy.optimize.differential_evolution(e.g.maxiter).
- Returns:
transform (dict
{'translate': (n_dims,), 'scale': (n_dims,)})cost (float) – KL cost of the mapped data.
Up to v1.0.0 the parameter vector was split with hard-coded
[0:3]/[3:6](assumes 3 dimensions) while the bounds had ``2 * n_dims`` entries (2-D data)
raised
IndexErrorand > 3-D data raised inscale(issue #57).
- fit_genetic_likelihood(x, y=None, popsize=15, model=None, seed=None, **de_kwargs)#
As
fit_genetic()but maximisesmodel’s summed class likelihood of the mapped data (cost1000 - sum(model._likelihood(x_mapped))).modelis a fitted KDE-family model andxmust already be in that model’s transformed feature space. Returns(transform, KL cost of the mapped data).
- fit_map(x, y=None, bias={'REM': 2})#
Random search over
self.Ntransforms (np.random.seed(0), deterministic).- Parameters:
x (np.ndarray, shape (n_samples, n_dims))
y (np.ndarray of str, shape (n_samples,), optional) – Labels (’’ = unlabelled). If given, classes are balanced and the semi-supervised cost (per-class KL to
TEMPLATE_CLASS+ global KL) is used; requirescreate_template(x, y).bias (dict class -> float) – Weight of each class’s KL term.
- Returns:
costs (np.ndarray, shape (N,))
transforms (list of dict
{'translate': (n_dims,), 'scale': (n_dims,)})Up to v1.0.0 this did
y = deepcopy(x), so the labels were replaced by thefeature matrix and the semi-supervised cost was computed on nonsense “labels”
without any error (issue #57).
- get_probabilities(x)#
- map(x, name, model=None, popsize=15, seed=None, **de_kwargs)#
Map
xonto the template, fitting the map once pername.If
nameis not inself.MAPSa transform is fitted withfit_genetic_likelihood()whenmodelis given, otherwise withfit_genetic(), and stored asself.MAPS[name] = {'transformation': tr, 'cost': cost}(this restores the behaviour of the commented-outfit_map(x, name, y, model)this method was written against). Later calls with the samenamereuse it.Up to v1.0.0 this called
fit_map(x, name=..., model=...), which accepts neither keyword (TypeError) and does not store a map (issue #57).- Returns:
The mapped copy of
x.- Return type:
np.ndarray, shape (n_samples, n_dims)
- class brainmaze_eeg.classifiers.MultiChannelMVGaussBayesClassifier(fbands=[[0.5, 5], [4, 9], [8, 14], [11, 16], [14, 20], [20, 30]], segm_size=30, fs=200, bands_to_erase=[], filter_bands=True, nfft=12000, window_smooth_n=3, window_std=1, cat_bias={'AWAKE': 1, 'N2': 1, 'N3': 1, 'REM': 1}, Selector2=True, log_lik_floor='auto', log_lik_quantile=0.001, log_lik_margin=30.0)#
Gaussian naive Bayes (
sklearn.naive_bayes.GaussianNB) on a feature matrix.“Multi-channel” refers to the intended use: concatenate the per-channel feature vectors (
extract_featuresis single-channel) into one row per epoch. There is no feature selection, PCA or normalisation;transformis the identity. Unlike the KDE family,scoresare not smoothed across epochs.Constructor parameters as for
KDEBayesianModel(window_*,cat_bias,Selector2are accepted but unused). Out-of-distribution epochs are flagged as in the KDE family (log_lik_flooretc., see_EpochClassifierMixin), using the class-conditional log-likelihood of the naive-Bayes model (joint log-likelihood minus log prior; in-sample training values).__name__was"KDEBayesianModel"up to v1.0.0 (copy-paste); it is now"MultiChannelMVGaussBayesClassifier".- extract_features(signal, return_names=False)#
- extract_features_bulk(list_of_signals, fsamp_list, return_names=False)#
Extract features from a list of epochs.
Unlike
KDEBayesianModel.extract_features_bulk()this does not resample: every entry offsamp_listmust equalself.fs(Hz), otherwiseValueError(up to v1.0.0 other rates were silently processed as if sampled atself.fs, mislabelling every frequency band).Returns
(features (n_epochs, n_features), fs), or(features, feature_names)ifreturn_names.
- fit(X, y)#
- Parameters:
X (np.ndarray, shape (n_epochs, n_features))
y (array-like, shape (n_epochs,))
- max_log_likelihood(X)#
Best-class log-likelihood
max_c log p(x | c)per row, shape (n_epochs,).
- scores(X, start_time=None)#
Class probabilities.
start_timeis accepted for API symmetry and unused (no smoothing).- Returns:
Columns are
self.classifier.classes_; rows sum to 1, except out-of-distribution rows (best class log-likelihood belowself.log_lik_floor_), which are all-NaN.self.max_log_lik_is set. Up to v1.0.0 this returned an (n_classes, n_classes) frame whose cells held whole arrays, andpredictraisedValueError(issue #57).- Return type:
pd.DataFrame, shape (n_epochs, n_classes)
- transform(X)#
Identity (this model has no feature transform); returns
np.asarray(X).
- class brainmaze_eeg.classifiers.SleepClassifierWrapper(n_jobs=10)#
One
KDEBayesianModelper stimulation frequency (0, 2, 7, 72.5 Hz) at 250 Hz.train(X, df, fs):Xis an (n_epochs, n_samples) array of 30-s epochs sampled atfsHz (resampled to 250 Hz);dfhas one row per epoch with columnsannotation(str) andfreq(stimulation frequency, Hz). Epochs with datarate <= 0.85,N1andUNKNOWNare dropped. Models: 0 Hz <- freq 0; 2 Hz <- freq 0 or 2 (bands around 2 Hz harmonics erased); 7 Hz <- freq 0 or 7; 72.5 Hz <- freq 72.5. A model with no training epochs is skipped with a warning.predict_signal(X, fs, stim_freq): anyfs; returns the annotation table of the matching model. Training and prediction resample each 30-s epoch to 250 Hz the same way (KDEBayesianModel.extract_features_bulk->unify_sampling_frequency, checked by_resample_epochs). Up to round 1 of PR #69predict_signalaccepted only 250 or 500 Hz and used a different anti-alias path (Butterworth 40 Hz +[::2]) fromtrain(review R9).Every
traincall starts from an emptyMODELdict: a model skipped for lack of epochs is absent, never a stale one from an earlier call.- Parameters:
n_jobs (int) – Passed to every model’s
RFECV(default 10, as before).
- predict_signal(X, fs, stim_freq, datarate_threshold=0.85)#
Classify a continuous signal with the model of
stim_freq.X: np.ndarray (n_samples,), NaN = missing;fs: any rate in Hz (epochs are resampled to 250 Hz exactly as intrain). Returns the model’spredict_signalannotation table ('UNKNOWN'= out-of-distribution epoch).ValueErrorif no model was trained forstim_freq.
- train(X, df, fs=250)#
- Parameters:
X (np.ndarray, shape (n_epochs, n_samples)) – Epochs (30 s each).
df (pd.DataFrame, n_epochs rows, columns
annotation,freq)fs (float, Hz) – Sampling rate of
X(default 250, the models’ rate). Up to v1.0.0 there was nofsandextract_features_bulkwas called without its requiredfsamp_list(TypeError; issue #57); its(features, fs)tuple was also passed tofitas if it were the feature matrix.
- class brainmaze_eeg.classifiers.SleepStageProbabilityMarkovChainFilter#
- correct_certainty(certainty: dict)#
- fit(scores, y)#
Adapt the transition matrix to the trained states and set per-state stability.
- Parameters:
scores (pd.DataFrame, shape (n_epochs, n_states)) – Training-set class probabilities (columns = state names).
y (array-like of str, shape (n_epochs,)) – Training labels. Must be a subset of
AWAKE, N1, N2, N3, REM; other labels raiseValueError(up to v1.0.0 they were silently ignored by the filter). Everyfitstarts again from all five states and the canonical transition matrix, then removes the states absent fromy.
Notes
Stability of a state =
1 - t*, wheret*is the ROC threshold on that state’s score maximisingsqrt((1 - FPR)^2 + TPR^2). scikit-learn >= 1.3 puts+infas the first ROC threshold; if that point won (a weak score), the stability became-infand the whole transition matrix NaN. Only finite thresholds are now considered and the stability is clipped to [0, 1].
- fit_optimize(scores, y, Niter=200, popsize=10)#
- get_changing_state_posteriors(state, x_prob)#
- get_changing_state_priors(state)#
- get_changing_state_probabilities(state, x_prob)#
- get_prob_to_change(state, x_prob)#
- get_state_change_posterior(state, x_prob)#
- get_state_change_prior(state)#
- get_state_idx(state)#
- get_state_priors(state)#
- predict(scores, state='AWAKE')#
- remove_class(class_name)#
- reset_probabilities()#
- weight_probabilities(stability: ndarray)#
- class brainmaze_eeg.classifiers.SleepStructureClassifier(states=('AWAKE', 'N1', 'N2', 'N3', 'REM'))#
Transductive classifier based on the structure of a recording: the distribution of pairwise difference vectors between epochs.
fit: for every ordered pair of distinct training epochs(i, j)the differencev = x_i - x_jis represented as[||v||, v / ||v||](shaped + 1) and onegaussian_kdeis fitted per state pair(s_i, s_j).scores(x): the same pairwise vectors are formed among the epochs of ``x`` (e.g. one test recording; no labels are used). For epochk, the score of statesissum_{i != k} sum_{s'} KDE[s'][s](x_i - x_k), i.e. how well the epoch’s position relative to every other epoch matches states(the other epoch’s state marginalised), normalised over states.Cost: O(n^2) pairs for both fitting and scoring (
n= epochs); keepnmodest (a few hundred).- Parameters:
states (list of str) – States the model may use (default
['AWAKE', 'N1', 'N2', 'N3', 'REM'], the same wake label as the other classifiers in this module; up to v1.0.0 the default was'WAKE', so passstates=[...]explicitly to keep that label). Everyfitstarts again from this list: states without training epochs are dropped for that fit with a warning (self.STATES= the states actually fitted), and training labels that are not in the list raiseValueError. Each fitted state pair needs enough pairs for a non-singular(d + 1)-dimensional KDE (seefit).
- fit(x, y)#
- Parameters:
x (np.ndarray, shape (n_epochs, n_features))
y (array-like of str, shape (n_epochs,))
Notes
Up to v1.0.0 the pairs included
i == j(zero vectors, direction 0/0 = NaN), sogaussian_kderaised on NaN input; self-pairs are now excluded.
- predict(x)#
Arg-max state per epoch, np.ndarray of str, shape (n_epochs,). An epoch whose densities all underflow (all-NaN score row) gets
UNKNOWN_LABEL.
- scores(x)#
- Parameters:
x (np.ndarray, shape (n_epochs, n_features), n_epochs >= 2)
- Returns:
Columns
self.STATES; rows sum to 1.- Return type:
pd.DataFrame, shape (n_epochs, n_states)
Notes
Up to v1.0.0 this (a) called
get_mutual_vectors(x)without labels, which raisesUnboundLocalErrorin brainmaze_utils, (b) evaluated the KDEs on raw difference vectors although they were fitted on[norm, direction](dimension mismatch), and (c) aggregated withN = x.shape[1](n_features) instead of the number of epochs, returning n_features rows. All three are fixed.
- brainmaze_eeg.classifiers.UNKNOWN_LABEL = 'UNKNOWN'#
out-of-distribution epochs (best state log-likelihood below the model’s floor, see
log_lik_floor). Same string the rest of the package uses for unscored epochs (e.g.SleepClassifierWrapper.traindrops'UNKNOWN'training labels).- Type:
Label given to epochs that cannot be classified
- class brainmaze_eeg.classifiers.multivariate_normal_(X, allow_singular=False, seed=None)#
Frozen multivariate normal fitted to data (sample mean and covariance).
- Parameters:
X (np.ndarray, shape (n_features, n_samples)) – Training data, features in rows (same convention as
scipy.stats.gaussian_kde).allow_singular (see
scipy.stats.multivariate_normal.)seed (see
scipy.stats.multivariate_normal.)(n_features (pdf / logpdf take the same)
shape (n_samples) layout and return)
(n_samples
).
- logpdf(X)#
Log of the multivariate normal probability density function.
- Parameters:
x (array_like) – Quantiles, with the last axis of x denoting the components.
- Returns:
pdf – Log of the probability density function evaluated at x
- Return type:
ndarray or scalar
Notes
See class definition for a detailed description of parameters.
- pdf(X)#
Multivariate normal probability density function.
- Parameters:
x (array_like) – Quantiles, with the last axis of x denoting the components.
- Returns:
pdf – Probability density function evaluated at x
- Return type:
ndarray or scalar
Notes
See class definition for a detailed description of parameters.