Scikit Modules#
- class brainmaze_eeg.scikit_modules.FeatureAugmentorModule#
Feature augmentation using an ‘augment_features’ function from the ‘PiesUtils’ package. See the code for additional details.
- fit(X=None, Y=None)#
Fit method (no-op for this transformer).
- fit_transform(X, Y=None)#
Fit and transform in one step (same as transform for this module).
- transform(X)#
Transform features by applying mutual and standalone operations.
- Parameters:
X (np.ndarray) – Input feature matrix of shape [n_samples, n_features].
- Returns:
Augmented feature matrix.
- Return type:
np.ndarray
- class brainmaze_eeg.scikit_modules.Log10Module#
Base-10 logarithm transformation module compatible with scikit-learn pipelines.
- fit(X, Y=None)#
Fit method (no-op for this transformer).
- fit_transform(X, Y=None)#
Fit and transform in one step.
- transform(X, Y=None)#
Apply base-10 logarithm transformation to features.
- class brainmaze_eeg.scikit_modules.LogModule#
Natural logarithm transformation module compatible with scikit-learn pipelines.
- fit(X, Y=None)#
Fit method (no-op for this transformer).
- fit_transform(X, Y=None)#
Fit and transform in one step.
- transform(X, Y=None)#
Apply natural logarithm transformation to features.
- class brainmaze_eeg.scikit_modules.PCAModule(var_threshold=0.98)#
PCA module using scikit-learn’s PCA with automatic component selection. Extends sklearn.decomposition.PCA.
- fit(X, y=None)#
Fit the model with X.
- Parameters:
X ({array-like, sparse matrix} of shape (n_samples, n_features)) – Training data, where n_samples is the number of samples and n_features is the number of features.
y (Ignored) – Ignored.
- Returns:
self – Returns the instance itself.
- Return type:
object
- fit_transform(X, y=None)#
Fit and transform data in one step.
- Parameters:
X (np.ndarray) – Input data matrix.
y (array-like, optional) – Target values (ignored, for scikit-learn compatibility).
- Returns:
Transformed data in PCA space.
- Return type:
np.ndarray
- class brainmaze_eeg.scikit_modules.PCAModuleSVD(var_threshold=0.98)#
PCA via eigendecomposition of the covariance matrix, with automatic component selection by an explained-variance threshold. Compatible with scikit-learn pipelines.
Note
Despite the name, this does not use an SVD: it eigendecomposes the sample covariance
C = X.T @ X / (n - 1). It does not centreX; it assumes the input is already zero-mean per feature (e.g. the output of a z-score step). On non-centred data the “covariance” is the second-moment matrix and the components are not principal components. UsePCAModule(scikit-learn, centres the data) if that assumption does not hold.- eigen_vals#
Set by
fit. Eigenvalues ofC(variance along each component), real and sorted in descending order. Tiny negative values from round-off are clipped to 0.- Type:
np.ndarray, shape (n_features,), float
- eigen_vecs#
Matching unit eigenvectors in columns (
eigen_vecs[:, i]<->eigen_vals[i]). Sign convention (like scikit-learn’ssvd_flip): the largest-magnitude loading of each column is positive (the first one on ties), so the projections do not depend on the LAPACK build.- Type:
np.ndarray, shape (n_features, n_features), float
- explained_variance_ratio#
eigen_vals / eigen_vals.sum().- Type:
np.ndarray, shape (n_features,)
- n#
Number of retained components: the smallest
nsuch that the cumulative explained-variance ratio of the firstncomponents is>= var_threshold(clipped ton_features).- Type:
int
- fit(X, Y=None)#
Fit the components on
X.- Parameters:
X (np.ndarray, shape (n_samples, n_features)) – Zero-mean data (see the class note;
Xis not centred here).Y (ignored)
- Return type:
self
Notes
Up to v1.0.0 this used
np.linalg.eig(general, non-symmetric solver) and took the firstneigenpairs as if they were in descending order.eigdoes not guarantee any order (e.g. it returned[.., 0.896, 0.292, 0.589, 0.449]for an 8-feature covariance) and returns complex dtype, so the selected components andncould be wrong. It now usesnp.linalg.eigh(symmetric solver, real output) and sorts explicitly in descending order (issue #39).
- fit_transform(X, Y=None)#
Fit and transform in one step.
- transform(X, Y=None)#
Transform data using fitted PCA components.
- Parameters:
X (np.ndarray) – Input data matrix.
Y (array-like, optional) – Ignored parameter for scikit-learn compatibility.
- Returns:
Transformed data in PCA space.
- Return type:
np.ndarray
- class brainmaze_eeg.scikit_modules.ZScoreModule(trainable=False, continuous_learning=False, multi_class=False)#
Z-score normalization compatible with scikit.pipeline.Pipeline Enables continuous learning - enabling continuous adaptation.
- Modes
Zscore normalization
- Zscore normalization with fixed mean and std values based on the initial training dataset
Possible category-wise normalization with mean and std values estimated from the training dataset - number of features is multiplied by number of categories
- Zscore normalization with an initial mean and std values trained on the training dataset - adaptation during inference
https://stats.stackexchange.com/questions/211837/variance-of-subsample
- continuous_learning#
If true - An instance updates mean and variance values during each prediction step. Initial outlier filtering is recommended
- Type:
bool
- trainable#
If false - An instance normalizes inference data based on their current mean value and std If true - An instance remembers mean and variance values of training data
- Type:
bool
- multi_class#
If true - An instance performs normalization for each training class separately Number of output features is multiplied by a number of training categories
- Type:
bool
- mean#
Trained mean values for each feature. In case multi_class == True -> list of numpy ndarrays for each category
- Type:
numpy ndarray / list
- std#
- Type:
numpy ndarray
- N#
- Type:
int
- fit(X=None, Y=None)#
- Parameters:
X (numpy ndarray) – shape[n_samples, n_features]
Y (list or numpy array, optional) – category reference for each sample - required only for option with multi_class normalization
- Return type:
None
- fit_transform(X=None, Y=None)#
- Parameters:
X (numpy ndarray) – shape[n_samples, n_features]
Y (list or numpy array, optional) – category reference for each sample - required only for option with multi_class normalization
- Returns:
transformed_data – shape[n_samples, n_features]
- Return type:
numpy ndarray
- transform(X=None)#
- Parameters:
X (numpy ndarray) – shape[n_samples, n_features]
- Returns:
transformed_data – shape[n_samples, n_features]
- Return type:
numpy ndarray