Annotations#

brainmaze_utils.annotations.create_day_indexes(df: DataFrame, hour: int | float = 12, tzinfo=None)#

Add a day column: an integer day index for each epoch, where days are cut at the given wall-clock hour (default noon, so one night falls within one day).

The index of an epoch is the number of calendar “days” (hour to hour) between the day containing the earliest epoch’s start and the day containing its own start; the earliest epoch is day 0. Day boundaries are computed on the local wall clock of the chosen timezone, so DST transitions are handled.

Parameters:
  • df (pandas.DataFrame) – Annotations with a start column (and end). start can be POSIX timestamps in seconds (int/float) or timezone-aware datetime/Timestamp values (all in the same timezone).

  • hour (int or float) – Hour of day (0 <= hour < 24, fractions allowed) at which days are split.

  • tzinfo (datetime.tzinfo, optional) –

    Timezone whose wall clock defines the day boundaries. Default: the timezone of the data if start is timezone-aware, otherwise the machine’s local timezone (dateutil.tz.tzlocal()).

    Warning

    For UTC data (e.g. the output of load_CyberPSG) the default cuts days at hour UTC: noon UTC is 6 or 7 AM in Chicago and 1 or 2 PM in Prague, not local noon. Pass the recording site’s timezone, e.g. tzinfo=dateutil.tz.gettz('America/Chicago'), to cut at local hour. (Same as v2.0.0.)

Returns:

A copy of df sorted by start (index reset) with an integer day column. start/end are returned unchanged.

Return type:

pandas.DataFrame

Notes

Note

Changed after v2.0.0: Rewritten. The previous implementation assigned via chained indexing, which is a silent no-op under pandas copy-on-write (pandas >= 3: every epoch got day 0), could miss the last day, and ignored the tzinfo argument.

brainmaze_utils.annotations.create_duration(dfHyp)#

Add a duration column (end - start in seconds).

Works for numeric timestamps and for timezone-aware datetimes.

Returns:

A copy with the duration column; the input frame is not modified.

Return type:

pandas.DataFrame

brainmaze_utils.annotations.filter_by_duration(dfAnnotations: DataFrame, duration: int | float)#

Keeps only epochs of the duration given by the input.

brainmaze_utils.annotations.filter_by_key(dfAnnotations: DataFrame, key: str, value: int | float | str)#

Removes annotations whose key column equals value, keeping all others.

Note

This drops the rows matching value (the inverse of what the name may suggest); e.g. filter_by_key(df, 'annotation', 'Arrousal') returns the frame with arousals removed.

brainmaze_utils.annotations.load_CyberPSG(path, tile=None, verbose=True, strip_suffixes=True)#

Load annotations from CyberPSG XML file(s).

Parameters:
  • path (str or list) – Path to CyberPSG XML file or list of paths for multiple files

  • tile (float, optional) – Time duration in seconds to tile annotations into fixed-length segments

  • verbose (bool, optional) – If True, display progress bar for multiple files (default: True)

  • strip_suffixes (bool, optional) – If True (default), one of the suffixes _bm (written by save_CyberPSG()), _best (written by v2.0.0 and earlier), _aisc or _PiesPro is removed from every label, so files written by any version load to the same label set (N2_bm, N2_best -> N2). Set to False to get the names exactly as stored, e.g. to tell a scorer’s N2 from a model’s N2_best in the same file (with True both load as N2).

Returns:

DataFrame with annotation columns (annotation, start, end, duration, and channel for channel annotations) or list of DataFrames if multiple paths provided. start/end are timezone-aware UTC datetimes; fractional seconds beyond microseconds (.NET writes 7 digits) are truncated. Label suffixes are handled as described for strip_suffixes.

Return type:

pandas.DataFrame or list

Notes

Note

Changed after v2.0.0: _bm and _best are stripped as well (v2.0.0 stripped only _aisc and _PiesPro, so its own files loaded as N2_best, IED_best, …).

brainmaze_utils.annotations.load_NSRR(path)#

Load sleep stage annotations from NSRR (National Sleep Research Resource) XML file.

Parameters:

path (str) – Path to NSRR XML annotation file

Returns:

DataFrame with sleep stage annotations

Return type:

pandas.DataFrame

brainmaze_utils.annotations.merge_annotations(df: DataFrame)#

Merge epochs with the same annotation that touch in time (end == start of the next one). Reverse of tile_annotations().

Epochs are grouped by annotation and by the values of every extra column (e.g. channel; missing values - None, NaN, pd.NA - are equal to each other). Within a group, epochs are sorted by start and chains of touching epochs are merged; the merged epoch keeps the group’s values. Epochs of other groups in between (e.g. interleaved channels, or an overlapping label from a second scorer) do not prevent merging. A duration column, if present, is recomputed.

Parameters:

df (pd.DataFrame) – DataFrame with numeric (timestamp) ‘start’, ‘end’ and ‘annotation’ columns, optionally more.

Returns:

merged annotations sorted by start (stable), columns start, end, annotation, then the extra columns in their original order (dtypes kept), then duration if it was present. The input frame is not modified.

Return type:

pd.DataFrame

Notes

Note

Changed after v2.0.0: v2.0.0 merged only rows that were consecutive in the input order, dropped every extra column and raised on some of them. Now the frame is grouped as described above, so unsorted input, interleaved channels and overlapping labels merge as well; for sorted, non-overlapping single-channel input the result is the same as before.

brainmaze_utils.annotations.save_CyberPSG(path, df)#

Save annotations to CyberPSG XML file format.

Parameters:
  • path (str) – Output path for the CyberPSG XML file

  • df (pandas.DataFrame) – DataFrame containing annotations with columns: annotation, start, end Optional column: channel (for channel-specific annotations)

Notes

Label names on disk: standard labels (AWAKE, N1, N2, N3, REM, UNKNOWN, Arousal, N, SLP, IED, seizure, seizure_05, seizure_08) are written with the suffix _bm ('N2' -> 'N2_bm'), with exactly the type UUIDs that v2.0.0 wrote for them (v2.0.0 wrote 'N2_best', same UUID). Input labels 'N2_best' or 'N2_bm' are the same type as 'N2'. Every other label is written as given. load_CyberPSG() strips _bm and _best, so load_CyberPSG(save_CyberPSG(df)) returns the labels of df (with any _best/_bm suffix removed), and files written by v2.0.0 load to the same labels. Times are written in UTC with microsecond precision. df is not modified.

Note

Changed after v2.0.0: The suffix written to standard labels is _bm instead of _best. Type UUIDs are unchanged. 'N2' and 'N2_best' in the same frame no longer raise (they are one type).

brainmaze_utils.annotations.tile_annotations(df: DataFrame, dur_threshold: int | float = 30)#

Optimized tile_annotations using vectorized operations. Tiles epochs to the max duration given by dur_threshold in seconds. Reverse to the ‘merge annotations’.

Parameters:
  • df (pd.DataFrame) – DataFrame with numeric (timestamp) ‘start’, ‘end’, and ‘annotation’ columns representing merged annotations. Extra columns (e.g. ‘channel’) are allowed and copied to every tile of their epoch.

  • dur_threshold (int) – The desired size of each chunk in seconds.

Returns:

DataFrame with tiled annotations (start, end, annotation, extra columns, then duration if present in the input).

Return type:

pd.DataFrame

brainmaze_utils.annotations.time_to_local(dfHyp)#

Convert start/end to timezone-aware datetimes in the machine’s local timezone (dateutil.tz.tzlocal()).

Accepted inputs per cell: POSIX timestamps in seconds (int/float) or timezone-aware datetime/Timestamp (naive datetimes raise).

Returns:

A copy; the input frame is not modified.

Return type:

pandas.DataFrame

brainmaze_utils.annotations.time_to_timestamp(dfHyp)#

Convert start/end to POSIX timestamps in seconds (float).

Accepted inputs per cell: timezone-aware datetime/Timestamp or numbers (passed through).

Returns:

A copy; the input frame is not modified.

Return type:

pandas.DataFrame

brainmaze_utils.annotations.time_to_timezone(dfHyp, tzinfo)#

Convert start/end to timezone-aware datetimes in tzinfo (a datetime.tzinfo, e.g. from dateutil.tz).

Returns:

A copy; the input frame is not modified.

Return type:

pandas.DataFrame

brainmaze_utils.annotations.time_to_utc(dfHyp)#

Convert start/end to timezone-aware UTC datetimes.

Accepted inputs per cell: POSIX timestamps in seconds (int/float) or timezone-aware datetime/Timestamp (naive datetimes raise).

Returns:

A copy; the input frame is not modified.

Return type:

pandas.DataFrame