utils

nomad.stop_detection.utils.applyParallel(groups, algorithm, algorithm_kwargs, reset_index=False, n_jobs=1, print_progress=False)[source]

Apply a callable over grouped data, optionally in parallel.

Parameters:
  • groups (DataFrameGroupBy) – Grouped dataframe iterator (e.g. data.groupby(user_id)).

  • algorithm (callable) – Function applied to each group DataFrame.

  • algorithm_kwargs (dict) – Arguments passed unchanged to algorithm.

  • reset_index (bool, default False) – Reset each group DataFrame before calling algorithm.

  • n_jobs (int, default 1) – Number of parallel jobs. 1 executes sequentially.

  • print_progress (bool, default False) – Whether to show a progress bar.

Returns:

List with one result per group.

Return type:

list

nomad.stop_detection.utils.explode_stops(stops, agg_freq='d', start_col='start_datetime', end_col='end_datetime', use_datetime=True)[source]

Explode each stop into one row per time bucket (day or week) based on agg_freq.

Parameters:
  • stops (pd.DataFrame)

  • agg_freq (str) – “d” for daily, “w” for weekly.

  • start_col (str) – Column names for start and end times.

  • end_col (str) – Column names for start and end times.

  • use_datetime (bool) – If False, converts start/end columns from unix seconds to datetime.

Returns:

Exploded table with updated start/end/duration per bucket.

Return type:

pd.DataFrame

nomad.stop_detection.utils.has_overlapping_stops(stop_data, traj_cols=None, **kwargs)[source]

Return True when any stop interval overlaps with the previous interval.

Intervals are interpreted from start/end columns when available, otherwise end times are reconstructed from duration in minutes.

nomad.stop_detection.utils.pad_short_stops(stop_data, pad=5, dur_min=None, traj_cols=None, **kwargs)[source]

Helper that pads stops shorter or equal than dur_min minutes extending the duration by pad minutes, but avoiding overlap. stop_data must be sorted chronologically and not overlap.

nomad.stop_detection.utils.summarize_stop(grouped_data, method='medoid', complete_output=False, keep_col_names=True, passthrough_cols=None, passthrough_agg=None, traj_cols=None, **kwargs)[source]
nomad.stop_detection.utils.summarize_stop_grid(grouped_data, complete_output=False, keep_col_names=True, passthrough_cols=None, traj_cols=None, passthrough_agg=None, **kwargs)[source]

Summarize index/grid‐based stop clusters (location_id only).

Parameters:
  • grouped_data (pd.DataFrame or gpd.GeoDataFrame) – All pings sharing the same location_id.

  • complete_output (bool) – If True, include n_pings, duration, and max_gap.

  • keep_col_names (bool) – If False, output start_/end_ keys are DEFAULT_SCHEMA ones; if True, they use the user’s time‐column name.

  • passthrough_cols (list[str], optional) – Additional columns (e.g. ‘user_id’) to carry through.

  • passthrough_agg (dict, optional) – Pandas-compatible aggregation function for each passthrough column that should not use its first value.

  • traj_cols (dict, optional) – Column‐name overrides.

Returns:

One‐row summary with:
  • start_[timestamp|datetime],

  • end_[timestamp|datetime] (if complete_output),

  • duration, n_pings, max_gap (if complete_output),

  • location_id, geometry (if present), plus passthrough_cols.

Return type:

pd.Series

nomad.stop_detection.utils.summarize_stops(data, labels, complete_output=False, dur_min=None, passthrough_cols=None, passthrough_agg=None, keep_col_names=True, traj_cols=None, **kwargs)[source]

Summarize non-noise point labels into one row per stop.

Parameters:
  • data (pd.DataFrame) – Input trajectory.

  • labels (pd.Series) – Cluster label for each trajectory row, with -1 denoting noise.

  • complete_output (bool) – Whether to include extended stop statistics.

  • dur_min (number, optional) – Minimum summarized stop duration to retain.

  • passthrough_cols (list, optional) – Columns copied into each stop summary. The first value is used by default.

  • passthrough_agg (dict, optional) – Pandas-compatible aggregation function for each passthrough column that should not use its first value.

  • keep_col_names (bool) – Whether to retain input coordinate and time column names.

  • traj_cols (dict, optional) – Canonical-to-actual column mapping.

  • **kwargs – Additional column mappings.

Returns:

One row per non-noise cluster.

Return type:

pd.DataFrame