visit_attribution

nomad.visit_attribution.visit_attribution.clip_stays_date(traj, dates, dawn_hour=6, dusk_hour=19)[source]
nomad.visit_attribution.visit_attribution.count_nights(usr_polygon, dawn_hour=6, dusk_hour=19, min_dwell=10)[source]
nomad.visit_attribution.visit_attribution.dawn_time(day_part, dawn_hour=6)[source]
nomad.visit_attribution.visit_attribution.detect_locations(data, epsilon=100, min_pts=1, return_locations=False, method='dbscan', algorithm=None, algorithm_kwargs=None, traj_cols=None, **kwargs)[source]

Assign recurring-location IDs to coordinate points.

Parameters:
  • data (pd.DataFrame) – Coordinate points or stop rows with x/y or longitude/latitude columns.

  • epsilon (float, default 100) – Maximum distance between rows in the same location. Units match projected coordinates or are meters for longitude/latitude coordinates.

  • min_pts (int, default 1) – Minimum number of stops required to form a location

  • return_locations (bool, default False) – If True, also return a GeoDataFrame summarizing each location. Location summaries currently require projected x/y coordinates.

  • method ({'dbscan', 'custom', 'infomap'}, default 'dbscan') – Location-detection method. Infomap is reserved for future multi-user stop-table support and currently raises NotImplementedError.

  • algorithm (callable, optional) – Location-detection callable required when method='custom'. It must return one label per input row.

  • algorithm_kwargs (dict, optional) – Arguments for a custom location-detection callable.

  • traj_cols (dict, optional) – Column name mappings for coordinates and location_id.

  • **kwargs – Additional arguments passed to column detection

Returns:

Location IDs aligned with data.index and, when requested, location centers, extents, and row counts.

Return type:

pd.Series or tuple of (pd.Series, gpd.GeoDataFrame)

Notes

Each noise row is retained as its own nonnegative location.

nomad.visit_attribution.visit_attribution.duration_at_night_fast(start, end, dawn_hour=6, dusk_hour=19)[source]
nomad.visit_attribution.visit_attribution.dusk_time(day_part, dusk_hour=19)[source]
nomad.visit_attribution.visit_attribution.night_stops(stop_table, user='user', dawn_hour=6, dusk_hour=19, min_dwell=10)[source]
nomad.visit_attribution.visit_attribution.oracle_map(data, true_visits, traj_cols=None, **kwargs)[source]

Map elements in traj to ground truth location based solely on time.

Parameters:
  • data (pd.DataFrame) – The trajectory DataFrame containing x and y coordinates.

  • true_visits (pd.DataFrame) – A visitation table containing location IDs, start times, and durations/end times.

  • traj_cols (list) – The columns in the trajectory DataFrame to be used for mapping.

  • **kwargs (dict) – Additional keyword arguments.

Returns:

A Series containing the location IDs corresponding to the pings in the trajectory.

Return type:

pd.Series

nomad.visit_attribution.visit_attribution.poi_map(data, poi_table, max_distance=0, data_crs=None, location_id=None, traj_cols=None, **kwargs)[source]

Assign each point or H3 containment area in data to a polygon in poi_table.

Points use geometric containment when max_distance==0 or the nearest neighbor within max_distance otherwise. H3 cells use intersecting POI cells or the POI cell with the smallest H3 grid distance.

Parameters:
  • data (pd.DataFrame or gpd.GeoDataFrame) – Input points, either as a DataFrame with coordinate columns or a GeoDataFrame, or a table containing H3 cells.

  • poi_table (gpd.GeoDataFrame) – Polygons to match against, indexed or with location_id column.

  • traj_cols (list of str, optional) – Names of the coordinate columns in data when it is a DataFrame.

  • max_distance (float, default 0) – Maximum search radius for nearest-neighbor matching. For H3 input, this is the maximum grid distance in cells.

  • data_crs (str or pyproj.CRS, optional) – CRS for data if it is a DataFrame; ignored for GeoDataFrames.

  • location_id (str, optional) – Name of the geometry ID column in poi_table; uses the GeoDataFrame index if not provided.

  • **kwargs – Passed to trajectory‐column parsing helper.

Returns:

Indexed like data, with each entry set to the matching polygon’s ID (from location_id or poi_table.index). Points or cells not contained or beyond max_distance yield NaN. Ties retain the first POI in poi_table.

Return type:

pd.Series

nomad.visit_attribution.visit_attribution.point_in_polygon(data, poi_table, method='centroid', data_crs=None, max_distance=0, cluster_label=None, location_id=None, recompute_location=True, traj_cols=None, **kwargs)[source]

Assign each stop or cluster of pings in data to a polygon in poi_table, either by the cluster’s centroid location or by the most frequent polygon hit.

Parameters:
  • data (pd.DataFrame or gpd.GeoDataFrame) – A table of pings (with optional stop/duration columns) or stops, indexed by observation or cluster.

  • poi_table (gpd.GeoDataFrame) – Polygons to match against, with CRS set and optional ID column.

  • method ({'centroid', 'majority'}, default 'centroid') – ‘centroid’ uses each cluster’s mean point; ‘majority’ picks the polygon most often visited within each cluster (only for ping data).

  • data_crs (str or pyproj.CRS, optional) – CRS for data when it is a plain DataFrame; ignored if data is a GeoDataFrame.

  • max_distance (float, default 0) – Search radius for nearest‐neighbor fall-back; zero triggers strict point-in-polygon matching.

  • cluster_label (str, optional) – Column name holding cluster IDs in ping data; inferred from data if absent.

  • location_id (str, optional) – Column in poi_table containing the output ID; uses the GeoDataFrame index if None.

  • recompute_location (bool, default True) – For labeled ping data, ignored for stop data. If False and a location column (as determined by location_id) already exists in data, it will be reused instead of overwritten.

  • traj_cols (list of str, optional) – Names of the coordinate columns in data when it is a DataFrame.

  • **kwargs – Passed through to poi_map or the trajectory-column parser.

Returns:

Indexed like data, giving the matched polygon ID for each stop or ping. Points or clusters that fall outside every polygon or beyond max_distance are set to NaN.

Return type:

pd.Series

nomad.visit_attribution.visit_attribution.slice_datetimes_interval_fast(start, end)[source]