visit_attribution
- nomad.visit_attribution.visit_attribution.clip_stays_date(traj, dates, dawn_hour=6, dusk_hour=19)[source]
- nomad.visit_attribution.visit_attribution.count_nights(usr_polygon, dawn_hour=6, dusk_hour=19, min_dwell=10)[source]
- nomad.visit_attribution.visit_attribution.detect_locations(data, epsilon=100, min_pts=1, return_locations=False, method='dbscan', algorithm=None, algorithm_kwargs=None, traj_cols=None, **kwargs)[source]
Assign recurring-location IDs to coordinate points.
- Parameters:
data (pd.DataFrame) – Coordinate points or stop rows with x/y or longitude/latitude columns.
epsilon (float, default 100) – Maximum distance between rows in the same location. Units match projected coordinates or are meters for longitude/latitude coordinates.
min_pts (int, default 1) – Minimum number of stops required to form a location
return_locations (bool, default False) – If True, also return a GeoDataFrame summarizing each location. Location summaries currently require projected x/y coordinates.
method ({'dbscan', 'custom', 'infomap'}, default 'dbscan') – Location-detection method. Infomap is reserved for future multi-user stop-table support and currently raises
NotImplementedError.algorithm (callable, optional) – Location-detection callable required when
method='custom'. It must return one label per input row.algorithm_kwargs (dict, optional) – Arguments for a custom location-detection callable.
traj_cols (dict, optional) – Column name mappings for coordinates and location_id.
**kwargs – Additional arguments passed to column detection
- Returns:
Location IDs aligned with
data.indexand, when requested, location centers, extents, and row counts.- Return type:
pd.Series or tuple of (pd.Series, gpd.GeoDataFrame)
Notes
Each noise row is retained as its own nonnegative location.
- nomad.visit_attribution.visit_attribution.duration_at_night_fast(start, end, dawn_hour=6, dusk_hour=19)[source]
- nomad.visit_attribution.visit_attribution.night_stops(stop_table, user='user', dawn_hour=6, dusk_hour=19, min_dwell=10)[source]
- nomad.visit_attribution.visit_attribution.oracle_map(data, true_visits, traj_cols=None, **kwargs)[source]
Map elements in traj to ground truth location based solely on time.
- Parameters:
data (pd.DataFrame) – The trajectory DataFrame containing x and y coordinates.
true_visits (pd.DataFrame) – A visitation table containing location IDs, start times, and durations/end times.
traj_cols (list) – The columns in the trajectory DataFrame to be used for mapping.
**kwargs (dict) – Additional keyword arguments.
- Returns:
A Series containing the location IDs corresponding to the pings in the trajectory.
- Return type:
pd.Series
- nomad.visit_attribution.visit_attribution.poi_map(data, poi_table, max_distance=0, data_crs=None, location_id=None, traj_cols=None, **kwargs)[source]
Assign each point or H3 containment area in data to a polygon in poi_table.
Points use geometric containment when max_distance==0 or the nearest neighbor within max_distance otherwise. H3 cells use intersecting POI cells or the POI cell with the smallest H3 grid distance.
- Parameters:
data (pd.DataFrame or gpd.GeoDataFrame) – Input points, either as a DataFrame with coordinate columns or a GeoDataFrame, or a table containing H3 cells.
poi_table (gpd.GeoDataFrame) – Polygons to match against, indexed or with location_id column.
traj_cols (list of str, optional) – Names of the coordinate columns in data when it is a DataFrame.
max_distance (float, default 0) – Maximum search radius for nearest-neighbor matching. For H3 input, this is the maximum grid distance in cells.
data_crs (str or pyproj.CRS, optional) – CRS for data if it is a DataFrame; ignored for GeoDataFrames.
location_id (str, optional) – Name of the geometry ID column in poi_table; uses the GeoDataFrame index if not provided.
**kwargs – Passed to trajectory‐column parsing helper.
- Returns:
Indexed like data, with each entry set to the matching polygon’s ID (from location_id or poi_table.index). Points or cells not contained or beyond max_distance yield NaN. Ties retain the first POI in poi_table.
- Return type:
pd.Series
- nomad.visit_attribution.visit_attribution.point_in_polygon(data, poi_table, method='centroid', data_crs=None, max_distance=0, cluster_label=None, location_id=None, recompute_location=True, traj_cols=None, **kwargs)[source]
Assign each stop or cluster of pings in data to a polygon in poi_table, either by the cluster’s centroid location or by the most frequent polygon hit.
- Parameters:
data (pd.DataFrame or gpd.GeoDataFrame) – A table of pings (with optional stop/duration columns) or stops, indexed by observation or cluster.
poi_table (gpd.GeoDataFrame) – Polygons to match against, with CRS set and optional ID column.
method ({'centroid', 'majority'}, default 'centroid') – ‘centroid’ uses each cluster’s mean point; ‘majority’ picks the polygon most often visited within each cluster (only for ping data).
data_crs (str or pyproj.CRS, optional) – CRS for data when it is a plain DataFrame; ignored if data is a GeoDataFrame.
max_distance (float, default 0) – Search radius for nearest‐neighbor fall-back; zero triggers strict point-in-polygon matching.
cluster_label (str, optional) – Column name holding cluster IDs in ping data; inferred from data if absent.
location_id (str, optional) – Column in poi_table containing the output ID; uses the GeoDataFrame index if None.
recompute_location (bool, default True) – For labeled ping data, ignored for stop data. If False and a location column (as determined by location_id) already exists in data, it will be reused instead of overwritten.
traj_cols (list of str, optional) – Names of the coordinate columns in data when it is a DataFrame.
**kwargs – Passed through to poi_map or the trajectory-column parser.
- Returns:
Indexed like data, giving the matched polygon ID for each stop or ping. Points or clusters that fall outside every polygon or beyond max_distance are set to NaN.
- Return type:
pd.Series