home_attribution
- nomad.visit_attribution.home_attribution.compute_candidate_homes(stops_table, dusk_hour=19, dawn_hour=6, traj_cols=None, **kwargs)[source]
Aggregate night-time stop statistics that serve as features for home-location inference.
Internally this calls
nocturnal_stops()to keep only night portions of the stops, then counts how many distinct nights and ISO-8601 calendar weeks each user spent at each location.- Parameters:
stops_table (pandas.DataFrame) –
A stop table with at least
temporal columns (
start_*and eitherend_*orduration),one user identifier, and
one location identifier.
Column names can be supplied/overridden via traj_cols or kwargs.
dusk_hour (int, optional) – Same semantics as in
nocturnal_stops().dawn_hour (int, optional) – Same semantics as in
nocturnal_stops().traj_cols (dict, optional) – Mapping from canonical names (
"user_id","location_id","start_timestamp"/"start_datetime", etc.) to the actual column names in stops_table.**kwargs – Column-overrides written as
<canonical>=<actual>; passed through to the NOMAD I/O helpers.
- Returns:
Columns:
user_id– user identifierlocation_id– candidate home locationnum_nights– unique nights present at the locationnum_weeks– unique ISO weeks present at the locationtotal_duration– aggregated night-time minutes
- Return type:
pandas.DataFrame
- Raises:
ValueError – If neither an end column nor a duration column is present.
- nomad.visit_attribution.home_attribution.compute_candidate_workplaces(stops_table, work_start_hour=9, work_end_hour=17, include_weekends=False, traj_cols=None, **kwargs)[source]
Build per‑location daytime presence features for workplace inference.
Internally this calls
workday_stops()to keep only weekday work‑hour portions of the stops, then counts how many distinct workdays and ISO weeks each user spent at each location.Returns a DataFrame with the columns
['user_id', 'location_id', 'num_work_days', 'num_weeks', 'total_duration'].
- nomad.visit_attribution.home_attribution.nocturnal_stops(stops_table, dusk_hour=19, dawn_hour=6, start_datetime='start_datetime', end_datetime='end_datetime', duration='duration')[source]
Slice each stop so that only its night-time portion—defined by a daily window from
dusk_hourtodawn_hour—is retained.The helper assumes that
stops_tablealready contains fully parsed timezone-aware datetimes. For every stop it constructs the set of candidate night windows it might intersect, clips the stop to those windows, recomputes the duration, and drops rows whose clipped duration is zero.- Parameters:
stops_table (pandas.DataFrame) –
Output of a stop-detection algorithm with at least
a start column named
start_datetime(default"start_datetime") andan end column named
end_datetime(default"end_datetime").
Both must be
datetime64[ns, tz].dusk_hour (int, default 19) – Local hour (0–23) that marks the beginning of the nocturnal window on the same calendar day.
dawn_hour (int, default 6) – Local hour (0–23) that marks the end of the nocturnal window on the following calendar day.
start_datetime (str, optional) – Column names if different from the defaults.
end_datetime (str, optional) – Column names if different from the defaults.
- Returns:
Same columns as the input plus an updated
duration(integer minutes) and only those rows whose clipped duration is positive. Temporary helper columns are removed.- Return type:
pandas.DataFrame
Notes
A stop that spans several nights is exploded so that one row per night is produced.
The duration is integer-divided by 60 s → minutes.
Time-zone information is preserved.
- nomad.visit_attribution.home_attribution.select_home(candidate_homes, min_days, min_weeks, stops_table=None, last_date=None, traj_cols=None, **kwargs)[source]
Choose the single most plausible home location for each user filtering candidate_homes by minimum presence thresholds, and ranking the remaining locations by
number of nights (descending) and
total night-time duration (descending),
and finally returns the top-ranked location together with the date of the last stop observation.
- Parameters:
candidate_homes (pandas.DataFrame) – Output of
compute_candidate_homes().stops_table (pandas.DataFrame) – Full stop table (not night-clipped) used only to determine the most recent observation date.
min_days (int) – Minimum number of distinct nights required for a location to qualify as home.
min_weeks (int) – Minimum number of distinct ISO weeks required for a location to qualify as home.
traj_cols (dict, optional) – Column mapping overrides, as in other NOMAD helpers.
**kwargs – Additional parameters
- Returns:
Columns
['user_id', 'location_id', 'home_date']with exactly one row per user. home_date equals the date (YYYY-MM-DD) of the most recent stop in stops_table.- Return type:
pandas.DataFrame
Notes
Ties beyond the ranking rules are broken by first occurrence
Users with no location meeting the thresholds are omitted
- nomad.visit_attribution.home_attribution.select_workplace(candidate_workplaces, min_days, min_weeks, stops_table=None, last_date=None, traj_cols=None, **kwargs)[source]
Choose the single most plausible workplace for each user.
- Parameters:
candidate_workplaces (pandas.DataFrame) – Output of :pyfunc:`compute_candidate_workplaces`.
stops_table (pandas.DataFrame) – Full stop table used only to derive the date of the last observation.
min_days (int) – Presence thresholds.
min_weeks (int) – Presence thresholds.
- Returns:
Columns
['user_id', 'location_id', 'work_date']with exactly one row per user. work_date equals the date (YYYY-MM-DD) of the most recent stop in stops_table.- Return type:
pandas.DataFrame
Notes
candidate_workplacesmust contain['user_id', 'location_id', 'num_work_days', 'num_weeks', 'total_duration']as produced bycompute_candidate_workplaces().
- nomad.visit_attribution.home_attribution.workday_stops(stops_table, work_start_hour=9, work_end_hour=17, include_weekends=False, start_datetime='start_datetime', end_datetime='end_datetime', duration='duration')[source]
Clip each stop to the daily work-hour window
work_start_hour→work_end_hourand return only the portions that fall on workdays.Stops that span several calendar days are exploded into one row per day (exactly like the night-time helper) so long multi-day artefacts do not crash the logic.
- Parameters:
stops_table (pandas.DataFrame) – At minimum the two datetime columns given in start_datetime and end_datetime (timezone-aware).
work_start_hour (int, optional) – Local hours (0–23) delimiting the daily work block.
work_start_hourmust be strictly <work_end_hour; otherwise aValueErroris raised.work_end_hour (int, optional) – Local hours (0–23) delimiting the daily work block.
work_start_hourmust be strictly <work_end_hour; otherwise aValueErroris raised.include_weekends (bool, default False) – If False (default) rows whose clipped interval falls entirely on Saturday or Sunday are dropped.
start_datetime (str, optional) – Column names overriding the defaults.
end_datetime (str, optional) – Column names overriding the defaults.
- Returns:
Same schema as the input plus an updated integer-minute
durationand only rows whose clipped duration is positive.- Return type:
pandas.DataFrame
Notes
Night‑shift workplace detection (where work_start_hour ≥ work_end_hour) is not implemented because it conflicts conceptually with the night‑ time logic used for home inference. A dedicated routine should be implemented for such cases.