home_attribution

nomad.visit_attribution.home_attribution.compute_candidate_homes(stops_table, dusk_hour=19, dawn_hour=6, traj_cols=None, **kwargs)[source]

Aggregate night-time stop statistics that serve as features for home-location inference.

Internally this calls nocturnal_stops() to keep only night portions of the stops, then counts how many distinct nights and ISO-8601 calendar weeks each user spent at each location.

Parameters:
  • stops_table (pandas.DataFrame) –

    A stop table with at least

    • temporal columns (start_* and either end_* or duration),

    • one user identifier, and

    • one location identifier.

    Column names can be supplied/overridden via traj_cols or kwargs.

  • dusk_hour (int, optional) – Same semantics as in nocturnal_stops().

  • dawn_hour (int, optional) – Same semantics as in nocturnal_stops().

  • traj_cols (dict, optional) – Mapping from canonical names ("user_id", "location_id", "start_timestamp"/"start_datetime", etc.) to the actual column names in stops_table.

  • **kwargs – Column-overrides written as <canonical>=<actual>; passed through to the NOMAD I/O helpers.

Returns:

Columns:

  • user_id – user identifier

  • location_id – candidate home location

  • num_nights – unique nights present at the location

  • num_weeks – unique ISO weeks present at the location

  • total_duration – aggregated night-time minutes

Return type:

pandas.DataFrame

Raises:

ValueError – If neither an end column nor a duration column is present.

nomad.visit_attribution.home_attribution.compute_candidate_workplaces(stops_table, work_start_hour=9, work_end_hour=17, include_weekends=False, traj_cols=None, **kwargs)[source]

Build per‑location daytime presence features for workplace inference.

Internally this calls workday_stops() to keep only weekday work‑hour portions of the stops, then counts how many distinct workdays and ISO weeks each user spent at each location.

Returns a DataFrame with the columns ['user_id', 'location_id', 'num_work_days', 'num_weeks', 'total_duration'].

nomad.visit_attribution.home_attribution.nocturnal_stops(stops_table, dusk_hour=19, dawn_hour=6, start_datetime='start_datetime', end_datetime='end_datetime', duration='duration')[source]

Slice each stop so that only its night-time portion—defined by a daily window from dusk_hour to dawn_hour—is retained.

The helper assumes that stops_table already contains fully parsed timezone-aware datetimes. For every stop it constructs the set of candidate night windows it might intersect, clips the stop to those windows, recomputes the duration, and drops rows whose clipped duration is zero.

Parameters:
  • stops_table (pandas.DataFrame) –

    Output of a stop-detection algorithm with at least

    • a start column named start_datetime (default "start_datetime") and

    • an end column named end_datetime (default "end_datetime").

    Both must be datetime64[ns, tz].

  • dusk_hour (int, default 19) – Local hour (0–23) that marks the beginning of the nocturnal window on the same calendar day.

  • dawn_hour (int, default 6) – Local hour (0–23) that marks the end of the nocturnal window on the following calendar day.

  • start_datetime (str, optional) – Column names if different from the defaults.

  • end_datetime (str, optional) – Column names if different from the defaults.

Returns:

Same columns as the input plus an updated duration (integer minutes) and only those rows whose clipped duration is positive. Temporary helper columns are removed.

Return type:

pandas.DataFrame

Notes

  • A stop that spans several nights is exploded so that one row per night is produced.

  • The duration is integer-divided by 60 s → minutes.

  • Time-zone information is preserved.

nomad.visit_attribution.home_attribution.select_home(candidate_homes, min_days, min_weeks, stops_table=None, last_date=None, traj_cols=None, **kwargs)[source]

Choose the single most plausible home location for each user filtering candidate_homes by minimum presence thresholds, and ranking the remaining locations by

  1. number of nights (descending) and

  2. total night-time duration (descending),

and finally returns the top-ranked location together with the date of the last stop observation.

Parameters:
  • candidate_homes (pandas.DataFrame) – Output of compute_candidate_homes().

  • stops_table (pandas.DataFrame) – Full stop table (not night-clipped) used only to determine the most recent observation date.

  • min_days (int) – Minimum number of distinct nights required for a location to qualify as home.

  • min_weeks (int) – Minimum number of distinct ISO weeks required for a location to qualify as home.

  • traj_cols (dict, optional) – Column mapping overrides, as in other NOMAD helpers.

  • **kwargs – Additional parameters

Returns:

Columns ['user_id', 'location_id', 'home_date'] with exactly one row per user. home_date equals the date (YYYY-MM-DD) of the most recent stop in stops_table.

Return type:

pandas.DataFrame

Notes

  • Ties beyond the ranking rules are broken by first occurrence

  • Users with no location meeting the thresholds are omitted

nomad.visit_attribution.home_attribution.select_workplace(candidate_workplaces, min_days, min_weeks, stops_table=None, last_date=None, traj_cols=None, **kwargs)[source]

Choose the single most plausible workplace for each user.

Parameters:
  • candidate_workplaces (pandas.DataFrame) – Output of :pyfunc:`compute_candidate_workplaces`.

  • stops_table (pandas.DataFrame) – Full stop table used only to derive the date of the last observation.

  • min_days (int) – Presence thresholds.

  • min_weeks (int) – Presence thresholds.

Returns:

Columns ['user_id', 'location_id', 'work_date'] with exactly one row per user. work_date equals the date (YYYY-MM-DD) of the most recent stop in stops_table.

Return type:

pandas.DataFrame

Notes

candidate_workplaces must contain ['user_id', 'location_id', 'num_work_days', 'num_weeks', 'total_duration'] as produced by compute_candidate_workplaces().

nomad.visit_attribution.home_attribution.workday_stops(stops_table, work_start_hour=9, work_end_hour=17, include_weekends=False, start_datetime='start_datetime', end_datetime='end_datetime', duration='duration')[source]

Clip each stop to the daily work-hour window work_start_hour → work_end_hour and return only the portions that fall on workdays.

Stops that span several calendar days are exploded into one row per day (exactly like the night-time helper) so long multi-day artefacts do not crash the logic.

Parameters:
  • stops_table (pandas.DataFrame) – At minimum the two datetime columns given in start_datetime and end_datetime (timezone-aware).

  • work_start_hour (int, optional) – Local hours (0–23) delimiting the daily work block. work_start_hour must be strictly < work_end_hour; otherwise a ValueError is raised.

  • work_end_hour (int, optional) – Local hours (0–23) delimiting the daily work block. work_start_hour must be strictly < work_end_hour; otherwise a ValueError is raised.

  • include_weekends (bool, default False) – If False (default) rows whose clipped interval falls entirely on Saturday or Sunday are dropped.

  • start_datetime (str, optional) – Column names overriding the defaults.

  • end_datetime (str, optional) – Column names overriding the defaults.

Returns:

Same schema as the input plus an updated integer-minute duration and only rows whose clipped duration is positive.

Return type:

pandas.DataFrame

Notes

  • Night‑shift workplace detection (where work_start_hour ≥ work_end_hour) is not implemented because it conflicts conceptually with the night‑ time logic used for home inference. A dedicated routine should be implemented for such cases.