filters_spark

nomad.filters_spark.filter_users(traj: DataFrame, start_time, end_time, timezone=None, polygon=None, min_active_days=1, min_pings_per_day=1, traj_cols=None, crs='EPSG:3857', spark_session=None, **kwargs)[source]

Subsets to users who have at least min_pings_per_day pings on min_active_days distinct days in the polygon within the timeframe start_time to start_time.

Parameters:
  • traj (pd.DataFrame) – Trajectory DataFrame with latitude and longitude columns.

  • start_time – Start of the timeframe for filtering.

  • end_time – End of the timeframe for filtering.

  • polygon (shapely.geometry.Polygon) – Polygon defining the area to retain points within. If None, no spatial filtering is applied.

  • min_active_days (int) – User is retained if they have at least min_pings_per_day pings on min_active_days distinct days. Defaults to 1.

  • min_pings_per_day (int) – User is retained if they have at least min_pings_per_day pings on min_active_days distinct days. Defaults to 1.

  • traj_cols (dict, optional) – A dictionary defining column mappings for ‘x’, ‘y’, ‘longitude’, ‘latitude’, ‘timestamp’, or ‘datetime’. If not provided, the function will attempt to use default column names or those provided in kwargs.

  • crs (str, optional) – Coordinate Reference System (CRS) for the polygon. Defaults to “EPSG:3857”.

  • spark_session (SparkSession, optional) – Spark session for distributed computation, if needed.

  • **kwargs – Additional parameters like ‘user_id’, ‘latitude’, ‘longitude’, or ‘datetime’ column names.

Returns:

Filtered DataFrame with points inside the polygon’s bounds.

Return type:

pd.DataFrame

nomad.filters_spark.to_projection(traj: DataFrame, input_crs=None, output_crs=None, spark_session: pyspark.sql.SparkSession = None, **kwargs)[source]

Projects coordinate columns from one Coordinate Reference System (CRS) to another.

This function takes a DataFrame containing coordinate columns and projects them from one CRS to another specified CRS. It supports both local and distributed computation using Spark. (TODO: SPARK)

If columns names are not specified in kwargs, the function will attempt to use default column names.

Parameters:
  • traj (pd.DataFrame) – Trajectory DataFrame containing coordinate columns.

  • input_crs (str, optional) – EPSG code for the original CRS.

  • output_crs (str, optional) – EPSG code for the target CRS.

  • spark_session (SparkSession, optional.) – Spark session for distributed computation, if needed.

  • **kwargs – Additional parameters to specify names of spatial columns to project from. E.g., ‘latitude’, ‘longitude’ or ‘x’, ‘y’.

Returns:

A pair of Series with the new projected coordinates.

Return type:

(pd.Series, pd.Series)

Raises:

ValueError – If expected coordinate columns are missing.