filters_spark
- nomad.filters_spark.filter_users(traj: DataFrame, start_time, end_time, timezone=None, polygon=None, min_active_days=1, min_pings_per_day=1, traj_cols=None, crs='EPSG:3857', spark_session=None, **kwargs)[source]
Subsets to users who have at least min_pings_per_day pings on min_active_days distinct days in the polygon within the timeframe start_time to start_time.
- Parameters:
traj (pd.DataFrame) – Trajectory DataFrame with latitude and longitude columns.
start_time – Start of the timeframe for filtering.
end_time – End of the timeframe for filtering.
polygon (shapely.geometry.Polygon) – Polygon defining the area to retain points within. If None, no spatial filtering is applied.
min_active_days (int) – User is retained if they have at least min_pings_per_day pings on min_active_days distinct days. Defaults to 1.
min_pings_per_day (int) – User is retained if they have at least min_pings_per_day pings on min_active_days distinct days. Defaults to 1.
traj_cols (dict, optional) – A dictionary defining column mappings for ‘x’, ‘y’, ‘longitude’, ‘latitude’, ‘timestamp’, or ‘datetime’. If not provided, the function will attempt to use default column names or those provided in kwargs.
crs (str, optional) – Coordinate Reference System (CRS) for the polygon. Defaults to “EPSG:3857”.
spark_session (SparkSession, optional) – Spark session for distributed computation, if needed.
**kwargs – Additional parameters like ‘user_id’, ‘latitude’, ‘longitude’, or ‘datetime’ column names.
- Returns:
Filtered DataFrame with points inside the polygon’s bounds.
- Return type:
pd.DataFrame
- nomad.filters_spark.to_projection(traj: DataFrame, input_crs=None, output_crs=None, spark_session: pyspark.sql.SparkSession = None, **kwargs)[source]
Projects coordinate columns from one Coordinate Reference System (CRS) to another.
This function takes a DataFrame containing coordinate columns and projects them from one CRS to another specified CRS. It supports both local and distributed computation using Spark. (TODO: SPARK)
If columns names are not specified in kwargs, the function will attempt to use default column names.
- Parameters:
traj (pd.DataFrame) – Trajectory DataFrame containing coordinate columns.
input_crs (str, optional) – EPSG code for the original CRS.
output_crs (str, optional) – EPSG code for the target CRS.
spark_session (SparkSession, optional.) – Spark session for distributed computation, if needed.
**kwargs – Additional parameters to specify names of spatial columns to project from. E.g., ‘latitude’, ‘longitude’ or ‘x’, ‘y’.
- Returns:
A pair of Series with the new projected coordinates.
- Return type:
(pd.Series, pd.Series)
- Raises:
ValueError – If expected coordinate columns are missing.