base
- nomad.io.base.from_df(df, parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, filters=None, sort_times=True, traj_cols=None, **kwargs)[source]
Converts a DataFrame into a standardized trajectory format by validating and casting specified spatial and temporal columns.
- Parameters:
df (pd.DataFrame or gpd.GeoDataFrame) – The input DataFrame containing trajectory data.
traj_cols (dict, optional) – Mapping of expected trajectory column names (e.g., ‘latitude’, ‘longitude’, ‘datetime’, ‘user_id’, etc.) to actual column names in df. If None, kwargs is used for inference.
parse_dates (bool, default=True) – Whether to parse datetime columns as pandas datetime objects.
mixed_timezone_behavior ({'utc', 'naive', 'object'}, default='naive') – Controls how datetime columns with mixed time zones are handled: - ‘utc’: Convert all datetimes to UTC. - ‘naive’: Strip time zone information and store offsets separately. - ‘object’: Keep timestamps as pd.Timestamp objects with mixed time zones.
fixed_format (str, optional) – Format string for faster parsing of datetime columns if known.
**kwargs (dict) – Additional parameters for column inference when traj_cols is not provided.
- Returns:
The processed DataFrame with validated and correctly typed trajectory columns.
- Return type:
pd.DataFrame
Notes
Any specified traj_cols that do not exist in df will trigger a warning.
If traj_cols is not provided, missing trajectory columns are inferred from kwargs or filled with default schema values when possible.
Spatial columns are validated, and datetime columns are processed based on parse_dates and mixed_timezone_behavior.
If mixed_timezone_behavior=’naive’, a separate column storing UTC offsets (in seconds) is added.
- nomad.io.base.from_file(filepath, format='csv', parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, sep=',', filters=None, sort_times=True, traj_cols=None, **kwargs)[source]
Load and cast trajectory data from a specified file path or list of paths.
- Parameters:
filepath (str or list of str) – Path or list of paths to the file(s) or directories containing the data.
format (str, optional) – The format of the data files, either ‘csv’ or ‘parquet’.
traj_cols (dict, optional) – Mapping from trajectory fields (e.g. ‘latitude’, ‘timestamp’) to column names in the input.
parse_dates (bool, default True) – Whether to parse timestamp columns as datetime.
mixed_timezone_behavior ({'utc', 'naive', 'object'}, default='naive') – Controls how datetime columns with mixed time zones are handled: - ‘utc’: Convert all datetimes to UTC. - ‘naive’: Strip time zone information and store offsets separately. - ‘object’: Keep timestamps as pd.Timestamp objects with mixed time zones.
fixed_format (str, optional) – strftime format string for datetime parsing.
sep (str, default ',') – Field delimiter for CSV input.
filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read‐time filter for Parquet. Accepts a PyArrow Expression, a (column, operator, value) tuple, or a list of such tuples (AND‐chained).
**kwargs (dict) – Additional parameters for column inference when traj_cols is not provided.
- Returns:
DataFrame with trajectory columns cast and dates parsed.
- Return type:
pd.DataFrame
- nomad.io.base.sample_from_file(filepath, users=None, format='csv', frac_users=1.0, frac_records=1.0, seed=None, parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, sep=',', filters=None, within=None, poly_crs=None, data_crs=None, sort_times=True, traj_cols=None, **kwargs)[source]
Read and sample trajectory data from a file.
- Parameters:
filepath (str or Path) – Path to the input file.
users (list of hashable or None, default None) – If provided, only include these user IDs; if None, include all users.
format (str, default "csv") – Data format, e.g. “csv” or “parquet”.
traj_cols (dict or None, default None) – Mapping of trajectory column names (e.g. {“uid”: “user_id”}), or None to use defaults.
frac_users (float, default 1.0) – Fraction of users to sample (0 < frac_users ≤ 1).
frac_records (float, default 1.0) – Fraction of each user’s records to sample (0 < frac_records ≤ 1).
seed (int or None, default None) – Random seed for reproducibility.
parse_dates (bool, default True) – Whether to parse date/time columns as datetime objects.
mixed_timezone_behavior (str, default "naive") – How to handle mixed‐timezone timestamps; options might include “naive”, “utc”, etc.
fixed_format (str or None, default None) – If specified, enforce this input format (overrides autodetection).
sep (str) – Separator character for reader. Defaults to “,”.
filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read-time filter(s) for Parquet inputs. Ignored for single CSV files.
**kwargs – Passed through to the underlying reader (e.g. pandas.read_csv).
- Returns:
Sampled trajectory data.
- Return type:
pandas.DataFrame
- nomad.io.base.sample_users(filepath, format='csv', size=1.0, seed=None, sep=',', filters=None, within=None, poly_crs=None, data_crs=None, traj_cols=None, **kwargs)[source]
Sample users from a dataset, with optional read-time filtering for Parquet.
- Parameters:
filepath (str or Path) – Path to the data file or directory.
format ({'csv','parquet'}, default 'csv') – Input format.
size (float or int, default 1.0) – Fraction (0–1] or absolute number of users to sample.
seed (int, optional) – Random seed for reproducibility.
sep (str, default ',') – CSV delimiter.
filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read-time filter(s) for Parquet inputs. Ignored for single CSV files.
within (shapely Polygon/MultiPolygon or WKT str, default None) – If supplied, keep only points whose coordinates fall inside this polygon.
data_crs (str or pyproj.CRS, optional) – CRS for data when it is a plain DataFrame; ignored if data is a GeoDataFrame.
traj_cols (dict, optional) – Mapping of logical names (‘user_id’, etc.) to actual column names.
**kwargs – Passed through to the underlying reader.
- Returns:
Sampled user IDs.
- Return type:
pd.Series
- nomad.io.base.table_columns(filepath, format='csv', include_schema=False, sep=',')[source]
Return column names or the full schema of a data source.
The ‘sep’ argument specifies the delimiter and is only used for ‘csv’ format; it is ignored when reading ‘parquet’ files.