base

nomad.io.base.from_df(df, parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, filters=None, sort_times=True, traj_cols=None, **kwargs)[source]

Converts a DataFrame into a standardized trajectory format by validating and casting specified spatial and temporal columns.

Parameters:
  • df (pd.DataFrame or gpd.GeoDataFrame) – The input DataFrame containing trajectory data.

  • traj_cols (dict, optional) – Mapping of expected trajectory column names (e.g., ‘latitude’, ‘longitude’, ‘datetime’, ‘user_id’, etc.) to actual column names in df. If None, kwargs is used for inference.

  • parse_dates (bool, default=True) – Whether to parse datetime columns as pandas datetime objects.

  • mixed_timezone_behavior ({'utc', 'naive', 'object'}, default='naive') – Controls how datetime columns with mixed time zones are handled: - ‘utc’: Convert all datetimes to UTC. - ‘naive’: Strip time zone information and store offsets separately. - ‘object’: Keep timestamps as pd.Timestamp objects with mixed time zones.

  • fixed_format (str, optional) – Format string for faster parsing of datetime columns if known.

  • **kwargs (dict) – Additional parameters for column inference when traj_cols is not provided.

Returns:

The processed DataFrame with validated and correctly typed trajectory columns.

Return type:

pd.DataFrame

Notes

  • Any specified traj_cols that do not exist in df will trigger a warning.

  • If traj_cols is not provided, missing trajectory columns are inferred from kwargs or filled with default schema values when possible.

  • Spatial columns are validated, and datetime columns are processed based on parse_dates and mixed_timezone_behavior.

  • If mixed_timezone_behavior=’naive’, a separate column storing UTC offsets (in seconds) is added.

nomad.io.base.from_file(filepath, format='csv', parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, sep=',', filters=None, sort_times=True, traj_cols=None, **kwargs)[source]

Load and cast trajectory data from a specified file path or list of paths.

Parameters:
  • filepath (str or list of str) – Path or list of paths to the file(s) or directories containing the data.

  • format (str, optional) – The format of the data files, either ‘csv’ or ‘parquet’.

  • traj_cols (dict, optional) – Mapping from trajectory fields (e.g. ‘latitude’, ‘timestamp’) to column names in the input.

  • parse_dates (bool, default True) – Whether to parse timestamp columns as datetime.

  • mixed_timezone_behavior ({'utc', 'naive', 'object'}, default='naive') – Controls how datetime columns with mixed time zones are handled: - ‘utc’: Convert all datetimes to UTC. - ‘naive’: Strip time zone information and store offsets separately. - ‘object’: Keep timestamps as pd.Timestamp objects with mixed time zones.

  • fixed_format (str, optional) – strftime format string for datetime parsing.

  • sep (str, default ',') – Field delimiter for CSV input.

  • filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read‐time filter for Parquet. Accepts a PyArrow Expression, a (column, operator, value) tuple, or a list of such tuples (AND‐chained).

  • **kwargs (dict) – Additional parameters for column inference when traj_cols is not provided.

Returns:

DataFrame with trajectory columns cast and dates parsed.

Return type:

pd.DataFrame

nomad.io.base.localize_from_offset(naive_dt, timezone_offset)[source]
nomad.io.base.naive_datetime_from_unix_and_offset(utc_timestamps, timezone_offset)[source]
nomad.io.base.sample_from_file(filepath, users=None, format='csv', frac_users=1.0, frac_records=1.0, seed=None, parse_dates=True, mixed_timezone_behavior='naive', fixed_format=None, sep=',', filters=None, within=None, poly_crs=None, data_crs=None, sort_times=True, traj_cols=None, **kwargs)[source]

Read and sample trajectory data from a file.

Parameters:
  • filepath (str or Path) – Path to the input file.

  • users (list of hashable or None, default None) – If provided, only include these user IDs; if None, include all users.

  • format (str, default "csv") – Data format, e.g. “csv” or “parquet”.

  • traj_cols (dict or None, default None) – Mapping of trajectory column names (e.g. {“uid”: “user_id”}), or None to use defaults.

  • frac_users (float, default 1.0) – Fraction of users to sample (0 < frac_users ≤ 1).

  • frac_records (float, default 1.0) – Fraction of each user’s records to sample (0 < frac_records ≤ 1).

  • seed (int or None, default None) – Random seed for reproducibility.

  • parse_dates (bool, default True) – Whether to parse date/time columns as datetime objects.

  • mixed_timezone_behavior (str, default "naive") – How to handle mixed‐timezone timestamps; options might include “naive”, “utc”, etc.

  • fixed_format (str or None, default None) – If specified, enforce this input format (overrides autodetection).

  • sep (str) – Separator character for reader. Defaults to “,”.

  • filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read-time filter(s) for Parquet inputs. Ignored for single CSV files.

  • **kwargs – Passed through to the underlying reader (e.g. pandas.read_csv).

Returns:

Sampled trajectory data.

Return type:

pandas.DataFrame

nomad.io.base.sample_users(filepath, format='csv', size=1.0, seed=None, sep=',', filters=None, within=None, poly_crs=None, data_crs=None, traj_cols=None, **kwargs)[source]

Sample users from a dataset, with optional read-time filtering for Parquet.

Parameters:
  • filepath (str or Path) – Path to the data file or directory.

  • format ({'csv','parquet'}, default 'csv') – Input format.

  • size (float or int, default 1.0) – Fraction (0–1] or absolute number of users to sample.

  • seed (int, optional) – Random seed for reproducibility.

  • sep (str, default ',') – CSV delimiter.

  • filters (pyarrow.dataset.Expression or tuple or list of tuples, optional) – Read-time filter(s) for Parquet inputs. Ignored for single CSV files.

  • within (shapely Polygon/MultiPolygon or WKT str, default None) – If supplied, keep only points whose coordinates fall inside this polygon.

  • data_crs (str or pyproj.CRS, optional) – CRS for data when it is a plain DataFrame; ignored if data is a GeoDataFrame.

  • traj_cols (dict, optional) – Mapping of logical names (‘user_id’, etc.) to actual column names.

  • **kwargs – Passed through to the underlying reader.

Returns:

Sampled user IDs.

Return type:

pd.Series

nomad.io.base.table_columns(filepath, format='csv', include_schema=False, sep=',')[source]

Return column names or the full schema of a data source.

The ‘sep’ argument specifies the delimiter and is only used for ‘csv’ format; it is ignored when reading ‘parquet’ files.

nomad.io.base.to_file(df, path, format='csv', traj_cols=None, output_traj_cols=None, partition_by=None, filesystem=None, use_offset=False, **kwargs)[source]
nomad.io.base.zoned_datetime_from_ts_and_offset(utc_timestamps, timezone_offset)[source]