Skip to content

Individual measures

Individual measures describe each person's movement, place use, regularity, and inferred behavior. Most accept a trajectory-like dataframe and return one row, scalar, or variable-length result per user. Pass explicit column names when auto-detection does not match your schema.

Question Representative APIs
How far and how often does someone move? jump_lengths, radius_of_gyration, waiting_times
Which places matter? home_location, work_location, location_frequency
Is their movement regular or predictable? regularity, entropy, predictability, motifs
Does route distance matter? jump_lengths_road, radius_of_gyration_road

For a task-focused introduction, see Measure people and populations.

API

fastmob.measures.individual.activity_transition_matrix(visits, activity_col=None, user_id_col=None, timestamp_col=None, day_col=None, day_filter=None, unknown_label='UNKNOWN')

Compute the activity transition matrix for a visits DataFrame.

Counts how often each activity-type transition (from -> to) occurs across all users, then normalises to percentages. Returns a DataFrame with an activity label column plus one column per target activity.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with at least an activity column and, optionally, a user-ID and timestamp column.

required
activity_col str or None

Column name for the activity/purpose type. Auto-detected if None.

None
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
timestamp_col str or None

Column name for the visit timestamp used for sorting. Auto-detected if None (rows are used in their current order when no timestamp is found).

None
day_col str or None

Column name for the day-of-week string. Auto-detected if None.

None
day_filter str or None

One of None (all days), "weekdays" (Mon–Fri only), or "weekends" (Sat–Sun only). When not None, day_col must be resolvable.

None

Returns:

Type Description
DataFrame

Transition matrix as percentages, in the same backend as visits, with an activity label column plus one column per target activity.

Raises:

Type Description
ValueError

When day_filter is requested but no day-of-week column can be found.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import activity_transition_matrix
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u2", "u2", "u2"],
...         "start_timestamp": pd.to_datetime(
...             [
...                 "2020-01-01 08:00",
...                 "2020-01-01 09:00",
...                 "2020-01-01 18:00",
...                 "2020-01-01 07:30",
...                 "2020-01-01 12:00",
...                 "2020-01-01 19:00",
...             ]
...         ),
...         "purpose": ["HOME", "WORK", "HOME", "HOME", "SHOP", "HOME"],
...     }
... )
>>> result = activity_transition_matrix(visits)
>>> print(result.set_index("activity").round(1).to_string())
          HOME  SHOP  WORK
activity
HOME       0.0  25.0  25.0
SHOP      25.0   0.0   0.0
WORK      25.0   0.0   0.0

fastmob.measures.individual.compute_profiles(visits, *, user_id_col=None, location_id_col=None, start_col=None, end_col=None, purpose_col=None, n_clusters=3, random_state=0)

Compute per-user mobility-profile metrics and assign profile labels.

For each user, computes intermittency, degree of return, regularity, diversity, (normalized) entropy, and stationarity from a stay-level visit table, then clusters users into Routiner/Regular/Scouter profiles (Amichi et al. 2020) on standardized [intermittency, degree_of_return].

Parameters:

Name Type Description Default
visits Any

Stay-level table; any Narwhals-compatible eager backend. Expected columns (auto-detected when not given explicitly): a user ID, a location ID, a stay start timestamp, a stay end timestamp, and optionally a "purpose"/activity type column.

required
user_id_col str | None

Explicit column name overrides; auto-detected when None. purpose_col may be absent entirely (no activity-type column).

None
location_id_col str | None

Explicit column name overrides; auto-detected when None. purpose_col may be absent entirely (no activity-type column).

None
start_col str | None

Explicit column name overrides; auto-detected when None. purpose_col may be absent entirely (no activity-type column).

None
end_col str | None

Explicit column name overrides; auto-detected when None. purpose_col may be absent entirely (no activity-type column).

None
purpose_col str | None

Explicit column name overrides; auto-detected when None. purpose_col may be absent entirely (no activity-type column).

None
n_clusters int

Number of clusters to form; must be <= 3 to get named profiles ("routiners"/"regulars"/"scouters"), otherwise clusters are named "cluster_0", "cluster_1", ...

3
random_state int

Seed for KMeans, for reproducible cluster assignments.

0

Returns:

Type Description
DataFrame

One row per clustered user with columns [user_id_col, "intermittency", "degree_of_return", "regularity", "diversity", "entropy", "stationarity", "profile"], in the same backend as input.

Raises:

Type Description
ValueError

If fewer than n_clusters users have finite profiling metrics.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import compute_profiles
>>> rows = []
>>> patterns = {
...     "u1": ["a", "b", "c", "d", "a", "b", "c", "d"],
...     "u2": ["a", "b", "a", "c", "a", "b", "a", "d"],
...     "u3": ["a", "a", "b", "a", "a", "a", "b", "a"],
... }
>>> for uid, locations in patterns.items():
...     for i, location in enumerate(locations):
...         rows.append(
...             {
...                 "uid": uid,
...                 "start_timestamp": pd.Timestamp("2020-01-01") + pd.Timedelta(hours=i),
...                 "end_timestamp": pd.Timestamp("2020-01-01") + pd.Timedelta(hours=i, minutes=30),
...                 "location_id": location,
...             }
...         )
>>> visits = pd.DataFrame(rows)
>>> result = compute_profiles(visits, n_clusters=3, random_state=7)
>>> sorted(result["profile"].unique())
['regulars', 'routiners', 'scouters']

fastmob.measures.individual.daily_activity_distribution(visits, activity_col=None, start_time_col=None, end_time_col=None, bin_size_minutes=10, day_col=None, day_filter=None, unknown_label='UNKNOWN')

Compute a daily activity distribution matrix over fixed time bins.

fastmob.measures.individual.daily_motifs(visits, *, uid_col=None, location_col=None, purpose_col=None, datetime_col=None, end_datetime_col=None, duration_col=None)

Compute one home-anchored mobility motif per user and day.

The returned dataframe matches the input backend and contains the resolved user-ID column, date, and motif_id. Row ordering is detected automatically before selecting the native kernel.

fastmob.measures.individual.distance_straight_line

distance_straight_line(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the total trajectory length (km) for each user.

The distance straight line \(d_{SL}\) travelled by an individual \(u\) is the sum of Haversine distances between consecutive GPS fixes in the time-ordered trajectory [WTDED2015]_:

\[ d_{SL}(u) = \sum_{j=2}^{n_u} dist(r_{j-1}, r_j) \]

where \(n_u\) is the number of recorded points for \(u\), \(r_j\) is the \(j\)-th point as a \((lat, lng)\) pair, and \(dist\) is the Haversine distance between two points.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "distance_straight_line"]. Distance values are in kilometres. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import distance_straight_line
>>> result = distance_straight_line(df)
>>> print(result.round({"distance_straight_line": 3}).head().to_string(index=False))
 uid  distance_straight_line
   0              374531.472
   1              774347.886
   2               86036.423
References
  • [WTDED2015] Williams, N. E., Thomas, T. A., Dunbar, M., Eagle, N. & Dobra, A. (2015) Measures of Human Mobility Using Mobile Phone Records Enhanced with GIS Data. PLOS ONE 10(7): e0133630. https://doi.org/10.1371/journal.pone.0133630
See Also

jump_lengths : Individual jump distances between consecutive points. maximum_distance : Largest single jump length per user.

fastmob.measures.individual.diversity

Trajectory diversity measure using suffix-array entropy.

diversity(visits, user_id_col=None, location_id_col=None)

Compute trajectory diversity per user using suffix-array entropy.

Per-user: factorize the location column, then call the Rust suffix-array kernel on the integer-coded token sequence to compute the ratio of distinct substrings to total substrings.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "diversity"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import diversity
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u1", "u2", "u2", "u2"],
...         "location_id": ["home", "work", "home", "gym", "home", "shop", "home"],
...     }
... )
>>> result = diversity(visits)
>>> print(result.round(3).to_string(index=False))
user_id  diversity
     u1      0.900
     u2      0.833

fastmob.measures.individual.exploration_profiling(visits, user_id_col=None, location_id_col=None, datetime_col=None, cold_start_strategy='frequency', known_suffixes=('_HOME', '_WORK'), clustering_method='kmeans', random_seed=42, n_iterations=300, impute_gaps=False)

Compute intermittency and degree of return, then cluster users into mobility profiles.

Calls :func:intermittance_and_degree_of_return to get per-user statistics, then applies a Rust clustering kernel on the degree_of_return values to assign each user to one of three profiles:

  • routiners — high degree of return; mostly revisit familiar places.
  • regulars — balanced mix of exploration and return.
  • scouters — low degree of return; mostly explore new places.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None
datetime_col str or None

Column name for visit timestamps. Auto-detected if None. Required when cold_start_strategy="baseline".

None
cold_start_strategy str

How to initialize known places. Passed directly to :func:intermittance_and_degree_of_return.

'frequency'
known_suffixes tuple of str

Location ID suffixes treated as known (only used when cold_start_strategy="suffix").

('_HOME', '_WORK')
clustering_method str

"kmeans" for K-Means or "gmm" for Gaussian Mixture Model.

'kmeans'
random_seed int

Seed for the clustering algorithm — ensures reproducible results.

42
n_iterations int

Maximum iterations for the clustering algorithm.

300
impute_gaps bool

If True, pass through to :func:intermittance_and_degree_of_return to fill eligible missing 5-minute trajectory slices before computing return/exploration statistics.

False

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "intermittency", "degree_of_return", "mean_return", "mean_exploration", "profile"]. The "profile" column contains one of "routiners", "regulars", or "scouters". The return type matches the input backend (pandas, polars, …).

Raises:

Type Description
ValueError

If fewer than 3 users are present (cannot form 3 clusters), or if an unknown clustering_method is given.

ImportError

If the compiled fastmob._core extension is not available (run maturin develop first).

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import exploration_profiling
>>> rows = []
>>> patterns = {
...     "u1": ["a", "b", "c", "d", "a"],
...     "u2": ["a", "b", "a", "c", "a", "d"],
...     "u3": ["a", "a", "b", "a", "c", "a"],
...     "u4": ["a", "b", "c", "a", "b", "c"],
...     "u5": ["a", "b", "a", "b", "a", "b"],
... }
>>> for uid, locations in patterns.items():
...     for i, location in enumerate(locations):
...         rows.append(
...             {
...                 "user_id": uid,
...                 "start_timestamp": pd.Timestamp("2020-01-01") + pd.Timedelta(hours=i),
...                 "location_id": location,
...             }
...         )
>>> visits = pd.DataFrame(rows)
>>> result = exploration_profiling(visits, cold_start_strategy="none", random_seed=7)
>>> print(result.round(3).to_string(index=False))
user_id  intermittency  degree_of_return  mean_return  mean_exploration   profile
     u1          5.000             0.245          1.0             4.000  scouters
     u2          2.333             0.644          1.0             1.333  regulars
     u3          2.000             0.785          1.0             1.000  regulars
     u4          6.000             0.785          3.0             3.000  regulars
     u5          6.000             1.107          4.0             2.000 routiners

fastmob.measures.individual.frequency_rank

frequency_rank(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the frequency rank of each distinct location for every user.

The frequency rank K_f(r_i) of location r_i is 1 if it is the most frequently visited location, 2 if it is the second-most frequently visited, and so on. Ties in visit count are broken by the order they appear after sorting (stable), matching the skmob reference implementation.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per (user, location) pair with columns [uid_col, lat_col, lng_col, "frequency_rank"]. Rank 1 is the most frequently visited location. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import frequency_rank
>>> result = frequency_rank(df)
>>> print(result.head().to_string(index=False))
 uid       lat         lng  frequency_rank
   0 39.762146 -104.982480               1
   0 39.891077 -105.068532               2
   0 39.739154 -104.984703               3
   0 39.891586 -105.068463               4
   0 39.827022 -105.143191               5
References
  • [BDEM2015] Barbosa, H., de Lima-Neto, F. B., Evsukoff, A., Menezes, R. (2015) The effect of recency to human mobility, EPJ Data Science 4(21), https://epjdatascience.springeropen.com/articles/10.1140/epjds/s13688-015-0059-8

fastmob.measures.individual.home_location

home_location(traj, *, start_night=22, end_night=7, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the most-visited nighttime location for each user.

The home location \(h(u)\) of an individual \(u\) is the location visited most often during the nighttime window [start_night, 24) ∪ [0, end_night) [CBTDHVSB2012] [PSO2012]:

\[ h(u) = \arg\max_{i} \bigl|\{r_i \mid t(r_i) \in [t_{\text{start}}, t_{\text{end}}]\}\bigr| \]

where \(r_i\) is a location visited by \(u\), \(t(r_i)\) is the time of the visit, and \(t_{\text{start}}\) / \(t_{\text{end}}\) bound the nighttime window. When a user has no nighttime records, the most-visited location across all hours is used as a fallback.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
start_night int

Hour (0–23) at which the nighttime window begins. Default: 22.

22
end_night int

Hour (0–23) at which the nighttime window ends (exclusive). Default: 7.

7
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, lat_col, lng_col] (using the detected column names). The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import home_location
>>> result = home_location(df)
>>> print(result.round({"lat": 3, "lng": 3}).head().to_string(index=False))
 uid    lat      lng
   0 39.891 -105.069
   1 37.630 -122.411
   2 39.739 -104.985
References
  • [CBTDHVSB2012] Csaji, B. C., Browet, A., Traag, V. A., Delvenne, J.-C., Huens, E., Van Dooren, P., Smoreda, Z. & Blondel, V. D. (2012) Exploring the Mobility of Mobile Phone Users. Physica A: Statistical Mechanics and its Applications 392(6), 1459-1473, https://www.sciencedirect.com/science/article/pii/S0378437112010059
  • [PSO2012] Phithakkitnukoon, S., Smoreda, Z. & Olivier, P. (2012) Socio-geography of human mobility: A study using longitudinal mobile phone data. PLOS ONE 7(6): e39253. https://doi.org/10.1371/journal.pone.0039253
See Also

max_distance_from_home : Maximum distance from the inferred home location.

fastmob.measures.individual.individual_mobility_network

individual_mobility_network(traj, *, self_loops=False, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the individual mobility network as a directed edge-list DataFrame.

An Individual Mobility Network (IMN) of an individual \(u\) is a directed weighted graph \(G_u = (V, E)\) where \(V\) is the set of distinct visited locations and \(E\) is the set of directed trips between locations [RGNPPG2014] [BL2012] [SQBB2010]_. The edge weight function

\[ \omega: E \to \mathbb{N} \]

returns the number of times \(u\) travelled that edge.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
self_loops bool

When False (default), consecutive visits to the same location are skipped so no self-loop edges are created. When True, a step that stays at the same location contributes to a (loc, loc) edge.

False
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per directed edge with columns [uid_col, "lat_origin", "lng_origin", "lat_dest", "lng_dest", "n_trips"]. When uid_col is None the uid column is omitted. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import individual_mobility_network
>>> result = individual_mobility_network(df)
>>> print(result.round({"lat_origin": 3, "lng_origin": 3, "lat_dest": 3, "lng_dest": 3}).head().to_string(index=False))
 uid  lat_origin  lng_origin  lat_dest  lng_dest  n_trips
   0      37.775    -122.419    37.601  -122.382        1
   0      37.601    -122.382    37.615  -122.390        1
   0      37.615    -122.390    39.879  -104.682        1
   0      39.879    -104.682    39.739  -104.985        1
   0      39.739    -104.985    39.762  -104.982       19
References
  • [RGNPPG2014] Rinzivillo, S., Gabrielli, L., Nanni, M., Pappalardo, L., Pedreschi, D. & Giannotti, F. (2012) The purpose of motion: Learning activities from Individual Mobility Networks. Proceedings of the 2014 IEEE International Conference on Data Science and Advanced Analytics, 312-318, https://ieeexplore.ieee.org/document/7058090
  • [BL2012] Bagrow, J. P. & Lin, Y.-R. (2012) Mesoscopic Structure and Social Aspects of Human Mobility. PLOS ONE 7(5): e37676. https://doi.org/10.1371/journal.pone.0037676
  • [SQBB2010] Song, C., Qu, Z., Blumm, N. & Barabasi, A. L. (2010) Limits of Predictability in Human Mobility. Science 327(5968), 1018-1021, https://science.sciencemag.org/content/327/5968/1018

fastmob.measures.individual.intermittance_and_degree_of_return(visits, user_id_col=None, location_id_col=None, datetime_col=None, cold_start_strategy='frequency', known_suffixes=('_HOME', '_WORK'), use_trajectory=True, impute_gaps=False)

Compute intermittancy and degree of return per user using vectorized operations.

For each user, partitions the visit sequence into alternating blocks of explorations (new places) and returns (revisits or home/work visits). Then computes summary statistics over those blocks. By default, stay intervals are reconstructed into 5-minute trajectory slices when start and end timestamp columns are available.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None
datetime_col str or None

Column name for visit timestamps. Auto-detected if None. Required when cold_start_strategy="baseline" (used to count distinct active days). Also used as the stay start timestamp when use_trajectory=True.

None
cold_start_strategy str

How to initialize "known" places.

  • "frequency" — locations visited at or above mean frequency (×0.8).
  • "baseline" — Papandrea et al. Algorithm 1: k-means (k=3) on the per-location relevance score R_u = d_visit / d_total; MVP cluster (most frequently visited) is pre-labelled as known. Requires a datetime column.
  • "max_frequency" — top 10 % by visit count.
  • "suffix" — locations whose ID ends with any of known_suffixes.
  • "none" — no cold-start; every first visit is an exploration.
'frequency'
known_suffixes tuple of str

Used only when cold_start_strategy="suffix".

('_HOME', '_WORK')
use_trajectory bool

If True, and both start and end timestamp columns are available, expand each stay into observed 5-minute slices from ceil(start, "5min") to floor(end, "5min"). Duplicate (user, timestamp) slices keep the first input occurrence. If no end timestamp column is available, the function falls back to treating each input row as one sequence event. If False, each input row is always one sequence event.

True
impute_gaps bool

If True, and trajectory reconstruction is possible, fill missing 5-minute slices between each user's first and last observed slice using per-user anchors inferred from observed locations: hours 2-5 use the most frequent nighttime location, hour 10 uses the most frequent location at 10, and hours 14-16 use the most frequent afternoon location. Missing slices outside those windows, or inside a window with no inferred anchor, remain absent. Ignored when use_trajectory=False or when no end timestamp column is available.

False

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "intermittency", "degree_of_return", "mean_return", "mean_exploration"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import intermittance_and_degree_of_return
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u1", "u2", "u2", "u2"],
...         "start_timestamp": pd.to_datetime(
...             [
...                 "2020-01-01 08:00",
...                 "2020-01-01 09:00",
...                 "2020-01-01 18:00",
...                 "2020-01-02 08:00",
...                 "2020-01-01 07:30",
...                 "2020-01-01 12:00",
...                 "2020-01-01 19:00",
...             ]
...         ),
...         "location_id": ["home", "work", "home", "gym", "home", "shop", "home"],
...     }
... )
>>> result = intermittance_and_degree_of_return(visits, cold_start_strategy="none")
>>> print(result.round(3).to_string(index=False))
user_id  intermittency  degree_of_return  mean_return  mean_exploration
     u1            2.5             0.588          1.0               1.5
     u2            3.0             0.464          1.0               2.0

fastmob.measures.individual.jump_lengths

jump_lengths(traj, merge=False, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Compute jump lengths (km) for each user in the trajectory.

A jump length is the Haversine distance (in km) between consecutive GPS fixes for the same user, sorted by datetime.

Input that is already grouped by user and ordered by datetime within each user is detected automatically and computed on a contiguous fast path, roughly halving the work; there is nothing to opt into.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory data; any Narwhals-compatible eager dataframe (pandas, polars, …). Must have columns for datetime, latitude, and longitude. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
merge bool

When True, return flat jump lengths across all users using a backend-appropriate array object. When False (default), return a per-user dataframe.

False
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame or array - like

When merge=False: a dataframe with columns [uid_col, "jump_lengths"] where each row holds an array-like sequence of jump lengths for one user. The returned backend matches the input backend. When merge=True: flat jump lengths as a NumPy array for NumPy-backed inputs or a PyArrow array for Arrow-backed inputs.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import jump_lengths
>>> result = jump_lengths(df)
>>> preview = result.assign(n_jumps=result["jump_lengths"].str.len())
>>> print(preview[["uid", "n_jumps"]].head().to_string(index=False))
 uid  n_jumps
   0     2098
   1     1209
   2     1690
References
  • [BHG2006] Brockmann, D., Hufnagel, L. & Geisel, T. (2006) The scaling laws of human travel. Nature 439, 462-465, https://www.nature.com/articles/nature04292
  • [GHB2008] Gonzalez, M. C., Hidalgo, C. A. & Barabasi, A. L. (2008) Understanding individual human mobility patterns. Nature, 453, 779-782, https://www.nature.com/articles/nature06958.
  • [PRQPG2013] Pappalardo, L., Rinzivillo, S., Qu, Z., Pedreschi, D. & Giannotti, F. (2013) Understanding the patterns of car travel. European Physics Journal Special Topics 215(1), 61-73, https://link.springer.com/article/10.1140%2Fepjst%2Fe2013-01715-5
See Also

maximum_distance : Largest single jump length per user. distance_straight_line : Sum of all jump lengths per user.

fastmob.measures.individual.jump_lengths_road(traj, *, network, uid_col=None, lat_col=None, lng_col=None, datetime_col=None, snap_max_distance_m=750.0)

Road-network jump lengths (km): distance between consecutive stops for the same user, sorted by datetime -- mirrors :func:~fastmob.measures.individual.jump_lengths.jump_lengths's sort key and its inclusion of zero-length jumps, but measures along the network instead of straight-line, falling back to Haversine per-pair when unsnapped or disconnected.

Parameters:

Name Type Description Default
traj Any

Trajectory dataframe; any Narwhals-compatible eager backend.

required
network RoadNetwork

A prepared :class:fastmob.network.RoadNetwork.

required
uid_col str | None

Explicit column name overrides; auto-detected when None.

None
lat_col str | None

Explicit column name overrides; auto-detected when None.

None
lng_col str | None

Explicit column name overrides; auto-detected when None.

None
datetime_col str | None

Explicit column name overrides; auto-detected when None.

None
snap_max_distance_m float

Maximum distance (metres) to snap a stop to the network; farther stops fall back to Haversine entirely for any jump touching them.

750.0

Returns:

Type Description
ndarray

One value per consecutive same-user pair (length: len(traj) - n_users).

fastmob.measures.individual.k_radius_of_gyration

k_radius_of_gyration(traj, k=2, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Compute the k-radius of gyration (km) for each user in the trajectory.

The k-radius of gyration is the radius of gyration computed using only the k most-visited locations for each user. Formally:

\[ r_g^{(k)}(u) = \sqrt{ \frac{\sum_{i \in \mathrm{top}\text{-}k} w_i d_\mathrm{haversine}(r_i, r_\mathrm{cm})^2} {\sum_{i \in \mathrm{top}\text{-}k} w_i} } \]

where \(r_\mathrm{cm}\) is the weighted center of mass over the top-k locations and \(w_i\) is the visit count of location \(i\).

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory data; any Narwhals-compatible eager dataframe (pandas, polars, ...). Must have columns for datetime, latitude, and longitude. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
k int

Number of most-visited locations to consider. Defaults to 2. Valid range: \([2, +\infty)\).

2
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "k_radius_of_gyration"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import k_radius_of_gyration
>>> result = k_radius_of_gyration(df, k=2)
>>> print(result.round({"k_radius_of_gyration": 3}).head().to_string(index=False))
 uid  k_radius_of_gyration
   0                 7.859
   1                 4.069
   2                 5.794
References
  • [PSRPGB2015] Pappalardo, L., Simini, F. Rinzivillo, S., Pedreschi, D. Giannotti, F. & Barabasi, A. L. (2015) Returners and Explorers dichotomy in human mobility. Nature Communications 6, https://www.nature.com/articles/ncomms9166
See Also

radius_of_gyration : Radius of gyration over all visited locations.

fastmob.measures.individual.location_frequency

location_frequency(traj, normalize=True, as_ranks=False, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return visit frequency for each distinct location per user.

The visitation frequency \(f(r_i)\) of location \(r_i\) for individual \(u\) is the probability of visiting that location [SKWB2010] [PF2018]:

\[ f(r_i) = \frac{n(r_i)}{n_u} \]

where \(n(r_i)\) is the number of visits to location \(r_i\) by \(u\), and \(n_u\) is the total number of data points in \(u\)'s trajectory. When normalize=False, raw visit counts \(n(r_i)\) are returned instead.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
normalize bool

When True (default), the "location_frequency" column contains the probability of visiting that location (count / total_visits), so each user's frequencies sum to 1.0. When False, raw visit counts are returned.

True
as_ranks bool

When True, return a Python list where element i is the mean visit frequency of the i-th most-visited location across all users (rank-1 = most visited). The list length equals the maximum number of distinct locations visited by any single user. When False (default), return a DataFrame.

False
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame or list

When as_ranks=False: one row per (user, location) pair with columns [uid_col, lat_col, lng_col, "location_frequency"]. When as_ranks=True: a flat Python list[float] of mean per-rank frequencies. The returned DataFrame backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import location_frequency
>>> result = location_frequency(df)
>>> print(result.round({"location_frequency": 3}).head().to_string(index=False))
 uid       lat         lng  location_frequency
   0 39.762146 -104.982480               0.102
   0 39.891077 -105.068532               0.065
   0 39.739154 -104.984703               0.060
   0 39.891586 -105.068463               0.034
   0 39.827022 -105.143191               0.025
References
  • [SKWB2010] Song, C., Koren, T., Wang, P. & Barabasi, A.L. (2010) Modelling the scaling properties of human mobility. Nature Physics 6, 818-823, https://www.nature.com/articles/nphys1760
  • [PF2018] Pappalardo, L. & Simini, F. (2018) Data-driven generation of spatio-temporal routines in human mobility. Data Mining and Knowledge Discovery 32, 787-829, https://link.springer.com/article/10.1007/s10618-017-0548-4
See Also

frequency_rank : Rank locations by visit frequency (1 = most visited). visits_per_location : Total visits per location across all users (collective measure).

fastmob.measures.individual.max_distance_from_home

max_distance_from_home(traj, *, start_night=22, end_night=7, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the maximum Haversine distance (km) from each user's home location.

The maximum distance from home \(dh_{max}(u)\) of an individual \(u\) is defined as [CM2015]_:

\[ dh_{max}(u) = \max_{1 \leq i \leq n_u} dist(r_i, h(u)) \]

where \(n_u\) is the number of recorded points for \(u\), \(r_i\) is a location as a \((lat, lng)\) pair, \(h(u)\) is the home location of \(u\) (see :func:home_location), and \(dist\) is the Haversine distance.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
start_night int

Hour (0–23) at which the nighttime window begins. Forwarded to :func:home_location. Default: 22.

22
end_night int

Hour (0–23) at which the nighttime window ends (exclusive). Forwarded to :func:home_location. Default: 7.

7
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "max_distance_from_home"]. Distance values are in kilometres. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import max_distance_from_home
>>> result = max_distance_from_home(df)
>>> print(result.round({"max_distance_from_home": 3}).head().to_string(index=False))
 uid  max_distance_from_home
   0               11286.959
   1               12800.565
   2               11282.764
References
  • [CM2015] Canzian, L. & Musolesi, M. (2015) Trajectories of depression: unobtrusive monitoring of depressive states by means of smartphone mobility traces analysis. Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing, 1293-1304, https://dl.acm.org/citation.cfm?id=2805845
See Also

home_location : Inferred home location from nighttime visits. maximum_distance : Largest single jump length per user.

fastmob.measures.individual.maximum_distance

maximum_distance(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the maximum distance (km) covered in a single movement for each user.

The maximum distance \(d_{max}\) travelled by an individual \(u\) is the largest Haversine distance between any two consecutive GPS fixes in the time-ordered trajectory [WTDED2015] [LBH2012]:

\[ d_{max}(u) = \max_{1 \leq i < n_u} dist(r_i, r_{i+1}) \]

where \(n_u\) is the number of recorded points for \(u\), \(r_i\) and \(r_{i+1}\) are two consecutive points as \((lat, lng)\) pairs, and \(dist\) is the Haversine distance.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "maximum_distance"]. Distance values are in kilometres. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import maximum_distance
>>> result = maximum_distance(df)
>>> print(result.round({"maximum_distance": 3}).head().to_string(index=False))
 uid  maximum_distance
   0         11294.452
   1         12804.913
   2         11286.761
References
  • [WTDED2015] Williams, N. E., Thomas, T. A., Dunbar, M., Eagle, N. & Dobra, A. (2015) Measures of Human Mobility Using Mobile Phone Records Enhanced with GIS Data. PLOS ONE 10(7): e0133630. https://doi.org/10.1371/journal.pone.0133630
  • [LBH2012] Lu, X., Bengtsson, L. & Holme, P. (2012) Predictability of population displacement after the 2010 haiti earthquake. Proceedings of the National Academy of Sciences 109 (29) 11576-11581; https://doi.org/10.1073/pnas.1203882109
See Also

jump_lengths : All jump distances between consecutive points. distance_straight_line : Sum of all jump lengths per user.

fastmob.measures.individual.motif_distribution(daily, motif_id_col='motif_id')

Aggregate daily motif IDs into count and percentage columns.

fastmob.measures.individual.number_of_locations

number_of_locations(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the number of distinct locations visited by each user.

A distinct location is a unique exact (lat, lng) pair — matching the skmob convention of float equality without spatial clustering.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "number_of_locations"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import number_of_locations
>>> result = number_of_locations(df)
>>> print(result.head().to_string(index=False))
 uid  number_of_locations
   0                  542
   1                   97
   2                  427
References
  • [GHB2008] Gonzalez, M. C., Hidalgo, C. A. & Barabasi, A. L. (2008) Understanding individual human mobility patterns. Nature, 453, 779-782, https://www.nature.com/articles/nature06958.

fastmob.measures.individual.number_of_visits

number_of_visits(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the total number of trajectory points (visits) for each user.

A "visit" is defined as one row in the trajectory dataframe after null removal. The result is the row count per user — identical to the skmob number_of_visits individual measure.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "number_of_visits"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import number_of_visits
>>> result = number_of_visits(df)
>>> print(result.head().to_string(index=False))
 uid  number_of_visits
   0              2099
   1              1210
   2              1691

fastmob.measures.individual.radius_of_gyration

radius_of_gyration(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Compute the radius of gyration (km) for each user in the trajectory.

The radius of gyration captures how far a user typically roams from their center of mass. Formally:

\[ r_g(u) = \sqrt{\frac{1}{n_u} \sum_i d_\mathrm{haversine}(r_i, r_\mathrm{cm})^2} \]

where \(r_\mathrm{cm}\) is the arithmetic mean of the user's lat/lng coordinates.

Radius of gyration is order-independent, so chronological sorting is not required for correctness; grouping rows by user is what matters. Input that already groups each user's rows together is detected automatically and computed on a contiguous fast path, roughly halving the work; there is nothing to opt into.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory data; any Narwhals-compatible eager dataframe (pandas, polars, …). Must have columns for datetime, latitude, and longitude. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "radius_of_gyration"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import radius_of_gyration
>>> result = radius_of_gyration(df)
>>> print(result.round({"radius_of_gyration": 3}).head().to_string(index=False))
 uid  radius_of_gyration
   0            1564.439
   1            2467.777
   2            1600.452
References
  • [GHB2008] Gonzalez, M. C., Hidalgo, C. A. & Barabasi, A. L. (2008) Understanding individual human mobility patterns. Nature, 453, 779-782, https://www.nature.com/articles/nature06958.
  • [PRQPG2013] Pappalardo, L., Rinzivillo, S., Qu, Z., Pedreschi, D. & Giannotti, F. (2013) Understanding the patterns of car travel. European Physics Journal Special Topics 215(1), 61-73, https://link.springer.com/article/10.1140%2Fepjst%2Fe2013-01715-5
See Also

k_radius_of_gyration : Radius of gyration restricted to the k most-visited locations.

fastmob.measures.individual.radius_of_gyration_road(traj, *, network, uid_col=None, lat_col=None, lng_col=None, snap_max_distance_m=750.0)

Road-network radius of gyration (km) per user: RMS network distance from each of a user's stops to the arithmetic-mean centroid of their stops -- mirrors the unweighted-centroid formula r_g(u) = sqrt(mean(d(r_i, r_cm)^2)) used by :func:~fastmob.measures.individual.radius_of_gyration.radius_of_gyration, but measures d along the network instead of straight-line.

Parameters:

Name Type Description Default
traj Any

Trajectory dataframe; any Narwhals-compatible eager backend.

required
network RoadNetwork

A prepared :class:fastmob.network.RoadNetwork.

required
uid_col str | None

Explicit column name overrides; auto-detected when None.

None
lat_col str | None

Explicit column name overrides; auto-detected when None.

None
lng_col str | None

Explicit column name overrides; auto-detected when None.

None
snap_max_distance_m float

Maximum distance (metres) to snap a stop to the network.

750.0

Returns:

Type Description
DataFrame

[uid_col, "radius_of_gyration"], one row per user, in the same backend as input.

fastmob.measures.individual.random_entropy

random_entropy(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the random entropy of mobility for each user.

Random entropy is defined as \(\log_2(n)\), where \(n\) is the number of distinct locations visited by the user. A location is a unique exact (lat, lng) pair — matching the skmob convention (float equality, no spatial clustering).

This is the maximum possible entropy for a user who visits \(n\) distinct places, assuming all locations are equally likely.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "random_entropy"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import random_entropy
>>> result = random_entropy(df)
>>> print(result.round({"random_entropy": 3}).head().to_string(index=False))
 uid  random_entropy
   0           9.082
   1           6.600
   2           8.738
References
  • [EP2009] Eagle, N. & Pentland, A. S. (2009) Eigenbehaviors: identifying structure in routine. Behavioral Ecology and Sociobiology 63(7), 1057-1066, https://link.springer.com/article/10.1007/s00265-009-0830-6
  • [SQBB2010] Song, C., Qu, Z., Blumm, N. & Barabasi, A. L. (2010) Limits of Predictability in Human Mobility. Science 327(5968), 1018-1021, https://science.sciencemag.org/content/327/5968/1018
See Also

uncorrelated_entropy : Entropy weighted by visit frequency (ignores temporal order). real_entropy : Entropy that captures temporal order and frequency of visits.

fastmob.measures.individual.real_entropy

Real entropy of individual mobility trajectories (Kontoyiannis estimator).

real_entropy(traj, *, datetime_col=None, location_id_col=None, uid_col=None)

Return the real (true) entropy of mobility for each user.

Real entropy is estimated using the Kontoyiannis (1998) Lempel-Ziv entropy rate estimator applied to the sequence of visited locations. Each location is represented by its exact location_id value.

The estimator captures both the frequency and the order of visits, unlike random entropy (which ignores order) and uncorrelated entropy (which ignores temporal correlations).

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime and location-ID columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
location_id_col str or None

Explicit location-ID column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "real_entropy"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import real_entropy
>>> result = real_entropy(df)
>>> print(result.round({"real_entropy": 3}).head().to_string(index=False))
 uid  real_entropy
   0         4.905
   1         2.200
   2         4.683
References
  • [SQBB2010] Song, C., Qu, Z., Blumm, N. & Barabasi, A. L. (2010) Limits of Predictability in Human Mobility. Science 327(5968), 1018-1021, https://science.sciencemag.org/content/327/5968/1018
See Also

random_entropy : Maximum possible entropy assuming uniform visitation. uncorrelated_entropy : Entropy weighted by visit frequency (ignores temporal order).

fastmob.measures.individual.recency_rank

recency_rank(traj, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the recency rank of each distinct location for every user.

The recency rank K_s(r_i) of location r_i is 1 if it is the most recently visited location, 2 if it is the second-most recently visited, and so on. Ties (multiple visits to the same (lat, lng) pair) are resolved by keeping only the latest visit for each location before ranking.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per (user, location) pair with columns [uid_col, lat_col, lng_col, "recency_rank"]. Rank 1 is the most recently visited location. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import recency_rank
>>> result = recency_rank(df)
>>> print(result.head().to_string(index=False))
 uid       lat         lng  recency_rank
   0 39.891383 -105.070814             1
   0 39.891077 -105.068532             2
   0 39.750469 -104.999073             3
   0 39.752713 -104.996337             4
   0 39.752508 -104.996637             5
References
  • [BDEM2015] Barbosa, H., de Lima-Neto, F. B., Evsukoff, A., Menezes, R. (2015) The effect of recency to human mobility, EPJ Data Science 4(21), https://epjdatascience.springeropen.com/articles/10.1140/epjds/s13688-015-0059-8

fastmob.measures.individual.regularity

Regularity measure for visit trajectories.

regularity(visits, user_id_col=None, location_id_col=None, location_type_col=None)

Compute regularity per user.

Regularity measures how repetitively a user visits the same places. It is defined as:

regularity = 1 - (unique_locations / total_visits)

where unique_locations is the count of distinct (location_id, location_type) pairs (or just distinct location_id values when location_type_col is None), and total_visits is the row count for the user.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None
location_type_col str or None

Column name for the location type / activity purpose. Auto-detected if None; set explicitly to None to disable.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "regularity"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import regularity
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u1", "u2", "u2", "u2"],
...         "location_id": ["home", "work", "home", "gym", "home", "shop", "home"],
...     }
... )
>>> result = regularity(visits)
>>> print(result.round(3).to_string(index=False))
user_id  regularity
     u1       0.250
     u2       0.333

fastmob.measures.individual.trajectory_entropy(visits, user_id_col=None, location_id_col=None, timestamp_col=None, normalized=True)

Compute Kontoyiannis entropy of mobility trajectories per user.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None
timestamp_col str or None

Column name for ordering visits. Auto-detected if None; when no timestamp column is found the row order is preserved.

None
normalized bool

When True (default), divide raw entropy by log2(n) and clip to [0, 1]. When False, return raw Kontoyiannis bits.

True

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "entropy"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import trajectory_entropy
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u1", "u2", "u2", "u2"],
...         "start_timestamp": pd.to_datetime(
...             [
...                 "2020-01-01 08:00",
...                 "2020-01-01 09:00",
...                 "2020-01-01 18:00",
...                 "2020-01-02 08:00",
...                 "2020-01-01 07:30",
...                 "2020-01-01 12:00",
...                 "2020-01-01 19:00",
...             ]
...         ),
...         "location_id": ["home", "work", "home", "gym", "home", "shop", "home"],
...     }
... )
>>> result = trajectory_entropy(visits)
>>> print(result.round(3).to_string(index=False))
user_id  entropy
     u1      0.8
     u2      1.0

fastmob.measures.individual.trajectory_predictability(visits, user_id_col=None, location_id_col=None, timestamp_col=None)

Compute per-user maximum predictability via Fano's inequality.

Follows the Song et al. (2010) approach: estimate real entropy using Kontoyiannis (1998), then solve Fano's inequality for the maximum predictability upper bound.

Parameters:

Name Type Description Default
visits DataFrame - like

A DataFrame (any Narwhals-compatible backend) with visit rows.

required
user_id_col str or None

Column name for the user ID. Auto-detected if None.

None
location_id_col str or None

Column name for the location ID. Auto-detected if None.

None
timestamp_col str or None

Column name for ordering visits. Auto-detected if None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [user_id_col, "real_entropy", "predictability", "n_unique_locations", "n_steps"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> from fastmob.measures.individual import trajectory_predictability
>>> visits = pd.DataFrame(
...     {
...         "user_id": ["u1", "u1", "u1", "u1", "u2", "u2", "u2"],
...         "start_timestamp": pd.to_datetime(
...             [
...                 "2020-01-01 08:00",
...                 "2020-01-01 09:00",
...                 "2020-01-01 18:00",
...                 "2020-01-02 08:00",
...                 "2020-01-01 07:30",
...                 "2020-01-01 12:00",
...                 "2020-01-01 19:00",
...             ]
...         ),
...         "location_id": ["home", "work", "home", "gym", "home", "shop", "home"],
...     }
... )
>>> result = trajectory_predictability(visits)
>>> print(result.round(3).to_string(index=False))
user_id  real_entropy  predictability  n_unique_locations  n_steps
     u1         1.600           0.333                   3        4
     u2         1.585           0.500                   2        3

fastmob.measures.individual.uncorrelated_entropy

uncorrelated_entropy(traj, *, normalize=False, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the uncorrelated entropy of mobility for each user.

Uncorrelated entropy is the Shannon entropy over the distribution of visit probabilities across distinct locations:

\[ S_\mathrm{unc} = -\sum_i p_i \log_2(p_i) \]

where \(p_i\) is the fraction of visits to location \(i\) out of the user's total visits. A location is a unique exact (lat, lng) pair.

When normalize=True the result is divided by \(\log_2(n)\) (the random entropy) so the output lies in [0, 1]. If the user visits only one distinct location the entropy is 0; dividing by 0 is avoided by returning 0 directly.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
normalize bool

When True, divide the Shannon entropy by \(\log_2(n_\mathrm{distinct})\) to normalise into [0, 1]. Default False.

False
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame

One row per user with columns [uid_col, "uncorrelated_entropy"]. The returned backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import uncorrelated_entropy
>>> result = uncorrelated_entropy(df)
>>> print(result.round({"uncorrelated_entropy": 3}).head().to_string(index=False))
 uid  uncorrelated_entropy
   0                 7.442
   1                 3.650
   2                 6.908
References
  • [EP2009] Eagle, N. & Pentland, A. S. (2009) Eigenbehaviors: identifying structure in routine. Behavioral Ecology and Sociobiology 63(7), 1057-1066, https://link.springer.com/article/10.1007/s00265-009-0830-6
  • [SQBB2010] Song, C., Qu, Z., Blumm, N. & Barabasi, A. L. (2010) Limits of Predictability in Human Mobility. Science 327(5968), 1018-1021, https://science.sciencemag.org/content/327/5968/1018
  • [PVGSPG2016] Pappalardo, L., Vanhoof, M., Gabrielli, L., Smoreda, Z., Pedreschi, D. & Giannotti, F. (2016) An analytical framework to nowcast well-being using mobile phone data. International Journal of Data Science and Analytics 2(75), 75-92, https://link.springer.com/article/10.1007/s41060-016-0013-2
See Also

random_entropy : Maximum possible entropy assuming uniform visitation. real_entropy : Entropy that captures temporal order and frequency of visits.

fastmob.measures.individual.visit_purpose_distribution(visits, activity_col=None, normalize=True, day_col=None, day_filter=None, unknown_label='UNKNOWN')

Compute visit-purpose counts and percentages.

Missing activity columns and null activity values are recoverable: the affected visits are labelled with unknown_label and a warning is emitted.

fastmob.measures.individual.waiting_times

waiting_times(traj, merge=False, *, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return the waiting times (seconds) between consecutive GPS fixes for each user.

A waiting time (or inter-time) \(\Delta t\) is the elapsed time between two consecutive trajectory points of individual \(u\) [SKWB2010] [PF2018]:

\[ \Delta t = |t(r_i) - t(r_{i+1})| \]

where \(r_i\) and \(r_{i+1}\) are two consecutive points in the time-ordered trajectory and \(t(r)\) is the time when \(u\) visited point \(r\). Users with fewer than 2 points receive an empty list.

Parameters:

Name Type Description Default
traj DataFrame - like

Trajectory dataframe; any Narwhals-compatible eager backend (pandas, polars, …). Must have datetime, latitude, and longitude columns. A user-ID column is optional; when absent the whole frame is treated as a single individual.

required
merge bool

When True, return flat waiting times across all users using a backend-appropriate array object. When False (default), return a DataFrame with one row per user.

False
datetime_col str or None

Explicit datetime column name. Auto-detected when None.

None
lat_col str or None

Explicit latitude column name. Auto-detected when None.

None
lng_col str or None

Explicit longitude column name. Auto-detected when None.

None
uid_col str or None

Explicit user-ID column name. Auto-detected when None.

None

Returns:

Type Description
DataFrame or DataFrame or array - like

When merge=False: one row per user with columns [uid_col, "waiting_times"]; each cell is an array-like sequence of floats (seconds). When merge=True: flat waiting times as a NumPy array for NumPy-backed inputs or a PyArrow array for Arrow-backed inputs. The returned DataFrame backend matches the input backend.

Examples:

>>> import pandas as pd
>>> import fastmob
>>> url = fastmob.data.BRIGHTKITE_SAMPLE
>>> df = pd.read_csv(
...     url,
...     sep="\t",
...     header=0,
...     nrows=5000,
...     names=["uid", "datetime", "lat", "lng", "location id"],
... )
>>> df["datetime"] = pd.to_datetime(df["datetime"], errors="coerce")
>>> df["location_id"] = df["location id"].astype("string")
>>> df = df.dropna(subset=["uid", "datetime", "lat", "lng"])[
...     ["uid", "datetime", "lat", "lng", "location_id"]
... ]
>>> print(df.head().to_string(index=False))
 uid                  datetime       lat         lng                              location_id
   0 2010-10-16 06:02:04+00:00 39.891383 -105.070814         7a0f88982aa015062b95e3b4843f9ca2
   0 2010-10-16 03:48:54+00:00 39.891077 -105.068532         dd7cd3d264c2d063832db506fba8bf79
   0 2010-10-14 18:25:51+00:00 39.750469 -104.999073 9848afcc62e500a01cf6fbf24b797732f8963683
   0 2010-10-14 00:21:47+00:00 39.752713 -104.996337         2ef143e12038c870038df53e0478cefc
   0 2010-10-13 23:31:51+00:00 39.752508 -104.996637         424eb3dd143292f9e013efa00486c907
>>> from fastmob import waiting_times
>>> result = waiting_times(df)
>>> preview = result.assign(n_waiting_times=result["waiting_times"].str.len())
>>> print(preview[["uid", "n_waiting_times"]].head().to_string(index=False))
 uid  n_waiting_times
   0             2098
   1             1209
   2             1690
References
  • [SKWB2010] Song, C., Koren, T., Wang, P. & Barabasi, A.L. (2010) Modelling the scaling properties of human mobility. Nature Physics 6, 818-823, https://www.nature.com/articles/nphys1760
  • [PF2018] Pappalardo, L. & Simini, F. (2018) Data-driven generation of spatio-temporal routines in human mobility. Data Mining and Knowledge Discovery 32, 787-829, https://link.springer.com/article/10.1007/s10618-017-0548-4

fastmob.measures.individual.work_location

Work-location inference from weekday daytime observations.

work_location(traj, *, start_work=8, end_work=18, datetime_col=None, lat_col=None, lng_col=None, uid_col=None)

Return each user's most-visited weekday daytime location.

Uses Monday--Friday observations whose local hour lies in [start_work, end_work). Users without an observation in that window are omitted: unlike home inference, treating arbitrary visits as work is not a useful fallback.