Flowcast

Public model card

How Flowcast predicts demand.

Flowcast predicts hourly pickup demand at the NYC taxi-zone level using recent demand history, rolling behavior, calendar signals, weather, and zone context. Evaluation is time-based so the model is tested on periods that occur after the training data.

Run: Full-scale SageMaker Pipeline run.

Model facts

Model family
XGBoost · one global model
Target
Hourly yellow-taxi pickups per TLC zone
Horizons
1h / 6h / 24h (horizon is a model input)
Active model version
rsx1twqvdz9b
Training data end date
Dec 31, 2025
Evaluation data window
Mar 1, 2026 – Jul 31, 2026
NYC zone coverage
262 / 262 five-borough TLC zones

Evaluation metrics

Measured on data the model never saw.

WAPE
18.6%
Baseline 24.7%
MAE (pickups)
3.72
Baseline 4.94
RMSE (pickups)
11.52
Baseline 15.63
R²
0.958
Baseline 0.922
Baseline improvement
+24.8%
relative WAPE reduction
Accuracy by forecast horizon
HorizonFlowcast WAPEBaseline WAPEImprovementMAEZone-hours scored
1h17.9%24.7%+27.7%3.57961,540
6h19.1%24.7%+22.8%3.81961,540
24h18.9%24.7%+23.8%3.77961,540

Validation

Time moves forward. The split does too.

TRAIN

Historical period used to learn model parameters.

Jan 1, 2024 – Dec 31, 2025

13,779,890 zone-hour-horizon examples

VALIDATE

Used for tuning, feature decisions, and uncertainty calibration.

Jan 1, 2026 – Feb 28, 2026

1,112,976 zone-hour-horizon examples

TEST

Held out until the final evaluation.

Mar 1, 2026 – Jul 31, 2026

2,885,406 zone-hour-horizon examples

The test period is scored once, after training and calibration are finished. The quality gate and the published metrics both come from that single evaluation.

What the model sees

Five kinds of signal, all known at the forecast origin.

Recent demand

Lagged hourly pickup counts capture short-term momentum and weekly seasonality.

Rolling behavior

Shifted rolling averages and volatility summarize demand patterns without leaking the target hour.

Calendar

Hour, weekday, weekend/holiday, month, and cyclical time encodings describe recurring patterns.

Weather

Temperature, precipitation, wind, visibility, and pressure provide environmental context.

Location

Taxi-zone and borough context allow the model to learn persistent geographic differences.

Technical details: all 48 model inputs
Recent demand
  • lag_1
  • lag_2
  • lag_3
  • lag_6
  • lag_12
  • lag_24
  • lag_48
  • lag_168
  • trend_recent_vs_daily
  • seasonal_naive
Rolling behavior
  • rolling_mean_3
  • rolling_mean_6
  • rolling_mean_24
  • rolling_mean_168
  • rolling_std_24
  • rolling_std_168
  • trend_daily_vs_weekly
Calendar
  • horizon_hours
  • target_hour_of_day
  • target_day_of_week
  • target_month
  • target_is_weekend
  • hour_sin
  • hour_cos
  • dow_sin
  • dow_cos
  • month_sin
  • month_cos
  • target_is_holiday
Weather
  • temperature_c
  • temperature_c_at_origin
  • dewpoint_c
  • dewpoint_c_at_origin
  • wind_speed_mps
  • wind_speed_mps_at_origin
  • visibility_km
  • visibility_km_at_origin
  • precip_1h_mm
  • precip_1h_mm_at_origin
  • pressure_hpa
  • pressure_hpa_at_origin
Location
  • zone_id
  • borough_Bronx
  • borough_Brooklyn
  • borough_Manhattan
  • borough_Queens
  • borough_Staten Island
  • zone_avg_demand_train

Lags and rolling windows are computed per zone on the series shifted one hour, so no window includes the origin hour or anything after it. The origin hour’s own in-progress count is never an input. Target-hour weather in backtests is the observed value (see limitations).

What drives the model

Mean absolute SHAP contribution by input group on 2,000 held-out test rows (TreeExplainer, log scale). Longer bars move forecasts more.

  • Daily and weekly pattern
  • Rolling averages
  • Time of day and calendar
  • Zone and borough
  • Recent pickup volume
  • Weather

Baseline & quality gate

The model has to beat a real baseline.

The seasonal-naive baseline predicts that a zone will behave like the same hour one week earlier. Flowcast compares against that baseline on the same held-out data before a model version is accepted.

In the SageMaker Pipeline, a condition step registers the model in the SageMaker Model Registry only when every check below passes.

Quality gate for rsx1twqvdz9b

Passed: WAPE 18.6% vs. baseline 24.7% (24.8% better; required ≥ 5% and WAPE ≤ 35%).

  • Every prediction is a finite number
  • No negative pickup predictions
  • Time-leakage checks passed
  • Lower WAPE than the one-week baseline
  • Beats the baseline by the required margin
  • WAPE under the maximum allowed

Uncertainty

Every forecast comes with a range.

Intervals are calibrated on the validation period only, from how far real demand landed from the forecast, separately for each horizon and for low- and high-volume zones. They are multiplicative, so a quiet zone gets a narrow range and a busy one a wider range.

Observed coverage on the test period

Share of held-out zone-hours whose actual value fell inside the band. Nominal targets: 80% and 95%.

Interval coverage by horizon
Horizon80% band95% band
1h82.5%95.9%
6h82.8%96.1%
24h83.0%96.1%

The Forecast page shows the 80% band.

Limitations

What this model does not know.

Zone-level geography

TLC trip records identify taxi zones, not exact pickup GPS coordinates. Every map color is a zone total.

Data freshness

Public TLC trip data is not a real-time feed; it is published with a delay of about two months. The current data runs through Jul 31, 2026.

Weather backtests

Historical evaluation uses the observed weather at the target hour. Operational forecasting should use weather forecasts available at the origin time; the “forecast from latest data” mode carries the origin-hour observation forward instead.

Unexpected events

Events are not modeled. Concerts, parades, disruptions and sudden regime shifts can be hard to predict from recent history alone.

Sparse zones

Yellow-taxi pickups concentrate in Manhattan and at the airports. Many outer-borough zones see only a few pickups per hour, where percentage errors swing widely. 2 zones had no pickups at all in training.

Yellow taxi only

Flowcast models yellow-taxi pickups. Green taxis and app-based for-hire vehicles are not included, so it is not total ride demand.

Coverage: 262 of 262 five-borough TLC zones. Not NYC zones, so not modeled: Newark Airport (EWR, ID 1); N/A (Unknown, ID 264); Outside of NYC (N/A, ID 265).