Flowcast

Model performance

See where the model is right, wrong, and improving.

Replay held-out periods, compare Flowcast predictions with observed demand and the one-week seasonal baseline, and inspect error by horizon, borough, time, and taxi zone.

How to read this dashboard

Everything here is held-out test data: hours the model never saw while training or tuning. The Forecast page answers “what does the model expect?”; this page answers “how well has it done?”.

Actual
Observed yellow-taxi pickups from the public TLC trip records.
Predicted
What Flowcast forecast for the same zone and hour, made from information available at the forecast origin.
Baseline
The seasonal-naive forecast: the same zone and hour one week earlier. Flowcast has to beat it.
WAPE
Total absolute error ÷ total actual pickups. 15% means misses add up to 15% of real demand.
MAE / RMSE
Average miss per zone-hour, in pickups. RMSE weighs large misses more heavily.
Residual
Actual minus predicted. Positive = Flowcast underpredicted; negative = overpredicted.

Use the filter bar to switch between the full test period and a replay of its last seven days, pick a horizon, and narrow to a borough or a single zone. Every chart, card and table updates together.

Evaluation window
Horizon
  • Full held-out test period
  • 1h horizon
  • Metric: WAPE
  • Held-out test period
Actual pickups

Observed TLC pickups in the selected frame/filter.

Predicted pickups

Flowcast’s predicted pickups for the same frame.

WAPE

Total absolute error ÷ total actual demand. Lower is better.

MAE

Average miss per zone-hour, in pickups. Lower is better.

RMSE

Like MAE, but large misses count more. Lower is better.

Baseline improvement

Error reduction vs. same zone and hour one week earlier. Positive is better.

Backtest frame

Observed and predicted demand for one held-out timestamp. Scrub or play through the last 7 days of the test period; click a zone to filter the dashboard to it.

Citywide demand over time

Observed pickups vs. Flowcast predictions and the one-week seasonal baseline (daily totals, 1h horizon).

Loading evaluation artifacts…

Error by forecast horizon

Lower values indicate more accurate forecasts. Compare 1h, 6h, and 24h on the same evaluation window (WAPE, All boroughs).

Loading evaluation artifacts…

Performance by borough

Forecast error across the five boroughs for the current filters (WAPE, 1h).

Loading evaluation artifacts…

Yellow-taxi pickups are concentrated in Manhattan and at the airports; outer-borough zones are sparse, so a few pickups can swing their percentage error.

Residual distribution

Residual = actual minus predicted. Values near zero indicate smaller errors; positive values mean Flowcast underpredicted and negative values mean it overpredicted.

Loading evaluation artifacts…

Error by hour of day

WAPE by the target hour’s time of day. Shows when in the day the model and the baseline miss most.

Loading evaluation artifacts…

Zone performance

Every NYC taxi zone in the current filters, 1h horizon, full held-out test period. Sort any column; select a zone to focus the dashboard and map on it.

Largest forecast misses

Zone-hours with the largest absolute prediction errors in the selected window.

Loading evaluation artifacts…

Big misses cluster around unusual hours (events, weather, holidays) that recent history and the weekly pattern do not anticipate.