Flowcast

AWS / MLOps

From public trip data to deployed forecasts.

Flowcast connects ingestion, feature engineering, model training, evaluation, model registration, inference, and a geospatial product in one reproducible system.

Stage 1

Data

Purpose
Collect the public inputs once and keep them unchanged, so every model version can be rebuilt from the same raw files.
Why this stage exists
A stable raw layer separates “what the city published” from every downstream decision, which keeps experiments reproducible.
Input
Monthly NYC TLC Yellow Taxi trip Parquet files (2024-01 → 2026-07, 31 months), NOAA GHCNh hourly observations for Central Park, and the official TLC taxi-zone lookup and shapefile.
Output
Immutable objects under s3://<project-bucket>/raw/{tlc,noaa,zones}/.
Service / implementation
Ingestion scripts (scripts/ingest_tlc.py, scripts/ingest_noaa.py) download from TLC’s CloudFront host and NOAA NCEI into a private Amazon S3 bucket.
Artifact or payload
Raw Parquet (pickup timestamp + pickup LocationID), pipe-delimited GHCNh station files, zone lookup CSV.
Cost behavior
S3 storage and requests only; reading public sources is free.

Request path

What happens when you pick a time.

01
User selects time + horizon
Forecast page controls
02
API request
POST /forecast via API Gateway
03
Feature assembly
Lambda loads the origin’s feature rows for every zone from S3
04
SageMaker inference
One batched call to the serverless endpoint
05
Zone predictions
Intervals applied, normalized citywide response
06
Map / dashboard
Joined to taxi-zone polygons by LocationID

When the live API is not configured or the endpoint is switched off, the Forecast page reads the same model version’s precomputed frames and says so in its status strip.