Skip to content

Repository files navigation

forge-pdm-mlops

forge-pdm-mlops

A predictive-maintenance ML pipeline with the production spine closed end to end — train → registry → serve → drift → retrain → cloud.

Try it live ROC-AUC ~0.82 205 offline tests 100% synthetic data Python 3.11+ MIT License


▶ Try it — no signup, no key

Interactive demo — score machine telemetry for failure risk, upload your own CSV/Parquet, or generate a synthetic fleet and get one risk score per vehicle. Runs on Google Cloud Run with a managed Neon Postgres behind it, at $0.

Fleet generation is a web + worker system: the page kicks off a run and the API answers in ~16 ms, because the generator runs in a separate Cloud Run Job — not in a background thread on the serving container. A vehicle is flagged on sustained risk (the peak of a 1-hour rolling mean), not on its single worst reading: the generator injects sensor outliers on purpose, and ranking on the max flags healthy vehicles. Both choices are measured, not asserted — ADR-026.

Also live on Hugging Face Spaces: /health · /model-info · /docs

The honesty boundary. The data is 100% synthetic. The served model is a fixture-trained demo, labelled as such by /model-info — the ≈0.82 figure is the full-data model trained locally, and it is the only number ever reported. Nothing here claims a live production deployment; the drift→retrain loop is a demonstrated closed loop on synthetic data.


Quickstart

pip install -e .[dev]
pytest -q                 # offline — no network or generator needed (205 with the [serve,ops,cloud] extras)
pdm --version

That runs against a committed smoke fixture. To train on real (full, regenerated) data:

pip install -e .[generate]     # the pinned generator
pdm train                      # train both models → MLflow → register the winner
pdm serve                      # FastAPI over the production-aliased model
pdm flow --season heatwave     # the marquee: detect drift → retrain → promote-or-hold
pdm train train both models, track to MLflow, register the winner
pdm promote / pdm rollback metric-gated promotion to production; deterministic rollback
pdm serve FastAPI /predict, /health, /model-info
pdm monitor Evidently drift report + the share-of-features decision
pdm flow the closed detect → retrain → promote-or-hold loop
pdm detect / pdm tune / pdm sequence / pdm ceiling the modelling arc (see below)

What this is

The MLOps half of a two-repo story. Its companion can-telemetry-forge is a clean-room generator of synthetic, SAE J1939-grounded heavy-equipment telemetry. This repo is the ML system in production on top of it.

Built the data engine, then the ML-in-production system over it.

Nothing about the model is clever — that's the point. The dataset is diverse, statistically credible and fully reproducible, so the pipeline around it is the thing on display.


The four things worth your time

🚦 A model that scores worse cannot reach production

Promotion to the production alias is metric-gated: pdm promote reads the candidate's and the incumbent's ROC-AUC from their MLflow source runs and moves the alias only if the candidate clears the gate.

What the gate does and does not decide. It compares two point estimates of one metric on one fixed grouped split (~27 held-out units). That is enough to stop a visibly worse model and to make promotion reproducible; it is not a confidence interval, and a candidate that regresses on calibration, latency or a subgroup while holding ROC-AUC passes it untouched. The mechanism is the deliverable here — the statistical power is bounded by a synthetic fleet of 134 units. A worse candidate does not promote — asserted by test — and that rejection is a governed, structured outcome, not an exception. Rollback restores the prior version deterministically.

The load-bearing part: the auto-retrain loop routes through that same gate, unchanged. A retrained model that doesn't beat the incumbent is held, not shipped (proven by a test that sets an impossible bar). So auto-retrain can never quietly mean auto-degrade — the guarantee isn't hoping the retrain is good, it's the gate for when it isn't.

🐛 The bug was in the data, not the model

The classifier scored ≈ 0.55 — chance. Rather than tune it, I measured why: a failing unit's pre-failure rows were statistically identical to its healthy rows. There was no signal to learn.

The root cause was upstream, in the generator — failures had a when, but the sensors had no path toward it. I fixed it there (progressive pre-failure degradation, no label leak), and the same model reached ≈ 0.82 ROC-AUC. Two attempts to also rebalance the failure hazard were measured and rejected.

Finding that my own showcase was measuring at chance — and saying so — is the point.

🧭 Then I measured that 0.82 is the data's ceiling

Instead of chasing a better model forever, I characterized the ceiling from three independent angles: an AUC decomposition by time-to-failure horizon, a deliberately label-leaking upper bound (a fenced diagnostic, never a reported metric, asserted by test), and an out-of-fold stacking redundancy probe.

All three converged: ≈0.82 is the data's limit, not the pipeline's. One probe refuted my own hypothesis — stacking beat the best base model by ~+0.007, leaving a little combinable signal — and I reported that instead of what I expected.

Two more negatives, measured and reported: grouped-CV Optuna HPO moves the number by +0.003 (LightGBM) / +0.000 (LogReg), and a causal TCN lands below the cheap temporal-feature LightGBM. Tuning wasn't the lever; the data was. Knowing the ceiling is the licence to stop optimizing and go build the spine.

🌩️ Operated on managed cloud — not just containerized

The same image runs on Google Cloud Run (a managed serverless container runtime, not a VM) with a managed Neon Postgres behind it. The /demo page scores your inputs and logs each prediction to the managed database, then reads it back — so the managed resource has a real job, not a decorative one. Secrets live in Secret Manager; the image builds via Cloud Build. The whole thing runs at $0 (scale-to-zero + Neon free tier).

store_pg.open_log() accepts any SQLAlchemy URL, so the same code runs on tmp SQLite in tests and Postgres in production — and graceful degradation is a hard invariant: with no DATABASE_URL the demo simply doesn't persist. Adding a managed resource cannot break any other deploy.


The drift → retrain loop

The generator exposes a season knob that shifts the whole fleet's operating distribution (a heatwave runs the machines hotter). That is the drift stimulus:

baseline model  ──serve──►  production
        │
   --season heatwave  ──►  the data distribution drifts
        │
   drift monitor FIRES (Evidently)  ──►  retrain on the new distribution
        │
   re-run the SAME model comparison  ──►  register + promote-or-HOLD (the F3 gate)

Drift is decided by a share-of-features policy, not a single column — one noisy signal doesn't trigger a retrain.

pdm monitor --season heatwave     # the drift report + the decision
pdm flow    --season heatwave     # the full loop, in-process

Honest evaluation, by construction

  • A leakage guard that fails the build if a label-side column reaches the features.
  • Era-gated missingness preserved as signal — a sensor an older CAN bus never reported is NULL, not zero; LightGBM consumes the NaN natively. No blind imputation.
  • Unit-grouped splits, so no machine's autocorrelated series straddles the train/test line.
  • One seed threads data → split → metrics. Same seed, same numbers.
  • Reported models always train on the full regenerated dataset. The committed data/sample_readings.parquet is a smoke fixture so clone && pytest runs offline — it is never a training set for a reported number (ADR-001).

The stack (and why two orchestration layers)

Concern Tool Note
Tracking + model registry MLflow Local SQLite backend — no server, no cost. Built on version aliases (the current API; the classic stages are deprecated).
Models scikit-learn + LightGBM A LogReg baseline and a LightGBM contender, compared through MLflow — model selection as a recorded process, not folklore.
Serving FastAPI Serves the production-aliased model; a promotion or rollback changes /predict with no redeploy.
Drift Evidently Baseline vs. a season-shifted distribution, under an auditable share-of-features policy.
Orchestration Prefect Authors detect → retrain → promote-or-hold; runs in-process for tests.
Scheduled execution GitHub Actions Triggers the flow on a cron, on free runners.
Managed cloud Cloud Run + Neon Managed runtime + managed Postgres, in production, at $0.

Prefect and Actions sit at different layers — Actions is the scheduler, Prefect is the flow author — and each closes a distinct gap. Rationale in ADR-002.


Going deeper

The README is the tour. The substance lives in:

  • docs/ARCHITECTURE.md — module design and data flow.
  • docs/DECISIONS.md — the ADRs: why aliases over stages, why a fixture is never a training set, why serving loads from a client-resolved path, and the rest.
  • docs/ROADMAP.md — every phase (F0–F9, plus Epoch 2) with objective and definition of done, including the two deliberately deferred ones (RUL reframing, cross-dataset validation on NASA C-MAPSS).
  • docs/EPOCH2_PLAN.md — the current epoch: from one risk score to a per-vehicle, multi-fault maintenance report. Its §1 locked decisions carry the reasoning that a later session could otherwise quietly reverse.
  • docs/DEPLOY.md — the hosted and managed-cloud deploys.

License

MIT © 2026 Jorge Ribeiro

About

MLOps pipeline over synthetic predictive-maintenance telemetry: MLflow tracking + model registry, FastAPI serving, and a drift -> auto-retrain loop (Evidently + Prefect). Clean-room, 100% synthetic.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages