End-to-end ML application that predicts machine failure probability, flags unusual machine behaviour, explains predictions with SHAP, and serves everything through a FastAPI service and a Streamlit dashboard.
Status: on the author's Windows machine (Python 3.12) the full training run, the 34-test pytest suite and the API
/healthendpoint have been executed successfully. MLflow logging, the dashboard,/predictthrough the running server and Docker Compose are not yet verified. Numbers below come from one run (seed 42, 2,000-row held-out test set) and can shift slightly with library versions.
Unplanned machine downtime is costly. Given sensor readings, the system answers: is the machine behaving abnormally, what is its failure probability, which factors drive that prediction, and which machines should be serviced first. Failure probability and anomaly status are separate signals: anomaly score is not a failure probability.
Sensors/CSV -> Ingestion -> Validation & Preprocessing -> Feature Engineering
-> Isolation Forest (anomaly) XGBoost (failure probability)
-> SHAP explainability -> FastAPI -> Streamlit dashboard / PostgreSQL
Training (src/train.py) and inference (src/predict.py, api/) are separate. The API loads models/model_bundle.joblib; it never trains.
Python 3.11+, pandas, NumPy, scikit-learn, XGBoost, SHAP, MLflow, FastAPI, Uvicorn, Pydantic, Streamlit, Plotly, Matplotlib, SQLAlchemy + PostgreSQL, Docker / Compose, pytest.
AI4I 2020 Predictive Maintenance Dataset (UCI Machine Learning Repository, id 601). Columns: UDI, Product ID, Type (L/M/H), Air temperature [K], Process temperature [K], Rotational speed [rpm], Torque [Nm], Tool wear [min], target Machine failure, and failure-mode flags TWF/HDF/PWF/OSF/RNF.
Measured on the downloaded file: 10,000 records; no missing values, duplicate rows, invalid values or duplicate machine IDs; failure rate about 3.4% (train 3.39%, test 3.40%; 8,000 train / 2,000 test rows, 68 test failures). Full statistics: models/reports/dataset_summary.json.
Deviations from the original brief (dataset limits): AI4I has no timestamps, no repeated machines, and no vibration, pressure, voltage or current columns. Therefore: lag/rolling/rate-of-change features exist as a tested utility but are not used by the model; the API/dashboard use the AI4I sensor set; the dashboard has no date/time filter; Product ID is used as machine_id (one reading per machine). Nothing was fabricated to fill these gaps.
src/preprocessing.py: column validation, type coercion, duplicate detection (ignoring the row id), invalid-value removal, stratified train/test split, and a ColumnTransformer (median impute + standard scale; mode impute + one-hot for Type) that lives inside the sklearn pipelines, so it is fitted on training folds only. The failure-mode flags are components of the target and are excluded as features (leakage guard in assert_no_leakage).
| Feature | Definition |
|---|---|
temperature_diff_k |
process temperature - air temperature |
power_w |
torque x rotational speed x 2*pi/60 |
wear_torque_product |
tool wear x torque |
These are physically motivated; the dataset documentation describes failure modes tied to temperature difference, power and wear-torque, so tree models may find them easy to use. add_time_series_features (lag/rolling/change) is provided for future time-stamped data.
Isolation Forest fitted without labels on training features (contamination configurable, default 0.03). anomaly_score = -score_samples (higher = more unusual); ANOMALY/NORMAL comes from the model's cutoff. On the test set the model flagged 2.8% of machines; 26.8% of flagged machines actually failed versus 2.7% of unflagged ones (about 10x enrichment), but flagged machines contain only roughly 15 of the 68 failures. Anomaly status is a complementary signal, not a failure detector. Source: models/reports/anomaly_summary.json.
Baselines: Logistic Regression and Random Forest (class-weighted). Final model: XGBoost tuned with RandomizedSearchCV (12 iterations, 3-fold stratified CV, scored by PR-AUC, seed 42). All models are evaluated on the same held-out test set with accuracy, precision, recall, F1, ROC-AUC, PR-AUC and the confusion matrix. Under class imbalance, accuracy is misleading; recall, precision, F1 and PR-AUC matter most (explanations are stored in metrics.json).
Test-set results (generated by python -m src.train, copied from models/reports/model_comparison.md):
| model | accuracy | precision | recall | f1 | roc_auc | pr_auc |
|---|---|---|---|---|---|---|
| Logistic Regression | 0.8585 | 0.1791 | 0.8824 | 0.2978 | 0.9385 | 0.4297 |
| Random Forest | 0.9920 | 0.9333 | 0.8235 | 0.8750 | 0.9690 | 0.8687 |
| XGBoost | 0.9865 | 0.7662 | 0.8676 | 0.8138 | 0.9879 | 0.8979 |
At the default 0.5 threshold, Random Forest has the best precision and F1; XGBoost has the best recall among the tree models and the best ROC-AUC / PR-AUC. With only 68 test failures, a few predictions move these numbers noticeably, so the models are close and the ranking is not conclusive. XGBoost's lower precision comes partly from scale_pos_weight; threshold selection is a listed improvement. Plots: models/reports/*.png.
SHAP TreeExplainer on the XGBoost model. Per prediction: probability, risk level, top_factors (features pushing risk up) and factor_details (signed contributions). Global importance is saved to models/reports/shap_global_importance.*. These are model-derived contributions, not causal explanations.
| Endpoint | Purpose |
|---|---|
GET /health |
status and whether a model is loaded (degraded if not) |
GET /model/info |
version, params, test metrics, thresholds |
POST /predict |
single prediction with SHAP factors |
POST /predict/batch?explain=false |
batch prediction (max MAX_BATCH_SIZE) |
{"machine_id": "M14860", "machine_type": "M", "air_temperature_k": 298.1,
"process_temperature_k": 308.6, "rotational_speed_rpm": 1551, "torque_nm": 42.8, "tool_wear_min": 0}Response fields: prediction, failure_probability, risk_level (LOW/MEDIUM/HIGH, thresholds configurable via RISK_MEDIUM_THRESHOLD / RISK_HIGH_THRESHOLD), anomaly_score, anomaly_status, top_factors, factor_details, model_version. Interactive docs at /docs. If DATABASE_URL is set, predictions for known machine IDs are logged to PostgreSQL (failures there never break a prediction).
Streamlit pages: Overview, Machine Monitoring, Sensor Analysis, Prediction, Explainability, Model Performance. All scores come from the API; performance pages read the saved reports. Default scope is the held-out test split to avoid in-sample optimism. Screenshots: TODO after first run (none captured yet).
database/schema.sql (machines, sensor_readings, predictions, maintenance_records), seed.sql (loads real dataset rows only), queries.sql (7 analytics queries with CTEs and window functions; two are adapted/inapplicable to AI4I and noted in comments). maintenance_records stays empty because the dataset has none. The DB layer (src/database.py) is SQLAlchemy-based; Oracle is a design goal, not implemented or tested.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # edit secrets; never commit .env
python -m src.data_loader download # or place ai4i2020.csv in data/raw/
python -m src.data_loader summary
python -m src.train # trains, evaluates, logs to MLflow, saves the bundle
pytest # tests train a tiny model on a synthetic fixture
uvicorn api.main:app --reload # terminal 2
streamlit run dashboard/app.py # terminal 3
mlflow ui --backend-store-uri sqlite:///mlruns/mlflow.dbDocker: train first (the API loads the saved model), then docker compose up --build. Services: db (5432), api (8000), dashboard (8501), mlflow (5000). The compose file has not been run yet.
tests/ covers preprocessing, feature engineering (including per-machine/no-future-leak checks), model loading, prediction format and risk thresholds, and the API (health, validation, responses, degraded mode). Fixtures use a small synthetic frame only for tests; it is never used for the project model.
See repository tree: src/ (pipeline), api/, dashboard/, database/, tests/, notebooks/ (thin, unexecuted; outputs generate when you run them), models/, mlruns/.
- One reading per machine, no time axis: no true temporal forecasting or time-based split.
- The engineered features mirror documented failure mechanisms, so high scores may partly reflect that; interpret accordingly.
- MLflow uses a SQLite backend (
mlruns/mlflow.db); artifact paths are stored as absolute host paths, so artifact links may not resolve inside the Docker MLflow UI. requirements.txthas lower bounds only; freeze exact versions after your first successful run.
Time-series data with lag/rolling features and time-based splits, probability calibration and cost-based threshold selection, drift monitoring, authentication on the API, Oracle support, model registry promotion in MLflow, CI pipeline.