A machine-learning pipeline for detecting illicit Bitcoin transactions and explaining predictions with SHAP.
This project uses the Elliptic Bitcoin transaction dataset to classify transactions as licit or illicit.
The pipeline processes 165 anonymised transaction features, compares several tree-based classifiers, selects a decision threshold using validation data, and evaluates the final model on later, unseen time steps. SHAP is then used to explain both global model behaviour and individual predictions.
The project is an educational exploration of cryptocurrency transaction analysis, imbalanced classification, temporal evaluation, and explainable machine learning. It is not intended as a production fraud-detection system.
| Feature | Implementation |
|---|---|
| Transaction classification | Detects licit and illicit Bitcoin transactions |
| Temporal evaluation | Trains and tests on separate chronological time periods |
| Model comparison | Evaluates Random Forest, Extra Trees, and Histogram Gradient Boosting configurations |
| Imbalance-aware metrics | Reports precision, recall, F1, balanced accuracy, average precision, and ROC-AUC |
| Threshold selection | Selects the classification threshold using validation data only |
| Explainability | Produces global SHAP summaries and individual waterfall explanations |
| Reproducibility | Uses fixed random seeds and saves model metadata and validation results |
-
Transaction features and class labels are loaded from the Elliptic dataset.
-
Transactions labelled
unknownare excluded. -
Original labels are converted into binary classes:
licitandillicit. -
Data is divided chronologically:
- Training: time steps 1-29
- Validation: time steps 30-34
- Testing: time steps 35-49
-
Multiple tree-based classifiers and class-weight configurations are evaluated on the validation period.
-
The decision threshold is selected using validation predictions.
-
The selected model is retrained on the combined training and validation data.
-
Final performance is measured once on the untouched test period.
-
SHAP explanations are generated for global behaviour and representative predictions.
The selected model was an unweighted Random Forest using a decision threshold of 0.5267.
| Metric | Test result |
|---|---|
| Accuracy | 98.08% |
| Majority-class baseline | 93.50% |
| Balanced accuracy | 85.91% |
| Illicit precision | 97.99% |
| Illicit recall | 71.93% |
| Illicit F1-score | 82.96% |
| Macro F1-score | 90.97% |
| Average precision | 79.38% |
| ROC-AUC | 93.02% |
The high illicit precision means that very few licit transactions were incorrectly flagged. Recall is lower, meaning the model did not detect every illicit transaction.
SHAP measures how strongly each feature influences the model's output. The summary plot shows both the magnitude and direction of feature effects across a sample of test transactions.
The Elliptic dataset intentionally anonymises its transaction features. Therefore, explanations identify influential feature indices rather than named financial attributes.
View additional SHAP explanations
View dataset, installation, and usage instructions
Download the Elliptic Bitcoin transaction dataset from Kaggle.
Extract these files into a directory named dataset at the repository root:
dataset/
├── elliptic_txs_classes.csv
├── elliptic_txs_edgelist.csv
└── elliptic_txs_features.csv
The dataset is not included in this repository. Refer to its Kaggle page for its documentation and usage terms.
Clone the repository:
git clone https://github.com/maxfroggatt/Explainable-Bitcoin-Transaction-Detection.git
cd Explainable-Bitcoin-Transaction-DetectionCreate a virtual environment:
python -m venv .venvActivate it using the appropriate command.
Windows PowerShell
.\.venv\Scripts\Activate.ps1Windows Command Prompt
.venv\Scripts\activate.batmacOS/Linux
source .venv/bin/activateInstall the dependencies:
python -m pip install -r requirements.txtTrain and evaluate the classifiers:
python src/train_model.pyGenerate the SHAP explanations:
python src/explain_shap.pyGenerated results are written to outputs/, including:
- Evaluation metrics and confusion matrix
- Precision-recall curve
- Validation results and model metadata
- Trained model artefact
- Test data used for explanation
- Global and local SHAP visualisations
- The dataset contains anonymised features, limiting the semantic interpretation of individual SHAP explanations.
- The current pipeline uses the supplied transaction features rather than directly modelling the edge list with a graph neural network.
- The dataset is highly imbalanced, so accuracy alone is not an adequate measure of performance.
- Although illicit precision is high, the model detects approximately 72% of illicit transactions and therefore still produces false negatives.
- Predictions indicate patterns learned from historical labels and do not prove that a transaction is criminal.






