So far, our journey has taken us from raw data to machine learning:
Collect → Clean → Transform → Visualize → Model → Evaluate
But there is one important question left:
A machine learning model may tell us:
“This customer is likely to be a High-Value Customer.”
But as data scientists, we often need to know:
Which features influenced that prediction, and in what direction?
This is where Explainable AI (XAI) comes in.
In this final phase, we will explore two popular explainability techniques:
SHAP — SHapley Additive exPlanations
import pandas as pd
customer_features = pd.read_csv('Online Retail Phase 4 Output.csv',index_col = 0)
customer_features.head()
features = [
"NumberOfOrders",
"TotalQuantity",
"AvgOrderValue"
]
X = customer_features[features]
y = customer_features["HighValueCustomer"]
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.20,
random_state=42,
stratify=y
)
We already trained our models in Phase 5.
For consistency, let's train Gradient Boosting again using the training data.
from sklearn.ensemble import GradientBoostingClassifier
gb_model = GradientBoostingClassifier(
random_state=42
)
gb_model.fit(X_train, y_train)
GradientBoostingClassifier(random_state=42)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
GradientBoostingClassifier(random_state=42)
X_train ↓ Gradient Boosting ↓ Trained Model ↓ Predictions
Before jumping into SHAP and LIME, let's start with a simple question:
Which features does the model consider important overall?
Gradient Boosting provides a built-in featureimportances attribute.
feature_importance = pd.DataFrame({
"Feature": X_train.columns,
"Importance": gb_model.feature_importances_
})
feature_importance = feature_importance.sort_values(
"Importance",
ascending=False
)
feature_importance
| Feature | Importance | |
|---|---|---|
| 1 | TotalQuantity | 0.860538 |
| 2 | AvgOrderValue | 0.072451 |
| 0 | NumberOfOrders | 0.067010 |
import matplotlib.pyplot as plt
plt.figure(figsize=(7, 4))
plt.barh(
feature_importance["Feature"],
feature_importance["Importance"]
)
plt.xlabel("Importance")
plt.ylabel("Feature")
plt.title("Gradient Boosting Feature Importance")
plt.gca().invert_yaxis()
plt.show()
Feature importance tells us:
Which features are important?
It doesn't necessarily tell us:
Why did a particular customer receive this prediction?
And it doesn't clearly show whether a feature pushed the prediction higher or lower.
That's where SHAP becomes useful.
SHAP is based on Shapley values from game theory.
The basic idea is:
Each feature gets a contribution toward the model's prediction.
For example:
Prediction
↓
High Value Customer
↑
│
┌──┼───────────────┐
│ │ │
Orders Quantity Avg Order Value
+ + +
import shap
For tree-based models such as Gradient Boosting, SHAP provides a tree-based explainer.
explainer = shap.TreeExplainer(gb_model)
shap_values = explainer.shap_values(X_test)
Let's start with a global view.
shap.summary_plot(
shap_values,
X_test
)
The plot helps us understand:
Which features have the greatest overall influence
Whether high or low feature values push predictions higher/lower
The distribution of feature effects across observations
This is more informative than simply saying:
“NumberOfOrders is important.”
We can start asking:
How does NumberOfOrders influence the prediction?
For a simpler summary:
shap.summary_plot(
shap_values,
X_test,
plot_type="bar"
)
This gives us an overall ranking of feature importance based on the average magnitude of SHAP contributions.
The difference:
Traditional Feature Importance
How important is the feature to the tree model?
SHAP
How much does the feature contribute to predictions?
Global explanations tell us what happens across the dataset.
But suppose we want to answer:
Why was this particular customer classified as High Value?
Let's select one test observation.
customer_index = 0
customer = X_test.iloc[[customer_index]]
customer
| NumberOfOrders | TotalQuantity | AvgOrderValue | |
|---|---|---|---|
| 1206 | 3 | 249 | 114.613333 |
prediction = gb_model.predict(customer)
probability = gb_model.predict_proba(customer)[:, 1]
print("Prediction:", prediction[0])
print("Probability:", probability[0])
Prediction: 0 Probability: 0.0009248888730795335
shap.waterfall_plot(
shap.Explanation(
values=shap_values[customer_index],
base_values=explainer.expected_value,
data=customer.iloc[0],
feature_names=X_test.columns
)
)
This SHAP waterfall explains why this particular customer was predicted as a low-value customer.
E[f(X)] = -3.512👉 Key takeaway: All three features push the prediction downward, with AvgOrderValue having the largest influence.
SHAP answers: “Why did the model make this prediction?”
This gives us a customer-level explanation.
We can interpret it as:
Base prediction
+
Feature contribution
+
Feature contribution
+
Feature contribution
↓
Final prediction
And now we can close the entire series.
Phase 1 — Data Collection
We started with raw transactional data.
↓
Phase 2 — Data Cleaning
We handled missing values, duplicates, inconsistencies and prepared the dataset.
↓
Phase 3 — Transformation & Visualization
We engineered useful features and explored patterns, relationships and distributions.
↓
Phase 4 — Logistic Regression
We built our first predictive model and understood the relationship between features and the target.
↓
Phase 5 — Machine Learning
We introduced a Train-Test Split and compared multiple ML algorithms on unseen data.
↓
Phase 6 — Explainability
We used SHAP to understand why our models make their predictions.