The Data Journey: Messy to Meaningful - Phase 6¶

Introduction — From Prediction to Explanation¶

So far, our journey has taken us from raw data to machine learning:

Collect → Clean → Transform → Visualize → Model → Evaluate

But there is one important question left:

Why did the model make this prediction?¶

A machine learning model may tell us:

“This customer is likely to be a High-Value Customer.”

But as data scientists, we often need to know:

Which features influenced that prediction, and in what direction?

This is where Explainable AI (XAI) comes in.

In this final phase, we will explore two popular explainability techniques:

SHAP — SHapley Additive exPlanations

In [1]:
import pandas as pd

customer_features = pd.read_csv('Online Retail Phase 4 Output.csv',index_col = 0)
customer_features.head()

features = [
    "NumberOfOrders",
    "TotalQuantity",
    "AvgOrderValue"
]

X = customer_features[features]
y = customer_features["HighValueCustomer"]

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42,
    stratify=y
)

Before Explainability — Load Our Final Model¶

We already trained our models in Phase 5.

For consistency, let's train Gradient Boosting again using the training data.

In [3]:
from sklearn.ensemble import GradientBoostingClassifier

gb_model = GradientBoostingClassifier(
    random_state=42
)

gb_model.fit(X_train, y_train)
Out[3]:
GradientBoostingClassifier(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=42)

Now we have a trained model:¶

X_train ↓ Gradient Boosting ↓ Trained Model ↓ Predictions

First: Understand Feature Importance¶

Before jumping into SHAP and LIME, let's start with a simple question:

Which features does the model consider important overall?

Gradient Boosting provides a built-in featureimportances attribute.

In [5]:
feature_importance = pd.DataFrame({
    "Feature": X_train.columns,
    "Importance": gb_model.feature_importances_
})

feature_importance = feature_importance.sort_values(
    "Importance",
    ascending=False
)

feature_importance
Out[5]:
Feature Importance
1 TotalQuantity 0.860538
2 AvgOrderValue 0.072451
0 NumberOfOrders 0.067010

We can visualize it:¶

In [6]:
import matplotlib.pyplot as plt

plt.figure(figsize=(7, 4))

plt.barh(
    feature_importance["Feature"],
    feature_importance["Importance"]
)

plt.xlabel("Importance")
plt.ylabel("Feature")
plt.title("Gradient Boosting Feature Importance")

plt.gca().invert_yaxis()

plt.show()

But there is a limitation.¶

Feature importance tells us:

Which features are important?

It doesn't necessarily tell us:

Why did a particular customer receive this prediction?

And it doesn't clearly show whether a feature pushed the prediction higher or lower.

That's where SHAP becomes useful.

SHAP — Explaining the Model¶

SHAP is based on Shapley values from game theory.

The basic idea is:

Each feature gets a contribution toward the model's prediction.

For example:

In [ ]:
Prediction
    ↓
High Value Customer
    ↑
    │
 ┌──┼───────────────┐
 │  │               │
Orders       Quantity       Avg Order Value
  +              +                +

A positive SHAP contribution pushes the prediction toward one class, while a negative contribution pushes it in the opposite direction.¶

Install and Import SHAP¶

In [7]:
import shap

Create the SHAP Explainer¶

For tree-based models such as Gradient Boosting, SHAP provides a tree-based explainer.

In [9]:
explainer = shap.TreeExplainer(gb_model)

shap_values = explainer.shap_values(X_test)

Now SHAP has calculated the contribution of each feature for the test observations.¶

Global Explanation — Which Features Matter Most?¶

Let's start with a global view.

In [10]:
shap.summary_plot(
    shap_values,
    X_test
)

How to read this plot¶

The plot helps us understand:

  • Which features have the greatest overall influence

  • Whether high or low feature values push predictions higher/lower

  • The distribution of feature effects across observations

This is more informative than simply saying:

“NumberOfOrders is important.”

We can start asking:

How does NumberOfOrders influence the prediction?

SHAP Bar Plot¶

For a simpler summary:

In [11]:
shap.summary_plot(
    shap_values,
    X_test,
    plot_type="bar"
)

This gives us an overall ranking of feature importance based on the average magnitude of SHAP contributions.

The difference:

Traditional Feature Importance

How important is the feature to the tree model?

SHAP

How much does the feature contribute to predictions?

Local Explanation — Explain One Customer¶

Global explanations tell us what happens across the dataset.

But suppose we want to answer:

Why was this particular customer classified as High Value?

Let's select one test observation.

In [12]:
customer_index = 0

customer = X_test.iloc[[customer_index]]

customer
Out[12]:
NumberOfOrders TotalQuantity AvgOrderValue
1206 3 249 114.613333
In [13]:
prediction = gb_model.predict(customer)

probability = gb_model.predict_proba(customer)[:, 1]

print("Prediction:", prediction[0])
print("Probability:", probability[0])
Prediction: 0
Probability: 0.0009248888730795335
Now we want to understand why the model arrived at that prediction.¶

SHAP Waterfall Plot¶

In [14]:
shap.waterfall_plot(
    shap.Explanation(
        values=shap_values[customer_index],
        base_values=explainer.expected_value,
        data=customer.iloc[0],
        feature_names=X_test.columns
    )
)

This SHAP waterfall explains why this particular customer was predicted as a low-value customer.

  • Baseline: E[f(X)] = -3.512
  • AvgOrderValue = 114.613: strongest negative contribution (-1.94)
  • TotalQuantity = 249: negative contribution (-0.88)
  • NumberOfOrders = 3: negative contribution (-0.65)
  • These contributions move the prediction to f(x) = -6.985.

👉 Key takeaway: All three features push the prediction downward, with AvgOrderValue having the largest influence.

SHAP answers: “Why did the model make this prediction?”

This gives us a customer-level explanation.

We can interpret it as:

In [ ]:
Base prediction
      +
Feature contribution
      +
Feature contribution
      +
Feature contribution
      ↓
Final prediction

The Complete Data Journey¶

And now we can close the entire series.

Phase 1 — Data Collection

We started with raw transactional data.

↓

Phase 2 — Data Cleaning

We handled missing values, duplicates, inconsistencies and prepared the dataset.

↓

Phase 3 — Transformation & Visualization

We engineered useful features and explored patterns, relationships and distributions.

↓

Phase 4 — Logistic Regression

We built our first predictive model and understood the relationship between features and the target.

↓

Phase 5 — Machine Learning

We introduced a Train-Test Split and compared multiple ML algorithms on unseen data.

↓

Phase 6 — Explainability

We used SHAP to understand why our models make their predictions.