logo icods

Explainable AI (SHAP) on XGBoost for E-Commerce Customer Churn: An Exploratory Case Study of Interpretability Under Data-Quality Constraints

Authors

  • Tikno

    Universitas Internasional Semen Indonesia
    Author

DOI:

https://doi.org/10.62201/tn7hra16

Keywords:

Customer Churn, XGBoost, SHAP, Explainable AI, CRM, E-commerce, Data Quality

Abstract

Customer churn remains a critical challenge in e-commerce, yet predictive performance is fundamentally bounded by the quality of the underlying data and its labels. This study applies the XGBoost algorithm together with SHAP (SHapley Additive exPlanations) to build a churn prediction model that balances predictive capability with managerial interpretability, using the secondary public dataset "E-Commerce Customer Behavior Analysis" from Kaggle, consisting of 250,000 transaction records aggregated into 49,673 unique customer profiles. Class imbalance was addressed through XGBoost's built-in class-weighting mechanism, with hyperparameters optimized using Optuna. The best configuration achieved a Recall of 0.9909 at the F1-optimal threshold of 0.5733, but this figure is misleading in isolation: the accompanying ROC-AUC of 0.5070 and PR-AUC of 0.2028 confirm near-random discrimination. The root cause is a critical data validity limitation — the churn label carries no documented derivation methodology, and all features exhibited Pearson correlations below |r| = 0.01 against the target, consistent with suspected synthetic or randomized label generation. Despite near-random model performance, SHAP was applied in an exploratory manner to surface dominant behavioral patterns: Tenure, Recency, and Age emerged as the most influential features, directionally consistent with RFM-based churn theory. These findings are reported as exploratory hypotheses rather than confirmed causal drivers, and were translated into a set of segmented, hypothesis-driven CRM retention ideas pending empirical validation. The primary contribution of this study is methodological: demonstrating that interpretability tools retain value for hypothesis generation in data-constrained settings, while affirming that high-quality, operationally derived churn labels are a prerequisite for actionable predictive output.

Downloads

Published

2026-09-07