An Explainable Stacking Ensemble for Early Prediction of Subnational Year-End Spending in Kalimantan
DOI:
https://doi.org/10.62201/87wq1717Keywords:
year-end expenditure realization, stacking ensemble, explainable machine learning, fiscal decentralization, cross-validation designAbstract
Indonesian subnational governments concentrate a disproportionate share of annual expenditure in the final quarter of the fiscal year, while fiscal oversight remains largely retrospective. This study asks whether each jurisdiction's December expenditure realization can be predicted from mid-year information, and whether such a model can be trusted where an early-warning system would add the most value. We assemble a panel of 27 subnational budget targets spanning five Kalimantan provinces, two government tiers, and three fiscal years (2023–2025), observed at up to six monthly checkpoints, yielding 160 target–checkpoint observations. A stacking ensemble of a random forest and a histogram-based gradient boosting regressor under a ridge meta-learner is evaluated under two leave-one-group-out regimes, with all preprocessing fitted on training folds only and inner folds grouped by target so that the six checkpoints sharing one label cannot leak into the meta-learner. Under leave-one-year-out validation the ensemble attains a target-level R² of 0.910 against 0.859 for a naïve seasonal-scaling baseline, although on percentage error the two are effectively tied (20.5% versus 21.0%). Under leave-one-province-out validation the ensemble is outperformed by that same baseline on all four metrics (R² 0.574 versus 0.832) and fails on the largest fiscal unit (−0.416) and the smallest (−2.908). Grouped-permutation importance assigns the within-year dynamics family an error increase of 2 IDR billion against a base error near 3,900, while the revenue–expenditure gap and jurisdiction identity dominate. A direct ablation confirms the diagnosis: deleting jurisdiction identity costs 0.019 R² temporally but recovers 0.129 R² spatially, and an absorption-rate target drives R² to −1.485. In small subnational panels, temporal validation alone materially overstates what a predictive fiscal-monitoring model knows.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Proceeding of International Conference on Digital, Social, and Science

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.







