Variational Inference of Extremes for Rare Event Modeling: Theory and Applications
Pub. online: 4 August 2026
Type: Methodology Article
Open Access
Area: Engineering Science
Accepted
31 July 2024
31 July 2024
Published
4 August 2026
4 August 2026
Abstract
Severe event class imbalance is observed in many fields, such as insurance fraud, severe weather, traffic safety, and rare disease, and has far-reaching significance when the generalization of minority event classes is of primary interest. While existing solutions mostly focus on differential sampling or sample re-weighting approaches to alleviate the imbalance issue, we take a novel alternative view on promoting generalization. Our proposal is formulated under the generative Bayesian framework, positing that predictors are the stochastic proxies of latent causes with limited samples, whose exceedance leads to extreme events. To accurately capture the extended tail, our solution adopts the Gumbel distribution as prior and is modeled with a variational inference framework. Following the assertion that exceedance leads to extremes, we devise a disentangled additive monotonic neural architecture to predict the risk. The proposed model acknowledges representation uncertainties while embracing improved interpretability, generalization, and robustness. We provide theoretical insights to show the merits of the proposed approach. To verify the effectiveness in empirical settings, we conducted studies on various real-world data against the state-of-the-art counterparts, with encouraging results reported.
References
Alemi, A. A., Fischer, I., Dillon, J. V. and Murphy, K. (2016). Deep variational information bottleneck. arXiv:1612.00410.
Byrd, J. and Lipton, Z. C. (2018). What is the effect of importance weighting in deep learning? arXiv:1812.03372.
Coles, S., Bawa, J., Trenner, L. and Dorazio, P. (2001). An introduction to statistical modeling of extreme values 208. Springer. https://doi.org/10.1007/978-1-4471-3675-0. MR1932132
Fan, J., Song, R. et al. (2010). Sure independence screening in generalized linear models with NP-dimensionality. The Annals of Statistics 38(6) 3567–3604. https://doi.org/10.1214/10-AOS798. MR2766861
Firth, D. (1993). Bias reduction of maximum likelihood estimates. Biometrika 27–38. https://doi.org/10.1093/biomet/80.1.27. MR1225212
Fortuin, V. (2022). Priors in bayesian deep learning: A review. International Statistical Review 90(3) 563–591. https://doi.org/10.1111/insr.12502. MR4524825
Guo, F. (2019). Statistical methods for naturalistic driving studies. Annual Review of Statistics and its Application 6 309–328. https://doi.org/10.1146/annurev-statistics-030718-105153. MR3939523
Han, Q., Wang, T., Chatterjee, S., Samworth, R. J. et al. (2019). Isotonic regression in general dimensions. The Annals of Statistics 47(5) 2440–2471. https://doi.org/10.1214/18-AOS1753. MR3988762
Hastie, T. J. and Tibshirani, R. J. (1990). Generalized additive models 43. CRC press. MR1082147
Kingma, D. P. and Welling, M. (2013). Auto-encoding variational bayes. arXiv:1312.6114.
Li, R., Zhong, W. and Zhu, L. (2012). Feature screening via distance correlation learning. Journal of the American Statistical Association 107(499) 1129–1139. https://doi.org/10.1080/01621459.2012.695654. MR3010900
Li, X., Li, R., Xia, Z. and Xu, C. (2020). Distributed feature screening via componentwise debiasing. Journal of Machine Learning Research 21. MR4071207
Lusa, L. et al. (2017). Gradient boosting for high-dimensional prediction of rare events. Computational Statistics & Data Analysis 113 19–37. https://doi.org/10.1016/j.csda.2016.07.016. MR3662388
Massoli, F. V., Falchi, F., Kantarci, A., Akti, ., Ekenel, H. K. and Amato, G. (2021). MOCCA: Multilayer one-class classification for anomaly detection. IEEE Transactions on Neural Networks and Learning Systems. https://doi.org/10.1109/tnnls.2021.3130074. MR4442308
Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C. and Joulin, A. (2017). Advances in pre-training distributed word representations. arXiv:1712.09405.
Ren, M., Zeng, W., Yang, B. and Urtasun, R. (2018). Learning to reweight examples for robust deep learning. arXiv:1803.09050.
Shen, D., Wang, G., Wang, W., Min, M. R., Su, Q., Zhang, Y., Li, C., Henao, R. and Carin, L. (2018). Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms. arXiv:1805.09843.
Sur, P. and Candès, E. J. (2019). A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences 116(29) 14516–14525. https://doi.org/10.1073/pnas.1810420116. MR3984492
Székely, G. J., Rizzo, M. L., Bakirov, N. K. et al. (2007). Measuring and testing dependence by correlation of distances. The Annals of Statistics 35(6) 2769–2794. https://doi.org/10.1214/009053607000000505. MR2382665
Zhang, F., Gao, C. et al. (2020). Convergence rates of variational posterior distributions. Annals of Statistics 48(4) 2180–2207. https://doi.org/10.1214/19-AOS1883. MR4134791