The New England Journal of Statistics in Data Science logo


  • Help
Login Register

  1. Home
  2. To appear
  3. Variational Inference of Extremes for Ra ...

The New England Journal of Statistics in Data Science

Submit your article Information Become a Peer-reviewer
  • Article info
  • Full article
  • More
    Article info Full article

Variational Inference of Extremes for Rare Event Modeling: Theory and Applications
Chen Qian   Chenyang Tao   Jingbin Xu     All authors (4)

Authors

 
Placeholder
https://doi.org/10.51387/26-NEJSDS071
Pub. online: 4 August 2026      Type: Methodology Article      Open accessOpen Access
Area: Engineering Science

Accepted
31 July 2024
Published
4 August 2026

Abstract

Severe event class imbalance is observed in many fields, such as insurance fraud, severe weather, traffic safety, and rare disease, and has far-reaching significance when the generalization of minority event classes is of primary interest. While existing solutions mostly focus on differential sampling or sample re-weighting approaches to alleviate the imbalance issue, we take a novel alternative view on promoting generalization. Our proposal is formulated under the generative Bayesian framework, positing that predictors are the stochastic proxies of latent causes with limited samples, whose exceedance leads to extreme events. To accurately capture the extended tail, our solution adopts the Gumbel distribution as prior and is modeled with a variational inference framework. Following the assertion that exceedance leads to extremes, we devise a disentangled additive monotonic neural architecture to predict the risk. The proposed model acknowledges representation uncertainties while embracing improved interpretability, generalization, and robustness. We provide theoretical insights to show the merits of the proposed approach. To verify the effectiveness in empirical settings, we conducted studies on various real-world data against the state-of-the-art counterparts, with encouraging results reported.

References

[1] 
Alemi, A. A., Fischer, I., Dillon, J. V. and Murphy, K. (2016). Deep variational information bottleneck. arXiv:1612.00410.
[2] 
Bengio, Y., Courville, A. and Vincent, P. (2013). Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence 35(8) 1798–1828.
[3] 
Byrd, J. and Lipton, Z. C. (2018). What is the effect of importance weighting in deep learning? arXiv:1812.03372.
[4] 
Cao, K., Wei, C., Gaidon, A., Arechiga, N. and Ma, T. (2019). Learning imbalanced datasets with label-distribution-aware margin loss. In Advances in Neural Information Processing Systems 1565–1576.
[5] 
Carlini, N. and Wagner, D. (2017). Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy 39–57. IEEE.
[6] 
Casale, F. P., Dalca, A., Saglietti, L., Listgarten, J. and Fusi, N. (2018). Gaussian process prior variational autoencoders. Advances in neural information processing systems 31.
[7] 
Chen, R. T., Li, X., Grosse, R. B. and Duvenaud, D. K. (2018). Isolating sources of disentanglement in variational autoencoders. Advances in neural information processing systems 31.
[8] 
Chen, Y., Zhou, X. S. and Huang, T. S. (2001). One-class SVM for learning in image retrieval. In Proceedings 2001 International Conference on Image Processing (Cat. No. 01CH37205) 1 34–37. IEEE.
[9] 
Coles, S., Bawa, J., Trenner, L. and Dorazio, P. (2001). An introduction to statistical modeling of extreme values 208. Springer. https://doi.org/10.1007/978-1-4471-3675-0. MR1932132
[10] 
Dingus, T. A., Guo, F., Lee, S., Antin, J. F., Perez, M., Buchanan-King, M. and Hankey, J. (2016). Driver crash risk factors and prevalence evaluation using naturalistic driving data. Proceedings of the National Academy of Sciences 113(10) 2636–2641.
[11] 
Dong, Q., Gong, S. and Zhu, X. (2018). Imbalanced deep learning by minority class incremental rectification. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(6) 1367–1381.
[12] 
Fan, J., Song, R. et al. (2010). Sure independence screening in generalized linear models with NP-dimensionality. The Annals of Statistics 38(6) 3567–3604. https://doi.org/10.1214/10-AOS798. MR2766861
[13] 
Finn, C., Abbeel, P. and Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning 1126–1135. PMLR.
[14] 
Firth, D. (1993). Bias reduction of maximum likelihood estimates. Biometrika 27–38. https://doi.org/10.1093/biomet/80.1.27. MR1225212
[15] 
Fortuin, V. (2022). Priors in bayesian deep learning: A review. International Statistical Review 90(3) 563–591. https://doi.org/10.1111/insr.12502. MR4524825
[16] 
Guo, F. (2019). Statistical methods for naturalistic driving studies. Annual Review of Statistics and its Application 6 309–328. https://doi.org/10.1146/annurev-statistics-030718-105153. MR3939523
[17] 
Haixiang, G., Yijing, L., Shang, J., Mingyun, G., Yuanyue, H. and Bing, G. (2017). Learning from class-imbalanced data: Review of methods and applications. Expert Systems with Applications 73 220–239.
[18] 
Han, Q., Wang, T., Chatterjee, S., Samworth, R. J. et al. (2019). Isotonic regression in general dimensions. The Annals of Statistics 47(5) 2440–2471. https://doi.org/10.1214/18-AOS1753. MR3988762
[19] 
Hastie, T. J. and Tibshirani, R. J. (1990). Generalized additive models 43. CRC press. MR1082147
[20] 
He, H. and Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21(9) 1263–1284.
[21] 
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S. and Lerchner, A. (2017). Beta-VAE: learning basic visual concepts with a constrained variational framework. The International Conference on Learning Representations 2(5) 6.
[22] 
Jin, D., Jin, Z., Zhou, J. T. and Szolovits, P. (2020). Is Bert really robust? A strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI Conference on Artificial Intelligence 34 8018–8025.
[23] 
King, G. and Zeng, L. (2001). Logistic regression in rare events data. Political Analysis 9(2) 137–163.
[24] 
Kingma, D. P. and Welling, M. (2013). Auto-encoding variational bayes. arXiv:1312.6114.
[25] 
Kingma, D. P., Welling, M. et al. (2019). An introduction to variational autoencoders. Foundations and Trendső in Machine Learning 12(4) 307–392.
[26] 
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I. and Welling, M. (2016). Improved variational inference with inverse autoregressive flow. Advances in neural information processing systems 29.
[27] 
Li, R., Zhong, W. and Zhu, L. (2012). Feature screening via distance correlation learning. Journal of the American Statistical Association 107(499) 1129–1139. https://doi.org/10.1080/01621459.2012.695654. MR3010900
[28] 
Li, X., Li, R., Xia, Z. and Xu, C. (2020). Distributed feature screening via componentwise debiasing. Journal of Machine Learning Research 21. MR4071207
[29] 
Lin, T. -Y., Goyal, P., Girshick, R., He, K. and Dollár, P. (2017). Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision 2980–2988.
[30] 
Liu, L., Racz, D., Vaillancourt, K., Michelman, J., Barnes, M., Mellem, S., Eastham, P., Green, B., Armstrong, C., Bal, R. et al. (2023). Smartphone-based hard-braking event detection at scale for road safety services. Transportation research part C: emerging technologies 146 103949.
[31] 
Lusa, L. et al. (2017). Gradient boosting for high-dimensional prediction of rare events. Computational Statistics & Data Analysis 113 19–37. https://doi.org/10.1016/j.csda.2016.07.016. MR3662388
[32] 
Massoli, F. V., Falchi, F., Kantarci, A., Akti, ., Ekenel, H. K. and Amato, G. (2021). MOCCA: Multilayer one-class classification for anomaly detection. IEEE Transactions on Neural Networks and Learning Systems. https://doi.org/10.1109/tnnls.2021.3130074. MR4442308
[33] 
Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C. and Joulin, A. (2017). Advances in pre-training distributed word representations. arXiv:1712.09405.
[34] 
O’Kelly, M., Sinha, A., Namkoong, H., Tedrake, R. and Duchi, J. C. (2018). Scalable end-to-end autonomous vehicle testing via rare-event simulation. Advances in Neural Information Processing Systems 31 9827–9838.
[35] 
Ren, M., Zeng, W., Yang, B. and Urtasun, R. (2018). Learning to reweight examples for robust deep learning. arXiv:1803.09050.
[36] 
Rezende, D. and Mohamed, S. (2015). Variational inference with normalizing flows. In International conference on machine learning 1530–1538. PMLR.
[37] 
Ruff, L., Vandermeulen, R., Goernitz, N., Deecke, L., Siddiqui, S. A., Binder, A., Müller, E. and Kloft, M. (2018). Deep one-class classification. In International Conference on Machine Learning 4393–4402.
[38] 
Shen, D., Wang, G., Wang, W., Min, M. R., Su, Q., Zhang, Y., Li, C., Henao, R. and Carin, L. (2018). Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms. arXiv:1805.09843.
[39] 
Shi, L., Qian, C. and Guo, F. (2022). Real-time driving risk assessment using deep learning with XGBoost. Accident Analysis & Prevention 178 106836.
[40] 
Snoek, J., Ovadia, Y., Fertig, E., Lakshminarayanan, B., Nowozin, S., Sculley, D., Dillon, J., Ren, J. and Nado, Z. (2019). Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems 13969–13980.
[41] 
Sohn, K., Lee, H. and Yan, X. (2015). Learning structured output representation using deep conditional generative models. Advances in Neural Information Processing Systems 28.
[42] 
Sønderby, C. K., Raiko, T., Maaløe, L., Sønderby, S. K. and Winther, O. (2016). Ladder variational autoencoders. Advances in neural information processing systems 29.
[43] 
Sur, P. and Candès, E. J. (2019). A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences 116(29) 14516–14525. https://doi.org/10.1073/pnas.1810420116. MR3984492
[44] 
Székely, G. J., Rizzo, M. L., Bakirov, N. K. et al. (2007). Measuring and testing dependence by correlation of distances. The Annals of Statistics 35(6) 2769–2794. https://doi.org/10.1214/009053607000000505. MR2382665
[45] 
Tomczak, J. and Welling, M. (2018). VAE with a VampPrior. In International Conference on Artificial Intelligence and Statistics 1214–1223. PMLR.
[46] 
Wang, H. (2020). Logistic regression for massive data with rare events. In International Conference on Machine Learning 9829–9836. PMLR.
[47] 
Wehenkel, A. and Louppe, G. (2019). Unconstrained monotonic neural networks. Advances in Neural Information Processing Systems 32.
[48] 
Wu, T., Liu, Z., Huang, Q., Wang, Y. and Lin, D. (2021). Adversarial robustness under long-tailed distribution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 8659–8668.
[49] 
Zhang, F., Gao, C. et al. (2020). Convergence rates of variational posterior distributions. Annals of Statistics 48(4) 2180–2207. https://doi.org/10.1214/19-AOS1883. MR4134791
[50] 
Zhang, Q., Yu, S., Xin, J. and Chen, B. (2022). Multi-view information bottleneck without variational approximation. In ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing 4318–4322. IEEE.
[51] 
Zhang, W. E., Sheng, Q. Z., Alhazmi, A. and Li, C. (2020). Adversarial attacks on deep-learning models in natural language processing: A survey. ACM Transactions on Intelligent Systems and Technology 11(3) 1–41.
[52] 
Zheng, L., Sayed, T. and Mannering, F. (2021). Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions. Analytic Methods in Accident Research 29 100142.

Full article PDF XML
Full article PDF XML

Copyright
© 2026 New England Statistical Society
by logo by logo
Open access article under the CC BY license.

Keywords
Rare event modeling Variational inference Generative Bayesian models Extreme value theory

Metrics
since December 2021
18

Article info
views

9

Full article
views

9

PDF
downloads

5

XML
downloads

Export citation

Copy and paste formatted citation
Placeholder

Download citation in file


Share


RSS

The New England Journal of Statistics in Data Science

  • ISSN: 2693-7166
  • Copyright © 2021 New England Statistical Society

About

  • About journal

For contributors

  • Submit
  • OA Policy
  • Become a Peer-reviewer
Powered by PubliMill  •  Privacy policy