<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "JATS-journalpublishing1.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">NEJSDS</journal-id>
<journal-title-group><journal-title>The New England Journal of Statistics in Data Science</journal-title></journal-title-group>
<issn pub-type="ppub">2693-7166</issn><issn-l>2693-7166</issn-l>
<publisher>
<publisher-name>New England Statistical Society</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">NEJSDS071</article-id>
<article-id pub-id-type="doi">10.51387/26-NEJSDS071</article-id>
<article-categories><subj-group subj-group-type="heading">
<subject>Methodology Article</subject></subj-group><subj-group subj-group-type="area">
<subject>Engineering Science</subject></subj-group></article-categories>
<title-group>
<article-title>Variational Inference of Extremes for Rare Event Modeling: Theory and Applications</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Qian</surname><given-names>Chen</given-names></name><email xlink:href="mailto:chenqian@vt.edu">chenqian@vt.edu</email><xref ref-type="aff" rid="j_nejsds071_aff_001"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Tao</surname><given-names>Chenyang</given-names></name><email xlink:href="mailto:chenyang.tao@duke.edu">chenyang.tao@duke.edu</email><xref ref-type="aff" rid="j_nejsds071_aff_002"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Xu</surname><given-names>Jingbin</given-names></name><email xlink:href="mailto:jingbin@vt.edu">jingbin@vt.edu</email><xref ref-type="aff" rid="j_nejsds071_aff_003"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Guo</surname><given-names>Feng</given-names></name><email xlink:href="mailto:feng.guo@vt.edu">feng.guo@vt.edu</email><xref ref-type="aff" rid="j_nejsds071_aff_004"/><xref ref-type="corresp" rid="cor1">∗</xref>
</contrib>
<aff id="j_nejsds071_aff_001">Department of Statistics, <institution>Virginia Tech</institution>, <country>U.S.A</country>. E-mail address: <email xlink:href="mailto:chenqian@vt.edu">chenqian@vt.edu</email></aff>
<aff id="j_nejsds071_aff_002"><institution>Amazon</institution>, <country>U.S.A</country>. E-mail address: <email xlink:href="mailto:chenyang.tao@duke.edu">chenyang.tao@duke.edu</email></aff>
<aff id="j_nejsds071_aff_003">Department of Statistics, <institution>Virginia Tech</institution>, <country>U.S.A</country>. E-mail address: <email xlink:href="mailto:jingbin@vt.edu">jingbin@vt.edu</email></aff>
<aff id="j_nejsds071_aff_004">Department of Statistics, <institution>Virginia Tech</institution> and Virginia Tech Transportation Institute, <country>U.S.A</country>. E-mail address: <email xlink:href="mailto:feng.guo@vt.edu">feng.guo@vt.edu</email></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>∗</label>Corresponding author.</corresp>
</author-notes>
<pub-date pub-type="ppub"><year>2026</year></pub-date><pub-date pub-type="epub"><day>4</day><month>8</month><year>2026</year></pub-date><volume content-type="ahead-of-print">0</volume><issue>0</issue><fpage>1</fpage><lpage>19</lpage><history><date date-type="accepted"><day>31</day><month>7</month><year>2024</year></date></history>
<permissions><copyright-statement>© 2026 New England Statistical Society</copyright-statement><copyright-year>2026</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>Open access article under the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">CC BY</ext-link> license.</license-p></license></permissions>
<abstract>
<p>Severe event class imbalance is observed in many fields, such as insurance fraud, severe weather, traffic safety, and rare disease, and has far-reaching significance when the generalization of minority event classes is of primary interest. While existing solutions mostly focus on differential sampling or sample re-weighting approaches to alleviate the imbalance issue, we take a novel alternative view on promoting generalization. Our proposal is formulated under the generative Bayesian framework, positing that predictors are the stochastic proxies of latent causes with limited samples, whose exceedance leads to extreme events. To accurately capture the extended tail, our solution adopts the Gumbel distribution as prior and is modeled with a variational inference framework. Following the assertion that exceedance leads to extremes, we devise a disentangled additive monotonic neural architecture to predict the risk. The proposed model acknowledges representation uncertainties while embracing improved interpretability, generalization, and robustness. We provide theoretical insights to show the merits of the proposed approach. To verify the effectiveness in empirical settings, we conducted studies on various real-world data against the state-of-the-art counterparts, with encouraging results reported.</p>
</abstract>
<kwd-group>
<label>Keywords and phrases</label>
<kwd>Rare event modeling</kwd>
<kwd>Variational inference</kwd>
<kwd>Generative Bayesian models</kwd>
<kwd>Extreme value theory</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="j_nejsds071_s_001">
<label>1</label>
<title>Introduction</title>
<p>Rare event modeling (REM) deals with situations with an extremely low prevalence of events but influential consequences. Common examples of high social-economic value events include vehicle collisions in traffic safety [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>], autonomous vehicle testing [<xref ref-type="bibr" rid="j_nejsds071_ref_034">34</xref>], life-threatening rare diseases in healthcare, and fraudulent financial activities [<xref ref-type="bibr" rid="j_nejsds071_ref_032">32</xref>]. Accurate and robust modeling of rare events is critical for the identification of the relevant risk factors so as to predict, and hopefully prevent, associated negative outcomes via intervention. As an example, it is estimated that an adult driver in the U.S. will, on average, travel approximately 6.8 million miles before experiencing a single traffic crash [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>]. Given that traffic crashes are relatively rare events compared to normal driving, the precise identification and prediction of such incidents are of utmost importance for reducing fatalities and enhancing overall safety [<xref ref-type="bibr" rid="j_nejsds071_ref_039">39</xref>].</p>
<p>Characterized by severe event class imbalance and the lack of minority labels, rare event modeling falls outside the comfort zone of standard statistical approaches [<xref ref-type="bibr" rid="j_nejsds071_ref_046">46</xref>]. Without explicit statistical adjustments, the imbalance drives a learning agent to bias toward the majority class; at the same time, the absence of adequate minority examples causes unprotected models to wrongly capture spurious features that do not generalize. Without adequate generalization ability, the implemented REM model would substantially diminish effectiveness in detecting rare events, which are usually associated with risks that people aim to avoid. In the context of traffic crashes, inaccurate predictions by the REM model could potentially result in severe accidents occurring.</p>
<p>To establish reliability for REMs, the implemented model should have sufficient robustness to handle future out-of-sample predictions for the extreme event, i.e. the data which has not been exposed to during the training of REM. [<xref ref-type="bibr" rid="j_nejsds071_ref_048">48</xref>] point out that learning from less represented events is vulnerable to data perturbations. These generated misjudgments might lead to safety-critical situations. For instance, if practitioners employ the REM model to assess the safety functions of autonomous vehicles, including decisions about whether to execute maneuvers to prevent traffic crashes based on crash predictions, a less robust REM model could potentially permit self-driving cars to take unforeseen actions, possibly resulting in collisions [<xref ref-type="bibr" rid="j_nejsds071_ref_005">5</xref>].</p>
<p>Pioneered by the work of [<xref ref-type="bibr" rid="j_nejsds071_ref_023">23</xref>], there has been extensive research to overcome the potential pitfalls of rare-event modeling in the classical statistical regression framework. Prominent examples include coefficient bias correction [<xref ref-type="bibr" rid="j_nejsds071_ref_023">23</xref>] and penalized maximum likelihood estimation [<xref ref-type="bibr" rid="j_nejsds071_ref_014">14</xref>]. However, difficulties arise when applying such correction techniques in real-world settings: (<italic>i</italic>) the estimation efficiency for rare event modeling is dictated by the number of events of interest given fixed dimensions, and high-dimensional inputs render the conclusions unreliable and ungeneralizable [<xref ref-type="bibr" rid="j_nejsds071_ref_046">46</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_043">43</xref>]; (<inline-formula id="j_nejsds071_ineq_001"><alternatives><mml:math>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mi mathvariant="italic">i</mml:mi></mml:math><tex-math><![CDATA[$ii$]]></tex-math></alternatives></inline-formula>) heavy tails are commonly expected in rare event data, and most existing solutions fail to account for such characteristics [<xref ref-type="bibr" rid="j_nejsds071_ref_048">48</xref>].</p>
<p>Statistical re-sampling and re-weighting are the two most popular sample adjustment strategies to mitigate severe event class unbalancing issues. Re-sampling approaches alter the learner’s exposure frequency by over-sampling the minority class and down-sampling the majority class. The re-weighting approach adjusts the relative importance by assigning more weight to the less-represented examples. In line with these practices, recent work represented by [<xref ref-type="bibr" rid="j_nejsds071_ref_004">4</xref>] and [<xref ref-type="bibr" rid="j_nejsds071_ref_029">29</xref>] also considers modifications to derive class-sensitive loss that properly penalizes the minority event class, allowing the model to learn efficient representation without getting trapped in the complications associated with sample adjustments.</p>
<p>The re-sampling and re-weighting statistical approaches have limited capability to handle the situation when inputs are outside the norm. Re-sampling typically involves either over-sampling or under-sampling [<xref ref-type="bibr" rid="j_nejsds071_ref_003">3</xref>]. The former is often subjected to the loss of estimation efficiency [<xref ref-type="bibr" rid="j_nejsds071_ref_046">46</xref>], the latter is often associated with compromised generalization [<xref ref-type="bibr" rid="j_nejsds071_ref_004">4</xref>]. The re-weighting schemes are often criticized for the numerical instability [<xref ref-type="bibr" rid="j_nejsds071_ref_035">35</xref>]. Recent works by [<xref ref-type="bibr" rid="j_nejsds071_ref_004">4</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_029">29</xref>] modify the hinge loss and entropy loss to capitalize the minority class prediction. However, research points out this might result in overfitting issues due to training bias and label noise [<xref ref-type="bibr" rid="j_nejsds071_ref_035">35</xref>].</p>
<p>An alternative way to overcome the difficulties of rare event modeling is to solicit inductive bias, thereby imposing structural constraints to suppress the spurious associations and identify robust predictive features. Many such models are framed under the few-shot learning setup, capitalizing on its superior ability to generalize in the small sample regime. These models often operate under relatively strong assumptions that the majority event classes sufficiently characterize the feature encoder and meta-predictors, leaving only a handful of parameters to tune for the minority classes, thereby boosting efficiency and generalization. Sharing an aggregating knowledge across different learning tasks defined in majority classes contributes significantly to the success of few-shot learning on generalization [<xref ref-type="bibr" rid="j_nejsds071_ref_013">13</xref>]. Few-shot learning relies on the assumption that all classes in the data share common features and leverages knowledge transfer from majority classes to enhance learning for minority classes. However, in reality, this assumption may not always hold true, as different classes can have varying data generation processes. For example, traffic crashes and normal driving segments often display distinct characteristics [<xref ref-type="bibr" rid="j_nejsds071_ref_016">16</xref>]. The failure to meet the assumptions of few-shot learning methods can potentially lead to a degradation in their performance [<xref ref-type="bibr" rid="j_nejsds071_ref_011">11</xref>].</p>
<p>Adding to the difficulty is the challenge of discerning spurious correlations as rare events do not always associate with stable, recognizable patterns. Instead, feature irregularities are common for rare events, which makes it unattainable to confidently model the rare event behavior. In such scenarios, rare event modeling can be formulated by “establishing the norm” using the majority of examples and to test if input is out of the ordinary, i.e., the anomaly detection approach with abundant tools developed based on various heuristics [<xref ref-type="bibr" rid="j_nejsds071_ref_008">8</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_037">37</xref>]. Unsupervised anomaly detection approaches often involve density modeling, which has been a long-standing challenge for high-dimensional inputs [<xref ref-type="bibr" rid="j_nejsds071_ref_008">8</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_037">37</xref>]</p>
<p>Despite the varying degree of empirical success of existing rare event modeling schemes, a few limitations have been largely overlooked. First, assumptions made by different modeling schemes are often at odds, implying performance depends on whether underlying assumptions have been satisfied. Second, representation has not been sufficiently valued in existing research, potentially taking a toll on the minority class generalization. In light of the above, it is imperative to develop robust and interpretable learning strategies that can fully recognize the unique characteristics of rare-event data.</p>
<p>In recognition of the limitations of current rare-event modeling discussed above, we present a novel learning framework, named <italic>Variational Inference for Rare Event Modeling</italic> (VI-REM), that explicitly addresses some of the weaknesses of existing schemes. Our main contributions include (<italic>i</italic>) formulation of a variational representation learning scheme based on information theoretic model, allowing disentangled extreme representations for rare-event prediction; (<inline-formula id="j_nejsds071_ineq_002"><alternatives><mml:math>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mi mathvariant="italic">i</mml:mi></mml:math><tex-math><![CDATA[$ii$]]></tex-math></alternatives></inline-formula>) development of a robust and interpretable prediction arm that joins the strength of a generalized additive model and an isotonic neural net; (<inline-formula id="j_nejsds071_ineq_003"><alternatives><mml:math>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mi mathvariant="italic">i</mml:mi></mml:math><tex-math><![CDATA[$iii$]]></tex-math></alternatives></inline-formula>) presentation of theoretical insights that justify the empirical gains through multiple numerical experiments of VI-REM.</p>
</sec>
<sec id="j_nejsds071_s_002">
<label>2</label>
<title>Background</title>
<p>Denote the input features as <inline-formula id="j_nejsds071_ineq_004"><alternatives><mml:math>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\mathbf{x}\in {\mathbb{R}^{d}}$]]></tex-math></alternatives></inline-formula>, the latent variables as <inline-formula id="j_nejsds071_ineq_005"><alternatives><mml:math>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\mathbf{z}\in {\mathbb{R}^{k}}$]]></tex-math></alternatives></inline-formula>, and class label <inline-formula id="j_nejsds071_ineq_006"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo fence="true" stretchy="false">}</mml:mo></mml:math><tex-math><![CDATA[$y\in \{0,1\}$]]></tex-math></alternatives></inline-formula> for which <inline-formula id="j_nejsds071_ineq_007"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$y=1$]]></tex-math></alternatives></inline-formula> represents the rare-event of interest with probability of occurrence <italic>p</italic>, and <inline-formula id="j_nejsds071_ineq_008"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$y=0$]]></tex-math></alternatives></inline-formula> as <italic>Normality</italic>. Training instances are given by <italic>n</italic> paired samples <inline-formula id="j_nejsds071_ineq_009"><alternatives><mml:math>
<mml:mi mathvariant="script">D</mml:mi>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo fence="true" stretchy="false">}</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[$\mathcal{D}={\{{\mathbf{x}^{(i)}},{y^{(i)}}\}_{i=1}^{n}}$]]></tex-math></alternatives></inline-formula> with sample size <inline-formula id="j_nejsds071_ineq_010"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${n_{1}}$]]></tex-math></alternatives></inline-formula> in class <italic>Event</italic> and <inline-formula id="j_nejsds071_ineq_011"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">n</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${n_{0}}=n-{n_{1}}$]]></tex-math></alternatives></inline-formula> in class <italic>Normality</italic>. The goal is to predict the conditional likelihood of an event given the features, i.e., <inline-formula id="j_nejsds071_ineq_012"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y|\mathbf{x})$]]></tex-math></alternatives></inline-formula> with accuracy and robustness. The rare-event problem arises when <inline-formula id="j_nejsds071_ineq_013"><alternatives><mml:math>
<mml:mi mathvariant="italic">p</mml:mi>
<mml:mo stretchy="false">≪</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$p\ll 1$]]></tex-math></alternatives></inline-formula>.</p>
<sec id="j_nejsds071_s_003">
<label>2.1</label>
<title>Information Theoretic Generative Model</title>
<p>The generative Bayesian model has gained popularity due to its scalability, interpretability, and generalizability [<xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>]. Under the full Bayesian setup, the generative process for the data likelihood of <italic>y</italic> given <bold>x</bold> can be obtained by marginalization over latent variables <bold>z</bold>: 
<disp-formula id="j_nejsds071_eq_001">
<label>(2.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∫</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mtext>y</mml:mtext>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mtext>y</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \Pr (y|\mathbf{x})=\frac{\textstyle\int \Pr (\text{y},\mathbf{x},\mathbf{z})d\mathbf{z}}{\Pr (\mathbf{x})}=\int \Pr (\text{y}|\mathbf{x},\mathbf{z})\Pr (\mathbf{z}|\mathbf{x})d\mathbf{z}\]]]></tex-math></alternatives>
</disp-formula> 
A common assumption in such generative process is <bold>z</bold> contains sufficient information of <bold>x</bold> for <italic>y</italic>, i.e <inline-formula id="j_nejsds071_ineq_014"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mtext>y</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mtext>y</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\text{y}|\mathbf{x},\mathbf{z})=\Pr (\text{y}|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. Similar assumption can be found in [<xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>]. With this, Equation <xref rid="j_nejsds071_eq_001">2.1</xref> can be simplified as 
<disp-formula id="j_nejsds071_eq_002">
<label>(2.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mtext>y</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \Pr (y|\mathbf{x})=\int \Pr (\text{y}|\mathbf{z})\Pr (\mathbf{z}|\mathbf{x})d\mathbf{z}\]]]></tex-math></alternatives>
</disp-formula> 
For REM which belongs to supervised learning tasks, the core of the generative Bayesian model is to reconstruct the conditional probability of response variable <italic>y</italic> given predictors <bold>x</bold>, which is <inline-formula id="j_nejsds071_ineq_015"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y|\mathbf{x})$]]></tex-math></alternatives></inline-formula>, with high accuracy [<xref ref-type="bibr" rid="j_nejsds071_ref_046">46</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_043">43</xref>].</p>
<p>For REM, maximizing over Equation <xref rid="j_nejsds071_eq_002">2.2</xref> is challenging because of the computational intractability of the integration as well as the difficulty to find the latent <bold>z</bold> given high-dimensional <bold>x</bold>.</p>
<p>To address the computational intractability issue for latent <bold>z</bold>, the Variational Information Bottleneck (VIB) [<xref ref-type="bibr" rid="j_nejsds071_ref_001">1</xref>] utilizes the information-theoretic methods to find a meaningful representation <bold>z</bold> from <bold>x</bold> by maximizing the following equations: 
<disp-formula id="j_nejsds071_eq_003">
<label>(2.3)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">arg</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">max</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi mathvariant="italic">I</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mi mathvariant="italic">I</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \arg \underset{\mathbf{z}}{\max }I(\mathbf{z},y)-\beta I(\mathbf{z},\mathbf{x})\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_016"><alternatives><mml:math>
<mml:mi mathvariant="italic">I</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$I(\mathbf{z},y)$]]></tex-math></alternatives></inline-formula> measures the mutual information between <bold>z</bold> and <italic>y</italic>, and the parameter <inline-formula id="j_nejsds071_ineq_017"><alternatives><mml:math>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\beta \gt 0$]]></tex-math></alternatives></inline-formula> is a hyper parameter.</p>
<p>Under the generative Bayesian setup, maximizing Equation <xref rid="j_nejsds071_eq_004">2.4</xref> is identical to minimize the following loss function based on Variational Inference (VI) [<xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>]: 
<disp-formula id="j_nejsds071_eq_004">
<label>(2.4)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mtext>Loss</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>VIB</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">[</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">]</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}{\text{Loss}_{\text{VIB}}}=& -{\mathcal{D}_{\mathrm{KL}}}\Big({q_{\phi }}(\mathbf{z}|\mathbf{x})\| \Pr (\mathbf{z})\Big)\\ {} & +\beta {\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}\Big[\log {p_{\theta }}(\mathbf{y}|\mathbf{z})\Big]\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_018"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})$]]></tex-math></alternatives></inline-formula> denotes a prior distribution on <bold>z</bold>, <inline-formula id="j_nejsds071_ineq_019"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\theta }}(\mathbf{y}|\mathbf{z})$]]></tex-math></alternatives></inline-formula> denotes a conditional distribution of <bold>y</bold> given <bold>z</bold> parameterized by <italic>θ</italic>, and <inline-formula id="j_nejsds071_ineq_020"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> is a family of distributions parameterized by <italic>ϕ</italic> to approximate the true posterior density [<xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>].</p>
<p>The first term <inline-formula id="j_nejsds071_ineq_021"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mathcal{D}_{\mathrm{KL}}}({q_{\phi }}(\mathbf{z}|\mathbf{x})\| \Pr (\mathbf{z}))$]]></tex-math></alternatives></inline-formula> is the KL divergence between <inline-formula id="j_nejsds071_ineq_022"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_023"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula>. It shows the complexity of the posterior distribution. Smaller values mean that <inline-formula id="j_nejsds071_ineq_024"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> is similar to <inline-formula id="j_nejsds071_ineq_025"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})$]]></tex-math></alternatives></inline-formula>, which implies a simpler posterior. The second term <inline-formula id="j_nejsds071_ineq_026"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[${\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}[\log {p_{\theta }}(\mathbf{y}|\mathbf{z})]$]]></tex-math></alternatives></inline-formula> shows how well latent <bold>z</bold> can reconstruct label <bold>y</bold>, i.e. the task performance.</p>
<p>The parameter <italic>β</italic> is used to keep a balance between the performance on the task and the complexity of the latent representation <bold>z</bold> [<xref ref-type="bibr" rid="j_nejsds071_ref_050">50</xref>]. If <inline-formula id="j_nejsds071_ineq_027"><alternatives><mml:math>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$\beta =1$]]></tex-math></alternatives></inline-formula>, then Variational Information Bottleneck in Equation <xref rid="j_nejsds071_eq_004">2.4</xref> is the same as maximizing the evidence lower bound (ELBO) of <inline-formula id="j_nejsds071_ineq_028"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y|\mathbf{x})$]]></tex-math></alternatives></inline-formula>, given latent variable <bold>z</bold> [<xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>]; if <inline-formula id="j_nejsds071_ineq_029"><alternatives><mml:math>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mo stretchy="false">≠</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$\beta \ne 1$]]></tex-math></alternatives></inline-formula> Variational Information Bottleneck will be the <italic>β</italic>-Variational Auto Encoder (VAE) [<xref ref-type="bibr" rid="j_nejsds071_ref_021">21</xref>] in the version of supervised learning.</p>
<p>For supervised learning tasks, maximizing over the VIB objective can improve model’s generalizability [<xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>], because the <inline-formula id="j_nejsds071_ineq_030"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[${\mathcal{D}_{\mathrm{KL}}}\Big({q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x},\mathbf{y})\| \Pr (\mathbf{z})\Big)$]]></tex-math></alternatives></inline-formula> regularizes the Kullback–Leibler (KL) divergence between approximated posterior <inline-formula id="j_nejsds071_ineq_031"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x},\mathbf{y})$]]></tex-math></alternatives></inline-formula> and prior distribution <inline-formula id="j_nejsds071_ineq_032"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})$]]></tex-math></alternatives></inline-formula> and helps to prevent <inline-formula id="j_nejsds071_ineq_033"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x},\mathbf{y})$]]></tex-math></alternatives></inline-formula> being overfitted because of simplicity in prior distribution. The generative Bayesian model also shows superior robustness over the deterministic models, especially in handling testing data with considerable noises [<xref ref-type="bibr" rid="j_nejsds071_ref_001">1</xref>].</p>
</sec>
<sec id="j_nejsds071_s_004">
<label>2.2</label>
<title>Challenges and Motivation</title>
<p>When applying the generative Bayesian model for REM, there are three challenges from statistical perspectives. The first challenge is setting an appropriate prior <inline-formula id="j_nejsds071_ineq_034"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})$]]></tex-math></alternatives></inline-formula> for the less represented events. Within the current state of the art (SOTA) approaches, most work uses the isotropic Gaussian [<xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>] or non-parametric based approaches [<xref ref-type="bibr" rid="j_nejsds071_ref_045">45</xref>]. However, neither of them could be a good fit in REM, where the events of interest are always associated with heavy tail and rarity of observations.</p>
<p>The second challenge in REM comes from the model’s generalization ability to unseen data [<xref ref-type="bibr" rid="j_nejsds071_ref_048">48</xref>], as the events of interest do not always associate with robust patterns. The disentanglement learning approach proves its effectiveness in representing information with stability, especially for the patterns that are less stable [<xref ref-type="bibr" rid="j_nejsds071_ref_002">2</xref>]. Promising work includes generalized additive models (GAM) [<xref ref-type="bibr" rid="j_nejsds071_ref_019">19</xref>] and <italic>β</italic>-Variational Auto Encoder (VAE) [<xref ref-type="bibr" rid="j_nejsds071_ref_021">21</xref>]. A feasible approach for disentanglement is to assume sets of statistically independent variables for modeling [<xref ref-type="bibr" rid="j_nejsds071_ref_007">7</xref>]. In this work, we follow the work in [<xref ref-type="bibr" rid="j_nejsds071_ref_021">21</xref>] to assume sets of statistically independent variables assuring the model’s disentanglement in REM. This setup helps to decouple the representations into individual features that each capture a unique aspect of data.</p>
<p>The third challenge arises from how we can build an interpretable prediction model. Isotonic regression, also known as monotonic regression, is an important class of constrained regression techniques that works under the assumption of monotonicity [<xref ref-type="bibr" rid="j_nejsds071_ref_018">18</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_047">47</xref>]. Research has verified its appropriateness in preventing overfitting and enhancing model interpretability [<xref ref-type="bibr" rid="j_nejsds071_ref_018">18</xref>]. Specifically, for a scalar predictor <italic>x</italic>, the relationship of <italic>x</italic> toward <inline-formula id="j_nejsds071_ineq_035"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$y=g(x)$]]></tex-math></alternatives></inline-formula> is assumed to be non-decreasing with regard to <italic>x</italic>; i.e., <inline-formula id="j_nejsds071_ineq_036"><alternatives><mml:math>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$g({x_{1}})\le g({x_{2}})$]]></tex-math></alternatives></inline-formula> if <inline-formula id="j_nejsds071_ineq_037"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${x_{1}}\le {x_{2}}$]]></tex-math></alternatives></inline-formula>. Such relation is ubiquitous in practice, and monotonicity allows extra flexibility compared to linear relations yet retains a similar level of interpretability of the predictive features and modeling robustness. In this paper, we leverage the recent introduction of the unconstrained monotonic neural net (UMNN) [<xref ref-type="bibr" rid="j_nejsds071_ref_047">47</xref>] to approximate flexible mapping functions in the generative Bayesian model.</p>
</sec>
</sec>
<sec id="j_nejsds071_s_005" sec-type="methods">
<label>3</label>
<title>Methodology</title>
<p>We formulate VI-REM under the latent variable model framework, in which the event class labels can be predicted from the latent <bold>z</bold>. The key idea of the VI-REM is to amortize the difficulty of direct prediction of rare-events to the representation learning stage. The premise is the extreme latent representations lead to extreme events. The approach partitions the representation space into normal and extreme regions, where the prevalence of target events in the latter far exceeds those in the former. This alleviates the majority bias issue that afflicts conventional schemes, as the event distribution is more balanced for the extreme region.</p>
<sec id="j_nejsds071_s_006">
<label>3.1</label>
<title>Extreme Prior Distribution</title>
<p>The Extreme Value Theory (EVT) studies the statistical behavior of extremes [<xref ref-type="bibr" rid="j_nejsds071_ref_009">9</xref>]. To model the extreme behavior of statistical features, <italic>maximum model</italic> within the domain of EVT describes the maximal value in a sequence of independent observations. Modeling extremes poses a challenge due to the limited number of observations. However, extreme value theory (EVT) asserts that, under mild conditions [<xref ref-type="bibr" rid="j_nejsds071_ref_009">9</xref>], extreme values actually belong to the Generalized Extreme Value (GEV) family of distributions. For a sequence of independent and identically distributed random variables, denoted as <inline-formula id="j_nejsds071_ineq_038"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">N</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${X_{1}},{X_{2}},\dots ,{X_{N}}$]]></tex-math></alternatives></inline-formula>, the asymptotic distribution of the maximum value <inline-formula id="j_nejsds071_ineq_039"><alternatives><mml:math>
<mml:mo movablelimits="false">max</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo movablelimits="false">…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\max ({X_{1}},{X_{2}},\dots ,{X_{N}})$]]></tex-math></alternatives></inline-formula> follows: 
<disp-formula id="j_nejsds071_eq_005">
<label>(3.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">lim</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">N</mml:mi>
<mml:mo stretchy="false">→</mml:mo>
<mml:mi>∞</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">max</mml:mo>
<mml:mfenced separators="" open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo movablelimits="false">…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">N</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mtext>GEV</mml:mtext>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">ξ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \underset{N\to \infty }{\lim }\Pr \big(\max \left\{{X_{1}},\dots ,{X_{N}}\right\}\le x\big)=\text{GEV}({\mu _{1}},{\sigma _{1}},\xi )\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_040"><alternatives><mml:math>
<mml:mtext>GEV</mml:mtext>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">ξ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\text{GEV}({\mu _{1}},{\sigma _{1}},\xi )$]]></tex-math></alternatives></inline-formula> denotes a GEV distribution with <inline-formula id="j_nejsds071_ineq_041"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mu _{1}}$]]></tex-math></alternatives></inline-formula> as the location, <inline-formula id="j_nejsds071_ineq_042"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\sigma _{1}}$]]></tex-math></alternatives></inline-formula> as the scale and <italic>ξ</italic> as the extreme value index. Equation <xref rid="j_nejsds071_eq_005">3.1</xref> is known as the Fisher–Tippett–Gnedenko theorem [<xref ref-type="bibr" rid="j_nejsds071_ref_009">9</xref>], allowing for accurate modeling of extreme values with very limited examples.</p>
<p>The GEV has three forms depending on the parameter <italic>ξ</italic>: Fréchet distribution (<inline-formula id="j_nejsds071_ineq_043"><alternatives><mml:math>
<mml:mi mathvariant="italic">ξ</mml:mi>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\xi \gt 0$]]></tex-math></alternatives></inline-formula>), Gumbel distribution(<inline-formula id="j_nejsds071_ineq_044"><alternatives><mml:math>
<mml:mi mathvariant="italic">ξ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\xi =0$]]></tex-math></alternatives></inline-formula>), and reversed Weibull distribution (<inline-formula id="j_nejsds071_ineq_045"><alternatives><mml:math>
<mml:mi mathvariant="italic">ξ</mml:mi>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\xi \lt 0$]]></tex-math></alternatives></inline-formula>). The Fréchet and Weibull distributions have bounded support, e.g. <inline-formula id="j_nejsds071_ineq_046"><alternatives><mml:math>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ξ</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle></mml:math><tex-math><![CDATA[$x\gt {\mu _{1}}-\frac{{\sigma _{1}}}{\xi }$]]></tex-math></alternatives></inline-formula> for Fréchet distribution. To cover the whole space, we use the Gumbel distribution, denoted as <inline-formula id="j_nejsds071_ineq_047"><alternatives><mml:math>
<mml:mi mathvariant="script">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{G}({\mu _{1}},{\sigma _{1}})$]]></tex-math></alternatives></inline-formula>, as the prior because the support for Gumbel is <inline-formula id="j_nejsds071_ineq_048"><alternatives><mml:math>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi></mml:math><tex-math><![CDATA[$x\in \mathbb{R}$]]></tex-math></alternatives></inline-formula>.</p>
<p>In VI-REM, we assume extreme latent in the representation space lead to the happening of <italic>Event</italic> cases. As these extreme representation usually deviate from the center of normality, we leverage <italic>maximum model</italic> in Equation <xref rid="j_nejsds071_eq_005">3.1</xref> to represent these extreme cases. We extend standard Gaussian prior [<xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>] via incorporating a Gumbel based prior to accommodate heavy tails. We model the normal regular representations with a Gaussian distribution, and the extreme representation with a Gumbel distribution. We assume the representation space of latent variable of <italic>z</italic> can be partitioned into the into normal and extreme regions. So that the probability density can be: 
<disp-formula id="j_nejsds071_eq_006">
<label>(3.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mtext mathvariant="italic">Normality</mml:mtext>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>×</mml:mo>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo stretchy="false">∉</mml:mo>
<mml:mtext mathvariant="italic">Normality</mml:mtext>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>×</mml:mo>
<mml:mi mathvariant="script">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\Pr (z)=\Pr (z\in & \textit{Normality})\times \mathcal{N}({\mu _{2}},{\sigma _{2}})\\ {} & +\Pr (z\notin \textit{Normality})\times \mathcal{G}({\mu _{1}},{\sigma _{1}})\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_049"><alternatives><mml:math>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{N}({\mu _{2}},{\sigma _{2}})$]]></tex-math></alternatives></inline-formula> is a Gaussian distribution with parameter <inline-formula id="j_nejsds071_ineq_050"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mu _{2}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_051"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\sigma _{2}}$]]></tex-math></alternatives></inline-formula>. For simplicity, we set <inline-formula id="j_nejsds071_ineq_052"><alternatives><mml:math>
<mml:mi mathvariant="italic">λ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mtext mathvariant="italic">Normality</mml:mtext>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\lambda =\Pr (z\in \textit{Normality})$]]></tex-math></alternatives></inline-formula>. We define the extreme prior (EP) as a weighted mixture of a Gumbel distribution and Gaussian distribution in equation <xref rid="j_nejsds071_eq_006">3.2</xref>, and its associated probability density function is formulated as: 
<disp-formula id="j_nejsds071_eq_007">
<label>(3.3)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">π</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true">{</mml:mo>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true">}</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true">{</mml:mo>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true">}</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}{p_{\boldsymbol{\kappa }}^{\text{EP}}}(z)=& \frac{\lambda }{\sqrt{2\pi {\sigma _{2}^{2}}}}\exp \Bigg\{-\frac{{(z-{\mu _{2}})^{2}}}{2{\sigma _{2}^{2}}}\Bigg\}\\ {} & +\frac{1-\lambda }{{\sigma _{1}}}\exp \Bigg\{-\frac{z-{\mu _{1}}}{{\sigma _{1}}}-\exp \Big(-\frac{z-{\mu _{1}}}{{\sigma _{1}}}\Big)\Bigg\}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<fig id="j_nejsds071_fig_001">
<label>Figure 1</label>
<caption>
<p>Illustration of the Extreme Prior distribution. <inline-formula id="j_nejsds071_ineq_053"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_054"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{2}}$]]></tex-math></alternatives></inline-formula> are simulated independently in the figure. (a) The scatter plot for different EP distributions and the red points come from the Gumbel distribution part and the green points come from the Normal distribution part. For case 1, both of the latent variables <inline-formula id="j_nejsds071_ineq_055"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∼</mml:mo>
<mml:mo movablelimits="false">EP</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>10</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.995</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${z_{1}},{z_{2}}\sim \operatorname{EP}(10,2,0.995,0,2)$]]></tex-math></alternatives></inline-formula>. (b) We set both the latent factors <inline-formula id="j_nejsds071_ineq_056"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∼</mml:mo>
<mml:mo movablelimits="false">EP</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>10</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.995</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${z_{1}},{z_{2}}\sim \operatorname{EP}(10,0.2,0.995,0,2)$]]></tex-math></alternatives></inline-formula> and as <inline-formula id="j_nejsds071_ineq_057"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_058"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{2}}$]]></tex-math></alternatives></inline-formula> comes from the same distribution, we only show the results for the probability density estimation for <inline-formula id="j_nejsds071_ineq_059"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula>. We consider the EP distribution, kernel density estimation (a non-parametric approach), and Normal distribution to fit the distribution of <inline-formula id="j_nejsds071_ineq_060"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula>. The black line represents the estimated entire distribution functions, and the blue displays the distribution function for the tailed part.</p>
</caption>
<graphic xlink:href="nejsds071_g001.jpg"/>
</fig>
<p>To induce more flexibility into the model, we set the parameters <inline-formula id="j_nejsds071_ineq_061"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">λ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\boldsymbol{\kappa }=(\lambda ,{\mu _{1}},{\sigma _{1}},{\mu _{2}},{\sigma _{2}})$]]></tex-math></alternatives></inline-formula> of EP prior are all learnable during the optimization process. The parameters <inline-formula id="j_nejsds071_ineq_062"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mu _{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_063"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mu _{2}}$]]></tex-math></alternatives></inline-formula> controls the center of Normal and Gumbel distribution. The parameter <inline-formula id="j_nejsds071_ineq_064"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\sigma _{1}}$]]></tex-math></alternatives></inline-formula> controls the shape of distribution tail. Figure <xref rid="j_nejsds071_fig_001">1</xref>a displays the scatter plot for different EP distributions. Theorem <xref rid="j_nejsds071_stat_001">3.1</xref> states that as long as <inline-formula id="j_nejsds071_ineq_065"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[${\sigma _{1}},{\sigma _{2}}\gt 0$]]></tex-math></alternatives></inline-formula>, the proposed EP prior is a proper prior. If the tailed behavior does not actually lead to rarity, the EP can fall back to a standard Gaussian distribution.</p><statement id="j_nejsds071_stat_001"><label>Proposition 3.1.</label>
<p>The proposed EP distribution is a proper prior, satisfying: 
<disp-formula id="j_nejsds071_eq_008">
<label>(3.4)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="italic">R</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\int _{z\in R}}{p_{\boldsymbol{\kappa }}^{\text{EP}}}(z)dz=1\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_066"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">λ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\boldsymbol{\kappa }=(\lambda ,{\mu _{1}},{\sigma _{1}},{\mu _{2}},{\sigma _{2}})$]]></tex-math></alternatives></inline-formula> are the parameters for the EP prior and <inline-formula id="j_nejsds071_ineq_067"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[${\sigma _{1}},{\sigma _{2}}\gt 0$]]></tex-math></alternatives></inline-formula>.</p></statement>
<p>Figure 1b compares the performance of different estimators in modeling the tail part of the distributions. The fitted Normal distribution considerably underestimates the tail portion, while non-parametric kernel density estimation tends to overfit the observed samples. In scenarios involving heavy-tailed distributions, which are commonly observed in risk-related analyses [<xref ref-type="bibr" rid="j_nejsds071_ref_009">9</xref>], traditional parametric methods (such as fitted Normal distribution) and non-parametric methods (like kernel density estimation) may struggle to accurately capture the tail probabilities due to the limited number of rare events. However, extreme value theory-based estimation (EP distribution) excels in handling heavy-tailed scenarios effectively. It not only addresses the challenges posed by limited rare events but also demonstrates substantial generalization ability in accurately describing the tails of distributions.</p>
<p>EP distribution can be generalized to multi-dimensional latent components by applying the above setup to each individual dimension. We assume mutual independence among multi-dimensional latents <inline-formula id="j_nejsds071_ineq_068"><alternatives><mml:math>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\mathbf{z}\in {\mathbb{R}^{k}}$]]></tex-math></alternatives></inline-formula>, so as the EP distribution for <bold>z</bold> with parameter <inline-formula id="j_nejsds071_ineq_069"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\boldsymbol{\omega }\in {\mathbb{R}^{5k}}$]]></tex-math></alternatives></inline-formula> can be extended as where <inline-formula id="j_nejsds071_ineq_070"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi></mml:math><tex-math><![CDATA[${\mathbf{z}_{i}}\in \mathbb{R}$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_009">
<label>(3.5)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∏</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {p_{\boldsymbol{\omega }}^{\text{EP}}}(\mathbf{z})={\prod \limits_{i=1}^{k}}{p_{{\boldsymbol{\kappa }_{i}}}^{\text{EP}}}({\mathbf{z}_{i}})\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_071"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\boldsymbol{\kappa }_{i}}=({\lambda _{i}},{\mu _{1,i}},{\sigma _{1,i}},{\mu _{2,i}},{\sigma _{2,i}})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_072"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\boldsymbol{\omega }=({\boldsymbol{\kappa }_{1}},\dots ,{\boldsymbol{\kappa }_{k}})$]]></tex-math></alternatives></inline-formula>.</p>
<p>The decomposition above assumes that the prior distribution is independent, which implies a possible disentanglement of the latent factor <bold>z</bold> [<xref ref-type="bibr" rid="j_nejsds071_ref_021">21</xref>]. One can also consider dependent priors for <bold>z</bold>. For example, a Gaussian process-based prior <inline-formula id="j_nejsds071_ineq_073"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn mathvariant="bold">0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="script">K</mml:mi>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})=\mathcal{N}(\mathbf{0},\mathcal{K}zz)$]]></tex-math></alternatives></inline-formula> with a kernel matrix <inline-formula id="j_nejsds071_ineq_074"><alternatives><mml:math>
<mml:mi mathvariant="script">K</mml:mi>
<mml:mi mathvariant="italic">z</mml:mi>
<mml:mi mathvariant="italic">z</mml:mi></mml:math><tex-math><![CDATA[$\mathcal{K}zz$]]></tex-math></alternatives></inline-formula> [<xref ref-type="bibr" rid="j_nejsds071_ref_006">6</xref>] to account for dependence in the prior distribution, or to introduce a hierarchical structure to the prior with <inline-formula id="j_nejsds071_ineq_075"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∏</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∣</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>−</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (\mathbf{z})=\Pr ({\mathbf{z}_{1}}){\textstyle\prod _{i=2}^{k}}\Pr ({\mathbf{z}_{i}}\mid {\mathbf{z}_{i-1}})$]]></tex-math></alternatives></inline-formula> [<xref ref-type="bibr" rid="j_nejsds071_ref_042">42</xref>]. However, more complex priors will increase the number of parameters to be optimized, thus to increase the overfitting risk in REM [<xref ref-type="bibr" rid="j_nejsds071_ref_045">45</xref>]. We focus on the independent prior in this work, and defer more flexible priors to future research.</p>
</sec>
<sec id="j_nejsds071_s_007">
<label>3.2</label>
<title>Monotonic Additive Neural Network</title>
<p>To facilitate interpretability and generalization for REM, it is beneficial to impose constraints on <inline-formula id="j_nejsds071_ineq_076"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. We combine two prominent techniques: the generalized additive model (GAM) [<xref ref-type="bibr" rid="j_nejsds071_ref_019">19</xref>] and isotonic regression for <inline-formula id="j_nejsds071_ineq_077"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>, and we name it as Monotonic Additive Neural Network (MANN). Let <italic>y</italic> come from a Bernoulli(<italic>p</italic>) distribution and <inline-formula id="j_nejsds071_ineq_078"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> be the Logit function and the <inline-formula id="j_nejsds071_ineq_079"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> can be the following: 
<disp-formula id="j_nejsds071_eq_010">
<label>(3.6)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {p_{\boldsymbol{\theta }}}(y|\mathbf{z})=\frac{1}{1+\exp \Big(-{\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})\Big)}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>We assume sets of statistically independent variables [<xref ref-type="bibr" rid="j_nejsds071_ref_007">7</xref>] in MANN to assure model’s disentanglement. Each <inline-formula id="j_nejsds071_ineq_080"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{j}}$]]></tex-math></alternatives></inline-formula> takes input from the <italic>j</italic>-th dimension of the latent to model the generation of <italic>y</italic>. This additive decomposition encourages disentangled representation. The Logit function <inline-formula id="j_nejsds071_ineq_081"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> is based on GAM: 
<disp-formula id="j_nejsds071_eq_011">
<label>(3.7)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:mfenced>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})=\log \left(\frac{{p_{\boldsymbol{\theta }}}(y|\mathbf{z})}{1-{p_{\boldsymbol{\theta }}}(y|\mathbf{z})}\right)=\alpha (\boldsymbol{\theta })+{\sum \limits_{j=1}^{k}}{f_{j}}({\mathbf{z}_{j}})\]]]></tex-math></alternatives>
</disp-formula> 
where each of the individual function is in Equation <xref rid="j_nejsds071_eq_012">3.8</xref>, <inline-formula id="j_nejsds071_ineq_082"><alternatives><mml:math>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>.</mml:mo>
<mml:mi mathvariant="italic">k</mml:mi></mml:math><tex-math><![CDATA[$j=1,..k$]]></tex-math></alternatives></inline-formula> and <italic>k</italic> is the dimension of latent variables. 
<disp-formula id="j_nejsds071_eq_012">
<label>(3.8)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">β</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">t</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:mi mathvariant="italic">t</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {f_{j}}({\mathbf{z}_{j}})={\beta _{j}}(\theta ){\int _{0}^{{\mathbf{z}_{j}}}}\exp \Big({h_{j}}(t;\boldsymbol{\theta })\Big)\mathrm{d}t\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>It is easy to see all <inline-formula id="j_nejsds071_ineq_083"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{j}}$]]></tex-math></alternatives></inline-formula> are monotonic, with the direction dictated by the sign of <inline-formula id="j_nejsds071_ineq_084"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">β</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\beta _{j}}$]]></tex-math></alternatives></inline-formula>. We model <inline-formula id="j_nejsds071_ineq_085"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">t</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${h_{j}}(t;\boldsymbol{\theta })$]]></tex-math></alternatives></inline-formula> above using MLP based deep neural networks, and compute <inline-formula id="j_nejsds071_ineq_086"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${f_{j}}({\mathbf{z}_{j}})$]]></tex-math></alternatives></inline-formula> via Clenshaw-Curtis numerical integration [<xref ref-type="bibr" rid="j_nejsds071_ref_047">47</xref>].</p>
<p>We use a toy examples to illustrate the advantages of monotonicity in dealing with REM. We set <italic>y</italic> from a Bernoulli distribution and the conditional probability follows 
<disp-formula id="j_nejsds071_eq_013">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>−</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \Pr (y|{z_{1}},{z_{2}})=\frac{1}{1+\exp (-{z_{1}^{3}}-0.5{z_{2}^{3}})}\]]]></tex-math></alternatives>
</disp-formula> 
where the random variables <inline-formula id="j_nejsds071_ineq_087"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∼</mml:mo>
<mml:mo movablelimits="false">EP</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.99</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${z_{1}}\sim \operatorname{EP}(2,0.5,0.99,-2,0.5)$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_088"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∼</mml:mo>
<mml:mo movablelimits="false">EP</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.99</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>−</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${z_{2}}\sim \operatorname{EP}(1,0.5,0.99,-1,0.5)$]]></tex-math></alternatives></inline-formula>. Through this, the probability for <inline-formula id="j_nejsds071_ineq_089"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mn>0.02</mml:mn></mml:math><tex-math><![CDATA[$\Pr (y=1)=0.02$]]></tex-math></alternatives></inline-formula>. We set a <inline-formula id="j_nejsds071_ineq_090"><alternatives><mml:math>
<mml:mn>2</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn></mml:math><tex-math><![CDATA[$2\times 500$]]></tex-math></alternatives></inline-formula> fully connected layers for both MANN and MLP. From Figure <xref rid="j_nejsds071_fig_002">2</xref>, the MANN models shows better fittings for the low data density region.</p>
<fig id="j_nejsds071_fig_002">
<label>Figure 2</label>
<caption>
<p>The marginal relationship of <inline-formula id="j_nejsds071_ineq_091"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_092"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{2}}$]]></tex-math></alternatives></inline-formula> with the Logit value of condition probability, that is <inline-formula id="j_nejsds071_ineq_093"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\log \Big(\frac{\Pr (y|{z_{1}})}{1-\Pr (y|{z_{1}})}\Big)$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_094"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\log \Big(\frac{\Pr (y|{z_{2}})}{1-\Pr (y|{z_{2}})}\Big)$]]></tex-math></alternatives></inline-formula> for <inline-formula id="j_nejsds071_ineq_095"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_096"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${z_{2}}$]]></tex-math></alternatives></inline-formula>. The black points represent the observed samples of the Gumbel part.</p>
</caption>
<graphic xlink:href="nejsds071_g002.jpg"/>
</fig>
</sec>
<sec id="j_nejsds071_s_008">
<label>3.3</label>
<title>Variational Inference Learning</title>
<p>Our VI-REM model combines the MANN and EP prior distribution. We establish VI-REM on the Multilayer Perceptron (MLP), which is flexible to handle different learning tasks. We adopt the setup from VIB [<xref ref-type="bibr" rid="j_nejsds071_ref_001">1</xref>], the loss function to be minimized for VI-REM is: 
<disp-formula id="j_nejsds071_eq_014">
<label>(3.9)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>−</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathcal{L}(\boldsymbol{\theta },\boldsymbol{\phi },\boldsymbol{\omega })={\mathcal{D}_{\mathrm{KL}}}& \Big({q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})||{p_{\boldsymbol{\omega }}^{\text{EP}}}(\mathbf{z})\Big)\\ {} -& \beta {\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}\Big(\log {p_{\boldsymbol{\theta }}}(y|\mathbf{z})\Big)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <italic>β</italic> is empirically set to be <inline-formula id="j_nejsds071_ineq_097"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${n_{0}}/{n_{1}}$]]></tex-math></alternatives></inline-formula> for models to be better optimized. <inline-formula id="j_nejsds071_ineq_098"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\omega }}^{\text{EP}}}(\mathbf{z})$]]></tex-math></alternatives></inline-formula> is the EP distribution with parameters <inline-formula id="j_nejsds071_ineq_099"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ω</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\omega }$]]></tex-math></alternatives></inline-formula>.</p>
<p>The <inline-formula id="j_nejsds071_ineq_100"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> is a family of distributions to approximate the true posterior density [<xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>]. Practitioners usually refer <inline-formula id="j_nejsds071_ineq_101"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> as Encoder structure as it reduces the original high-dimensional features <bold>x</bold> into lower dimensions <bold>z</bold>. The <inline-formula id="j_nejsds071_ineq_102"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> can be set as isotropic Gaussian distributions through: 
<disp-formula id="j_nejsds071_eq_015">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∣</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>×</mml:mo>
<mml:mi mathvariant="bold">I</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})=\mathcal{N}\Big(\mathbf{z}\mid {\mu _{\phi }}(\mathbf{x}),{\sigma _{\phi }^{2}}(\mathbf{x})\times \mathbf{I}\Big)\]]]></tex-math></alternatives>
</disp-formula> 
where <bold>I</bold> is the identity matrix. The parameters <inline-formula id="j_nejsds071_ineq_103"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mu _{\phi }}(\mathbf{x}),{\sigma _{\phi }^{2}}(\mathbf{x})={g_{\boldsymbol{\phi }}}(\mathbf{x})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_104"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${g_{\phi }}(\mathbf{x})$]]></tex-math></alternatives></inline-formula> is established based on MLP parameterized by <inline-formula id="j_nejsds071_ineq_105"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\phi }$]]></tex-math></alternatives></inline-formula>. This is a standard choice in literature of VI [<xref ref-type="bibr" rid="j_nejsds071_ref_025">25</xref>]. The usage of isotropic Gaussian distributions imply the individual components of the posterior <inline-formula id="j_nejsds071_ineq_106"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}({\mathbf{z}_{1}}|\mathbf{x})$]]></tex-math></alternatives></inline-formula>, ..., <inline-formula id="j_nejsds071_ineq_107"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">K</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}({\mathbf{z}_{K}}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> are mutually independent [<xref ref-type="bibr" rid="j_nejsds071_ref_025">25</xref>]. For more complex cases with dependent posterior, we can use distributions with tractable likelihood, such as normalizing flows [<xref ref-type="bibr" rid="j_nejsds071_ref_036">36</xref>] or inverse autoregressive flows [<xref ref-type="bibr" rid="j_nejsds071_ref_026">26</xref>]. Our results show that isotropic Gaussian distributions perform well. We use them in the following analysis. <inline-formula id="j_nejsds071_ineq_108"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\log {p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> is based on MANN parameterized by <inline-formula id="j_nejsds071_ineq_109"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_016">
<label>(3.10)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">{</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">}</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \log {p_{\boldsymbol{\theta }}}(y|\mathbf{z})=-\log \Big\{1+\exp \big(-\alpha (\boldsymbol{\theta })-{\sum \limits_{j=1}^{k}}{f_{j}}({\mathbf{z}_{j}})\big)\Big\}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Updating the parameters <inline-formula id="j_nejsds071_ineq_110"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_111"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\phi }$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_112"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ω</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\omega }$]]></tex-math></alternatives></inline-formula> in VI-REM is based on stochastic gradient descent. Updating the parameter <inline-formula id="j_nejsds071_ineq_113"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\phi }$]]></tex-math></alternatives></inline-formula> is the same as the strategies in [<xref ref-type="bibr" rid="j_nejsds071_ref_041">41</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_024">24</xref>]. As for parameters <inline-formula id="j_nejsds071_ineq_114"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\boldsymbol{\omega }\in {\mathbb{R}^{5k}}$]]></tex-math></alternatives></inline-formula> for <inline-formula id="j_nejsds071_ineq_115"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>EP</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\omega }}^{\text{EP}}}(\mathbf{z})$]]></tex-math></alternatives></inline-formula> and we calculate the derivatives with Equation <xref rid="j_nejsds071_eq_009">3.5</xref>: 
<disp-formula id="j_nejsds071_eq_017">
<label>(3.11)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo fence="true" stretchy="false">]</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}{\nabla _{\boldsymbol{\omega }}}\mathcal{L}(\boldsymbol{\theta },\boldsymbol{\phi },\boldsymbol{\omega })& ={\nabla _{\boldsymbol{\omega }}}{\mathcal{D}_{\mathrm{KL}}}[{q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})||{p_{\boldsymbol{\omega }}}(\mathbf{z})]\\ {} & =-{\mathbb{E}_{{q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})}}\Big({\sum \limits_{i=1}^{k}}{\nabla _{\boldsymbol{\omega }}}\log {p_{{\boldsymbol{\kappa }_{i}}}}({\mathbf{z}_{i}})\Big)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
The parameters <inline-formula id="j_nejsds071_ineq_116"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\lambda _{1}},\dots ,{\lambda _{k}}\in (0,1)$]]></tex-math></alternatives></inline-formula> determine the weight for the Gumbel distribution within the EP prior. Their derivatives can be calculated as follows: 
<disp-formula id="j_nejsds071_eq_018">
<label>(3.12)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">κ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">{</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">}</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}{\nabla _{{\lambda _{i}}}}\mathcal{L}(\boldsymbol{\theta },\boldsymbol{\phi },\boldsymbol{\omega })& =-{\mathbb{E}_{{q_{\boldsymbol{\phi }}}({\mathbf{z}_{i}}|\mathbf{x})}}\Big(\frac{{\nabla _{{\lambda _{i}}}}{p_{{\boldsymbol{\kappa }_{i}}}}({\mathbf{z}_{i}})}{{p_{{\boldsymbol{\kappa }_{i}}}}({\mathbf{z}_{i}})}\Big)\\ {} & =-{\mathbb{E}_{{q_{\boldsymbol{\phi }}}({\mathbf{z}_{i}}|\mathbf{x})}}\Big\{\frac{{h_{1,i}}-{h_{2,i}}}{{\lambda _{i}}{h_{1,i}}+(1-{\lambda _{i}}){h_{2,i}}}\Big\}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_117"><alternatives><mml:math>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">k</mml:mi></mml:math><tex-math><![CDATA[$i=1,\dots ,k$]]></tex-math></alternatives></inline-formula>, and <italic>k</italic> is the dimension of the latent variable <bold>z</bold>; <inline-formula id="j_nejsds071_ineq_118"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{i}}$]]></tex-math></alternatives></inline-formula> is the <italic>i</italic>th component of latent <bold>z</bold>; <inline-formula id="j_nejsds071_ineq_119"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\lambda _{i}}$]]></tex-math></alternatives></inline-formula> is the parameter <italic>λ</italic> for latent <inline-formula id="j_nejsds071_ineq_120"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{i}}$]]></tex-math></alternatives></inline-formula>; <inline-formula id="j_nejsds071_ineq_121"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">π</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[${h_{1,i}}=\frac{1}{\sqrt{2\pi {\sigma _{2,i}^{2}}}}\exp \Big(\frac{-{({\mathbf{z}_{i}}-{\mu _{2,i}})^{2}}}{2{\sigma _{2,i}^{2}}}\Big)$]]></tex-math></alternatives></inline-formula>; and <inline-formula id="j_nejsds071_ineq_122"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">μ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">σ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[${h_{2,i}}=\frac{1}{{\sigma _{1,i}}}\exp \Big(\frac{-({\mathbf{z}_{i}}-{\mu _{1,i}})}{{\sigma _{1,i}}}-\exp \big(\frac{-({\mathbf{z}_{i}}-{\mu _{1,i}})}{{\sigma _{1,i}}}\big)\Big)$]]></tex-math></alternatives></inline-formula>.</p>
<p>During the training process, parameter <inline-formula id="j_nejsds071_ineq_123"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\lambda _{i}}$]]></tex-math></alternatives></inline-formula> might collapse to zero due to sparse parameter regime. To avoid this issues, we set parameter <inline-formula id="j_nejsds071_ineq_124"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\lambda _{i}}$]]></tex-math></alternatives></inline-formula> through a logistic function as <inline-formula id="j_nejsds071_ineq_125"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle></mml:math><tex-math><![CDATA[${\lambda _{i}}=\frac{1}{1+\exp (-{\Lambda _{i}})}$]]></tex-math></alternatives></inline-formula>. This transformation enables us to update the weight for <inline-formula id="j_nejsds071_ineq_126"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\lambda _{i}}\in (0,1)$]]></tex-math></alternatives></inline-formula> by updating the real-valued parameter <inline-formula id="j_nejsds071_ineq_127"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi></mml:math><tex-math><![CDATA[${\Lambda _{i}}\in \mathbb{R}$]]></tex-math></alternatives></inline-formula>. The logistic function helps prevent the parameter <inline-formula id="j_nejsds071_ineq_128"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\lambda _{i}}$]]></tex-math></alternatives></inline-formula> from collapsing to zero. To obtain <inline-formula id="j_nejsds071_ineq_129"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\lambda _{i}}$]]></tex-math></alternatives></inline-formula>, we utilize the chain rule on Equation <xref rid="j_nejsds071_eq_018">3.12</xref> to compute the derivatives of <inline-formula id="j_nejsds071_ineq_130"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Λ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\nabla _{{\Lambda _{i}}}}\mathcal{L}(\boldsymbol{\theta },\boldsymbol{\phi },\boldsymbol{\omega })$]]></tex-math></alternatives></inline-formula>. The effectiveness of this approach is confirmed through numerical studies in Sections <xref rid="j_nejsds071_s_015">5</xref> and <xref rid="j_nejsds071_s_018">6</xref>.</p>
<p>The calculation of derivatives <inline-formula id="j_nejsds071_ineq_131"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula> involves the derivatives of each individual function <inline-formula id="j_nejsds071_ineq_132"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${f_{j}}({\mathbf{z}_{j}})$]]></tex-math></alternatives></inline-formula> in Equation <xref rid="j_nejsds071_eq_012">3.8</xref>. Based on Leibniz integral rule, we can have: 
<disp-formula id="j_nejsds071_eq_019">
<label>(3.13)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">t</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:mi mathvariant="italic">t</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">β</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\nabla _{\boldsymbol{\theta }}}{f_{j}}({\mathbf{z}_{j}})={\int _{{s_{0}}}^{{\mathbf{z}_{j}}}}{\nabla _{\boldsymbol{\theta }}}\exp \Big({h_{j}}(t;\boldsymbol{\theta })\Big)\mathrm{d}t+{\nabla _{\boldsymbol{\theta }}}{\beta _{j}}(\boldsymbol{\theta })\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Based on Equation <xref rid="j_nejsds071_eq_019">3.13</xref>, the derivatives of the Logit function <inline-formula id="j_nejsds071_ineq_133"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> with respect to <inline-formula id="j_nejsds071_ineq_134"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula> can be obtained through: 
<disp-formula id="j_nejsds071_eq_020">
<label>(3.14)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">β</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∫</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">t</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:mi mathvariant="italic">t</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}{\nabla _{\boldsymbol{\theta }}}{\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})={\nabla _{\boldsymbol{\theta }}}& \alpha (\boldsymbol{\theta })+{\sum \limits_{j=1}^{k}}{\nabla _{\boldsymbol{\theta }}}{\beta _{j}}(\boldsymbol{\theta })\\ {} & +{\sum \limits_{j=1}^{k}}{\int _{{s_{0}}}^{{\mathbf{z}_{j}}}}{\nabla _{\boldsymbol{\theta }}}\exp \Big({h_{j}}(t;\boldsymbol{\theta })\Big)\mathrm{d}t\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>With <inline-formula id="j_nejsds071_ineq_135"><alternatives><mml:math>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∼</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathbf{z}\sim {q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula>, the derivatives w.r.t parameters <inline-formula id="j_nejsds071_ineq_136"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula> are: 
<disp-formula id="j_nejsds071_eq_021">
<label>(3.15)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold-italic">ω</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">[</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">]</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">[</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∇</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">]</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}& {\nabla _{\boldsymbol{\theta }}}\mathcal{L}(\boldsymbol{\theta },\boldsymbol{\phi },\boldsymbol{\omega })={\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}\Big[{\nabla _{\boldsymbol{\theta }}}\log {p_{\theta }}(y|\mathbf{z})\Big]\\ {} & ={\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}\Big[\frac{\exp (-{\psi _{\theta }}(y|\mathbf{z}))}{1+\exp (-{\psi _{\theta }}(y|\mathbf{z}))}\times {\nabla _{\boldsymbol{\theta }}}{\psi _{\boldsymbol{\theta }}}(y|\mathbf{z})\Big]\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
With the optimized parameters <inline-formula id="j_nejsds071_ineq_137"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">θ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\theta }$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_138"><alternatives><mml:math>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi></mml:math><tex-math><![CDATA[$\boldsymbol{\phi }$]]></tex-math></alternatives></inline-formula> through the introduced learning process, we can generate the predictions for <inline-formula id="j_nejsds071_ineq_139"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y|\mathbf{x})$]]></tex-math></alternatives></inline-formula> through a Monte Carlo estimation where <italic>L</italic> is number of generated samples, where <inline-formula id="j_nejsds071_ineq_140"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">∼</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mathbf{z}^{(l)}}\sim {q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_022">
<label>(3.16)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≈</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">L</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">L</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \Pr (y|\mathbf{x})={\mathbb{E}_{{q_{\phi }}(\mathbf{z}|\mathbf{x})}}\Big(\log {p_{\boldsymbol{\theta }}}(y|\mathbf{z})\Big)\approx \frac{1}{L}{\sum \limits_{l=1}^{L}}{p_{\boldsymbol{\theta }}}(y|{\mathbf{z}^{(l)}})\]]]></tex-math></alternatives>
</disp-formula>
</p>
</sec>
<sec id="j_nejsds071_s_009">
<label>3.4</label>
<title>Theoretical Justifications of VI-REM</title>
<p>We theoretically discuss modeling the tail in REM can improve generalization power. The generalization power of <inline-formula id="j_nejsds071_ineq_141"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y|\mathbf{x})$]]></tex-math></alternatives></inline-formula> involve <inline-formula id="j_nejsds071_ineq_142"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_143"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\theta }}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> as in Equation <xref rid="j_nejsds071_eq_022">3.16</xref>. The asymptotic behavior of <inline-formula id="j_nejsds071_ineq_144"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> through VI has been extensively discussed in [<xref ref-type="bibr" rid="j_nejsds071_ref_049">49</xref>]. This paper focuses on <inline-formula id="j_nejsds071_ineq_145"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\theta }}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>.</p>
<p>We use <italic>generalization gap</italic> to measure model’s generalization ability, which is defined as the difference between the model’s performance on training data and on unseen testing data. Denote <inline-formula id="j_nejsds071_ineq_146"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[${\mathrm{P}_{n}}\big(g(\mathbf{z})\big)$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_147"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr \big(g(\mathbf{z})\big)$]]></tex-math></alternatives></inline-formula> respectively as the empirical measure and underlying ground truth for some function <inline-formula id="j_nejsds071_ineq_148"><alternatives><mml:math>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$g(\mathbf{z})$]]></tex-math></alternatives></inline-formula>. Define <inline-formula id="j_nejsds071_ineq_149"><alternatives><mml:math>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[$\mathcal{F}$]]></tex-math></alternatives></inline-formula> as the set of all monotonic functions, let <italic>f</italic> be the candidate function belonging to function space <inline-formula id="j_nejsds071_ineq_150"><alternatives><mml:math>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[$\mathcal{F}$]]></tex-math></alternatives></inline-formula>, and <inline-formula id="j_nejsds071_ineq_151"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[${f_{0}}\in \mathcal{F}$]]></tex-math></alternatives></inline-formula> is the ground truth. We use <inline-formula id="j_nejsds071_ineq_152"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mathrm{L}_{f}}(\mathbf{z},y)$]]></tex-math></alternatives></inline-formula> to denote the loss measured by function <italic>f</italic> taking value at <bold>z</bold> for the response <italic>y</italic>. We identify <inline-formula id="j_nejsds071_ineq_153"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">≜</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mi mathvariant="italic">η</mml:mi>
<mml:mo fence="true" stretchy="false">}</mml:mo></mml:math><tex-math><![CDATA[${\Omega _{\eta }}\triangleq \{\| \mathbf{z}\| \lt \eta \}$]]></tex-math></alternatives></inline-formula> as an <italic>η</italic>-ball where the bulk of the distribution mass resides, i.e., <inline-formula id="j_nejsds071_ineq_154"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">ϵ</mml:mi></mml:math><tex-math><![CDATA[$\Pr ({\Omega _{\eta }})\ge 1-\epsilon $]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_155"><alternatives><mml:math>
<mml:mi mathvariant="italic">ϵ</mml:mi>
<mml:mo stretchy="false">≪</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$\epsilon \ll 1$]]></tex-math></alternatives></inline-formula>, and consequently <inline-formula id="j_nejsds071_ineq_156"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">≜</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>∖</mml:mo>
<mml:mi mathvariant="normal">Ω</mml:mi></mml:math><tex-math><![CDATA[${\Omega _{\eta }^{C}}\triangleq {\mathbb{R}^{d}}\setminus \Omega $]]></tex-math></alternatives></inline-formula> contains the extremes.</p>
<p>We point out the theoretical generalization gap with the introduction of cutoff points in Theorem <xref rid="j_nejsds071_stat_002">3.2</xref>. The following regularized assumptions are needed to establish Theorem <xref rid="j_nejsds071_stat_002">3.2</xref>:</p>
<list>
<list-item id="j_nejsds071_li_001">
<label>C.1</label>
<p>The function <inline-formula id="j_nejsds071_ineq_157"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f(\mathbf{z})$]]></tex-math></alternatives></inline-formula> follows the Lipschitz property: 
<disp-formula id="j_nejsds071_eq_023">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ |f(\mathbf{z})-f({\mathbf{z}^{\prime }})|\le {k_{n}}\| \mathbf{z}-{\mathbf{z}^{\prime }}\| \]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_158"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${k_{n}}$]]></tex-math></alternatives></inline-formula> is some positive constant.</p>
</list-item>
<list-item id="j_nejsds071_li_002">
<label>C.2</label>
<p>The ground truth function <inline-formula id="j_nejsds071_ineq_159"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{0}}$]]></tex-math></alternatives></inline-formula> satisfies the condition: 
<disp-formula id="j_nejsds071_eq_024">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">arg</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">min</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {f_{0}}=\arg \underset{f\in \mathcal{F}}{\min }\mathbb{E}\Big(\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\Big)\]]]></tex-math></alternatives>
</disp-formula>
</p>
</list-item>
<list-item id="j_nejsds071_li_003">
<label>C.3</label>
<p>Function <inline-formula id="j_nejsds071_ineq_160"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f(\mathbf{z})$]]></tex-math></alternatives></inline-formula> satisfies the condition for nonzero <bold>z</bold>: 
<disp-formula id="j_nejsds071_eq_025">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \underset{f\in \mathcal{F}}{\sup }|f(\mathbf{z})-{f_{0}}(\mathbf{z})|\le \underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}\| \| \mathbf{z}\| \]]]></tex-math></alternatives>
</disp-formula>
</p>
</list-item>
</list>
<statement id="j_nejsds071_stat_002"><label>Theorem 3.2.</label>
<p><italic>Under the assumptions in C.1–C.3 described above, the following generalization gap holds</italic> 
<disp-formula id="j_nejsds071_eq_026">
<label>(3.17)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mi mathvariant="italic">δ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msqrt>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Big(& \underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\big|\Big)\\ {} & \le \sqrt{\frac{\log |\mathcal{F}|+\log (1/\delta )}{2n}}+\gamma {h^{1/2}}{\eta ^{1/2}}\underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }}}^{1/2}}\\ {} & +\gamma \underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\sup }\| \mathbf{z}{\| ^{1/2}}{(1-h)^{1/2}}\underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }^{C}}}^{1/2}}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
<italic>where</italic> <inline-formula id="j_nejsds071_ineq_161"><alternatives><mml:math>
<mml:mi mathvariant="italic">δ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">n</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\delta =|\mathcal{F}|{e^{-2n{\phi ^{2}}}}$]]></tex-math></alternatives></inline-formula> <italic>and</italic> <inline-formula id="j_nejsds071_ineq_162"><alternatives><mml:math>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
<mml:mo stretchy="false">|</mml:mo></mml:math><tex-math><![CDATA[$|\mathcal{F}|$]]></tex-math></alternatives></inline-formula> <italic>is the cardinality of space</italic> <inline-formula id="j_nejsds071_ineq_163"><alternatives><mml:math>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[$\mathcal{F}$]]></tex-math></alternatives></inline-formula><italic>. ϕ, γ are some constant,</italic> <inline-formula id="j_nejsds071_ineq_164"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${I_{\mathbf{z}\in {\Omega _{\eta }}:||\mathbf{z}||\le \eta }}$]]></tex-math></alternatives></inline-formula> <italic>is parameter λ in EP distribution, and h is the empirical value of</italic> <inline-formula id="j_nejsds071_ineq_165"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${I_{\mathbf{z}\in {\Omega _{\eta }}:||\mathbf{z}||\le \eta }}$]]></tex-math></alternatives></inline-formula><italic>.</italic></p></statement>
<p>Through Theorem <xref rid="j_nejsds071_stat_002">3.2</xref>, we decompose the generalization gap into estimation error, the approximation error for the non-tailed part, and the approximation error for the tailed part over a certain cutoff point as stated in the following Remark.</p><statement id="j_nejsds071_stat_003"><label>Remark 1.</label>
<p><italic>Assume</italic> <inline-formula id="j_nejsds071_ineq_166"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[${f_{1}},{f_{2}}\in \mathcal{F}$]]></tex-math></alternatives></inline-formula> <italic>are two functions belonging the space</italic> <inline-formula id="j_nejsds071_ineq_167"><alternatives><mml:math>
<mml:mi mathvariant="script">F</mml:mi></mml:math><tex-math><![CDATA[$\mathcal{F}$]]></tex-math></alternatives></inline-formula><italic>. We set</italic> <inline-formula id="j_nejsds071_ineq_168"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{1}}$]]></tex-math></alternatives></inline-formula> <italic>as the</italic> <inline-formula id="j_nejsds071_ineq_169"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\log {p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> <italic>which is obtained by maximizing Equation</italic> <xref rid="j_nejsds071_eq_014"><italic>3.9</italic></xref><italic>. Similarly, we set</italic> <inline-formula id="j_nejsds071_ineq_170"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{2}}$]]></tex-math></alternatives></inline-formula> <italic>as</italic> <inline-formula id="j_nejsds071_ineq_171"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\log {p_{{\boldsymbol{\theta }^{\prime }}}}(y|{\mathbf{z}^{\prime }})$]]></tex-math></alternatives></inline-formula><italic>is obtained by optimizing</italic> 
<disp-formula id="j_nejsds071_eq_027">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">KL</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">∣</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="italic">Gauss</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \beta {\mathbb{E}_{{q_{{\boldsymbol{\phi }^{\prime }}}}({\mathbf{z}^{\prime }}|\mathbf{x})}}\Big(\log {p_{{\boldsymbol{\theta }^{\prime }}}}(y|{\mathbf{z}^{\prime }})\Big)-{\mathcal{D}_{\mathrm{KL}}}\Big({q_{{\boldsymbol{\phi }^{\prime }}}}({\mathbf{z}^{\prime }}|\mathbf{x})\mid {p^{\textit{Gauss}}}({\mathbf{z}^{\prime }})\Big)\]]]></tex-math></alternatives>
</disp-formula> 
<italic>where</italic> <inline-formula id="j_nejsds071_ineq_172"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="italic">Gauss</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p^{\textit{Gauss}}}(\mathbf{z})$]]></tex-math></alternatives></inline-formula> <italic>is a isotropic Gaussian distribution. Denote</italic> <inline-formula id="j_nejsds071_ineq_173"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>−</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\Delta _{{\Omega _{\eta }^{C}}}}=||{f_{2}}-{f_{0}}|{|_{\mathbf{z}\in {\Omega _{\eta }^{C}}}^{1/2}}-||{f_{1}}-{f_{0}}|{|_{\mathbf{z}\in {\Omega _{\eta }^{C}}}^{1/2}}$]]></tex-math></alternatives></inline-formula> <italic>and</italic> <inline-formula id="j_nejsds071_ineq_174"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>−</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\Delta _{{\Omega _{\eta }}}}=||{f_{1}}-{f_{0}}|{|_{\mathbf{z}\in {\Omega _{\eta }}}^{1/2}}-||{f_{2}}-{f_{0}}|{|_{\mathbf{z}\in {\Omega _{\eta }}}^{1/2}}$]]></tex-math></alternatives></inline-formula><italic>. As long as the follows holds:</italic> 
<disp-formula id="j_nejsds071_eq_028">
<label>(3.18)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">&gt;</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">r</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\Delta _{{\Omega _{\eta }^{C}}}}/{\Delta _{{\Omega _{\eta }}}}\gt \frac{{\eta ^{1/2}}}{r{\sup _{\mathbf{z}\in {\Omega _{\eta }^{C}}}}\| \mathbf{z}{\| ^{1/2}}}{\Big(\frac{h}{1-h}\Big)^{1/2}}\]]]></tex-math></alternatives>
</disp-formula> 
<italic>where r is some constant. Then based on conditions C.1–C.3 and Theorem</italic> <xref rid="j_nejsds071_stat_002"><italic>3.2</italic></xref><italic>, the generalization gap over function</italic> <inline-formula id="j_nejsds071_ineq_175"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{1}}$]]></tex-math></alternatives></inline-formula><italic>,</italic> <inline-formula id="j_nejsds071_ineq_176"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{2}}$]]></tex-math></alternatives></inline-formula> <italic>can have the following:</italic> 
<disp-formula id="j_nejsds071_eq_029">
<label>(3.19)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Big(& \underset{{f_{1}}\in \mathcal{F}}{\sup }\big|{\mathrm{P}_{n}}\big({\mathrm{L}_{{f_{1}}}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\big|\Big)\\ {} & \lt \mathbb{E}\Big(\underset{{f_{2}}\in \mathcal{F}}{\sup }\big|{\mathrm{P}_{n}}\big({\mathrm{L}_{{f_{2}}}}({\mathbf{z}^{\prime }},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}({\mathbf{z}^{\prime }},y)\big)\big|\Big)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p></statement>
<p>Remark <xref rid="j_nejsds071_stat_003">1</xref> supports that the proposed extreme prior can improve model generalization. The misrepresented tailed part in REM amplifies the approximation error, thus restricting the model’s generalizability. Intuitively from Figure <xref rid="j_nejsds071_fig_001">1</xref>b, utilizing the heavy-tailed prior distribution remedies the misrepresentation problem, thus helps to tighten the generalization bound.</p>
</sec>
</sec>
<sec id="j_nejsds071_s_010">
<label>4</label>
<title>Model Implementations</title>
<sec id="j_nejsds071_s_011">
<label>4.1</label>
<title>Model Interpretation</title>
<p>The interpretation of VI-REM, i.e., the relationship between <bold>x</bold> and label <bold>y</bold>, needs to be made through latent factor <bold>z</bold>. The MANN component in recognizing the <inline-formula id="j_nejsds071_ineq_177"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula> makes the relationship between latent factors <bold>z</bold> and response <italic>y</italic> easy to assess. However, interpreting the relationship between <bold>x</bold> and <bold>z</bold> is challenging due to the black-box neural network. We frame the model interpretability as a variable selection process, i.e., identifying the most relevant <bold>x</bold> for each of the latent factors.</p>
<p>The core of this interpretation approach is the evaluation of the statistical dependence between <bold>x</bold> and <bold>z</bold>. Recent studies suggest that distance correlation [<xref ref-type="bibr" rid="j_nejsds071_ref_044">44</xref>] is a powerful non-parametric tool to measure the statistical correlation between two random variables. The MANN structure guarantees that the latent factors <bold>z</bold> in VI-REM are largely statistically independent. Based on this fact, we can use the Distance Correlation [<xref ref-type="bibr" rid="j_nejsds071_ref_044">44</xref>] to select the critical <bold>x</bold> in determining <inline-formula id="j_nejsds071_ineq_178"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_179"><alternatives><mml:math>
<mml:mn>...</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[$...{\mathbf{z}_{k}}$]]></tex-math></alternatives></inline-formula>.</p>
<p>Considering the high dimensionality of the input features <bold>x</bold>, we adopt the Distance Correlation Sure Independence Screening (DC-SIS) framework [<xref ref-type="bibr" rid="j_nejsds071_ref_027">27</xref>], which is the SOTA approach to use Distance Correlation to select high-dimensional features. The essence of a DC-SIS is to rank variables based on the Distance Correlation. Under our setup, we can use the Distance Correlation between <inline-formula id="j_nejsds071_ineq_180"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{i}}$]]></tex-math></alternatives></inline-formula> and each of the input features <inline-formula id="j_nejsds071_ineq_181"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{x}_{1}}$]]></tex-math></alternatives></inline-formula>, ..., <inline-formula id="j_nejsds071_ineq_182"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{x}_{d}}$]]></tex-math></alternatives></inline-formula> to rank the importance of input features and select the most relevant ones. However, when directly using the DC-SIS to the proposed VI-REM model, the <inline-formula id="j_nejsds071_ineq_183"><alternatives><mml:math>
<mml:mi mathvariant="script">O</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{O}({n^{2}})$]]></tex-math></alternatives></inline-formula> computational complexity for Distance Correlation is intolerable for a large sample size [<xref ref-type="bibr" rid="j_nejsds071_ref_028">28</xref>], especially when hundreds of thousands of training data. Therefore, we adopted a Divide and Conquer algorithm [<xref ref-type="bibr" rid="j_nejsds071_ref_028">28</xref>] to overcome the computation challenge.</p>
<p>For latent factor <inline-formula id="j_nejsds071_ineq_184"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}}$]]></tex-math></alternatives></inline-formula>, we equally spilt the <italic>n</italic> samples into G groups with <inline-formula id="j_nejsds071_ineq_185"><alternatives><mml:math>
<mml:mi mathvariant="italic">s</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">n</mml:mi>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mi mathvariant="italic">G</mml:mi></mml:math><tex-math><![CDATA[$s=n/G$]]></tex-math></alternatives></inline-formula> data points in each subgroup. For each subgroup <inline-formula id="j_nejsds071_ineq_186"><alternatives><mml:math>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mi mathvariant="italic">G</mml:mi>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[$g\in [1,\dots G]$]]></tex-math></alternatives></inline-formula> with <italic>s</italic> samples, we measure the Distance Correlation of <inline-formula id="j_nejsds071_ineq_187"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\mathbf{x}_{l}^{g}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_188"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}^{g}}$]]></tex-math></alternatives></inline-formula>, where <inline-formula id="j_nejsds071_ineq_189"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\mathbf{x}_{l}^{g}}$]]></tex-math></alternatives></inline-formula> represents the input feature <inline-formula id="j_nejsds071_ineq_190"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{x}_{l}}$]]></tex-math></alternatives></inline-formula> within subgroup <italic>g</italic> and <inline-formula id="j_nejsds071_ineq_191"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}^{g}}$]]></tex-math></alternatives></inline-formula> represents the latent factor <inline-formula id="j_nejsds071_ineq_192"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}}$]]></tex-math></alternatives></inline-formula> within subgroup <italic>g</italic>. For simplicity, denote <inline-formula id="j_nejsds071_ineq_193"><alternatives><mml:math>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[$\mathbf{u}={\mathbf{x}_{l}^{g}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_194"><alternatives><mml:math>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[$\mathbf{v}={\mathbf{z}_{j}^{g}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_195"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${d_{u}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_196"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${d_{v}}$]]></tex-math></alternatives></inline-formula> as the dimension for <bold>u</bold> and <bold>v</bold>, and <inline-formula id="j_nejsds071_ineq_197"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${\mathbf{u}^{(i)}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_198"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${\mathbf{v}^{(i)}}$]]></tex-math></alternatives></inline-formula> as the <italic>i</italic> th sample of <bold>u</bold> and <bold>v</bold> correspondingly. The detailed formula for estimation of distance correlation is as follows: 
<disp-formula id="j_nejsds071_eq_030">
<label>(4.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">DC</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo><mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcorr</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcov</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcov</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo><mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcov</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\widehat{\operatorname{DC}}_{l}^{g}}=\widehat{\operatorname{dcorr}}({\mathbf{x}_{l}^{g}},{\mathbf{z}_{j}^{g}})=\frac{\widehat{\operatorname{dcov}}(\mathbf{u},\mathbf{v})}{\sqrt{\widehat{\operatorname{dcov}}(\mathbf{u},\mathbf{u})\widehat{\operatorname{dcov}}(\mathbf{v},\mathbf{v})}}\]]]></tex-math></alternatives>
</disp-formula> 
where 
<disp-formula id="j_nejsds071_eq_031">
<label>(4.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcov</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\widehat{\operatorname{dcov}}^{2}}(\mathbf{u},\mathbf{v})={\widehat{S}_{1}^{g}}+{\widehat{S}_{2}^{g}}-2{\widehat{S}_{3}^{g}}\]]]></tex-math></alternatives>
</disp-formula> 
and <inline-formula id="j_nejsds071_ineq_199"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\widehat{S}_{1}^{g}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_200"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\widehat{S}_{2}^{g}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_201"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\widehat{S}_{3}^{g}}$]]></tex-math></alternatives></inline-formula> for <inline-formula id="j_nejsds071_ineq_202"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">dcov</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">u</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="bold">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\widehat{\operatorname{dcov}}^{2}}(\mathbf{u},\mathbf{v})$]]></tex-math></alternatives></inline-formula> are as follow: 
<disp-formula id="j_nejsds071_eq_032">
<label>(4.3)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable equalrows="false" equalcolumns="false" columnalign="left">
<mml:mtr>
<mml:mtd class="array">
<mml:mtable displaystyle="true" columnalign="right">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd class="array">
<mml:mtable displaystyle="true" columnalign="right">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd class="array">
<mml:mtable displaystyle="true" columnalign="right">
<mml:mtr>
<mml:mtd>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="italic">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mfenced separators="" open="‖" close="‖">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{array}{l}\displaystyle {\widehat{S}_{1}^{g}}=\displaystyle \frac{1}{{s^{2}}}{\displaystyle \sum \limits_{i=1}^{s}}{\displaystyle \sum \limits_{j=1}^{s}}{\left\| {\mathbf{u}^{(i)}}-{\mathbf{u}^{(j)}}\right\| _{{d_{u}}}}{\left\| {\mathbf{v}^{(i)}}-{\mathbf{v}^{(j}}\right\| _{{d_{v}}}}\\ {} \displaystyle {\widehat{S}_{3}^{g}}=\displaystyle \frac{1}{{s^{3}}}{\displaystyle \sum \limits_{i=1}^{s}}{\displaystyle \sum \limits_{j=1}^{s}}{\displaystyle \sum \limits_{l=1}^{s}}{\left\| {\mathbf{u}^{(i)}}-{\mathbf{u}^{(l)}}\right\| _{{d_{u}}}}{\left\| {\mathbf{v}^{(j)}}-{\mathbf{v}^{(l)}}\right\| _{{d_{v}}}}\\ {} \displaystyle {\widehat{S}_{2}^{g}}=\displaystyle \frac{1}{{s^{2}}}{\displaystyle \sum \limits_{i=1}^{s}}{\displaystyle \sum \limits_{j=1}^{s}}{\left\| {\mathbf{u}^{(i)}}-{\mathbf{u}^{(j)}}\right\| _{{d_{u}}}}\displaystyle \frac{1}{{s^{2}}}{\displaystyle \sum \limits_{i=1}^{s}}{\displaystyle \sum \limits_{j=1}^{s}}{\left\| {\mathbf{v}^{(i)}}-{\mathbf{v}^{(j)}}\right\| _{{d_{v}}}}\end{array}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>For each subgroup, we conduct the DC-SIS [<xref ref-type="bibr" rid="j_nejsds071_ref_027">27</xref>] to select the set with the most contribution of input features <bold>x</bold> toward the latent factor <inline-formula id="j_nejsds071_ineq_203"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}}$]]></tex-math></alternatives></inline-formula>. In practice for each of the subgroup <italic>g</italic>, the selection process for a given size <inline-formula id="j_nejsds071_ineq_204"><alternatives><mml:math>
<mml:mi mathvariant="italic">q</mml:mi>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mi mathvariant="italic">s</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">n</mml:mi>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mi mathvariant="italic">G</mml:mi></mml:math><tex-math><![CDATA[$q\lt s=n/G$]]></tex-math></alternatives></inline-formula> can be: 
<disp-formula id="j_nejsds071_eq_033">
<label>(4.4)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfenced separators="" open="{" close="}">
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo>:</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">DC</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="2.5pt"/>
<mml:mtext>is among top</mml:mtext>
<mml:mspace width="2.5pt"/>
<mml:mi mathvariant="italic">q</mml:mi>
<mml:mspace width="2.5pt"/>
<mml:mtext>largest of</mml:mtext>
<mml:mspace width="2.5pt"/>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mspace width="2.5pt"/>
</mml:mrow>
</mml:mfenced>
<mml:mo mathvariant="normal">,</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\widehat{\mathcal{D}}_{g}}=\left\{l:{\widehat{\operatorname{DC}}_{l}^{g}}\hspace{2.5pt}\text{is among top}\hspace{2.5pt}q\hspace{2.5pt}\text{largest of}\hspace{2.5pt}d\hspace{2.5pt}\right\},\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_205"><alternatives><mml:math>
<mml:mi mathvariant="italic">l</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[$l\in [1,\dots ,d]$]]></tex-math></alternatives></inline-formula> and <italic>d</italic> is the dimension of input features <bold>x</bold>.</p>
<p>After conducting these <italic>G</italic> (in total we have <italic>G</italic> subgroups) independent calculations and selections, the final component of <inline-formula id="j_nejsds071_ineq_206"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}}$]]></tex-math></alternatives></inline-formula> denoted by <inline-formula id="j_nejsds071_ineq_207"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\widehat{\mathcal{D}}_{{\mathbf{z}_{j}}}}$]]></tex-math></alternatives></inline-formula> is established by: <inline-formula id="j_nejsds071_ineq_208"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>∩</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mi mathvariant="italic">G</mml:mi>
<mml:mo fence="true" stretchy="false">}</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\widehat{\mathcal{D}}_{{\mathbf{z}_{j}}}}={\cap _{g\in \{1,\dots G\}}}{\widehat{\mathcal{D}}_{g}}$]]></tex-math></alternatives></inline-formula>. In practice, we can simply uses the most common selected relevant features <inline-formula id="j_nejsds071_ineq_209"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{x}_{1}}$]]></tex-math></alternatives></inline-formula>, ..., <inline-formula id="j_nejsds071_ineq_210"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">d</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{x}_{d}}$]]></tex-math></alternatives></inline-formula> in each of the subgroup procedures. Repeat the above process for all <inline-formula id="j_nejsds071_ineq_211"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{j}}$]]></tex-math></alternatives></inline-formula> and calculate the <inline-formula id="j_nejsds071_ineq_212"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\widehat{\mathcal{D}}_{{\mathbf{z}_{j}}}}$]]></tex-math></alternatives></inline-formula>; i.e., the identified component of latent factors.</p><statement id="j_nejsds071_stat_004"><label>Theorem 4.1.</label>
<p><italic>Denote</italic> <inline-formula id="j_nejsds071_ineq_213"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${\mathcal{D}_{\mathbf{{z_{j}}}}^{\ast }}$]]></tex-math></alternatives></inline-formula> <italic>as the true components associated with latent</italic> <inline-formula id="j_nejsds071_ineq_214"><alternatives><mml:math><mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle></mml:math><tex-math><![CDATA[$\mathbf{{z_{j}}}$]]></tex-math></alternatives></inline-formula><italic>, and</italic> <inline-formula id="j_nejsds071_ineq_215"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mstyle><mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\widehat{\mathcal{D}}_{\mathbf{{z_{j}}}}}$]]></tex-math></alternatives></inline-formula> <italic>as the set identified by the previously mentioned model’s interpretation strategy. Then:</italic> 
<disp-formula id="j_nejsds071_eq_034">
<label>(4.5)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true" columnalign="right">
<mml:mtr>
<mml:mtd class="eqnarray-1">
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">⊆</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mstyle><mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">{</mml:mo>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="script">O</mml:mi>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo fence="true" stretchy="false">]</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">α</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo fence="true" stretchy="false">}</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">}</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">G</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">→</mml:mo>
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\Pr \left({\mathcal{D}_{\mathbf{{z_{j}}}}^{\ast }}\subseteq {\widehat{\mathcal{D}}_{\mathbf{{z_{j}}}}}\right)\ge \Big\{1& -\mathcal{O}\{d\exp [-{c_{1}}{n^{\alpha -2\alpha (u+v)}}]\\ {} & +d{n^{\alpha }}\exp (-{c_{2}}{n^{v}})\}{\Big\}^{G}}\to 1\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
<italic>where</italic> <inline-formula id="j_nejsds071_ineq_216"><alternatives><mml:math>
<mml:mi mathvariant="italic">G</mml:mi>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">α</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$G={n^{1-\alpha }}$]]></tex-math></alternatives></inline-formula><italic>,</italic> <inline-formula id="j_nejsds071_ineq_217"><alternatives><mml:math>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\alpha \gt 0$]]></tex-math></alternatives></inline-formula><italic>, d is the dimension of input features</italic> <bold>x</bold><italic>,</italic> <inline-formula id="j_nejsds071_ineq_218"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${c_{1}}$]]></tex-math></alternatives></inline-formula><italic>,</italic> <inline-formula id="j_nejsds071_ineq_219"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${c_{2}}$]]></tex-math></alternatives></inline-formula><italic>, u, v are some constant and</italic> <inline-formula id="j_nejsds071_ineq_220"><alternatives><mml:math>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$u+v\lt 1/2$]]></tex-math></alternatives></inline-formula><italic>.</italic></p></statement>
<p>Theorem <xref rid="j_nejsds071_stat_004">4.1</xref> shows that as long as the sample size <italic>n</italic> is getting sufficiently large, the true risk modulators for the learnt latents have high confidence to be selected by the given procedure latent factor identification.</p>
</sec>
<sec id="j_nejsds071_s_012">
<label>4.2</label>
<title>Robustness Evaluations</title>
<p>A robust prediction model should be less vulnerable to attacks and remain valid when testing data is out of domain [<xref ref-type="bibr" rid="j_nejsds071_ref_005">5</xref>]. Building a robust model is critical for deploying deep learning models in the application of safety-critical situation [<xref ref-type="bibr" rid="j_nejsds071_ref_048">48</xref>]. Robustness measures the model’s ability to handle data perturbations and helps understand the boundary of successful applications. Mathematically for a model <italic>f</italic> and its predictions over <italic>x</italic> <inline-formula id="j_nejsds071_ineq_221"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f(x)$]]></tex-math></alternatives></inline-formula>, robustness evaluations aims to find how the model <italic>f</italic> can handle <inline-formula id="j_nejsds071_ineq_222"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f({x^{\prime }})$]]></tex-math></alternatives></inline-formula>, where <inline-formula id="j_nejsds071_ineq_223"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${x^{\prime }}$]]></tex-math></alternatives></inline-formula> has minor perturbations to <italic>x</italic>. if <inline-formula id="j_nejsds071_ineq_224"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f({x^{\prime }})$]]></tex-math></alternatives></inline-formula> does not show much difference with <inline-formula id="j_nejsds071_ineq_225"><alternatives><mml:math>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$f(x)$]]></tex-math></alternatives></inline-formula>, then model <italic>f</italic> is robust and trustworthy to deploy [<xref ref-type="bibr" rid="j_nejsds071_ref_048">48</xref>].</p>
<p>Researchers have proposed various approaches under different application scenarios to generate <inline-formula id="j_nejsds071_ineq_226"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${x^{\prime }}$]]></tex-math></alternatives></inline-formula> to evaluate robustness. Regarding natural language processing tasks, [<xref ref-type="bibr" rid="j_nejsds071_ref_051">51</xref>] summarizes the current well-established methods and suggests mixing the words from different documents [<xref ref-type="bibr" rid="j_nejsds071_ref_022">22</xref>, <xref ref-type="bibr" rid="j_nejsds071_ref_051">51</xref>] could be a simple and efficient approach. For computer vision studies, image blurring and image rotation are all useful in practice. In this work, we use the rotation operator to investigate whether the model can efficiently extract invariant features from rare events.</p>
</sec>
<sec id="j_nejsds071_s_013">
<label>4.3</label>
<title>Extension to General Setups</title>
<p>The previous VI-REM model mostly focuses on binary classification, and we generalize it to multi-label classification in risk analysis. Suppose we have <inline-formula id="j_nejsds071_ineq_227"><alternatives><mml:math>
<mml:mi mathvariant="normal">L</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$\mathrm{L}+1$]]></tex-math></alternatives></inline-formula> classified labels, and we treat one dominant class as class <italic>Normality</italic>, the rest L minority classes as <italic>Event</italic>. For example, traffic crashes are rare and can be classified from the most severe fatal crashes to injury crashes to minor property-damage only crashes. To mimic the general settings in risk analysis, we assume the severity of the L minority class ranges from low to high.</p>
<p>Without generality, we set 1, ..., <italic>L</italic> as labels for the L minority classes, and class 1 as the least severe <italic>Event</italic> and class <italic>L</italic> is the most severe one. As the response variable, the severity of an event increases from class 1 toward class <italic>L</italic> and belongs to the ordinal variables, we applied the setup from the Ordered Logit Model (OLM). Under the OLM setup, we want to model <inline-formula id="j_nejsds071_ineq_228"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr (y\ge m|\mathbf{x})$]]></tex-math></alternatives></inline-formula> given different event severity where <inline-formula id="j_nejsds071_ineq_229"><alternatives><mml:math>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>.</mml:mo>
<mml:mi mathvariant="italic">L</mml:mi></mml:math><tex-math><![CDATA[$m=1,,,.L$]]></tex-math></alternatives></inline-formula>. So that under the VI-REM, we modify the Logit function in Equation <xref rid="j_nejsds071_eq_010">3.6</xref> through: 
<disp-formula id="j_nejsds071_eq_035">
<label>(4.6)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\psi _{\boldsymbol{\theta }}}\big(y\ge m|\mathbf{z}\big)=\log \left(\frac{{p_{\boldsymbol{\theta }}}(y\ge m|\mathbf{z})}{1-{p_{\boldsymbol{\theta }}}(y\ge m|\mathbf{z})}\right)\]]]></tex-math></alternatives>
</disp-formula> 
and the Logit function in the OLM based VI-REM follows: 
<disp-formula id="j_nejsds071_eq_036">
<label>(4.7)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">ψ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">m</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">α</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo largeop="true" movablelimits="false">∑</mml:mo></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\psi _{\boldsymbol{\theta }}}\big(y\ge m|\mathbf{z}\big)={\boldsymbol{\alpha }_{\boldsymbol{m}}}(\boldsymbol{\theta })+{\sum \limits_{j=1}^{k}}{f_{j}}({\mathbf{z}_{j}})\]]]></tex-math></alternatives>
</disp-formula> 
where the coefficients <inline-formula id="j_nejsds071_ineq_230"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">α</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\boldsymbol{\alpha }_{\boldsymbol{m}}}(\boldsymbol{\theta })$]]></tex-math></alternatives></inline-formula> are intercepts for different event categories and <inline-formula id="j_nejsds071_ineq_231"><alternatives><mml:math>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">α</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mo stretchy="false">⋯</mml:mo>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">α</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">L</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$0\lt {\boldsymbol{\alpha }_{\mathbf{1}}}(\boldsymbol{\theta })\lt \cdots \lt {\boldsymbol{\alpha }_{\boldsymbol{L}}}(\boldsymbol{\theta })$]]></tex-math></alternatives></inline-formula>. The individual function <inline-formula id="j_nejsds071_ineq_232"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${f_{j}}({\mathbf{z}_{j}})$]]></tex-math></alternatives></inline-formula> function follows the same as Equation <xref rid="j_nejsds071_eq_012">3.8</xref>.</p>
</sec>
<sec id="j_nejsds071_s_014">
<label>4.4</label>
<title>Experiment Setup</title>
<p>We considered the following machine learning approaches for comparison: (1) Sampling MLP, short as <bold>S-MLP</bold>, is a re-sampling scheme based on Multilayer Perceptron (MLP) over-samples from the less-represented class [<xref ref-type="bibr" rid="j_nejsds071_ref_003">3</xref>]. The hyperparameter <italic>r</italic> in the S-MLP is used to control the over-sample ratio, which is equal to the expected <inline-formula id="j_nejsds071_ineq_233"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${n_{1}}$]]></tex-math></alternatives></inline-formula> over <inline-formula id="j_nejsds071_ineq_234"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${n_{0}}$]]></tex-math></alternatives></inline-formula> after over-sampling. We use grid search cross-validation to select the sampling ratio <inline-formula id="j_nejsds071_ineq_235"><alternatives><mml:math>
<mml:mi mathvariant="italic">r</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0.01</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.05</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.1</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$r\in (0.01,0.05,0.1,0.2)$]]></tex-math></alternatives></inline-formula>. (2) Focal MLP, short as <bold>F-MLP</bold>, uses a re-weighting based loss that adapts the cross entropy loss putting more penalization on the less-represented class [<xref ref-type="bibr" rid="j_nejsds071_ref_029">29</xref>] based on MLPs. We use grid search cross-validation to select the hyperparameters <inline-formula id="j_nejsds071_ineq_236"><alternatives><mml:math>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0.25</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.75</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\alpha \in (0.25,0.5,0.75)$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_237"><alternatives><mml:math>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>5</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\gamma \in (2,5)$]]></tex-math></alternatives></inline-formula>. (3) <bold>MAML</bold> is a few-shot learning scheme from [<xref ref-type="bibr" rid="j_nejsds071_ref_013">13</xref>] and it focuses on improving the generalization performance for less represented data samples and we use 5-shot learning.</p>
<p>We also consider the Deep-SVDD [<xref ref-type="bibr" rid="j_nejsds071_ref_037">37</xref>], short as <bold>D-SVDD</bold>, which is a unsupervised anomaly detection approach. D-SVDD will train a MLP based neural network to represent the majority of the data. The data points falling outside the majority will be considered as outliers. We also consider <bold>GBDT</bold>, a tree based classifier optimizing objective function through a multiple boosting process, and it is a good fit for REM [<xref ref-type="bibr" rid="j_nejsds071_ref_031">31</xref>].</p>
<p>In terms of the evaluation metrics, we use the F1 score to evaluate the model’s classification ability, and we set 0.5 as the decision threshold. Another adopted evaluation metric is the Precision-Recall AUC (PR-AUC), which focuses on the trade-off between precision and recall ability [<xref ref-type="bibr" rid="j_nejsds071_ref_020">20</xref>]. Besides the F1 and PR-AUC score for classification ability, we also care about the model’s ability with uncertainty prediction as the model results provide support for decision-making [<xref ref-type="bibr" rid="j_nejsds071_ref_040">40</xref>]. We use the negative log likelihood (NLL), Brier score [<xref ref-type="bibr" rid="j_nejsds071_ref_040">40</xref>]. Considering the extreme imbalanced label in REM, we modify the Brier score as the Brier Skill Score (BSS) in the following: 
<disp-formula id="j_nejsds071_eq_037">
<label>(4.8)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">BSS</mml:mo>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \operatorname{BSS}=1-\frac{\frac{1}{n}{\textstyle\textstyle\sum _{i=1}^{n}}{\Big(\Pr ({y^{(i)}}|{\mathbf{x}^{(i)}})-{y^{(i)}}\Big)^{2}}}{\frac{1}{n}{\textstyle\textstyle\sum _{i=1}^{n}}{\Big(\frac{1}{n}{\textstyle\textstyle\sum _{i}^{n}}{y^{(i)}}-{y^{(i)}}\Big)^{2}}}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_238"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr ({y^{(i)}}|{\mathbf{x}^{(i)}})$]]></tex-math></alternatives></inline-formula> is predicted probability of label <inline-formula id="j_nejsds071_ineq_239"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${y^{(i)}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_240"><alternatives><mml:math>
<mml:mo movablelimits="false">BSS</mml:mo>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[$\operatorname{BSS}\in [0,1]$]]></tex-math></alternatives></inline-formula>. Similar as the BSS, we modify the NLL with the relative NLL (short as R-NLL) given the imbalanced label: 
<disp-formula id="j_nejsds071_eq_038">
<label>(4.9)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">R‐NLL</mml:mo>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo movablelimits="false">NLL</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo movablelimits="false">NLL</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mo largeop="false" movablelimits="false">∑</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ R‐NLL=\frac{{\textstyle\textstyle\sum _{i=1}^{n}}\operatorname{NLL}\Big(\Pr ({y^{(i)}}|{\mathbf{x}^{(i)}}),{y^{(i)}}\Big)}{{\textstyle\textstyle\sum _{i=1}^{n}}\operatorname{NLL}\Big(\frac{1}{n}{\textstyle\textstyle\sum _{i}^{n}}{y^{(i)}},{y^{(i)}}\Big)}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_241"><alternatives><mml:math>
<mml:mo movablelimits="false">NLL</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\operatorname{NLL}\Big(\Pr ({y^{(i)}}|{\mathbf{x}^{(i)}}),{y^{(i)}}\Big)$]]></tex-math></alternatives></inline-formula> measure the NLL between <inline-formula id="j_nejsds071_ineq_242"><alternatives><mml:math>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\Pr ({y^{(i)}}|{\mathbf{x}^{(i)}})$]]></tex-math></alternatives></inline-formula> and label <inline-formula id="j_nejsds071_ineq_243"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">i</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${y^{(i)}}$]]></tex-math></alternatives></inline-formula>.</p>
</sec>
</sec>
<sec id="j_nejsds071_s_015">
<label>5</label>
<title>Benchmark Dataset</title>
<p>We investigated VI-REM’s generalization performance on unstructured image and natural language processing tasks and evaluated robustness. We conducted the experiments via Pytorch on an NVIDIA GeForce RTX 4080 Laptop GPU. To ensure fair comparisons and convenience of visualization, we set the number of latents to 2 in all experiments. As the two considered benchmark datasets are well established in practice, we design an ablation study to understand the functionality of each component in VI-REM. For model VI-REM-V1, we alternate the EP distribution with the Gaussian distribution. Model VI-REM-V2 further replaces the MANN component with the fully connected MLPs. Model VI-REM-V3 alternates the weight <inline-formula id="j_nejsds071_ineq_244"><alternatives><mml:math>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[$\beta ={n_{0}}/{n_{1}}$]]></tex-math></alternatives></inline-formula> in Equation <xref rid="j_nejsds071_eq_014">3.9</xref> with a small number. For simplicity, we set the weight <inline-formula id="j_nejsds071_ineq_245"><alternatives><mml:math>
<mml:mi mathvariant="italic">β</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>5</mml:mn></mml:math><tex-math><![CDATA[$\beta =5$]]></tex-math></alternatives></inline-formula> in VI-REM-V3.</p>
<sec id="j_nejsds071_s_016">
<label>5.1</label>
<title>Benchmark Dataset: Imbalanced Digit Recognition</title>
<fig id="j_nejsds071_fig_003">
<label>Figure 3</label>
<caption>
<p>Learned representations via VI-REM under different tasks on original testing. Red points represent class <italic>Event</italic>, green points represent class <italic>Normality</italic>. Red points represent the class <italic>Event</italic>, green points represent the class <italic>Normality</italic>.</p>
</caption>
<graphic xlink:href="nejsds071_g003.jpg"/>
</fig>
<p>We applied the VI-REM to MNIST<xref ref-type="fn" rid="j_nejsds071_fn_001">1</xref><fn id="j_nejsds071_fn_001"><label><sup>1</sup></label>
<p><uri>http://yann.lecun.com/exdb/mnist/</uri></p></fn> dataset. The MNIST dataset covers handwritten digit numbers 0 to 9, corresponding to class 0 to class 9 as the label. Following a similar setup in [<xref ref-type="bibr" rid="j_nejsds071_ref_003">3</xref>], in each task, one of the ten classes is set as <italic>Event</italic>, while the samples from the remaining nine classes represent <italic>Normality</italic>. In the model training part, 100 images are randomly selected from the <italic>Event</italic> class, and nearly 600,000 images are selected from <italic>Normality</italic> class. As the selection of <italic>Event</italic> cases involves randomness, we repeat the experiment ten times for each task. In total, we conduct 100 individual runs of experiments. For model testing part, we create balanced testings consist of 5,000 images from the class <italic>Event</italic> and 5,000 images from the class <italic>Normality</italic>. Regarding robustness evaluations, we applied the random rotation operator to the original testing data. Each of the original testing images will rotate at a different angle, and the rotated angle is a random value from a uniform distribution <inline-formula id="j_nejsds071_ineq_246"><alternatives><mml:math>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>60</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>∘</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>60</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>∘</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[$[-{60^{\circ }},{60^{\circ }}]$]]></tex-math></alternatives></inline-formula>.</p>
<p>For VI-REM, the experiment sets <inline-formula id="j_nejsds071_ineq_247"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times 2$]]></tex-math></alternatives></inline-formula> fully connected (FC) MLPs for <inline-formula id="j_nejsds071_ineq_248"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_249"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times 500$]]></tex-math></alternatives></inline-formula> for MANN in recognizing <inline-formula id="j_nejsds071_ineq_250"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. We used Leaky ReLU for the activation function for both <inline-formula id="j_nejsds071_ineq_251"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_252"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\boldsymbol{\theta }}}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. For S-MLP, F-MLP and MAML, we set the neural network to be <inline-formula id="j_nejsds071_ineq_253"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times \times 2\times 500\times 500\times 500\times 2$]]></tex-math></alternatives></inline-formula> FC layers; for D-SVDD, we set the neural network to be <inline-formula id="j_nejsds071_ineq_254"><alternatives><mml:math>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$1000\times 1000\times 1000\times 2$]]></tex-math></alternatives></inline-formula> FC layers. For the GBDT model, we set the maximum depth of GBDT to be 10 and number of estimators in GBDT to be 50.</p>
<table-wrap id="j_nejsds071_tab_001">
<label>Table 1</label>
<caption>
<p>Digit Number Detection Overall Performance Comparison.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1 Score</sc>↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="9" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Original</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.28</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.49</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.75</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.24</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>1.32</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.47</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.98</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.82</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.74</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.31</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.98</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.78</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.83</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.40</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.98</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.81</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.38</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.80</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.60</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.01</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>1.13</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.51</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.96</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.86</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.25</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.54</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.86</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.63</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.22</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.78</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.45</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>VI-REM-V3</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.55</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.43</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.70</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.32</sc></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="9" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Rotation</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.24</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.76</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.65</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.24</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.23</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.82</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.69</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.12</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.31</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.50</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.77</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.33</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.25</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.70</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.78</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.22</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.38</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.81</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.54</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.01</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.87</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.21</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.89</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.71</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.34</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.45</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.72</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.36</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.26</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.75</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.65</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.17</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>VI-REM-V3</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.23</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.92</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.57</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.05</sc></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
<p>Figure <xref rid="j_nejsds071_fig_003">3</xref> shows the learned representations via VI-REM under different tasks under the original testing. It’s clear that all the class <italic>Event</italic> deviate from the distribution and mostly lie on the tail part. This pattern is consistent with the EP prior from the VI-REM model. Table <xref rid="j_nejsds071_tab_001">1</xref> reports the average of the 100 experiment results for both original testings and robustness evaluations. The proposed VI-REM model consistently outperformed the competing models regarding model accuracy and robustness. Under random rotation tasks, the VI-REM shows significantly improved model robustness compared to other SOTAs. Through ablation study, the MANN structure largely increases the model’s generalization ability under original testings, while the EP distribution contributes much larger in robustness.</p>
<table-wrap id="j_nejsds071_tab_002">
<label>Table 2</label>
<caption>
<p>Digit Number Detection Overall Performance Comparison with Different Number of Latent Variables.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Dim</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1 Score</sc>↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Original</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.13</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.51</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.96</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.86</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>3</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.33</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.31</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.96</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.77</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>5</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.48</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.56</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.96</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.87</sc></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="3" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Rotation</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.87</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.21</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.89</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.71</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>3</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.13</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.42</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.77</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.38</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>5</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.13</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.35</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.77</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.43</sc></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
<p>The dimension of latent variable is an important hyper parameter in the VI-REM. To examine how the latent variable dimension affects the performance of our model, we vary it from 2 to 3 and 5. The results are reported in Table <xref rid="j_nejsds071_tab_002">2</xref>, and the column “Dim” displays the dimension for the latent variables. We find that setting the dimension of the latent variable can generate satisfying results on both original testing and rotation testing.</p>
</sec>
<sec id="j_nejsds071_s_017">
<label>5.2</label>
<title>Benchmark Dataset: Imbalanced News Detection</title>
<p>This experiment applied the VI-REM to the standard natural language processing (NLP) benchmark dataset <italic>AG News Classification</italic>.<xref ref-type="fn" rid="j_nejsds071_fn_002">2</xref><fn id="j_nejsds071_fn_002"><label><sup>2</sup></label>
<p><uri>https://www.kaggle.com/amananandrai/ag-news-classification-dataset</uri></p></fn> The AG News dataset covers various news articles from “Sports”, “World”, “Business”, and “Sci/Tech” these four different categories. For implementing the applications, we used pre-trained fastText [<xref ref-type="bibr" rid="j_nejsds071_ref_033">33</xref>] and SWEM [<xref ref-type="bibr" rid="j_nejsds071_ref_038">38</xref>] to extract raw text representations. For the original input text data, we firstly apply the pre-trained fastText model to transform raw word input into the word embedding, where we set the dimensionality to 300. Based on the extracted word embedding, we used the average operator from SWEM [<xref ref-type="bibr" rid="j_nejsds071_ref_038">38</xref>] to combine word embedding and generate a <inline-formula id="j_nejsds071_ineq_255"><alternatives><mml:math>
<mml:mn>300</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$300\times 1$]]></tex-math></alternatives></inline-formula> vector for an individual news article. In terms of robustness evaluations, we randomly select 20% words from class <italic>Event</italic> and class <italic>Normality</italic>, and exchange them. Figure <xref rid="j_nejsds071_fig_004">4</xref> gives an illustrative example of the data processing and robustness evaluations.</p>
<p>We created an extremely imbalanced text classification task. For the training task, we treated 30,000 “Sports” news articles as <italic>Normality</italic> and 100 randomly selected news articles from other categories as <italic>Event</italic> types. Within each task, we repeat the random samplings for <italic>Event</italic> types ten times to alleviate randomness. In task 1, we use the news articles from the “World” category as class <italic>Event</italic> while in task 2 and 3, we set “Business” and “Sci/Tech” separately as <italic>Event</italic> types. For out-of-sample model performance evaluation, 1,900 news articles for each category made up the testing sample. For VI-REM, we set a <inline-formula id="j_nejsds071_ineq_256"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times 2$]]></tex-math></alternatives></inline-formula> FC neural layer for <inline-formula id="j_nejsds071_ineq_257"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${q_{\phi }}(\mathbf{z}|\mathbf{x})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_258"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times 1$]]></tex-math></alternatives></inline-formula> for the individual function <inline-formula id="j_nejsds071_ineq_259"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${f_{j}}({\mathbf{z}_{j}})$]]></tex-math></alternatives></inline-formula> of MANN in <inline-formula id="j_nejsds071_ineq_260"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">θ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${p_{\theta }}(y|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. For S-MLP, F-MLP and MAML, we set the neural network as <inline-formula id="j_nejsds071_ineq_261"><alternatives><mml:math>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>500</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$500\times 500\times 2\times 500\times 500\times 1$]]></tex-math></alternatives></inline-formula> FC layers; for D-SVDD, we set a <inline-formula id="j_nejsds071_ineq_262"><alternatives><mml:math>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>1000</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$1000\times 1000\times 1000\times 2$]]></tex-math></alternatives></inline-formula> network.</p>
<p>Table <xref rid="j_nejsds071_tab_003">3</xref> reports average performance under the original testings and robustness evaluations, and VI-REM shows better generalization ability and more robust model performance in dealing with data perturbations. The MANN structure largely increases the model’s generalization ability through the ablation study in Table <xref rid="j_nejsds071_tab_003">3</xref>. We also check how the model performances will change when setting a different number of latent dimension in Table <xref rid="j_nejsds071_tab_004">4</xref>.</p>
<fig id="j_nejsds071_fig_004">
<label>Figure 4</label>
<caption>
<p>Illustration of Robustness Evaluations for Imbalanced News Detection.</p>
</caption>
<graphic xlink:href="nejsds071_g004.jpg"/>
</fig>
<table-wrap id="j_nejsds071_tab_003">
<label>Table 3</label>
<caption>
<p>News Detection Model performance Comparisons.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Method</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1-Score</sc> ↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="9" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Original</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center">0.06</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_263"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.61</mml:mn></mml:math><tex-math><![CDATA[$-0.61$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.81</td>
<td style="vertical-align: top; text-align: center">0.33</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center">0.60</td>
<td style="vertical-align: top; text-align: center">0.44</td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.99</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.82</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.34</td>
<td style="vertical-align: top; text-align: center">0.43</td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.99</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.83</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.35</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.43</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.98</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.83</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center">0.45</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_264"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.69</mml:mn></mml:math><tex-math><![CDATA[$-0.69$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.60</td>
<td style="vertical-align: top; text-align: center">0.02</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>1.64</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.58</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.98</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.87</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.43</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.52</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.98</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.84</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.68</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.20</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.63</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>VI-REM-V3</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.50</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.03</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.92</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.50</sc></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="9" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Mixture</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center">0.05</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_265"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.95</mml:mn></mml:math><tex-math><![CDATA[$-0.95$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.66</td>
<td style="vertical-align: top; text-align: center">0.06</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center">0.34</td>
<td style="vertical-align: top; text-align: center">0.02</td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.97</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.65</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.11</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_266"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.49</mml:mn></mml:math><tex-math><![CDATA[$-0.49$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.86</td>
<td style="vertical-align: top; text-align: center">0.39</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.12</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.44</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.85</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.43</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center">0.42</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_267"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.75</mml:mn></mml:math><tex-math><![CDATA[$-0.75$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.57</td>
<td style="vertical-align: top; text-align: center">0.01</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.92</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.24</sc></bold></td>
<td style="vertical-align: top; text-align: center"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.74</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.80</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.15</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.69</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.42</sc></td>
<td style="vertical-align: top; text-align: center">−<sc>0.18</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.87</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.47</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>VI-REM-V3</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.32</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.38</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.85</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.32</sc></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap id="j_nejsds071_tab_004">
<label>Table 4</label>
<caption>
<p>News Detection Overall Performance Comparison with Different Number of Latent Variables.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Dim</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1 Score</sc>↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Original</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.64</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.58</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.98</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.87</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>3</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.06</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.42</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.96</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.85</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>5</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.91</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.06</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.99</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.27</sc></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="3" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Mixture</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.92</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.24</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.74</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>3</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.65</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.04</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.91</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.70</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>5</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.69</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">−<sc>0.17</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.95</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.19</sc></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="j_nejsds071_s_018">
<label>6</label>
<title>Real Data Applications</title>
<p>Two real data applications are considered. First application is to detect the fraudulent credit card transactions. The second application identifies traffic crash events based on driving data.</p>
<sec id="j_nejsds071_s_019">
<label>6.1</label>
<title>Application Case 1: Fraud Detection</title>
<p>We implemented the VI-REM model on the fraud detection applications<xref ref-type="fn" rid="j_nejsds071_fn_003">3</xref><fn id="j_nejsds071_fn_003"><label><sup>3</sup></label>
<p><uri>https://www.kaggle.com/mlg-ulb/creditcardfraud</uri></p></fn> where fraudulent credit card transactions are the rare events of interest. The dataset to which the VI-REM was applied consists of about 284,000 financial transactions associated with 28 predictive features, among which only 492 examples are labeled as <italic>Event</italic> resulting in an extremely low prevalence of <inline-formula id="j_nejsds071_ineq_268"><alternatives><mml:math>
<mml:mn>0.2</mml:mn>
<mml:mi mathvariant="normal">%</mml:mi></mml:math><tex-math><![CDATA[$0.2\% $]]></tex-math></alternatives></inline-formula>. We adopted a 5:5 train test split. To further evaluate model robustness in handling massive data noise, we created an artificially noisy setting by injecting Gaussian noise <inline-formula id="j_nejsds071_ineq_269"><alternatives><mml:math>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{N}(0,1)$]]></tex-math></alternatives></inline-formula> to the testing data, labeled as “Noisy 1” in Table <xref rid="j_nejsds071_tab_005">5</xref>. Further, we impose Gaussian noise <inline-formula id="j_nejsds071_ineq_270"><alternatives><mml:math>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{N}(0,2)$]]></tex-math></alternatives></inline-formula> to testing data and we refer it as “Noisy 2” in Table <xref rid="j_nejsds071_tab_005">5</xref>. The neural network structure is the same as Imbalanced News Detection.</p>
<p>Figure <xref rid="j_nejsds071_fig_005">5</xref>a and Figure <xref rid="j_nejsds071_fig_005">5</xref>b display the learnt representations for the fraud detection dataset in the original testing and noisy testings. Figure <xref rid="j_nejsds071_fig_005">5</xref>a and Figure <xref rid="j_nejsds071_fig_005">5</xref>b confirm that for the VI-REM model, the learnt representation for the <italic>Event</italic> cases lie on the tailed part while <italic>Normality</italic> mostly lie in the center.</p>
<fig id="j_nejsds071_fig_005">
<label>Figure 5</label>
<caption>
<p>The latent representation by different model in testing Set for Fraud Detection. PCA is short as Principal Component Analysis. For VI-REM model, the latents comes from <inline-formula id="j_nejsds071_ineq_271"><alternatives><mml:math>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∼</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">ϕ</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathbf{z}\sim {q_{\boldsymbol{\phi }}}(\mathbf{z}|\mathbf{z})$]]></tex-math></alternatives></inline-formula>. The red points represent the fraud transactions, the green points represent the normal ones.</p>
</caption>
<graphic xlink:href="nejsds071_g005.jpg"/>
</fig>
<table-wrap id="j_nejsds071_tab_005">
<label>Table 5</label>
<caption>
<p>Fraud Detection Model Comparison.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Method</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1-Score</sc> ↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="6" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Original</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center">0.54</td>
<td style="vertical-align: top; text-align: center">0.42</td>
<td style="vertical-align: top; text-align: center">0.64</td>
<td style="vertical-align: top; text-align: center">0.71</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center">1.06</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_272"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.31</mml:mn></mml:math><tex-math><![CDATA[$-0.31$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.71</td>
<td style="vertical-align: top; text-align: center">0.49</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center">2.23</td>
<td style="vertical-align: top; text-align: center">0.63</td>
<td style="vertical-align: top; text-align: center">0.79</td>
<td style="vertical-align: top; text-align: center">0.79</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.89</td>
<td style="vertical-align: top; text-align: center">0.56</td>
<td style="vertical-align: top; text-align: center">0.76</td>
<td style="vertical-align: top; text-align: center">0.77</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center">1.09</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_273"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.01</mml:mn></mml:math><tex-math><![CDATA[$-0.01$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.03</td>
<td style="vertical-align: top; text-align: center">0.00</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>3.72</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.68</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.80</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.81</sc></bold></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="6" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Noisy 1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center">0.07</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_274"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>5.49</mml:mn></mml:math><tex-math><![CDATA[$-5.49$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.36</td>
<td style="vertical-align: top; text-align: center">0.17</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center">0.54</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_275"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>2.16</mml:mn></mml:math><tex-math><![CDATA[$-2.16$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.67</td>
<td style="vertical-align: top; text-align: center">0.26</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.76</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_276"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.61</mml:mn></mml:math><tex-math><![CDATA[$-0.61$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.68</td>
<td style="vertical-align: top; text-align: center">0.44</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.31</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_277"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>1.56</mml:mn></mml:math><tex-math><![CDATA[$-1.56$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.46</td>
<td style="vertical-align: top; text-align: center">0.32</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center">0.99</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_278"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.01</mml:mn></mml:math><tex-math><![CDATA[$-0.01$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.02</td>
<td style="vertical-align: top; text-align: center">0.00</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>2.83</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.57</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.72</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.73</sc></bold></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="6" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>Noisy 2</sc></td>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center">0.02</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_279"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>30.2</mml:mn></mml:math><tex-math><![CDATA[$-30.2$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.33</td>
<td style="vertical-align: top; text-align: center">0.04</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center">0.30</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_280"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>4.46</mml:mn></mml:math><tex-math><![CDATA[$-4.46$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.53</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.17</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.28</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_281"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>2.82</mml:mn></mml:math><tex-math><![CDATA[$-2.82$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.41</td>
<td style="vertical-align: top; text-align: center">0.22</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center">0.13</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_282"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>5.10</mml:mn></mml:math><tex-math><![CDATA[$-5.10$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.21</td>
<td style="vertical-align: top; text-align: center">0.14</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center">0.55</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_283"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.26</mml:mn></mml:math><tex-math><![CDATA[$-0.26$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.01</td>
<td style="vertical-align: top; text-align: center">0.00</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>1.32</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.16</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.53</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.56</sc></bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
<p>Table <xref rid="j_nejsds071_tab_005">5</xref> reports results for both original and noisy settings. The proposed VI-REM model consistently outperformed the competing models in terms of accuracy and robustness. Under noisy testing tasks, which sometimes can be associated with data quality issues in real-world settings, the VI-REM shows significantly improved model robustness compared to all SOTAs. VI-REM’s supreme performance on robustness to data noises indicates it can learn ubiquitous discriminative information. We notice that under “Noisy 1” testings, our model VI-REM still has reliable prediction performance while the other competitors are like a random guesses.</p>
<p>Next, we investigate the relationship between latents <bold>z</bold> and label <italic>y</italic>. Figure <xref rid="j_nejsds071_fig_006">6</xref>a plots the decision boundary and associated predicted probability for determining the <italic>Event</italic> cases under original testings. It is clear that the larger distance with the distribution center, the more likely it will be class <italic>Event</italic>. Figure <xref rid="j_nejsds071_fig_006">6</xref>b details the relationship between latent factors <inline-formula id="j_nejsds071_ineq_284"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_285"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> and label <italic>y</italic>. Latent <inline-formula id="j_nejsds071_ineq_286"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> has a negative relationship with the occurrences of the events while <inline-formula id="j_nejsds071_ineq_287"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> has a positive relationship.</p>
<p>We identify the relationship between input features <bold>x</bold> and latent <bold>z</bold>. We consider the 28 individual features and the interaction of any two of the input features. In total, <inline-formula id="j_nejsds071_ineq_288"><alternatives><mml:math>
<mml:mfenced separators="" open="(" close=")">
<mml:mfrac linethickness="0.0pt">
<mml:mrow>
<mml:mn>28</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mfenced>
<mml:mo>+</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mfrac linethickness="0.0pt">
<mml:mrow>
<mml:mn>28</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mfenced>
<mml:mo>=</mml:mo>
<mml:mn>406</mml:mn></mml:math><tex-math><![CDATA[$\left(\genfrac{}{}{0.0pt}{}{28}{1}\right)+\left(\genfrac{}{}{0.0pt}{}{28}{2}\right)=406$]]></tex-math></alternatives></inline-formula> combinations are considered for model interpretations. We implement the strategy discussed in Section <xref rid="j_nejsds071_s_011">4.1</xref> where we <inline-formula id="j_nejsds071_ineq_289"><alternatives><mml:math>
<mml:mi mathvariant="italic">q</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>5</mml:mn></mml:math><tex-math><![CDATA[$q=5$]]></tex-math></alternatives></inline-formula> for each round of selections. Table <xref rid="j_nejsds071_tab_006">6</xref> reports the major components of latent factor <inline-formula id="j_nejsds071_ineq_290"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_291"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> respectively, where <inline-formula id="j_nejsds071_ineq_292"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>22</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{22}}$]]></tex-math></alternatives></inline-formula> represents the 22th variable of the input features. The most contributed part of <inline-formula id="j_nejsds071_ineq_293"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_294"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> do not overlap, confirming the disentanglement. We also fit a linear regression between <inline-formula id="j_nejsds071_ineq_295"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_296"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> and their selected major components, and it is reported as column “Coefficients” in Table <xref rid="j_nejsds071_tab_006">6</xref>. For example, the fitted coefficient between the combination of <inline-formula id="j_nejsds071_ineq_297"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{16}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_298"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>24</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{24}}$]]></tex-math></alternatives></inline-formula> and latent <inline-formula id="j_nejsds071_ineq_299"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> is 0.024, indicating they have a positive relationship.</p>
<table-wrap id="j_nejsds071_tab_006">
<label>Table 6</label>
<caption>
<p>Model Interpretations for Latent Factor <inline-formula id="j_nejsds071_ineq_300"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_301"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> in Application Case 1 and Application Case 2.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"> Selected Variable  </td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin">  Coefficient  </td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5" style="vertical-align: middle; text-align: center; border-bottom: solid thin">Case 1</td>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_302"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_303"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>22</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{22}}\times {V_{26}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.001</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_304"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_305"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>24</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{16}}\times {V_{24}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.024</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_306"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_307"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>18</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{16}}\times {V_{18}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.041</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_308"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_309"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>24</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{13}}\times {V_{24}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.006</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">Latent Variable <inline-formula id="j_nejsds071_ineq_310"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_311"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>23</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{16}}\times {V_{23}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.011</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="5" style="vertical-align: middle; text-align: center; border-bottom: solid thin">Case 1</td>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_312"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_313"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{3}}\times {V_{13}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.010</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_314"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_315"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>7</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{7}}\times {V_{13}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.002</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_316"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_317"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>19</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{13}}\times {V_{19}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.040</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_318"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_319"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{3}}\times {V_{16}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.008</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">Latent Variable <inline-formula id="j_nejsds071_ineq_320"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_321"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${V_{13}}\times {V_{20}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.023</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="5" style="vertical-align: middle; text-align: center; border-bottom: solid thin">Case 2</td>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_322"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_323"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-x</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-y</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-x}_{26}}\times {\text{acc-y}_{26}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.747</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_324"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_325"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>29</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-z}_{26}}\times {\text{acc-z}_{29}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−2.042</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_326"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_327"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>28</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-z}_{26}}\times {\text{acc-z}_{28}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.133</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">Latent Variable <inline-formula id="j_nejsds071_ineq_328"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_329"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>25</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>27</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-z}_{25}}\times {\text{acc-z}_{27}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">−0.973</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">Latent Variable <inline-formula id="j_nejsds071_ineq_330"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_331"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>×</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-z</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>30</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-z}_{26}}\times {\text{acc-z}_{30}}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">1.896</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="j_nejsds071_fig_006">
<label>Figure 6</label>
<caption>
<p>Model interpretations for VI-REM in fraud detection. (a)The decision boundary and associated predicted probability for determining the <italic>Event</italic> cases under original testings. (b) Latent factors plot. Part (1) and part (3) of displays the probability density functions of latent variable 1 and latent variable 2, where the red lines represent the <italic>Event</italic> class and the green lines represent the <italic>Normality</italic> class.Part (3) and (4) presents the marginal relationship of the latent factors <inline-formula id="j_nejsds071_ineq_332"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_333"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> with the Logit value of condition probability, that is <inline-formula id="j_nejsds071_ineq_334"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\log \Big(\frac{\Pr (y|{\mathbf{z}_{1}})}{1-\Pr (y|{\mathbf{z}_{1}})}\Big)$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_335"><alternatives><mml:math>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\log \Big(\frac{\Pr (y|{\mathbf{z}_{2}})}{1-\Pr (y|{\mathbf{z}_{1}})}\Big)$]]></tex-math></alternatives></inline-formula> for <inline-formula id="j_nejsds071_ineq_336"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_337"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula>.</p>
</caption>
<graphic xlink:href="nejsds071_g006.jpg"/>
</fig>
</sec>
<sec id="j_nejsds071_s_020">
<label>6.2</label>
<title>Application Case 2: Traffic Crash Identification</title>
<p>Traffic crashes are rare events, with an average rate of one crash 6.8 per million vehicle miles traveled in the US [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>]. The naturalistic driving study (NDS) provides an unprecedented opportunity to evaluate crash risk [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>]. NDSs are characterized by continuously recording driving information, such as three-dimensional Inertial Measurement Unit acceleration and multi-channel video recordings. Accurately identifying crashes with robustness benefits further understanding of driving behavior with the overall goal of reducing traffic accidents [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>], identifying high risky segments at road network [<xref ref-type="bibr" rid="j_nejsds071_ref_052">52</xref>] or improving the validation efficiency for autonomous vehicles [<xref ref-type="bibr" rid="j_nejsds071_ref_034">34</xref>].</p>
<p>This experiment applied the VI-REM to the Second Strategic Highway Research Program (SHRP2) NDS [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>], which is largest NDS to-date, with more than 1 million hours of continuous driving data. SHRP2 Insight website<xref ref-type="fn" rid="j_nejsds071_fn_004">4</xref><fn id="j_nejsds071_fn_004"><label><sup>4</sup></label>
<p><uri>https://insight.shrp2nds.us/</uri></p></fn> gives a more detailed introduction. We applied the VI-REM to identify traffic crashes based on the high-frequency longitudinal, lateral and vertical tri-axial acceleration for the vehicles.</p>
<p>SHRP2 crash events have four different severity levels [<xref ref-type="bibr" rid="j_nejsds071_ref_010">10</xref>], ranging from level 4 to level 1. Level 4 (L4) crash events are low-risk tire strikes, level 3 (L3) crash events are minor crashes, level 2 (L2) crash events are police reportable crashes with at least 1,500 dollars worth of damage, and level 1 (L1) are fatal crash events. Our data consists of 100 level 1, 150 level 2, 578 level 3, and 588 level 4 crash events. After a random selection process, we also collected nearly 120,000 safe driving segments from SHRP2. These selected safe driving segments have similar maximum and minimum speeds with crashes. Table <xref rid="j_nejsds071_tab_007">7</xref> summarizes the detail of different events and the last column calculates the imbalance ratio.</p>
<table-wrap id="j_nejsds071_tab_007">
<label>Table 7</label>
<caption>
<p>Summary Description of Level 1 to Level 4 Crash Events and Normal Drivings.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Description</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Number</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Ratio</sc></td>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: top; text-align: center">L1 Crash</td>
<td style="vertical-align: top; text-align: center">Fatal Crash</td>
<td style="vertical-align: top; text-align: center">100</td>
<td style="vertical-align: top; text-align: center">1200:1</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">L2 Crash</td>
<td style="vertical-align: top; text-align: center">Police Reportable Crashes</td>
<td style="vertical-align: top; text-align: center">150</td>
<td style="vertical-align: top; text-align: center">800:1</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">L3 Crash</td>
<td style="vertical-align: top; text-align: center">Minor Crash and MC</td>
<td style="vertical-align: top; text-align: center">578</td>
<td style="vertical-align: top; text-align: center">208:1</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center">L4 Crash</td>
<td style="vertical-align: top; text-align: center">Low-risk Tire Strikes</td>
<td style="vertical-align: top; text-align: center">588</td>
<td style="vertical-align: top; text-align: center">204:1</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">Normal Driving</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">Routine Safe Driving</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">119,996</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">1:1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We applied the ordinal-based VI-REM model to the SHRP2 data, as the label of events has different severity levels. To standardize the data, we put the most volatile time in the middle of each driving segment. Each driving segment has 5.1 seconds and as the collection frequency for acceleration is 10 HZ, so the length for each segment is <inline-formula id="j_nejsds071_ineq_338"><alternatives><mml:math>
<mml:mn>5.1</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>10</mml:mn>
<mml:mo>=</mml:mo>
<mml:mn>51</mml:mn></mml:math><tex-math><![CDATA[$5.1\times 10=51$]]></tex-math></alternatives></inline-formula>. As the acceleration has three directions, the dimensionality for the input features will be <inline-formula id="j_nejsds071_ineq_339"><alternatives><mml:math>
<mml:mn>51</mml:mn>
<mml:mo>×</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>=</mml:mo>
<mml:mn>153</mml:mn></mml:math><tex-math><![CDATA[$51\times 3=153$]]></tex-math></alternatives></inline-formula>. Figure <xref rid="j_nejsds071_fig_007">7</xref>a shows several driving segments. Our experiment adopted a 5/5 split for training and testing, and the neural network structure for each model is the same as Imbalanced Digit Recognition.</p>
<fig id="j_nejsds071_fig_007">
<label>Figure 7</label>
<caption>
<p>Several plots for identifying traffic crash events in SHRP2 data (a) Several driving segments are associated with the three-dimensional accelerations under different event categories. (b) Part (1) and Part (2) represent the kernel density estimation of latent <inline-formula id="j_nejsds071_ineq_340"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_341"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> given different event categories. Part (3) and Part (4) display the associated probability of different event categories given different values on <inline-formula id="j_nejsds071_ineq_342"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_343"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula>.</p>
</caption>
<graphic xlink:href="nejsds071_g007.jpg"/>
</fig>
<p>Our work transforms the assessment of an ordinal model into assessing the performance of a classification model based on different selections on the severity level. For example: 
<disp-formula id="j_nejsds071_eq_039">
<label>(6.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mtext>PR-AUC</mml:mtext>
<mml:mspace width="2.5pt"/>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo><mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">Y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="script">I</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">Y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mtext>and</mml:mtext>
<mml:mspace width="1em"/>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mtext>L1, L2, L3, L4</mml:mtext>
<mml:mo fence="true" stretchy="false">]</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}& \text{PR-AUC}\hspace{2.5pt}\Big(\widehat{\Pr }(Y\ge j),\mathcal{I}(Y\ge j)\Big)\hspace{1em}\\ {} & \text{and}\hspace{1em}j\in [\text{L1, L2, L3, L4}]\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_344"><alternatives><mml:math>
<mml:mi mathvariant="italic">Y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>…</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$Y=({y_{1}},\dots ,{y_{N}})$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_345"><alternatives><mml:math><mml:mover accent="false">
<mml:mrow>
<mml:mo movablelimits="false">Pr</mml:mo>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">Y</mml:mi>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\widehat{\Pr }(Y\ge j)$]]></tex-math></alternatives></inline-formula> is the predicted probability from model in Equation <xref rid="j_nejsds071_eq_036">4.7</xref>. Similar transformations will also be applied to the other metrics. Table <xref rid="j_nejsds071_tab_008">8</xref> compares the performances of different models for identifying different severity traffic crashes, where we skip the Deep-SVDD model as it is difficult to extend to ordinal data conditions. Table <xref rid="j_nejsds071_tab_008">8</xref> shows VI-REM’s supreme performances over other competitive models, especially for severe level 1, level 2 and level 3 crash events. Overall for identifying crash events (level 4 up to level 1), our proposed VI-REM can have 0.77 in F1 score, while the most competitive F1 score from benchmark model is 0.68 by Sampling MLP.</p>
<p>Analytical modeling for identifying traffic crashes is sensitive to kinematic noise (such as signal connection issues or different driving environments), where model performance can largely deteriorate as noise increases [<xref ref-type="bibr" rid="j_nejsds071_ref_030">30</xref>]. Considering this, where model robustness is fundamentally important for decision making, we created an artificially noisy setting where Gaussian noise <inline-formula id="j_nejsds071_ineq_346"><alternatives><mml:math>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mn>0.05</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathcal{N}(0,0.05)$]]></tex-math></alternatives></inline-formula> was injected to the input accelerations. The model performance metrics are presented in Table <xref rid="j_nejsds071_tab_009">9</xref>. As the table shows, the proposed VI-REM significantly outperforms alternative models based on the evaluation metrics. Compared with benchmarks, the VI-REM shows much more robust disincentive power under noisy testing conditions.</p>
<p>As for the model interpretations, Figure <xref rid="j_nejsds071_fig_007">7</xref>b shows the mutual relationship between <bold>z</bold> and label <italic>y</italic>. <inline-formula id="j_nejsds071_ineq_347"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> has larger discriminative over different crash events, and larger values in <inline-formula id="j_nejsds071_ineq_348"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> are associated with more considerable severity. We implement the same strategy in application case 1 to identify the relationship between input features <bold>x</bold> and latent <bold>z</bold>. We report the identified major components of <inline-formula id="j_nejsds071_ineq_349"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{1}}$]]></tex-math></alternatives></inline-formula> in Table <xref rid="j_nejsds071_tab_006">6</xref>, as latent <inline-formula id="j_nejsds071_ineq_350"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathbf{z}_{2}}$]]></tex-math></alternatives></inline-formula> does not have strong distinguishing ability. In Table <xref rid="j_nejsds071_tab_006">6</xref>, <inline-formula id="j_nejsds071_ineq_351"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mtext>acc-x</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>26</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\text{acc-x}_{26}}$]]></tex-math></alternatives></inline-formula> represents the longitudinal acceleration at point 26, and <inline-formula id="j_nejsds071_ineq_352"><alternatives><mml:math>
<mml:mtext>acc-y</mml:mtext></mml:math><tex-math><![CDATA[$\text{acc-y}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_353"><alternatives><mml:math>
<mml:mtext>acc-z</mml:mtext></mml:math><tex-math><![CDATA[$\text{acc-z}$]]></tex-math></alternatives></inline-formula> represent the lateral and vertical acceleration respectively. All the major contributing input features are around the middle of the driving segment through the model interpretations. This phenomenon is consistent with our experiment setup. We also find that vertical accelerations could be the most influential factor for identifying traffic crashes and determining their associated severity level. This observation is consistent with various empirical studies [<xref ref-type="bibr" rid="j_nejsds071_ref_030">30</xref>]</p>
<table-wrap id="j_nejsds071_tab_008">
<label>Table 8</label>
<caption>
<p>Performance Comparisons in Traffic Crash Identifications.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Method</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1-Score</sc> ↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_354"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.58</td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.58</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.70</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.77</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_355"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>2.50</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.52</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.69</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.68</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_356"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>3.32</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.59</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.77</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.74</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_357"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>2.71</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.56</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.78</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.71</sc></bold></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_358"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.65</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_359"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.23</mml:mn></mml:math><tex-math><![CDATA[$-0.23$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.49</td>
<td style="vertical-align: top; text-align: center">0.52</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_360"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.48</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_361"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.86</mml:mn></mml:math><tex-math><![CDATA[$-0.86$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.27</td>
<td style="vertical-align: top; text-align: center">0.31</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_362"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.62</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_363"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.61</mml:mn></mml:math><tex-math><![CDATA[$-0.61$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.38</td>
<td style="vertical-align: top; text-align: center">0.36</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_364"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.94</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_365"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.21</mml:mn></mml:math><tex-math><![CDATA[$-0.21$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.50</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.42</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_366"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>1.91</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.42</td>
<td style="vertical-align: top; text-align: center">0.64</td>
<td style="vertical-align: top; text-align: center">0.56</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_367"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">2.16</td>
<td style="vertical-align: top; text-align: center">0.44</td>
<td style="vertical-align: top; text-align: center">0.67</td>
<td style="vertical-align: top; text-align: center">0.64</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_368"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.67</td>
<td style="vertical-align: top; text-align: center">0.40</td>
<td style="vertical-align: top; text-align: center">0.65</td>
<td style="vertical-align: top; text-align: center">0.61</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_369"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">1.42</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.30</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.53</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.49</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_370"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.43</td>
<td style="vertical-align: top; text-align: center">0.26</td>
<td style="vertical-align: top; text-align: center">0.57</td>
<td style="vertical-align: top; text-align: center">0.65</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_371"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.36</td>
<td style="vertical-align: top; text-align: center">0.08</td>
<td style="vertical-align: top; text-align: center">0.63</td>
<td style="vertical-align: top; text-align: center">0.66</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_372"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.68</td>
<td style="vertical-align: top; text-align: center">0.24</td>
<td style="vertical-align: top; text-align: center">0.54</td>
<td style="vertical-align: top; text-align: center">0.54</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_373"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">1.44</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.09</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.57</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.51</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_374"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.71</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_375"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.01</mml:mn></mml:math><tex-math><![CDATA[$-0.01$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.70</td>
<td style="vertical-align: top; text-align: center">0.68</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_376"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.25</td>
<td style="vertical-align: top; text-align: center">0.40</td>
<td style="vertical-align: top; text-align: center">0.65</td>
<td style="vertical-align: top; text-align: center">0.60</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_377"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.15</td>
<td style="vertical-align: top; text-align: center">0.55</td>
<td style="vertical-align: top; text-align: center">0.72</td>
<td style="vertical-align: top; text-align: center">0.68</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_378"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.66</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.49</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.62</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.58</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap id="j_nejsds071_tab_009">
<label>Table 9</label>
<caption>
<p>Robustness Evaluations in Traffic Crash Identifications.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"/>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Method</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>R-NLL</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>BSS</sc> ↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>PR-AUC</sc>↑</td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>F1-Score</sc> ↑</td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><bold><sc>VI-REM</sc></bold></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_379"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.48</td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.48</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.69</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.73</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_380"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>2.43</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.47</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.68</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.66</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_381"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>3.28</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.59</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.77</sc></bold></td>
<td style="vertical-align: top; text-align: center"><bold><sc>0.74</sc></bold></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_382"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>2.57</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.57</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.78</sc></bold></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><bold><sc>0.72</sc></bold></td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_383"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.38</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_384"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>1.09</mml:mn></mml:math><tex-math><![CDATA[$-1.09$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.42</td>
<td style="vertical-align: top; text-align: center">0.38</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_385"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.38</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_386"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>1.42</mml:mn></mml:math><tex-math><![CDATA[$-1.42$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.24</td>
<td style="vertical-align: top; text-align: center">0.23</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_387"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.38</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_388"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.84</mml:mn></mml:math><tex-math><![CDATA[$-0.84$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.31</td>
<td style="vertical-align: top; text-align: center">0.31</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_389"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.67</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_390"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.64</mml:mn></mml:math><tex-math><![CDATA[$-0.64$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.43</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.38</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_391"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center"><bold><sc>1.85</sc></bold></td>
<td style="vertical-align: top; text-align: center">0.40</td>
<td style="vertical-align: top; text-align: center">0.58</td>
<td style="vertical-align: top; text-align: center">0.56</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_392"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.99</td>
<td style="vertical-align: top; text-align: center">0.39</td>
<td style="vertical-align: top; text-align: center">0.64</td>
<td style="vertical-align: top; text-align: center">0.62</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_393"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.60</td>
<td style="vertical-align: top; text-align: center">0.38</td>
<td style="vertical-align: top; text-align: center">0.61</td>
<td style="vertical-align: top; text-align: center">0.60</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_394"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">1.43</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.30</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.52</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.49</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_395"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.87</td>
<td style="vertical-align: top; text-align: center">0.05</td>
<td style="vertical-align: top; text-align: center">0.58</td>
<td style="vertical-align: top; text-align: center">0.61</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_396"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.66</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_397"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.65</mml:mn></mml:math><tex-math><![CDATA[$-0.65$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.62</td>
<td style="vertical-align: top; text-align: center">0.63</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_398"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.87</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_399"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.42</mml:mn></mml:math><tex-math><![CDATA[$-0.42$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.45</td>
<td style="vertical-align: top; text-align: center">0.48</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_400"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.63</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_401"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>1.34</mml:mn></mml:math><tex-math><![CDATA[$-1.34$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.45</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.29</td>
</tr>
</tbody><tbody>
<tr>
<td rowspan="4" style="vertical-align: middle; text-align: center; border-bottom: solid thin"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_402"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L4</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L4}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.55</td>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_403"><alternatives><mml:math>
<mml:mo>−</mml:mo>
<mml:mn>0.53</mml:mn></mml:math><tex-math><![CDATA[$-0.53$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">0.64</td>
<td style="vertical-align: top; text-align: center">0.60</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_404"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L3</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L3}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.00</td>
<td style="vertical-align: top; text-align: center">0.27</td>
<td style="vertical-align: top; text-align: center">0.62</td>
<td style="vertical-align: top; text-align: center">0.60</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><inline-formula id="j_nejsds071_ineq_405"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L2</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L2}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center">1.14</td>
<td style="vertical-align: top; text-align: center">0.55</td>
<td style="vertical-align: top; text-align: center">0.71</td>
<td style="vertical-align: top; text-align: center">0.68</td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><inline-formula id="j_nejsds071_ineq_406"><alternatives><mml:math>
<mml:mtext mathvariant="normal">Y</mml:mtext>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mtext>L1</mml:mtext></mml:math><tex-math><![CDATA[$\text{Y}\ge \text{L1}$]]></tex-math></alternatives></inline-formula></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.65</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.48</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.61</td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin">0.61</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>↑ for higher is better and ↓ for lower is better.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="j_nejsds071_s_021">
<label>6.3</label>
<title>Discussion on the Running Time</title>
<p>Before concluding our paper, we present the average running time of our method VI-REM and other competitive methods across different tasks in Table <xref rid="j_nejsds071_tab_010">10</xref>. Column “CV” represents the task for imbalanced digit recognition in Section <xref rid="j_nejsds071_s_016">5.1</xref>, column “NLP” represents the task for imbalanced detection in Section <xref rid="j_nejsds071_s_017">5.2</xref>, and column “Fraud” represents the task for fraud detection in Section <xref rid="j_nejsds071_s_019">6.1</xref>, and “Crash” represents the task in Section <xref rid="j_nejsds071_s_020">6.2</xref>. Our experiments show that MAML is the most robust among the competitive methods. In terms of efficiency, VI-REM is faster than MAML, especially under the MLP and CV task. We also find that most of the running time for the VI-REM model comes from the MANN part. Designing a more efficient regularized neural networks could be an interesting topic to explore in the future. The GBDT method performs fast under CV and Fraud task, but may experience longer iteration (with an average of 2113 seconds) to converge when dealing with complex text data. Under the OLM setup in Section <xref rid="j_nejsds071_s_020">6.2</xref>, the VI-REM is even faster than S-MLP.</p>
<table-wrap id="j_nejsds071_tab_010">
<label>Table 10</label>
<caption>
<p>Overall Computation Time across Different Methods on Different Tasks. The unit is <inline-formula id="j_nejsds071_ineq_407"><alternatives><mml:math>
<mml:mo>×</mml:mo>
<mml:mn>100</mml:mn></mml:math><tex-math><![CDATA[$\times 100$]]></tex-math></alternatives></inline-formula> Second.</p>
</caption>
<table>
<thead>
<tr>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Model</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>CV</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>NLP</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Fraud</sc></td>
<td style="vertical-align: top; text-align: center; border-top: double; border-bottom: solid thin"><sc>Crash</sc></td>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: top; text-align: center"><sc>GBDT</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.05</sc></td>
<td style="vertical-align: top; text-align: center"><sc>21.13</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.01</sc></td>
<td style="vertical-align: top; text-align: center"><sc>6.44</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>MAML</sc></td>
<td style="vertical-align: top; text-align: center"><sc>4.82</sc></td>
<td style="vertical-align: top; text-align: center"><sc>1.53</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2.78</sc></td>
<td style="vertical-align: top; text-align: center"><sc>3.79</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>F-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.30</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.11</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.35</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.32</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>S-MLP</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.51</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.14</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.32</sc></td>
<td style="vertical-align: top; text-align: center"><sc>3.53</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>D-SVDD</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.25</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.19</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.32</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>-</sc></td>
</tr>
</tbody><tbody>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM</sc></td>
<td style="vertical-align: top; text-align: center"><sc>3.17</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.80</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2.40</sc></td>
<td style="vertical-align: top; text-align: center"><sc>3.04</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center"><sc>VI-REM-V1</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2.95</sc></td>
<td style="vertical-align: top; text-align: center"><sc>0.77</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2.10</sc></td>
<td style="vertical-align: top; text-align: center"><sc>2.89</sc></td>
</tr>
<tr>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>VI-REM-V2</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.28</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.20</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.16</sc></td>
<td style="vertical-align: top; text-align: center; border-bottom: solid thin"><sc>0.27</sc></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="j_nejsds071_s_022">
<label>7</label>
<title>Summary and Discussion</title>
<p>This study proposes a Variational Inference toward extreme approach for rare event modeling (short as VI-REM). Our model induces uncertainty to representation learning through a novel EP distribution and addresses the over-fitting issue through MANN, which leads to a more robust model for rare events. We provide theoretical properties on the efficacy of the proposed model. Extensive empirical experiments are conducted to confirm the model performance over various application scenarios with promising results, especially with respect to model generalization, robustness and interpretability.</p>
<p>The VI-REM framework greatly improves the generalizability and interpretability of rare event modeling, two challenging issues associated with limited numbers of events. The application of VI-REM could accurately depict the risk associated with rare-events and allow the general public and decision-makers to set realistic expectations for rare events. The identification of adverse events at an early stage may allow mitigation of the damage and loss associated with the events. The features identified through the model are crucial for researchers and practitioners to identify the causes of rare events and take proper countermeasures to prevent and reduce the occurrence of future adverse events.</p>
</sec>
</body>
<back>
<app-group>
<app id="j_nejsds071_app_001"><label>Appendix A</label>
<title>Proof of Theorem <xref rid="j_nejsds071_stat_002">3.2</xref></title><statement id="j_nejsds071_stat_005"><label>Proof.</label>
<p>As the label <italic>y</italic> is binary variable, we set a cross entropy loss function, and we can specify <inline-formula id="j_nejsds071_ineq_408"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mathrm{L}_{f}}(\mathbf{z},y)$]]></tex-math></alternatives></inline-formula> as: 
<disp-formula id="j_nejsds071_eq_040">
<label>(A.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnalign="right">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \displaystyle {\mathrm{L}_{f}}(\mathbf{z},y)=y\log f(\mathbf{z})+(1-y)\log (1-f(\mathbf{z}))\]]]></tex-math></alternatives>
</disp-formula> 
So as when <inline-formula id="j_nejsds071_ineq_409"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[$y=1$]]></tex-math></alternatives></inline-formula> or when <inline-formula id="j_nejsds071_ineq_410"><alternatives><mml:math>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$y=0$]]></tex-math></alternatives></inline-formula>, we can have for <bold>z</bold>, <inline-formula id="j_nejsds071_ineq_411"><alternatives><mml:math>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${\mathbf{z}^{\prime }}$]]></tex-math></alternatives></inline-formula> 
<disp-formula id="j_nejsds071_eq_041">
<label>(A.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
<mml:mo>−</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo><mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo movablelimits="false">min</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}({\mathbf{z}^{\prime }},y)\| & \le |\log f(\mathbf{z})-\log f({\mathbf{z}^{\prime }})|\\ {} & =\Big|\log \left(1+\left(\frac{f(\mathbf{z})}{f({\mathbf{z}^{\prime }})}-1\right)\right)\Big|\\ {} & \le \frac{|f(\mathbf{z})-f({\mathbf{z}^{\prime }})|}{\min \big(f(\mathbf{z}),f({\mathbf{z}^{\prime }})\big)}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
Based on condition C.1, the function <inline-formula id="j_nejsds071_ineq_412"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[${\mathrm{L}_{f}}(\mathbf{z},y)$]]></tex-math></alternatives></inline-formula> follows: 
<disp-formula id="j_nejsds071_eq_042">
<label>(A.3)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mo stretchy="false">‖</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo>−</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \| {\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}({\mathbf{z}^{\prime }},y)\| \le {k^{\prime }_{n}}\| \mathbf{z}-{\mathbf{z}^{\prime }}\| \]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_413"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${k^{\prime }_{n}}$]]></tex-math></alternatives></inline-formula> is some constant <inline-formula id="j_nejsds071_ineq_414"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo movablelimits="false">min</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mstyle></mml:math><tex-math><![CDATA[${k^{\prime }_{n}}=\frac{{k_{n}}}{\min \big(f(\mathbf{z}),f({\mathbf{z}^{\prime }})\big)}$]]></tex-math></alternatives></inline-formula></p>
<p>Decompose generalization bound with triangle inequality: 
<disp-formula id="j_nejsds071_eq_043">
<label>(A.4)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Bigg(& \underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\Big|\Bigg)\\ {} & \le \mathbb{E}\Bigg(\underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{f}}(\mathbf{z},y)\big)\Big|\Bigg)\\ {} & +\mathbb{E}\Bigg(\underset{f\in \mathcal{F}}{\sup }\Big|\Pr \big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\Big|\Bigg)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where we know that <inline-formula id="j_nejsds071_ineq_415"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{0}}$]]></tex-math></alternatives></inline-formula> minimize of the loss function form assumption C.2, then we can have the upper bound for: 
<disp-formula id="j_nejsds071_eq_044">
<label>(A.5)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Bigg(& \underset{f\in \mathcal{F}}{\sup }\Big|\Pr \big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\Big|\Bigg)\\ {} & \le \mathbb{E}\Bigg(\Big(\Pr -{\mathrm{P}_{n}}\Big)\Big(\underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}(\mathbf{z},y)\big|\Big)\Bigg)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where we denote this upper bound to be 
<disp-formula id="j_nejsds071_eq_045">
<label>(A.6)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \mathbb{E}\Big(\mathrm{G}(\mathbf{z})\Big)=\mathbb{E}\Big((\Pr -{\mathrm{P}_{n}})\underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}(\mathbf{z},y)\big|\Big)\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Thus the generalization gap can be transferred as: 
<disp-formula id="j_nejsds071_eq_046">
<label>(A.7)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}& \mathbb{E}\Bigg(\underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{{f_{0}}}}(\mathbf{z},y)\big)\Big|\Bigg)\\ {} & \le \mathbb{E}\Big(\underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{f}}(\mathbf{z},y)\big)\Big|\Big)+\mathbb{E}\Big(\mathrm{G}(\mathbf{z})\Big)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>And the function <inline-formula id="j_nejsds071_ineq_416"><alternatives><mml:math>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo></mml:math><tex-math><![CDATA[$\mathrm{G}(\mathbf{z})$]]></tex-math></alternatives></inline-formula> can be decomposed as: 
<disp-formula id="j_nejsds071_eq_047">
<label>(A.8)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo>×</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathrm{G}(\mathbf{z})=(\Pr -{\mathrm{P}_{n}})& \underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}(\mathbf{z},y)\big|\times ({I_{n}^{+}}+{I_{n}^{-}})\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_417"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${I_{n}^{-}}$]]></tex-math></alternatives></inline-formula> as the empirical value of <inline-formula id="j_nejsds071_ineq_418"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${I_{\mathbf{z}\in {\Omega _{\eta }}:\| \mathbf{z}\| \lt \eta }}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_419"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${I_{n}^{+}}$]]></tex-math></alternatives></inline-formula> as the empirical value of <inline-formula id="j_nejsds071_ineq_420"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>:</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${I_{\mathbf{z}\in {\Omega _{\eta }^{C}}:\| \mathbf{z}\| \ge \eta }}$]]></tex-math></alternatives></inline-formula>. It is obvious <inline-formula id="j_nejsds071_ineq_421"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn></mml:math><tex-math><![CDATA[${I_{n}^{+}}+{I_{n}^{-}}=1$]]></tex-math></alternatives></inline-formula>. Thus the expectation generalization bound can be formed as: 
<disp-formula id="j_nejsds071_eq_048">
<label>(A.9)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo fence="true" stretchy="false">}</mml:mo>
<mml:mo>×</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" stretchy="true">|</mml:mo>
<mml:mo fence="true" stretchy="false">}</mml:mo>
<mml:mo>×</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Big(\mathrm{G}(\mathbf{z})\Big)& =\mathbb{E}\Big((\Pr -{\mathrm{P}_{n}})\underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}(\mathbf{z},y)\big|\}\times {I_{n}^{+}}\Big)\\ {} & +\mathbb{E}\Big((\Pr -{\mathrm{P}_{n}})\underset{f\in \mathcal{F}}{\sup }\big|{\mathrm{L}_{f}}(\mathbf{z},y)-{\mathrm{L}_{f}}(\mathbf{z},y)\big|\}\times {I_{n}^{-}}\Big)\\ {} & =\mathbb{E}\Big({\mathrm{G}^{+}}(\mathbf{z})\Big)+\mathbb{E}\Big({\mathrm{G}^{-}}(\mathbf{z})\Big)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>From Lemma 2 and Lemma 3 in [<xref ref-type="bibr" rid="j_nejsds071_ref_012">12</xref>] and Equation <xref rid="j_nejsds071_eq_042">A.3</xref>, we can have the following bound, where <inline-formula id="j_nejsds071_ineq_422"><alternatives><mml:math>
<mml:mi mathvariant="italic">j</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo fence="true" stretchy="false">]</mml:mo></mml:math><tex-math><![CDATA[$j\in [+,-]$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_049">
<label>(A.10)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="italic">ϵ</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \mathbb{E}\Big({\mathrm{G}^{j}}(\mathbf{z})\Big)\le 4{k^{\prime }_{n}}\mathbb{E}\left(\underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\epsilon \big(f(\mathbf{z})-{f_{0}}(\mathbf{z})\big){I_{n}^{j}}\Big|\right)\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_423"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup></mml:math><tex-math><![CDATA[${k^{\prime }_{n}}$]]></tex-math></alternatives></inline-formula> is some positive constant satisfying Lipschitz property, <italic>ϵ</italic> is the Rademacher sequence.</p>
<p>Similar to the Lemma 5 proof in [<xref ref-type="bibr" rid="j_nejsds071_ref_012">12</xref>], Cauchy–Schwarz inequality can bound it further by: 
<disp-formula id="j_nejsds071_eq_050">
<label>(A.11)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mtd>
<mml:mtd>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="italic">ϵ</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="italic">ϵ</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}4{k^{\prime }_{n}}\mathbb{E}& \left(\underset{f\in \mathcal{F}}{\sup }\Big|{\mathrm{P}_{n}}\epsilon \big(f(\mathbf{z})-{f_{0}}(\mathbf{z})\big){I_{n}^{j}}\Big|\right)\\ {} & \le 4{k^{\prime }_{n}}\mathbb{E}\left({\big({\mathrm{P}_{n}}\epsilon {I_{n}^{j}}\big)^{\frac{1}{2}}}\underset{f\in \mathcal{F}}{\sup }|f(\mathbf{z})-{f_{0}}(\mathbf{z}){|^{\frac{1}{2}}}\right)\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Based on the property in Rademacher sequence, we can have the following bound where <italic>c</italic> is some positive constant: 
<disp-formula id="j_nejsds071_eq_051">
<label>(A.12)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="italic">ϵ</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \mathbb{E}\big({\mathrm{P}_{n}}\epsilon \big)\le {c^{2}}/n\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Denote <inline-formula id="j_nejsds071_ineq_424"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi></mml:math><tex-math><![CDATA[${I_{n}^{+}}=h$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_425"><alternatives><mml:math>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi></mml:math><tex-math><![CDATA[${I_{n}^{-}}=1-h$]]></tex-math></alternatives></inline-formula>. Follows the assumption C.3, we can have the bound toward <inline-formula id="j_nejsds071_ineq_426"><alternatives><mml:math>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\mathbb{E}\big({\mathrm{G}^{-}}(\mathbf{z})\big)$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_052">
<label>(A.13)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="italic">c</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \mathbb{E}\Big({\mathrm{G}^{-}}(\mathbf{z})\Big)\le 4{k^{\prime }_{n}}c{n^{-\frac{1}{2}}}\| \mathbf{z}{\| ^{1/2}}{(1-h)^{1/2}}\underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }^{C}}}^{1/2}}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Similarly, we can have the bound toward <inline-formula id="j_nejsds071_ineq_427"><alternatives><mml:math>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo></mml:math><tex-math><![CDATA[$\mathbb{E}\big({\mathrm{G}^{+}}(\mathbf{z})\big)$]]></tex-math></alternatives></inline-formula>: 
<disp-formula id="j_nejsds071_eq_053">
<label>(A.14)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>+</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="italic">c</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo><mml:mstyle displaystyle="false">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \mathbb{E}\Big({\mathrm{G}^{+}}(\mathbf{z})\Big)\le 4{k^{\prime }_{n}}c{n^{-\frac{1}{2}}}{h^{1/2}}{\eta ^{1/2}}\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }}}^{1/2}}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>For simplicity, we set <inline-formula id="j_nejsds071_ineq_428"><alternatives><mml:math>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>4</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="italic">k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mi mathvariant="italic">c</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\gamma =4{k^{\prime }_{n}}c{n^{-1/2}}$]]></tex-math></alternatives></inline-formula> and we can have: 
<disp-formula id="j_nejsds071_eq_054">
<label>(A.15)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="normal">G</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo stretchy="false">≤</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
</mml:mtd>
<mml:mtd>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Big(\mathrm{G}(\mathbf{z})\Big)\le \gamma & {h^{1/2}}{\eta ^{1/2}}\underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }}}^{1/2}}\\ {} & +\gamma \underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\sup }\| \mathbf{z}{\| ^{1/2}}{(1-h)^{1/2}}\underset{f\in \mathcal{F}}{\sup }\| f-{f_{0}}{\| _{\mathbf{z}\in {\Omega _{\eta }^{C}}}^{1/2}}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>And similar to the classification generalization bound: 
<disp-formula id="j_nejsds071_eq_055">
<label>(A.16)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">(</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>−</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mi mathvariant="italic">y</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" stretchy="true">|</mml:mo>
<mml:mo maxsize="2.45em" minsize="2.45em" fence="true" mathvariant="normal">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≤</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mfrac>
<mml:mrow>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo movablelimits="false">log</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mi mathvariant="italic">δ</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mstyle>
</mml:mrow>
</mml:msqrt>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\mathbb{E}\Bigg(\underset{f\in \mathcal{F}}{\sup }\Big|& {\mathrm{P}_{n}}\big({\mathrm{L}_{f}}(\mathbf{z},y)\big)-\Pr \big({\mathrm{L}_{f}}(\mathbf{z},y)\big)\Big|\Bigg)\\ {} & \le \sqrt{\frac{\log |\mathcal{F}|+\log (1/\delta )}{2n}}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_429"><alternatives><mml:math>
<mml:mi mathvariant="italic">δ</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:mi mathvariant="script">F</mml:mi>
<mml:mo stretchy="false">|</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="italic">n</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">ϕ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$\delta =|\mathcal{F}|{e^{-2n{\phi ^{2}}}}$]]></tex-math></alternatives></inline-formula> where <italic>ϕ</italic> is some constant. Combine Equation <xref rid="j_nejsds071_eq_054">A.15</xref> and Equation <xref rid="j_nejsds071_eq_055">A.16</xref> based on Equation <xref rid="j_nejsds071_eq_046">A.7</xref>, we can have Theorem <xref rid="j_nejsds071_stat_002">3.2</xref>.  □</p></statement></app>
<app id="j_nejsds071_app_002"><label>Appendix B</label>
<title>Proof of the Remark <xref rid="j_nejsds071_stat_003">1</xref></title><statement id="j_nejsds071_stat_006"><label>Proof.</label>
<p>To simplify the proof, we denote the generalization gap for function <italic>f</italic> as <inline-formula id="j_nejsds071_ineq_430"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\mathcal{E}_{f}}$]]></tex-math></alternatives></inline-formula>. Then, the difference of the generalization bound by <inline-formula id="j_nejsds071_ineq_431"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_432"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${f_{2}}$]]></tex-math></alternatives></inline-formula> can be expressed as: 
<disp-formula id="j_nejsds071_eq_056">
<label>(B.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">≃</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\mathcal{E}_{{f_{1}}}}-{\mathcal{E}_{{f_{2}}}}\simeq {\Delta _{2}}-{\Delta _{1}}\]]]></tex-math></alternatives>
</disp-formula> 
where <inline-formula id="j_nejsds071_ineq_433"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\Delta _{2}}=\gamma {h^{1/2}}{\eta ^{1/2}}{\Delta _{\Omega }}$]]></tex-math></alternatives></inline-formula>. and 
<disp-formula id="j_nejsds071_eq_057">
<alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">max</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">sup</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo movablelimits="false">sup</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\Delta _{1}}=\underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\max }\Big(\sup \| \mathbf{z}{\| ^{1/2}},\sup \| {\mathbf{z}^{\prime }}{\| ^{1/2}}\Big)\gamma {(1-h)^{1/2}}{\Delta _{{\Omega _{\eta }^{C}}}}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>As <inline-formula id="j_nejsds071_ineq_434"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${\sup _{\mathbf{z}\in {\Omega _{\eta }^{C}}}}\| \mathbf{z}{\| ^{1/2}}$]]></tex-math></alternatives></inline-formula> and <inline-formula id="j_nejsds071_ineq_435"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">‖</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[${\sup _{{\mathbf{z}^{\prime }}\in {\Omega _{\eta }^{C}}}}\| {\mathbf{z}^{\prime }}{\| ^{1/2}}$]]></tex-math></alternatives></inline-formula> are at the same magnitude, we assume the following condition: 
<disp-formula id="j_nejsds071_eq_058">
<label>(B.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">max</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mo movablelimits="false">sup</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal">,</mml:mo>
<mml:mo movablelimits="false">sup</mml:mo>
<mml:mo stretchy="false">‖</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>′</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true" mathvariant="normal">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="italic">r</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\max }\Big(\sup \| \mathbf{z}{\| ^{1/2}},\sup \| {\mathbf{z}^{\prime }}{\| ^{1/2}}\Big)=r\underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\sup }\| \mathbf{z}{\| ^{1/2}}\]]]></tex-math></alternatives>
</disp-formula> 
where <italic>r</italic> is some constant. So that: 
<disp-formula id="j_nejsds071_eq_059">
<label>(B.3)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">≃</mml:mo>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">r</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo movablelimits="false">sup</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:mo stretchy="false">∈</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">‖</mml:mo>
<mml:mi mathvariant="bold">z</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">‖</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="italic">γ</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">h</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">Ω</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">η</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">C</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ {\mathcal{E}_{{f_{1}}}}-{\mathcal{E}_{{f_{2}}}}\simeq \gamma {h^{1/2}}{\eta ^{1/2}}{\Delta _{\Omega }}-r\underset{\mathbf{z}\in {\Omega _{\eta }^{C}}}{\sup }\| \mathbf{z}{\| ^{1/2}}\gamma {(1-h)^{1/2}}{\Delta _{{\Omega _{\eta }^{C}}}}\]]]></tex-math></alternatives>
</disp-formula>
</p>
<p>Intuitively, as long as <inline-formula id="j_nejsds071_ineq_436"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">Δ</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${\Delta _{2}}\lt {\Delta _{1}}$]]></tex-math></alternatives></inline-formula>, the overall model performance will be improved. As long as the Equation <xref rid="j_nejsds071_eq_028">3.18</xref> satisfies, we can have the statement in Remark <xref rid="j_nejsds071_stat_003">1</xref>.  □</p></statement></app>
<app id="j_nejsds071_app_003"><label>Appendix C</label>
<title>Proof of Theorem <xref rid="j_nejsds071_stat_004">4.1</xref></title><statement id="j_nejsds071_stat_007"><label>Proof.</label>
<p>For each round of the selections for model interpretation as in Equation <xref rid="j_nejsds071_eq_033">4.4</xref>, we have the following from[<xref ref-type="bibr" rid="j_nejsds071_ref_027">27</xref>]: 
<disp-formula id="j_nejsds071_eq_060">
<label>(C.1)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mtd>
<mml:mtd>
<mml:mo stretchy="false">⊆</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo fence="true" stretchy="false">}</mml:mo>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="script">O</mml:mi>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">{</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo fence="true" stretchy="false">]</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo>×</mml:mo>
<mml:mi mathvariant="italic">s</mml:mi>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">γ</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">}</mml:mo>
<mml:mspace width="2.5pt"/>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\Pr \{{\mathcal{D}_{\mathbf{z}}^{\ast }}& \subseteq {\widehat{\mathcal{D}}_{g}}\}\ge 1-\mathcal{O}\Big\{d\exp [-{c_{1}}{s^{1-2(u+v)}}]\\ {} & +d\times s\exp (-{c_{2}}{s^{\gamma }})\Big\}\hspace{2.5pt}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
where <italic>s</italic> is the sample size group <italic>g</italic>, <italic>d</italic> is the dimension of input features, and <inline-formula id="j_nejsds071_ineq_437"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${c_{1}}$]]></tex-math></alternatives></inline-formula>, <inline-formula id="j_nejsds071_ineq_438"><alternatives><mml:math>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub></mml:math><tex-math><![CDATA[${c_{2}}$]]></tex-math></alternatives></inline-formula>, <italic>u</italic>, <italic>v</italic> are some constant and <inline-formula id="j_nejsds071_ineq_439"><alternatives><mml:math>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
<mml:mo mathvariant="normal">&lt;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="normal" stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn></mml:math><tex-math><![CDATA[$u+v\lt 1/2$]]></tex-math></alternatives></inline-formula>.</p>
<p>Thus: 
<disp-formula id="j_nejsds071_eq_061">
<label>(C.2)</label><alternatives><mml:math display="block">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtable displaystyle="true" columnspacing="0pt" columnalign="right left">
<mml:mtr>
<mml:mtd>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mfenced separators="" open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">⊆</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mstyle><mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover></mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub></mml:mstyle>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
<mml:mtd>
<mml:mo>=</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">⊆</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo fence="true" stretchy="false">}</mml:mo>
<mml:mo>×</mml:mo>
<mml:mo stretchy="false">⋯</mml:mo>
<mml:mo>×</mml:mo>
<mml:mo movablelimits="false">Pr</mml:mo>
<mml:mo fence="true" stretchy="false">{</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>∗</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">⊆</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="false">
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mo stretchy="true">ˆ</mml:mo></mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo fence="true" stretchy="false">}</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo stretchy="false">≥</mml:mo>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">{</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="script">O</mml:mi>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">(</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo fence="true" stretchy="false">[</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>−</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo fence="true" stretchy="false">]</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd/>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mi mathvariant="italic">d</mml:mi>
<mml:mo>×</mml:mo>
<mml:mi mathvariant="italic">s</mml:mi>
<mml:mo movablelimits="false">exp</mml:mo>
<mml:mo mathvariant="normal" fence="true" stretchy="false">(</mml:mo>
<mml:mo>−</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="italic">c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">γ</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo mathvariant="normal" fence="true" stretchy="false">)</mml:mo>
<mml:mo maxsize="1.19em" minsize="1.19em" fence="true" mathvariant="normal">)</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo maxsize="1.61em" minsize="1.61em" fence="true">}</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">G</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mspace width="2.5pt"/>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
</mml:mtable></mml:math><tex-math><![CDATA[\[ \begin{aligned}{}\Pr \left({\mathcal{D}_{\mathbf{{z_{j}}}}^{\ast }}\subseteq {\widehat{\mathcal{D}}_{\mathbf{{z_{j}}}}}\right)& =\Pr \{{\mathcal{D}_{\mathbf{z}}^{\ast }}\subseteq {\widehat{\mathcal{D}}_{1}}\}\times \cdots \times \Pr \{{\mathcal{D}_{\mathbf{z}}^{\ast }}\subseteq {\widehat{\mathcal{D}}_{G}}\}\\ {} & \ge \Big\{1-\mathcal{O}\big(d\exp [-{c_{1}}{s^{1-2(u+v)}}]\\ {} & +d\times s\exp (-{c_{2}}{s^{\gamma }})\big){\Big\}^{G}}\hspace{2.5pt}\end{aligned}\]]]></tex-math></alternatives>
</disp-formula> 
Replacing <inline-formula id="j_nejsds071_ineq_440"><alternatives><mml:math>
<mml:mi mathvariant="italic">s</mml:mi>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="italic">n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>−</mml:mo>
<mml:mi mathvariant="italic">α</mml:mi>
</mml:mrow>
</mml:msup></mml:math><tex-math><![CDATA[$s={n^{-\alpha }}$]]></tex-math></alternatives></inline-formula> where <inline-formula id="j_nejsds071_ineq_441"><alternatives><mml:math>
<mml:mi mathvariant="italic">α</mml:mi>
<mml:mo mathvariant="normal">&gt;</mml:mo>
<mml:mn>0</mml:mn></mml:math><tex-math><![CDATA[$\alpha \gt 0$]]></tex-math></alternatives></inline-formula>, then we have Theorem <xref rid="j_nejsds071_stat_004">4.1</xref>.  □</p></statement></app></app-group>
<ack id="j_nejsds071_ack_001">
<title>Acknowledgements</title>
<p>This study is partially funded by the Safety through Disruption (Safe-D) National University Transportation Center, a grant from the U.S. Department of Transportation – Office of the Assistant Secretary for Research and Technology, University Transportation Centers Program.</p></ack>
<ref-list id="j_nejsds071_reflist_001">
<title>References</title>
<ref id="j_nejsds071_ref_001">
<label>[1]</label><mixed-citation publication-type="other"><string-name><surname>Alemi</surname>, <given-names>A. A.</given-names></string-name>, <string-name><surname>Fischer</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Dillon</surname>, <given-names>J. V.</given-names></string-name> and <string-name><surname>Murphy</surname>, <given-names>K.</given-names></string-name> (2016). Deep variational information bottleneck. <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1612.00410"><italic>1612.00410</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_002">
<label>[2]</label><mixed-citation publication-type="journal"><string-name><surname>Bengio</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Courville</surname>, <given-names>A.</given-names></string-name> and <string-name><surname>Vincent</surname>, <given-names>P.</given-names></string-name> (<year>2013</year>). <article-title>Representation learning: A review and new perspectives</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source> <volume>35</volume>(<issue>8</issue>) <fpage>1798</fpage>–<lpage>1828</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_003">
<label>[3]</label><mixed-citation publication-type="other"><string-name><surname>Byrd</surname>, <given-names>J.</given-names></string-name> and <string-name><surname>Lipton</surname>, <given-names>Z. C.</given-names></string-name> (2018). What is the effect of importance weighting in deep learning? <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1812.03372"><italic>1812.03372</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_004">
<label>[4]</label><mixed-citation publication-type="chapter"><string-name><surname>Cao</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Gaidon</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Arechiga</surname>, <given-names>N.</given-names></string-name> and <string-name><surname>Ma</surname>, <given-names>T.</given-names></string-name> (<year>2019</year>). <chapter-title>Learning imbalanced datasets with label-distribution-aware margin loss</chapter-title>. In <source>Advances in Neural Information Processing Systems</source> <fpage>1565</fpage>–<lpage>1576</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_005">
<label>[5]</label><mixed-citation publication-type="chapter"><string-name><surname>Carlini</surname>, <given-names>N.</given-names></string-name> and <string-name><surname>Wagner</surname>, <given-names>D.</given-names></string-name> (<year>2017</year>). <chapter-title>Towards evaluating the robustness of neural networks</chapter-title>. In <source>2017 IEEE Symposium on Security and Privacy</source> <fpage>39</fpage>–<lpage>57</lpage>. <publisher-name>IEEE</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_006">
<label>[6]</label><mixed-citation publication-type="journal"><string-name><surname>Casale</surname>, <given-names>F. P.</given-names></string-name>, <string-name><surname>Dalca</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Saglietti</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Listgarten</surname>, <given-names>J.</given-names></string-name> and <string-name><surname>Fusi</surname>, <given-names>N.</given-names></string-name> (<year>2018</year>). <article-title>Gaussian process prior variational autoencoders</article-title>. <source>Advances in neural information processing systems</source> <volume>31</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_007">
<label>[7]</label><mixed-citation publication-type="journal"><string-name><surname>Chen</surname>, <given-names>R. T.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Grosse</surname>, <given-names>R. B.</given-names></string-name> and <string-name><surname>Duvenaud</surname>, <given-names>D. K.</given-names></string-name> (<year>2018</year>). <article-title>Isolating sources of disentanglement in variational autoencoders</article-title>. <source>Advances in neural information processing systems</source> <volume>31</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_008">
<label>[8]</label><mixed-citation publication-type="chapter"><string-name><surname>Chen</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>X. S.</given-names></string-name> and <string-name><surname>Huang</surname>, <given-names>T. S.</given-names></string-name> (<year>2001</year>). <chapter-title>One-class SVM for learning in image retrieval</chapter-title>. In <source>Proceedings 2001 International Conference on Image Processing (Cat. No. 01CH37205)</source> <volume>1</volume> <fpage>34</fpage>–<lpage>37</lpage>. <publisher-name>IEEE</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_009">
<label>[9]</label><mixed-citation publication-type="book"><string-name><surname>Coles</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Bawa</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Trenner</surname>, <given-names>L.</given-names></string-name> and <string-name><surname>Dorazio</surname>, <given-names>P.</given-names></string-name> (<year>2001</year>). <source>An introduction to statistical modeling of extreme values</source> <volume>208</volume>. <publisher-name>Springer</publisher-name>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1007/978-1-4471-3675-0" xlink:type="simple">https://doi.org/10.1007/978-1-4471-3675-0</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=1932132">MR1932132</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_010">
<label>[10]</label><mixed-citation publication-type="journal"><string-name><surname>Dingus</surname>, <given-names>T. A.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Antin</surname>, <given-names>J. F.</given-names></string-name>, <string-name><surname>Perez</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Buchanan-King</surname>, <given-names>M.</given-names></string-name> and <string-name><surname>Hankey</surname>, <given-names>J.</given-names></string-name> (<year>2016</year>). <article-title>Driver crash risk factors and prevalence evaluation using naturalistic driving data</article-title>. <source>Proceedings of the National Academy of Sciences</source> <volume>113</volume>(<issue>10</issue>) <fpage>2636</fpage>–<lpage>2641</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_011">
<label>[11]</label><mixed-citation publication-type="journal"><string-name><surname>Dong</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Gong</surname>, <given-names>S.</given-names></string-name> and <string-name><surname>Zhu</surname>, <given-names>X.</given-names></string-name> (<year>2018</year>). <article-title>Imbalanced deep learning by minority class incremental rectification</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source> <volume>41</volume>(<issue>6</issue>) <fpage>1367</fpage>–<lpage>1381</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_012">
<label>[12]</label><mixed-citation publication-type="journal"><string-name><surname>Fan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Song</surname>, <given-names>R.</given-names></string-name> <etal>et al.</etal> (<year>2010</year>). <article-title>Sure independence screening in generalized linear models with NP-dimensionality</article-title>. <source>The Annals of Statistics</source> <volume>38</volume>(<issue>6</issue>) <fpage>3567</fpage>–<lpage>3604</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1214/10-AOS798" xlink:type="simple">https://doi.org/10.1214/10-AOS798</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=2766861">MR2766861</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_013">
<label>[13]</label><mixed-citation publication-type="chapter"><string-name><surname>Finn</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Abbeel</surname>, <given-names>P.</given-names></string-name> and <string-name><surname>Levine</surname>, <given-names>S.</given-names></string-name> (<year>2017</year>). <chapter-title>Model-agnostic meta-learning for fast adaptation of deep networks</chapter-title>. In <source>International Conference on Machine Learning</source> <fpage>1126</fpage>–<lpage>1135</lpage>. <publisher-name>PMLR</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_014">
<label>[14]</label><mixed-citation publication-type="journal"><string-name><surname>Firth</surname>, <given-names>D.</given-names></string-name> (<year>1993</year>). <article-title>Bias reduction of maximum likelihood estimates</article-title>. <source>Biometrika</source> <fpage>27</fpage>–<lpage>38</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1093/biomet/80.1.27" xlink:type="simple">https://doi.org/10.1093/biomet/80.1.27</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=1225212">MR1225212</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_015">
<label>[15]</label><mixed-citation publication-type="journal"><string-name><surname>Fortuin</surname>, <given-names>V.</given-names></string-name> (<year>2022</year>). <article-title>Priors in bayesian deep learning: A review</article-title>. <source>International Statistical Review</source> <volume>90</volume>(<issue>3</issue>) <fpage>563</fpage>–<lpage>591</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1111/insr.12502" xlink:type="simple">https://doi.org/10.1111/insr.12502</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=4524825">MR4524825</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_016">
<label>[16]</label><mixed-citation publication-type="journal"><string-name><surname>Guo</surname>, <given-names>F.</given-names></string-name> (<year>2019</year>). <article-title>Statistical methods for naturalistic driving studies</article-title>. <source>Annual Review of Statistics and its Application</source> <volume>6</volume> <fpage>309</fpage>–<lpage>328</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1146/annurev-statistics-030718-105153" xlink:type="simple">https://doi.org/10.1146/annurev-statistics-030718-105153</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=3939523">MR3939523</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_017">
<label>[17]</label><mixed-citation publication-type="journal"><string-name><surname>Haixiang</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Yijing</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Shang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Mingyun</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Yuanyue</surname>, <given-names>H.</given-names></string-name> and <string-name><surname>Bing</surname>, <given-names>G.</given-names></string-name> (<year>2017</year>). <article-title>Learning from class-imbalanced data: Review of methods and applications</article-title>. <source>Expert Systems with Applications</source> <volume>73</volume> <fpage>220</fpage>–<lpage>239</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_018">
<label>[18]</label><mixed-citation publication-type="journal"><string-name><surname>Han</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Chatterjee</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Samworth</surname>, <given-names>R. J.</given-names></string-name> <etal>et al.</etal> (<year>2019</year>). <article-title>Isotonic regression in general dimensions</article-title>. <source>The Annals of Statistics</source> <volume>47</volume>(<issue>5</issue>) <fpage>2440</fpage>–<lpage>2471</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1214/18-AOS1753" xlink:type="simple">https://doi.org/10.1214/18-AOS1753</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=3988762">MR3988762</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_019">
<label>[19]</label><mixed-citation publication-type="book"><string-name><surname>Hastie</surname>, <given-names>T. J.</given-names></string-name> and <string-name><surname>Tibshirani</surname>, <given-names>R. J.</given-names></string-name> (<year>1990</year>). <source>Generalized additive models</source> <volume>43</volume>. <publisher-name>CRC press</publisher-name>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=1082147">MR1082147</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_020">
<label>[20]</label><mixed-citation publication-type="journal"><string-name><surname>He</surname>, <given-names>H.</given-names></string-name> and <string-name><surname>Garcia</surname>, <given-names>E. A.</given-names></string-name> (<year>2009</year>). <article-title>Learning from imbalanced data</article-title>. <source>IEEE Transactions on Knowledge and Data Engineering</source> <volume>21</volume>(<issue>9</issue>) <fpage>1263</fpage>–<lpage>1284</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_021">
<label>[21]</label><mixed-citation publication-type="journal"><string-name><surname>Higgins</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Matthey</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Pal</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Burgess</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Glorot</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Botvinick</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Mohamed</surname>, <given-names>S.</given-names></string-name> and <string-name><surname>Lerchner</surname>, <given-names>A.</given-names></string-name> (<year>2017</year>). <article-title>Beta-VAE: learning basic visual concepts with a constrained variational framework</article-title>. <source>The International Conference on Learning Representations</source> <volume>2</volume>(<issue>5</issue>) <fpage>6</fpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_022">
<label>[22]</label><mixed-citation publication-type="chapter"><string-name><surname>Jin</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Jin</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>J. T.</given-names></string-name> and <string-name><surname>Szolovits</surname>, <given-names>P.</given-names></string-name> (<year>2020</year>). <chapter-title>Is Bert really robust? A strong baseline for natural language attack on text classification and entailment</chapter-title>. In <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> <volume>34</volume> <fpage>8018</fpage>–<lpage>8025</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_023">
<label>[23]</label><mixed-citation publication-type="journal"><string-name><surname>King</surname>, <given-names>G.</given-names></string-name> and <string-name><surname>Zeng</surname>, <given-names>L.</given-names></string-name> (<year>2001</year>). <article-title>Logistic regression in rare events data</article-title>. <source>Political Analysis</source> <volume>9</volume>(<issue>2</issue>) <fpage>137</fpage>–<lpage>163</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_024">
<label>[24]</label><mixed-citation publication-type="other"><string-name><surname>Kingma</surname>, <given-names>D. P.</given-names></string-name> and <string-name><surname>Welling</surname>, <given-names>M.</given-names></string-name> (2013). Auto-encoding variational bayes. <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1312.6114"><italic>1312.6114</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_025">
<label>[25]</label><mixed-citation publication-type="journal"><string-name><surname>Kingma</surname>, <given-names>D. P.</given-names></string-name>, <string-name><surname>Welling</surname>, <given-names>M.</given-names></string-name> <etal>et al.</etal> (<year>2019</year>). <article-title>An introduction to variational autoencoders</article-title>. <source>Foundations and Trendső in Machine Learning</source> <volume>12</volume>(<issue>4</issue>) <fpage>307</fpage>–<lpage>392</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_026">
<label>[26]</label><mixed-citation publication-type="journal"><string-name><surname>Kingma</surname>, <given-names>D. P.</given-names></string-name>, <string-name><surname>Salimans</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Jozefowicz</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Sutskever</surname>, <given-names>I.</given-names></string-name> and <string-name><surname>Welling</surname>, <given-names>M.</given-names></string-name> (<year>2016</year>). <article-title>Improved variational inference with inverse autoregressive flow</article-title>. <source>Advances in neural information processing systems</source> <volume>29</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_027">
<label>[27]</label><mixed-citation publication-type="journal"><string-name><surname>Li</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Zhong</surname>, <given-names>W.</given-names></string-name> and <string-name><surname>Zhu</surname>, <given-names>L.</given-names></string-name> (<year>2012</year>). <article-title>Feature screening via distance correlation learning</article-title>. <source>Journal of the American Statistical Association</source> <volume>107</volume>(<issue>499</issue>) <fpage>1129</fpage>–<lpage>1139</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1080/01621459.2012.695654" xlink:type="simple">https://doi.org/10.1080/01621459.2012.695654</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=3010900">MR3010900</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_028">
<label>[28]</label><mixed-citation publication-type="journal"><string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Xia</surname>, <given-names>Z.</given-names></string-name> and <string-name><surname>Xu</surname>, <given-names>C.</given-names></string-name> (<year>2020</year>). <article-title>Distributed feature screening via componentwise debiasing</article-title>. <source>Journal of Machine Learning Research</source> <volume>21</volume>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=4071207">MR4071207</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_029">
<label>[29]</label><mixed-citation publication-type="chapter"><string-name><surname>Lin</surname>, <given-names>T. -Y.</given-names></string-name>, <string-name><surname>Goyal</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>K.</given-names></string-name> and <string-name><surname>Dollár</surname>, <given-names>P.</given-names></string-name> (<year>2017</year>). <chapter-title>Focal loss for dense object detection</chapter-title>. In <source>Proceedings of the IEEE International Conference on Computer Vision</source> <fpage>2980</fpage>–<lpage>2988</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_030">
<label>[30]</label><mixed-citation publication-type="journal"><string-name><surname>Liu</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Racz</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Vaillancourt</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Michelman</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Barnes</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Mellem</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Eastham</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Green</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Armstrong</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Bal</surname>, <given-names>R.</given-names></string-name> <etal>et al.</etal> (<year>2023</year>). <article-title>Smartphone-based hard-braking event detection at scale for road safety services</article-title>. <source>Transportation research part C: emerging technologies</source> <volume>146</volume> <fpage>103949</fpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_031">
<label>[31]</label><mixed-citation publication-type="journal"><string-name><surname>Lusa</surname>, <given-names>L.</given-names></string-name> <etal>et al.</etal> (<year>2017</year>). <article-title>Gradient boosting for high-dimensional prediction of rare events</article-title>. <source>Computational Statistics &amp; Data Analysis</source> <volume>113</volume> <fpage>19</fpage>–<lpage>37</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1016/j.csda.2016.07.016" xlink:type="simple">https://doi.org/10.1016/j.csda.2016.07.016</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=3662388">MR3662388</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_032">
<label>[32]</label><mixed-citation publication-type="journal"><string-name><surname>Massoli</surname>, <given-names>F. V.</given-names></string-name>, <string-name><surname>Falchi</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Kantarci</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Akti</surname>, <given-names>.</given-names></string-name>, <string-name><surname>Ekenel</surname>, <given-names>H. K.</given-names></string-name> and <string-name><surname>Amato</surname>, <given-names>G.</given-names></string-name> (<year>2021</year>). <article-title>MOCCA: Multilayer one-class classification for anomaly detection</article-title>. <source>IEEE Transactions on Neural Networks and Learning Systems</source>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1109/tnnls.2021.3130074" xlink:type="simple">https://doi.org/10.1109/tnnls.2021.3130074</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=4442308">MR4442308</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_033">
<label>[33]</label><mixed-citation publication-type="other"><string-name><surname>Mikolov</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Grave</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Bojanowski</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Puhrsch</surname>, <given-names>C.</given-names></string-name> and <string-name><surname>Joulin</surname>, <given-names>A.</given-names></string-name> (2017). Advances in pre-training distributed word representations. <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1712.09405"><italic>1712.09405</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_034">
<label>[34]</label><mixed-citation publication-type="journal"><string-name><surname>O’Kelly</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Sinha</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Namkoong</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Tedrake</surname>, <given-names>R.</given-names></string-name> and <string-name><surname>Duchi</surname>, <given-names>J. C.</given-names></string-name> (<year>2018</year>). <article-title>Scalable end-to-end autonomous vehicle testing via rare-event simulation</article-title>. <source>Advances in Neural Information Processing Systems</source> <volume>31</volume> <fpage>9827</fpage>–<lpage>9838</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_035">
<label>[35]</label><mixed-citation publication-type="other"><string-name><surname>Ren</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Zeng</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>B.</given-names></string-name> and <string-name><surname>Urtasun</surname>, <given-names>R.</given-names></string-name> (2018). Learning to reweight examples for robust deep learning. <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1803.09050"><italic>1803.09050</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_036">
<label>[36]</label><mixed-citation publication-type="chapter"><string-name><surname>Rezende</surname>, <given-names>D.</given-names></string-name> and <string-name><surname>Mohamed</surname>, <given-names>S.</given-names></string-name> (<year>2015</year>). <chapter-title>Variational inference with normalizing flows</chapter-title>. In <source>International conference on machine learning</source> <fpage>1530</fpage>–<lpage>1538</lpage>. <publisher-name>PMLR</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_037">
<label>[37]</label><mixed-citation publication-type="chapter"><string-name><surname>Ruff</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Vandermeulen</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Goernitz</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Deecke</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Siddiqui</surname>, <given-names>S. A.</given-names></string-name>, <string-name><surname>Binder</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Müller</surname>, <given-names>E.</given-names></string-name> and <string-name><surname>Kloft</surname>, <given-names>M.</given-names></string-name> (<year>2018</year>). <chapter-title>Deep one-class classification</chapter-title>. In <source>International Conference on Machine Learning</source> <fpage>4393</fpage>–<lpage>4402</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_038">
<label>[38]</label><mixed-citation publication-type="other"><string-name><surname>Shen</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Min</surname>, <given-names>M. R.</given-names></string-name>, <string-name><surname>Su</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Henao</surname>, <given-names>R.</given-names></string-name> and <string-name><surname>Carin</surname>, <given-names>L.</given-names></string-name> (2018). Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms. <italic>arXiv:</italic><ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1805.09843"><italic>1805.09843</italic></ext-link>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_039">
<label>[39]</label><mixed-citation publication-type="journal"><string-name><surname>Shi</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Qian</surname>, <given-names>C.</given-names></string-name> and <string-name><surname>Guo</surname>, <given-names>F.</given-names></string-name> (<year>2022</year>). <article-title>Real-time driving risk assessment using deep learning with XGBoost</article-title>. <source>Accident Analysis &amp; Prevention</source> <volume>178</volume> <fpage>106836</fpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_040">
<label>[40]</label><mixed-citation publication-type="chapter"><string-name><surname>Snoek</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ovadia</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Fertig</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Lakshminarayanan</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Nowozin</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sculley</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Dillon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>J.</given-names></string-name> and <string-name><surname>Nado</surname>, <given-names>Z.</given-names></string-name> (<year>2019</year>). <chapter-title>Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift</chapter-title>. In <source>Advances in Neural Information Processing Systems</source> <fpage>13969</fpage>–<lpage>13980</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_041">
<label>[41]</label><mixed-citation publication-type="journal"><string-name><surname>Sohn</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>H.</given-names></string-name> and <string-name><surname>Yan</surname>, <given-names>X.</given-names></string-name> (<year>2015</year>). <article-title>Learning structured output representation using deep conditional generative models</article-title>. <source>Advances in Neural Information Processing Systems</source> <volume>28</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_042">
<label>[42]</label><mixed-citation publication-type="journal"><string-name><surname>Sønderby</surname>, <given-names>C. K.</given-names></string-name>, <string-name><surname>Raiko</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Maaløe</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Sønderby</surname>, <given-names>S. K.</given-names></string-name> and <string-name><surname>Winther</surname>, <given-names>O.</given-names></string-name> (<year>2016</year>). <article-title>Ladder variational autoencoders</article-title>. <source>Advances in neural information processing systems</source> <volume>29</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_043">
<label>[43]</label><mixed-citation publication-type="journal"><string-name><surname>Sur</surname>, <given-names>P.</given-names></string-name> and <string-name><surname>Candès</surname>, <given-names>E. J.</given-names></string-name> (<year>2019</year>). <article-title>A modern maximum-likelihood theory for high-dimensional logistic regression</article-title>. <source>Proceedings of the National Academy of Sciences</source> <volume>116</volume>(<issue>29</issue>) <fpage>14516</fpage>–<lpage>14525</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1073/pnas.1810420116" xlink:type="simple">https://doi.org/10.1073/pnas.1810420116</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=3984492">MR3984492</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_044">
<label>[44]</label><mixed-citation publication-type="journal"><string-name><surname>Székely</surname>, <given-names>G. J.</given-names></string-name>, <string-name><surname>Rizzo</surname>, <given-names>M. L.</given-names></string-name>, <string-name><surname>Bakirov</surname>, <given-names>N. K.</given-names></string-name> <etal>et al.</etal> (<year>2007</year>). <article-title>Measuring and testing dependence by correlation of distances</article-title>. <source>The Annals of Statistics</source> <volume>35</volume>(<issue>6</issue>) <fpage>2769</fpage>–<lpage>2794</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1214/009053607000000505" xlink:type="simple">https://doi.org/10.1214/009053607000000505</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=2382665">MR2382665</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_045">
<label>[45]</label><mixed-citation publication-type="chapter"><string-name><surname>Tomczak</surname>, <given-names>J.</given-names></string-name> and <string-name><surname>Welling</surname>, <given-names>M.</given-names></string-name> (<year>2018</year>). <chapter-title>VAE with a VampPrior</chapter-title>. In <source>International Conference on Artificial Intelligence and Statistics</source> <fpage>1214</fpage>–<lpage>1223</lpage>. <publisher-name>PMLR</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_046">
<label>[46]</label><mixed-citation publication-type="chapter"><string-name><surname>Wang</surname>, <given-names>H.</given-names></string-name> (<year>2020</year>). <chapter-title>Logistic regression for massive data with rare events</chapter-title>. In <source>International Conference on Machine Learning</source> <fpage>9829</fpage>–<lpage>9836</lpage>. <publisher-name>PMLR</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_047">
<label>[47]</label><mixed-citation publication-type="journal"><string-name><surname>Wehenkel</surname>, <given-names>A.</given-names></string-name> and <string-name><surname>Louppe</surname>, <given-names>G.</given-names></string-name> (<year>2019</year>). <article-title>Unconstrained monotonic neural networks</article-title>. <source>Advances in Neural Information Processing Systems</source> <volume>32</volume>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_048">
<label>[48]</label><mixed-citation publication-type="chapter"><string-name><surname>Wu</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name> and <string-name><surname>Lin</surname>, <given-names>D.</given-names></string-name> (<year>2021</year>). <chapter-title>Adversarial robustness under long-tailed distribution</chapter-title>. In <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> <fpage>8659</fpage>–<lpage>8668</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_049">
<label>[49]</label><mixed-citation publication-type="journal"><string-name><surname>Zhang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>C.</given-names></string-name> <etal>et al.</etal> (<year>2020</year>). <article-title>Convergence rates of variational posterior distributions</article-title>. <source>Annals of Statistics</source> <volume>48</volume>(<issue>4</issue>) <fpage>2180</fpage>–<lpage>2207</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1214/19-AOS1883" xlink:type="simple">https://doi.org/10.1214/19-AOS1883</ext-link>. <ext-link ext-link-type="uri" xlink:href="https://mathscinet.ams.org/mathscinet-getitem?mr=4134791">MR4134791</ext-link></mixed-citation>
</ref>
<ref id="j_nejsds071_ref_050">
<label>[50]</label><mixed-citation publication-type="chapter"><string-name><surname>Zhang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Xin</surname>, <given-names>J.</given-names></string-name> and <string-name><surname>Chen</surname>, <given-names>B.</given-names></string-name> (<year>2022</year>). <chapter-title>Multi-view information bottleneck without variational approximation</chapter-title>. In <source>ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing</source> <fpage>4318</fpage>–<lpage>4322</lpage>. <publisher-name>IEEE</publisher-name>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_051">
<label>[51]</label><mixed-citation publication-type="journal"><string-name><surname>Zhang</surname>, <given-names>W. E.</given-names></string-name>, <string-name><surname>Sheng</surname>, <given-names>Q. Z.</given-names></string-name>, <string-name><surname>Alhazmi</surname>, <given-names>A.</given-names></string-name> and <string-name><surname>Li</surname>, <given-names>C.</given-names></string-name> (<year>2020</year>). <article-title>Adversarial attacks on deep-learning models in natural language processing: A survey</article-title>. <source>ACM Transactions on Intelligent Systems and Technology</source> <volume>11</volume>(<issue>3</issue>) <fpage>1</fpage>–<lpage>41</lpage>.</mixed-citation>
</ref>
<ref id="j_nejsds071_ref_052">
<label>[52]</label><mixed-citation publication-type="journal"><string-name><surname>Zheng</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Sayed</surname>, <given-names>T.</given-names></string-name> and <string-name><surname>Mannering</surname>, <given-names>F.</given-names></string-name> (<year>2021</year>). <article-title>Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions</article-title>. <source>Analytic Methods in Accident Research</source> <volume>29</volume> <fpage>100142</fpage>.</mixed-citation>
</ref>
</ref-list>
</back>
</article>
