The timing of longitudinal measurements may depend upon outcome or disease severity. In biomedical studies relying on clinical encounter data, patients often have dense, irregular collections of visit data when suffering a worse health condition. In parallel, the longitudinal measurements may be impacted by the period of irregular visiting. Ignoring the impact of the outcome-dependent visiting process when constructing a longitudinal disease progression model can produce biased results. We propose a Bayesian joint model linking a mixed-effects model for the longitudinal marker and Weibull proportional hazards model with a log frailty for the visiting process, adjusting both longitudinal marker and event processes with covariates. We examine different random effect structures and performance characterizing disease trajectory. Motivated by clinical data on cystic fibrosis lung disease, we estimate the longitudinal process for lung function decline. Individuals with lower lung function tend to have more frequent clinical visits than those with higher lung function. Simulation studies suggest that incorporating a time-dependent Gaussian process is more important for model fit than adding the survival model via joint modeling; the random intercepts model exhibits maximum bias, especially when there is an outcome-dependent visiting process.
Severe event class imbalance is observed in many fields, such as insurance fraud, severe weather, traffic safety, and rare disease, and has far-reaching significance when the generalization of minority event classes is of primary interest. While existing solutions mostly focus on differential sampling or sample re-weighting approaches to alleviate the imbalance issue, we take a novel alternative view on promoting generalization. Our proposal is formulated under the generative Bayesian framework, positing that predictors are the stochastic proxies of latent causes with limited samples, whose exceedance leads to extreme events. To accurately capture the extended tail, our solution adopts the Gumbel distribution as prior and is modeled with a variational inference framework. Following the assertion that exceedance leads to extremes, we devise a disentangled additive monotonic neural architecture to predict the risk. The proposed model acknowledges representation uncertainties while embracing improved interpretability, generalization, and robustness. We provide theoretical insights to show the merits of the proposed approach. To verify the effectiveness in empirical settings, we conducted studies on various real-world data against the state-of-the-art counterparts, with encouraging results reported.
In extreme value analysis, the impact of rounding in data, a form of quantization, on statistical inferences beyond point estimation has not been comprehensively studied. This paper addresses these challenges by considering rounded data as interval-censored. The maximum likelihood estimators of the model parameters tailored to account for interval censoring are asymptotically unbiased and efficient. Further, we adapt classic goodness-of-fit tests, such as the Anderson-Darling test, for rounded data based on the maximum likelihood estimator. The resulting tests have appropriate sizes and considerable power. One application of such tests is threshold selection for the peak over threshold approach in extreme value analysis. The efficacy of our estimation approach and the goodness-of-fit tests are demonstrated through a simulation study involving data rounded from generalized Pareto distributions. Applying this method to precipitation data from 18 stations in eastern Washington, an area with typically low precipitation and expecting a significant rounding effect, we observe narrower interval estimates of return levels.
Observations of groundwater pollutants, such as arsenic or Perfluorooctane sulfonate (PFOS), are riddled with left censoring. These measurements have an impact on the health and lifestyle of the populace. Left censoring of these spatially correlated observations is usually addressed by applying Gaussian processes (GPs), which have theoretical advantages. However, this comes with a challenging computational complexity of $\mathcal{O}({n^{3}})$, impractical for large datasets. Additionally, a sizable proportion of the left-censored data creates further bottlenecks since the likelihood computation now involves an intractable high-dimensional integral of the multivariate Gaussian density. In this article, we tackle these two problems simultaneously by approximating the GP with a Gaussian Markov random field (GMRF) approach that exploits an explicit link between a GP with Matérn correlation function and a GMRF using stochastic partial differential equations (SPDEs). We introduce a GMRF-based measurement error into the model, which alleviates the likelihood computation for the censored data, drastically improving the computational speed while maintaining admirable accuracy. Our approach demonstrates robustness and substantial computational scalability compared to state-of-the-art methods for censored spatial responses across various simulation settings. Finally, the fit of this fully Bayesian model to the concentration of PFOS in groundwater available at 24,959 sites across California, where 46.62% responses are censored, produces prediction surface and uncertainty quantification in real-time, thereby substantiating the applicability and scalability of the proposed method. Code for implementation is made available via GitHub.
We consider the class of inverse probability weight (IPW) estimators, including the popular Horvitz–Thompson and Hájek estimators used routinely in survey sampling, causal inference and for Bayesian computation. We focus on the ‘weak paradoxes’ for these estimators due to two counterexamples by Basu (1988) and Wasserman (2004) and investigate the two natural Bayesian answers to this problem: one based on binning and smoothing: a ‘Bayesian sieve’ and the other based on a conjugate hierarchical model that allows borrowing information via exchangeability. We compare the mean squared errors for the two Bayesian estimators with the IPW estimators for Wasserman’s example via simulation studies on a broad range of parameter configurations. We also prove posterior consistency for the Bayes estimators under missing-completely-at-random assumption and show that it requires fewer assumptions on the inclusion probabilities. We also revisit the connection between the different problems where improved or adaptive IPW estimators will be useful, including survey sampling, evidence estimation strategies such as Conditional Monte Carlo, Riemannian sum, Trapezoidal rules and vertical likelihood, as well as average treatment effect estimation in causal inference.
Social networks primarily focus on the phenomenon of the contagion effect when examining behavior patterns within specific social groups. However, the impact of peer effects is characterized by the tendency to imitate the behaviors of friends and the selection process, where individuals tend to affiliate with others sharing similar traits, significantly contributing to shaping social behaviors that are frequently interconnecting. This article presents a Bayesian approach that uses latent-space estimation methods to detect and examine contagion effects, considering the impact of social selection. The research provides a methodological explanation, followed by a sequence of simulation trials designed to explore operational functionalities and possible real-world applications. To illustrate the potential correlation between changes in alcohol use and the influence of social networks, this study concludes by presenting an example of adolescent drinking behavior.
This study explores the application of the t-link model in high-dimensional variable selection for binary regression. The t-link model provides flexibility in binary modeling and offers robust inference in the presence of outliers, making it a preferable alternative to the commonly used probit and logit links. To address the computational challenges posed by a large number of covariates, the skinny Gibbs algorithm is employed, and the consistency of variable selection under this approximate algorithm is established. These advancements in both computational and theoretical perspectives enhance the practicality and ease of implementing the t-link model. The performance of the t-link model, with a specified degrees of freedom, is compared to logit link and the probit link through simulation studies and an application to PCR data. The results demonstrate the robustness and computational efficiency of the proposed method.
We study local change point detection in variance using generalized likelihood ratio tests. Building on [24], we utilize the multiplier bootstrap to approximate the unknown, non-asymptotic distribution of the test statistic and introduce a multiplicative bias correction that improves upon the existing additive version. This proposed correction offers a clearer interpretation of the bootstrap estimators while significantly reducing computational costs. Simulation results demonstrate that our method performs comparably to the original approach. We apply it to the growth rates of U.S. inflation, industrial production, and Bitcoin returns.