Modern Statistics is made of the sensible combination of direct evidence (the data directly relevant or the “individual data”) and indirect evidence (the data and knowledge indirectly relevant or the “group data”). The admissible procedures combine the two sources of information, and technology advances make indirect evidence more substantial and ubiquitous. It has been pointed out, however, that an important problem of Statistics when “borrowing strength” is to treat in a fundamentally different way exceptional cases that do not adapt to the central “aurea mediocritas”. This is what has been coined as “the Clemente problem” in honor of R. Clemente, an exceptional batter [
6]. In this article, we argue that the problem is caused by the simultaneous use of square loss function and conjugate (light-tailed) priors, which is the usual procedure. We propose in their place to use robust penalties, in the form of losses that penalize more severely huge errors, or (equivalently) priors of heavy tails which make the exceptional more probable. Using heavy-tailed priors, we can reproduce, in a Bayesian structured way, Efron and Morris’ “limited translated estimators” (with Double Exponential Priors) and “discarding priors estimators” (with Cauchy-like priors), which discard the prior in the presence of prior-likelihood conflict. We show that both Empirical Bayes and Full Bayes approaches can alleviate the Clemente problem and beat the James-Stein estimator in terms of smaller square errors, for sensible Robust Bayes priors. We model in parallel Empirical Bayes and Fully Bayesian hierarchical models, illustrating that the differences among sensible versions of both are relatively small, as compared with the effect due to the robust assumptions. We follow [
16] in using a heavy-tailed (scaled) Beta2 distribution for (squared) scales that arise naturally as an alternative to the usual Inverted-Gamma distribution. The combination of a Cauchy Prior for location and Scaled Beta2 for square scales yields a novel closed-form prior for location, extremely suitable for Objective Robust Bayesian Analysis. Finally, we calculate the predictive intervals and the Robust models covers the Clemente average at
$80\% $ of probability, which the conjugate do not. Our approach has connections with [
25], which employs a completely different methodology. This is an instance of the convergence of the best frequentist and Bayesian analyses [see
3].