
| Pengarang | : | Bandoro Bantarto |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 3) |
| Halaman | : | 191-201 |
| Abstrak | : | The classical birthday problem considers the probability that at least two people in a group of size N share the same birthday. The inverse birthday problem considers the estimation of the size N of a group given the number of different birthdays in the group. In practice, this problem is analogous to estimating the size of a population from occurrence data only. The inverse problem can be solved via two simple approaches including the method of moments for a multinominal model and the maximum likelihood estimate of a Poisson model, which we present in this study. We investigate properties of both methods and show that they can yield asymptotically equivalent Wald-type interval estimators. Moreover, we show that these methods estimate a lower bound for the population size when birth rates are nonhomogenous or individuals in the population are aggregated. A simulation study was conducted to evaluate the performance of the point estimates arising from the two approaches and to compare the performance of seven interval estimators, including likelihood ratio and log-transformation methods. We illustrate the utility of these methods by estimating: (1) the abundance of tree species over a 50-hectare forest plot, (2) the number of Chlamydia infections when only the number of different birthdays of the patients is known, and (3) the number of rainy days when the number of rainy weeks is known. Supplementary materials for this article are available online. |
| Pengarang | : | Joseph B. Lang |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 354-368 |
| Abstrak | : | This article introduces mean-minimum (MM) exact confidence intervals for a binomial probability. These intervals guarantee that both the mean and the minimum frequentist coverage never drop below specified values. For example, an MM 95[93]% interval has mean coverage at least 95% and minimum coverage at least 93%. In the conventional sense, such an interval can be viewed as an exact 93% interval that has mean coverage at least 95% or it can be viewed as an approximate 95% interval that has minimum coverage at least 93%. Graphical and numerical summaries of coverage and expected length suggest that the Blaker-based MM exact interval is an attractive alternative to, even an improvement over, commonly recommended approximate and exact intervals, including the Agresti–Coull approximate interval, the Clopper–Pearson (CP) exact interval, and the more recently recommended CP-, Blaker-, and Sterne-based mean-coverage-adjusted approximate intervals.Correlated data are commonly analyzed using models constructed using population-averaged generalized estimating equations (GEEs). The specification of a population-averaged GEE model includes selection of a structure describing the correlation of repeated measures. Accurate specification of this structure can improve efficiency, whereas the finite-sample estimation of nuisance correlation parameters can inflate the variances of regression parameter estimates. Therefore, correlation structure selection criteria should penalize, or account for, correlation parameter estimation. In this article, we compare recently proposed penalties in terms of their impacts on correlation structure selection and regression parameter estimation, and give practical considerations for data analysts. Supplementary materials for this article are available online. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 344-353 |
| Abstrak | : | Correlated data are commonly analyzed using models constructed using population-averaged generalized estimating equations (GEEs). The specification of a population-averaged GEE model includes selection of a structure describing the correlation of repeated measures. Accurate specification of this structure can improve efficiency, whereas the finite-sample estimation of nuisance correlation parameters can inflate the variances of regression parameter estimates. Therefore, correlation structure selection criteria should penalize, or account for, correlation parameter estimation. In this article, we compare recently proposed penalties in terms of their impacts on correlation structure selection and regression parameter estimation, and give practical considerations for data analysts. Supplementary materials for this article are available online. |
| Pengarang | : | Luís Gustavo Esteves |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 336-343 |
| Abstrak | : | Teaching how to derive minimax decision rules can be challenging because of the lack of examples that are simple enough to be used in the classroom. Motivated by this challenge, we provide a new example that illustrates the use of standard techniques in the derivation of optimal decision rules under the Bayes and minimax approaches. We discuss how to predict the value of an unknown quantity, θ ∈ {0, 1}, given the opinions of n experts. An important example of such crowdsourcing problem occurs in modern cosmology, where θ indicates whether a given galaxy is merging or not, and Y1, …, Yn are the opinions from n astronomers regarding θ. We use the obtained prediction rules to discuss advantages and disadvantages of the Bayes and minimax approaches to decision theory. The material presented here is intended to be taught to first-year graduate students. |
| Pengarang | : | |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 326-335 |
| Abstrak | : | The six recommendations made by the Guidelines for Assessment and Instruction in Statistics Education (GAISE) committee were first communicated in 2005 and more formally in 2010. In this article, 25 introductory statistics textbooks are examined to assess how well these textbooks have incorporated the three GAISE recommendations most relevant to implementation in textbooks (statistical literacy and thinking; use of real data; stress concepts over procedures). The implementation of another recommendation (using technology) is described but not assessed. In general, most textbooks appear to be adopting the GAISE recommendations reasonably well in both exposition and exercises. The textbooks are particularly adept at using real data, using real data well, and promoting statistical literacy. Textbooks are less adept—but still rated reasonably well, in general—at explaining concepts over procedures and promoting statistical thinking. In contrast, few textbooks have easy-usable glossaries of statistical terms to assist with understanding of statistical language and literacy development. Supplementary materials for this article are available online. |
| Pengarang | : | Louis Leahy |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 317-325 |
| Abstrak | : | This article describes a free, open-source collection of templates for the popular Excel (2013, and later versions) spreadsheet program. These templates are spreadsheet files that allow easy and intuitive learning and the implementation of practical examples concerning descriptive statistics, random variables, confidence intervals, and hypothesis testing. Although they are designed to be used with Excel, they can also be employed with other free spreadsheet programs (changing some particular formulas). Moreover, we exploit some possibilities of the ActiveX controls of the Excel Developer Menu to perform interactive Gaussian density charts. Finally, it is important to note that they can be often embedded in a web page, so it is not necessary to employ Excel software for their use. These templates have been designed as a useful tool to teach basic statistics and to carry out data analysis even when the students are not familiar with Excel. Additionally, they can be used as a complement to other analytical software packages. They aim to assist students in learning statistics, within an intuitive working environment. Supplementary materials with the Excel templates are available online. |
| Pengarang | : | Dannemiller, Steve |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 310-316 |
| Abstrak | : | The coefficient of determination, a.k.a. R2, is well-defined in linear regression models, and measures the proportion of variation in the dependent variable explained by the predictors included in the model. To extend it for generalized linear models, we use the variance function to define the total variation of the dependent variable, as well as the remaining variation of the dependent variable after modeling the predictive effects of the independent variables. Unlike other definitions that demand complete specification of the likelihood function, our definition of R2 only needs to know the mean and variance functions, so applicable to more general quasi-models. It is consistent with the classical measure of uncertainty using variance, and reduces to the classical definition of the coefficient of determination when linear regression models are considered. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 305-309 |
| Abstrak | : | The interval between two prespecified order statistics of a sample provides a distribution-free confidence interval for a population quantile. However, due to discreteness, only a small set of exact coverage probabilities is available. Interpolated confidence intervals are designed to expand the set of available coverage probabilities. However, we show here that the infimum of the coverage probability for an interpolated confidence interval is either the coverage probability for the inner interval or the coverage probability obtained by removing the more likely of the two extreme subintervals from the outer interval. Thus, without additional assumptions, interpolated intervals do not expand the set of available guaranteed coverage probabilities. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | The American Statistician |
| Volume / Edisi | : | 71 (No. 4) |
| Halaman | : | 302-304 |
| Abstrak | : | It has long been asserted that in univariate location-scale models, when concerned with inference for either the location or scale parameter, the use of the inverse of the scale parameter as a Bayesian prior yields posterior credible sets that have exactly the correct frequentist confidence set interpretation. This claim dates to at least Peers, and has subsequently been noted by various authors, with varying degrees of justification. We present a simple, direct demonstration of the exact matching property of the posterior credible sets derived under use of this prior in the univariate location-scale model. This is done by establishing an equivalence between the conditional frequentist and posterior densities of the pivotal quantities on which conditional frequentist inferences are based. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Analisis CSIS |
| Volume / Edisi | : | 37 (No. 3) |
| Halaman | : | 424-443 |
| Abstrak | : | Globalisasi sanggup mendesakkan energi yang kuat untuk menembus batas-batas teritorial geografis yang menjadi sekat antara satu komunitas dengan yang lain. Pada satu sisi, globalisasi menanamkan kesadaran atas keterkaitan secara erat antara satu negarabangsa dan lainnya sehingga sebuah dampak yang ditimbulkan oleh salah satu diantara mereka akan mempengaruhi yang lain secara luas. Pada sisi lain, globalisasi telah menyebabkan leburnya sekat-sekat identitas etnis, agama dan bahkan pada titik ekstrem kebangsaan. |