
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 625-635 |
| Abstrak | : | We use topological methods to investigate the small-scale variation and local spatial characteristics of the interstellar medium (ISM) in three regions of the southern sky. We demonstrate that there are circumstances where topological methods can identify differences in distributions when conventional marginal or correlation analyses may not. We propose a nonparametric method for comparing two fields based on the counts of topological features and the geometry of the associated persistence diagrams. We investigate the expected distribution of topological structures quantified through Betti numbers under Gaussian random field (GRF) assumptions, which underlie many astrophysical models of the ISM. When we apply the methods to the astrophysical data, we find strong evidence that one of the three regions is both topologically dissimilar to the other two and not consistent with an underlying GRF model. This region is proximal to a region of recent star formation whereas the others are more distant. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Dipoyudo Kirdi |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 610-624 |
| Abstrak | : | An important task in microbiome studies is to test the existence of and give characterization to differences in the microbiome composition across groups of samples. Important challenges of this problem include the large within-group heterogeneities among samples and the existence of potential confounding variables that, when ignored, increase the chance of false discoveries and reduce the power for identifying true differences. We propose a probabilistic framework to overcome these issues by combining three ideas: (i) a phylogenetic tree-based decomposition of the cross-group comparison problem into a series of local tests, (ii) a graphical model that links the local tests to allow information sharing across taxa, and (iii) a Bayesian testing strategy that incorporates covariates and integrates out the within-group variation, avoiding potentially unstable point estimates. With the proposed method, we analyze the American Gut data to compare the gut microbiome composition of groups of participants with different dietary habits. Our analysis shows that (i) the frequency of consuming fruit, seafood, vegetable, and whole grain are closely related to the gut microbiome composition and (ii) the conclusion of the analysis can change drastically when different sets of relevant covariates are adjusted, indicating the necessity of carefully selecting and including possible confounders in the analysis when comparing microbiome compositions with data from observational studies. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Neal S. Grantham |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 599-609 |
| Abstrak | : | Recent advances in bioinformatics have made high-throughput microbiome data widely available, and new statistical tools are required to maximize the information gained from these data. For example, analysis of high-dimensional microbiome data from designed experiments remains an open area in microbiome research. Contemporary analyses work on metrics that summarize collective properties of the microbiome, but such reductions preclude inference on the fine-scale effects of environmental stimuli on individual microbial taxa. Other approaches model the proportions or counts of individual taxa as response variables in mixed models, but these methods fail to account for complex correlation patterns among microbial communities. In this article, we propose a novel Bayesian mixed-effects model that exploits cross-taxa correlations within the microbiome, a model we call microbiome mixed model (MIMIX). MIMIX offers global tests for treatment effects, local tests and estimation of treatment effects on individual taxa, quantification of the relative contribution from heterogeneous sources to microbiome variability, and identification of latent ecological subcommunities in the microbiome. MIMIX is tailored to large microbiome experiments using a combination of Bayesian factor analysis to efficiently represent dependence between taxa and Bayesian variable selection methods to achieve sparsity. We demonstrate the model using a simulation experiment and on a 2 × 2 factorial experiment of the effects of nutrient supplement and herbivore exclusion on the foliar fungal microbiome of Andropogon gerardii, a perennial bunchgrass, as part of the global Nutrient Network research initiative. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Antony M. Overstall |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 583-598 |
| Abstrak | : | Bayesian optimal design is considered for experiments where the response distribution depends on the solution to a system of nonlinear ordinary differential equations. The motivation is an experiment to estimate parameters in the equations governing the transport of amino acids through cell membranes in human placentas. Decision-theoretic Bayesian design of experiments for such nonlinear models is conceptually very attractive, allowing the formal incorporation of prior knowledge to overcome the parameter dependence of frequentist design and being less reliant on asymptotic approximations. However, the necessary approximation and maximization of the, typically analytically intractable, expected utility results in a computationally challenging problem. These issues are further exacerbated if the solution to the differential equations is not available in closed-form. This article proposes a new combination of a probabilistic solution to the equations embedded within a Monte Carlo approximation to the expected utility with cyclic descent of a smooth approximation to find the optimal design. A novel precomputation algorithm reduces the computational burden, making the search for an optimal design feasible for bigger problems. The methods are demonstrated by finding new designs for a number of common models derived from differential equations, and by providing optimal designs for the placenta experiment. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 570-582 |
| Abstrak | : | Nowadays, events are spread rapidly along social networks. We are interested in whether people’s responses to an event are affected by their friends’ characteristics. For example, how soon will a person start playing a game given that his/her friends like it? Studying social network dependence is an emerging research area. In this work, we propose a novel latent spatial autocorrelation Cox model to study social network dependence with time-to-event data. The proposed model introduces a latent indicator to characterize whether a person’s survival time might be affected by his or her friends’ features. We first propose a score-type test for detecting the existence of social network dependence. If it exists, we further develop an EM-type algorithm to estimate the model parameters. The performance of the proposed test and estimators are illustrated by simulation studies and an application to a time-to-event dataset about playing a popular mobile game from one of the largest online social network platforms. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Jean-Noël Bacro |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 555-569 |
| Abstrak | : | The statistical modeling of space-time extremes in environmental applications is key to understanding complex dependence structures in original event data and to generating realistic scenarios for impact models. In this context of high-dimensional data, we propose a novel hierarchical model for high threshold exceedances defined over continuous space and time by embedding a space-time Gamma process convolution for the rate of an exponential variable, leading to asymptotic independence in space and time. Its physically motivated anisotropic dependence structure is based on geometric objects moving through space-time according to a velocity vector. We demonstrate that inference based on weighted pairwise likelihood is fast and accurate. The usefulness of our model is illustrated by an application to hourly precipitation data from a study region in Southern France, where it clearly improves on an alternative censored Gaussian space-time random field model. While classical limit models based on threshold-stability fail to appropriately capture relatively fast joint tail decay rates between asymptotic dependence and classical independence, strong empirical evidence from our application and other recent case studies motivates the use of more realistic asymptotic independence models such as ours. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Trambak Banerjee |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 538-554 |
| Abstrak | : | We develop a constrained extremely zero inflated joint (CEZIJ) modeling framework for simultaneously analyzing player activity, engagement, and dropouts (churns) in app-based mobile freemium games. Our proposed framework addresses the complex interdependencies between a player’s decision to use a freemium product, the extent of her direct and indirect engagement with the product and her decision to permanently drop its usage. CEZIJ extends the existing class of joint models for longitudinal and survival data in several ways. It not only accommodates extremely zero-inflated responses in a joint model setting but also incorporates domain-specific, convex structural constraints on the model parameters. Longitudinal data from app-based mobile games usually exhibit a large set of potential predictors and choosing the relevant set of predictors is highly desirable for various purposes including improved predictability. To achieve this goal, CEZIJ conducts simultaneous, coordinated selection of fixed and random effects in high-dimensional penalized generalized linear mixed models. For analyzing such large-scale datasets, variable selection and estimation are conducted via a distributed computing based split-and-conquer approach that massively increases scalability and provides better predictive performance over competing predictive methods. Our results reveal codependencies between varied player characteristics that promote player activity and engagement. Furthermore, the predicted churn probabilities exhibit idiosyncratic clusters of player profiles over time based on which marketers and game managers can segment the playing population for improved monetization of app-based freemium games. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 521-537 |
| Abstrak | : | The National Birth Defects Prevention Study (NBDPS) is a case-control study of birth defects conducted across 10 U.S. states. Researchers are interested in characterizing the etiologic role of maternal diet, collected using a food frequency questionnaire. Because diet is multidimensional, dimension reduction methods such as cluster analysis are often used to summarize dietary patterns. In a large, heterogeneous population, traditional clustering methods, such as latent class analysis, used to estimate dietary patterns can produce a large number of clusters due to a variety of factors, including study size and regional diversity. These factors result in a loss of interpretability of patterns that may differ due to minor consumption changes. Based on adaptation of the local partition process, we propose a new method, robust profile clustering, to handle these data complexities. Here, participants may be clustered at two levels: (1) globally, where women are assigned to an overall population-level cluster via an overfitted finite mixture model, and (2) locally, where regional variations in diet are accommodated via a beta-Bernoulli process dependent on subpopulation differences. We use our method to analyze the NBDPS data, deriving prepregnancy dietary patterns for women in the NBDPS while accounting for regional variability. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 501-520 |
| Abstrak | : | Cortical surface functional magnetic resonance imaging (cs-fMRI) has recently grown in popularity versus traditional volumetric fMRI. In addition to offering better whole-brain visualization, dimension reduction, removal of extraneous tissue types, and improved alignment of cortical areas across subjects, it is also more compatible with common assumptions of Bayesian spatial models. However, as no spatial Bayesian model has been proposed for cs-fMRI data, most analyses continue to employ the classical general linear model (GLM), a “massive univariate” approach. Here, we propose a spatial Bayesian GLM for cs-fMRI, which employs a class of sophisticated spatial processes to model latent activation fields. We make several advances compared with existing spatial Bayesian models for volumetric fMRI. First, we use integrated nested Laplacian approximations, a highly accurate and efficient Bayesian computation technique, rather than variational Bayes. To identify regions of activation, we utilize an excursions set method based on the joint posterior distribution of the latent fields, rather than the marginal distribution at each location. Finally, we propose the first multi-subject spatial Bayesian modeling approach, which addresses a major gap in the existing literature. The methods are very computationally advantageous and are validated through simulation studies and two task fMRI studies from the Human Connectome Project. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 530) |
| Halaman | : | 491-500 |
| Abstrak | : | What does statistics have to offer science and society, in this age of massive data, machine learning algorithms, and multiple online sources of tools for data analysis? I recall a few situations where statistics made a real difference and reinforced the impact of our discipline on society. Sometimes the difference lay in the insightful analysis and inference enabled by ground-breaking methods in our field like hypothesis testing, likelihood ratios, Bayesian models, jackknife, and bootstrap. But perhaps more often, the impacts came from thoughtful analyses before data were collected, and the questions that arose after the statistical analysis. The impact of understanding the problem, designing the experiment and data collections, conducting the pilot surveys, and raising important questions, is substantial. Through sensible explorations following formal statistical procedures, statisticians have made contributions in many domains. In this presentation, I recall some examples which made a long-lasting impact. Some of them, like randomization in clinical trials, known and familiar to all, are so ingrained in our practice that the role of statistics has been forgotten. Others may be less familiar but nonetheless benefited greatly from the critical input of statisticians. All remind us that our field remains today not only relevant but critical to science and society. |