OPAC - Pencarian Artikel Jurnal & Majalah Library USD

Menampilkan semua artikel (Halaman 497 dari 34170, Total: 341698 data)

Generic Inference on Quantile and Quantile Effect Functions for Discrete Outcomes

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 123-137
Abstrak : Quantile and quantile effect (QE) functions are important tools for descriptive and causal analysis due to their natural and intuitive interpretation. Existing inference methods for these functions do not apply to discrete random variables. This article offers a simple, practical construction of simultaneous confidence bands for quantile and QE functions of possibly discrete random variables. It is based on a natural transformation of simultaneous confidence bands for distribution functions, which are readily available for many problems. The construction is generic and does not depend on the nature of the underlying problem. It works in conjunction with parametric, semiparametric, and nonparametric modeling methods for observed and counterfactual distributions, and does not depend on the sampling scheme. We apply our method to characterize the distributional impact of insurance coverage on health care utilization and obtain the distributional decomposition of the racial test score gap. We find that universal insurance coverage increases the number of doctor visits across the entire distribution, and that the racial test score gap is small at early ages but grows with age due to socio-economic factors especially at the top of the distribution. Supplementary materials (additional results, R package, replication files) for this article are available online.

Penalized and Constrained Optimization: An Application to High-Dimensional Website Advertising

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 102-122
Abstrak : Firms are increasingly transitioning advertising budgets to Internet display campaigns, but this transition poses new challenges. These campaigns use numerous potential metrics for success (e.g., reach or click rate), and because each website represents a separate advertising opportunity, this is also an inherently high-dimensional problem. Further, advertisers often have constraints they wish to place on their campaign, such as targeting specific sub-populations or websites. These challenges require a method flexible enough to accommodate thousands of websites, as well as numerous metrics and campaign constraints. Motivated by this application, we consider the general constrained high-dimensional problem, where the parameters satisfy linear constraints. We develop the Penalized and Constrained optimization method (PaC) to compute the solution path for high-dimensional, linearly constrained criteria. PaC is extremely general; in addition to internet advertising, we show it encompasses many other potential applications, such as portfolio estimation, monotone curve estimation, and the generalized lasso. Computing the PaC coefficient path poses technical challenges, but we develop an efficient algorithm over a grid of tuning parameters. Through extensive simulations, we show PaC performs well. Finally, we apply PaC to a proprietary dataset in an exemplar Internet advertising case study and demonstrate its superiority over existing methods in this practical setting. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

Quantile Function on Scalar Regression Analysis for Distributional Data

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 90-106
Abstrak : Radiomics involves the study of tumor images to identify quantitative markers explaining cancer heterogeneity. The predominant approach is to extract hundreds to thousands of image features, including histogram features comprised of summaries of the marginal distribution of pixel intensities, which leads to multiple testing problems and can miss out on insights not contained in the selected features. In this paper, we present methods to model the entire marginal distribution of pixel intensities via the quantile function as functional data, regressed on a set of demographic, clinical, and genetic predictors to investigate their effects of imaging-based cancer heterogeneity. We call this approach quantile functional regression, regressing subject-specific marginal distributions across repeated measurements on a set of covariates, allowing us to assess which covariates are associated with the distribution in a global sense, as well as to identify distributional features characterizing these differences, including mean, variance, skewness, heavy-tailedness, and various upper and lower quantiles. To account for smoothness in the quantile functions, account for intrafunctional correlation, and gain statistical power, we introduce custom basis functions we call quantlets that are sparse, regularized, near-lossless, and empirically defined, adapting to the features of a given dataset and containing a Gaussian subspace so non-Gaussianness can be assessed. We fit this model using a Bayesian framework that uses nonlinear shrinkage of quantlet coefficients to regularize the functional regression coefficients and provides fully Bayesian inference after fitting a Markov chain Monte Carlo. We demonstrate the benefit of the basis space modeling through simulation studies, and apply the method to Magnetic resonance imaging (MRI)-based radiomic dataset from Glioblastoma Multiforme to relate imaging-based quantile functions to various demographic, clinical, and genetic predictors, finding specific differences in tumor pixel intensity distribution between males and females and between tumors with and without DDIT3 mutations. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

Mapping Tumor-Specific Expression QTLs in Impure Tumor Samples

Pengarang : DouglasR. Wilson
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 78-89
Abstrak : The study of gene expression quantitative trait loci (eQTL) is an effective approach to illuminate the functional roles of genetic variants. Computational methods have been developed for eQTL mapping using gene expression data from microarray or RNA-seq technology. Application of these methods for eQTL mapping in tumor tissues is problematic because tumor tissues are composed of both tumor and infiltrating normal cells (e.g., immune cells) and eQTL effects may vary between tumor and infiltrating normal cells. To address this challenge, we have developed a new method for eQTL mapping using RNA-seq data from tumor samples. Our method separately estimates the eQTL effects in tumor and infiltrating normal cells using both total expression and allele-specific expression (ASE). We demonstrate that our method controls Type I error rate and has higher power than some alternative approaches. We applied our method to study RNA-seq data from The Cancer Genome Atlas and illustrated the similarities and differences of eQTL effects in tumor and normal cells. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

Modeling Bronchiolitis Incidence Proportions in the Presence of Spatio-Temporal Uncertainty

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 66-78
Abstrak : Bronchiolitis (inflammation of the lower respiratory tract) in infants is primarily due to viral infection and is the single most common cause of infant hospitalization in the United States. To increase epidemiological understanding of bronchiolitis (and, subsequently, develop better prevention strategies), this research analyzes data on infant bronchiolitis cases from the U.S. Military Health System between the years 2003–2013 in Norfolk, Virginia, USA. For privacy reasons, child home addresses, birth dates, and diagnosis dates were randomized (jittered) creating spatio-temporal uncertainty in the geographic location and timing of bronchiolitis incidents. Using spatio-temporal point patterns, we created a modeling strategy that accounts for the jittering to estimate and quantify the uncertainty for the incidence proportion (IP) of bronchiolitis. Additionally, we regress the IP onto key covariates including pollution where we adequately account for uncertainty in the pollution levels (i.e., covariate uncertainty) using a land use regression model. Our analysis results indicate that the IP is positively associated with sulfur dioxide and population density. Further, we demonstrate how scientific conclusions may change if various sources of uncertainty (either spatio-temporal or covariate uncertainty) are not accounted for. Code submitted with this article was checked by an Associate Editor for Reproducibility and is available as an online supplement.

Demand Models With Random Partitions

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 47-65
Abstrak : Many economic models of consumer demand require researchers to partition sets of products or attributes prior to the analysis. These models are common in applied problems when the product space is large or spans multiple categories. While the partition is traditionally fixed a priori, we let the partition be a model parameter and propose a Bayesian method for inference. The challenge is that demand systems are commonly multivariate models that are not conditionally conjugate with respect to partition indices, precluding the use of Gibbs sampling. We solve this problem by constructing a new location-scale partition distribution that can generate random-walk Metropolis–Hastings proposals and also serve as a prior. Our method is illustrated in the context of a store-level category demand model, where we find that allowing for partition uncertainty is important for preserving model flexibility, improving demand forecasts, and learning about the structure of demand. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

Prediction and Inference With Missing Data in Patient Alert Systems

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 32-46
Abstrak : We describe the Bedside Patient Rescue (BPR) project, the goal of which is risk prediction of adverse events for non-intensive care unit patients using ∼100 variables (vitals, lab results, assessments, etc.). There are several missing predictor values for most patients, which in the health sciences is the norm, rather than the exception. A Bayesian approach is presented that addresses many of the shortcomings to standard approaches to missing predictors: (i) treatment of the uncertainty due to imputation is straight-forward in the Bayesian paradigm, (ii) the predictor distribution is flexibly modeled as an infinite normal mixture with latent variables to explicitly account for discrete predictors (i.e., as in multivariate probit regression models), and (iii) certain missing not at random situations can be handled effectively by allowing the indicator of missingness into the predictor distribution only to inform the distribution of the missing variables. The proposed approach also has the benefit of providing a distribution for the prediction, including the uncertainty inherent in the imputation. Therefore, we can ask questions such as: is it possible this individual is at high risk but we are missing too much information to know for sure? How much would we reduce the uncertainty in our risk prediction by obtaining a particular missing value? This approach is applied to the BPR problem resulting in excellent predictive capability to identify deteriorating patients. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

A Bayesian Approach to Multistate Hidden Markov Models: Application to Dementia Progression

Pengarang : Nicolas Debarsy
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 16-31
Abstrak : People are living longer than ever before, and with this arises new complications and challenges for humanity. Among the most pressing of these challenges is of understanding the role of aging in the development of dementia. This article is motivated by the Mayo Clinic Study of Aging data for 4742 subjects since 2004, and how it can be used to draw inference on the role of aging in the development of dementia. We construct a hidden Markov model (HMM) to represent progression of dementia from states associated with the buildup of amyloid plaque in the brain, and the loss of cortical thickness. A hierarchical Bayesian approach is taken to estimate the parameters of the HMM with a truly time-inhomogeneous infinitesimal generator matrix, and response functions of the continuous-valued biomarker measurements are cut-point agnostic. A Bayesian approach with these features could be useful in many disease progression models. Additionally, an approach is illustrated for correcting a common bias in delayed enrollment studies, in which some or all subjects are not observed at baseline. Standard software is incapable of accounting for this critical feature, so code to perform the estimation of the model described below is made available online. Code submitted with this article was checked by an Associate Editor for Reproducibility and is available as an online supplement.

A Hierarchical Model of Nonhomogeneous Poisson Processes for Twitter Retweets

Pengarang : -
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 529)
Halaman : 1-15
Abstrak : We present a hierarchical model of nonhomogeneous Poisson processes (NHPP) for information diffusion on online social media, in particular Twitter retweets. The retweets of each original tweet are modelled by a NHPP, for which the intensity function is a product of time-decaying components and another component that depends on the follower count of the original tweet author. The latter allows us to explain or predict the ultimate retweet count by a network centrality-related covariate. The inference algorithm enables the Bayes factor to be computed, to facilitate model selection. Finally, the model is applied to the retweet datasets of two hashtags. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement

Prediction, Estimation, and Attribution

Pengarang : Pengsheng Ji
Nama Majalah/Jurnal : Journal of the American Statistical Association
Volume / Edisi : 115 (No. 530)
Halaman : 636-655
Abstrak : The scientific needs and computational limitations of the twentieth century fashioned classical statistical methodology. Both the needs and limitations have changed in the twenty-first, and so has the methodology. Large-scale prediction algorithms—neural nets, deep learning, boosting, support vector machines, random forests—have achieved star status in the popular press. They are recognizable as heirs to the regression tradition, but ones carried out at enormous scale and on titanic datasets. How do these algorithms compare with standard regression techniques such as ordinary least squares or logistic regression? Several key discrepancies will be examined, centering on the differences between prediction and estimation or prediction and attribution (significance testing). Most of the discussion is carried out through small numerical examples.
← Back to HOME