
| Pengarang | : | Nabarun Deb |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 192-207 |
| Abstrak | : | In this article, we propose a general framework for distribution-free nonparametric testing in multi-dimensions, based on a notion of multivariate ranks defined using the theory of measure transportation. Unlike other existing proposals in the literature, these multivariate ranks share a number of useful properties with the usual one-dimensional ranks; most importantly, these ranks are distribution-free. This crucial observation allows us to design nonparametric tests that are exactly distribution-free under the null hypothesis. We demonstrate the applicability of this approach by constructing exact distribution-free tests for two classical nonparametric problems: (I) testing for mutual independence between random vectors, and (II) testing for the equality of multivariate distributions. In particular, we propose (multivariate) rank versions of distance covariance and energy statistic for testing scenarios (I) and (II), respectively. In both these problems, we derive the asymptotic null distribution of the proposed test statistics. We further show that our tests are consistent against all fixed alternatives. Moreover, the proposed tests are computationally feasible and are well-defined under minimal assumptions on the underlying distributions (e.g., they do not need any moment assumptions). We also demonstrate the efficacy of these procedures via extensive simulations. In the process of analyzing the theoretical properties of our procedures, we end up proving some new results in the theory of measure transportation and in the limit theory of permutation statistics using Stein’s method for exchangeable pairs, which may be of independent interest. |
| Pengarang | : | Zhenhua Lin |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 177-191 |
| Abstrak | : | We propose a new approach to the problem of high-dimensional multivariate ANOVA via bootstrapping max statistics that involve the differences of sample mean vectors. The proposed method proceeds via the construction of simultaneous confidence regions for the differences of population mean vectors. It is suited to simultaneously test the equality of several pairs of mean vectors of potentially more than two populations. By exploiting the variance decay property that is a natural feature in relevant applications, we are able to provide dimension-free and nearly parametric convergence rates for Gaussian approximation, bootstrap approximation, and the size of the test. We demonstrate the proposed approach with ANOVA problems for functional data and sparse count data. The proposed methodology is shown to work well in simulations and several real data applications. |
| Pengarang | : | Eugene Katsevich |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 165-176 |
| Abstrak | : | Scientific hypotheses in a variety of applications have domain-specific structures, such as the tree structure of the international classification of diseases (ICD), the directed acyclic graph structure of the gene ontology (GO), or the spatial structure in genome-wide association studies. In the context of multiple testing, the resulting relationships among hypotheses can create redundancies among rejections that hinder interpretability. This leads to the practice of filtering rejection sets obtained from multiple testing procedures, which may in turn invalidate their inferential guarantees. We propose Focused BH, a simple, flexible, and principled methodology to adjust for the application of any prespecified filter. We prove that Focused BH controls the false discovery rate under various conditions, including when the filter satisfies an intuitive monotonicity property and the p-values are positively dependent. We demonstrate in simulations that Focused BH performs well across a variety of settings, and illustrate this method’s practical utility via analyses of real datasets based on ICD and GO. |
| Pengarang | : | Jie Chen |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 147-164 |
| Abstrak | : | Gaussian random fields (GRF) are a fundamental stochastic model for spatiotemporal data analysis. An essential ingredient of GRF is the covariance function that characterizes the joint Gaussian distribution of the field. Commonly used covariance functions give rise to fully dense and unstructured covariance matrices, for which required calculations are notoriously expensive to carry out for large data. In this work, we propose a construction of covariance functions that result in matrices with a hierarchical structure. Empowered by matrix algorithms that scale linearly with the matrix dimension, the hierarchical structure is proved to be efficient for a variety of random field computations, including sampling, kriging, and likelihood evaluation. Specifically, with n scattered sites, sampling and likelihood evaluation has an O(n) cost and kriging has an ?????( log ????) cost after preprocessing, particularly favorable for the kriging of an extremely large number of sites (e.g., predicting on more sites than observed). We demonstrate comprehensive numerical experiments to show the use of the constructed covariance functions and their appealing computation time. Numerical examples on a laptop include simulated data of size up to one million, as well as a climate data product with over two million observations. |
| Pengarang | : | Kris H. Timotius [et al.] |
| Nama Majalah/Jurnal | : | Bina Darma |
| Volume / Edisi | : | 10-39 (No. 39) |
| Halaman | : | 100-106 |
| Abstrak | : | - |
| Pengarang | : | Wenxuan Zhong |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 135-146 |
| Abstrak | : | With rapid advances in information technology, massive datasets are collected in all fields of science, such as biology, chemistry, and social science. Useful or meaningful information is extracted from these data often through statistical learning or model fitting. In massive datasets, both sample size and number of predictors can be large, in which case conventional methods face computational challenges. Recently, an innovative and effective sampling scheme based on leverage scores via singular value decompositions has been proposed to select rows of a design matrix as a surrogate of the full data in linear regression. Analogously, variable screening can be viewed as selecting rows of the design matrix. However, effective variable selection along this line of thinking remains elusive. In this article, we bridge this gap to propose a weighted leverage variable screening method by using both the left and right singular vectors of the design matrix. We show theoretically and empirically that the predictors selected using our method can consistently include true predictors not only for linear models but also for complicated general index models. Extensive simulation studies show that the weighted leverage screening method is highly computationally efficient and effective. We also demonstrate its success in identifying carcinoma related genes using spatial transcriptome data. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Bina Darma |
| Volume / Edisi | : | 10-39 (No. 39) |
| Halaman | : | 95-99 |
| Abstrak | : | - |
| Pengarang | : | Tao Zhang |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 118 (No. 541) |
| Halaman | : | 122-134 |
| Abstrak | : | In this article, we develop uniform inference methods for the conditional mode based on quantile regression. Specifically, we propose to estimate the conditional mode by minimizing the derivative of the estimated conditional quantile function defined by smoothing the linear quantile regression estimator, and develop two bootstrap methods, a novel pivotal bootstrap and the nonparametric bootstrap, for our conditional mode estimator. Building on high-dimensional Gaussian approximation techniques, we establish the validity of simultaneous confidence rectangles constructed from the two bootstrap methods for the conditional mode. We also extend the preceding analysis to the case where the dimension of the covariate vector is increasing with the sample size. Finally, we conduct simulation experiments and a real data analysis using the U.S. wage data to demonstrate the finite sample performance of our inference method. The supplemental materials include the wage dataset, R codes and an appendix containing proofs of the main results, additional simulation results, discussion of model misspecification and quantile crossing, and additional details of the numerical implementation. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Bina Darma |
| Volume / Edisi | : | 10-39 (No. 39) |
| Halaman | : | 84-94 |
| Abstrak | : | - |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Bina Darma |
| Volume / Edisi | : | 10-39 (No. 39) |
| Halaman | : | 77-83 |
| Abstrak | : | - |