
| Pengarang | : | Lorin Crawford |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 531) |
| Halaman | : | 1139-1150 |
| Abstrak | : | Glioblastoma multiforme (GBM) is an aggressive form of human brain cancer that is under active study in the field of cancer biology. Its rapid progression and the relative time cost of obtaining molecular data make other readily available forms of data, such as images, an important resource for actionable measures in patients. Our goal is to use information given by medical images taken from GBM patients in statistical settings. To do this, we design a novel statistic—the smooth Euler characteristic transform (SECT)—that quantifies magnetic resonance images of tumors. Due to its well-defined inner product structure, the SECT can be used in a wider range of functional and nonparametric modeling approaches than other previously proposed topological summary statistics. When applied to a cohort of GBM patients, we find that the SECT is a better predictor of clinical outcomes than both existing tumor shape quantifications and common molecular assays. Specifically, we demonstrate that SECT features alone explain more of the variance in GBM patient survival than gene expression, volumetric features, and morphometric features. The main takeaways from our findings are thus 2-fold. First, they suggest that images contain valuable information that can play an important role in clinical prognosis and other medical decisions. Second, they show that the SECT is a viable tool for the broader study of medical imaging informatics. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Dipierro, Serena,Giacomin, Giovanni,Valdinoci, Enrico |
| Nama Majalah/Jurnal | : | Siam Journal On Applied Mathematics |
| Volume / Edisi | : | 83 (No. 5) |
| Halaman | : | 1935-1968 |
| Abstrak | : | We analyze the searching strategies of a forager diffusing in the whole space via an equation of fractional type. Specifically, the diffusion of the forager is regulated by a L\'evy flight whose exponent can be chosen in order to optimize a suitable foraging efficiency functional. Here, the dimension of the space is arbitrary. On the one hand, we show that the exponent s=0 corresponding to the limit case of heavy-tailed L\'evy flights is a pessimizer for the efficiency functional. On the other hand, we prove that, in situations of biological interest, one finds the most rewarding strategies arbitrarily close to s = 0. The combination of these results gives that the most rewarding searching option may turn out to be unfeasible, or at least unreliable, in practice, since small perturbations of the optimal searching exponent lead to pessimal patterns. The cases analyzed specifically are those of a target located in the proximity of the forager and that of sparse prey modeled by a target infinitely far from the initial position of the seeker. The efficiency functionals taken into account are either of pointwise type (in which the predator and the prey are modeled by moving points) or of set-dependent type (in which the predator and the prey correspond to regions of space with uniform density, thus modeling also the case of a sight range of the biological individuals involved). To implement our analysis, we also provide a number of structural results about finiteness, continuity, and asymptotic behaviors of the efficiency functionals. It is suggestive to relate the adoption of the most rewarding searching pattern close to pessimizers to a ``high-risk/high-gain"" strategy, in which the forager aims at high-energy content prey to mitigate the risk of failure. This setting is also connected to foraging modes of ``ambush"" type. |
| Pengarang | : | Li, Jialei,Liu, Xiaodong,Shi, Qingxiang |
| Nama Majalah/Jurnal | : | Siam Journal On Applied Mathematics |
| Volume / Edisi | : | 83 (No. 5) |
| Halaman | : | 1915-1934 |
| Abstrak | : | We consider the inverse time-harmonic elastic source problems with multifrequency sparse far field patterns. The unknown multiscale object is a combination of point sources and an extended source with compact support. Even if the extended source is unknown, we prove that the locations and the polarization strengths of the point sources can be uniquely determined by the multifrequency far field patterns at sparse observations. The least number of observation directions is given in terms of the number of the point sources. Having identified the point sources, under certain conditions, we further show that the multifrequency far field patterns at a fixed observation direction uniquely determine the narrowest strip containing the source support. With the increase of the number of observation directions, a convex hull of the source support can be obtained. Based on the constructive uniqueness proof, we introduce a novel direct sampling method both for locating the point sources and for reconstructing the support of the extended source. A formula for computing the polarization strengths is also given. Numerical examples in two dimensions are presented to verify the accuracy and robustness of the proposed methods for multiscale elastic sources. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 531) |
| Halaman | : | 1125-1138 |
| Abstrak | : | In the genomic era, the identification of gene signatures associated with disease is of significant interest. Such signatures are often used to predict clinical outcomes in new patients and aid clinical decision-making. However, recent studies have shown that gene signatures are often not replicable. This occurrence has practical implications regarding the generalizability and clinical applicability of such signatures. To improve replicability, we introduce a novel approach to select gene signatures from multiple datasets whose effects are consistently nonzero and account for between-study heterogeneity. We build our model upon some rank-based quantities, facilitating integration over different genomic datasets. A high-dimensional penalized generalized linear mixed model is used to select gene signatures and address data heterogeneity. We compare our method to some commonly used strategies that select gene signatures ignoring between-study heterogeneity. We provide asymptotic results justifying the performance of our method and demonstrate its advantage in the presence of heterogeneity through thorough simulation studies. Lastly, we motivate our method through a case study subtyping pancreatic cancer patients from four gene expression studies. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 531) |
| Halaman | : | 1111-1124 |
| Abstrak | : | People are increasingly concerned with understanding their personal environment, including possible exposure to harmful air pollutants. To make informed decisions on their day-to-day activities, they are interested in real-time information on a localized scale. Publicly available, fine-scale, high-quality air pollution measurements acquired using mobile monitors represent a paradigm shift in measurement technologies. A methodological framework utilizing these increasingly fine-scale measurements to provide real-time air pollution maps and short-term air quality forecasts on a fine-resolution spatial scale could prove to be instrumental in increasing public awareness and understanding. The Google Street View study provides a unique source of data with spatial and temporal complexities, with the potential to provide information about commuter exposure and hot spots within city streets with high traffic. We develop a computationally efficient spatiotemporal model for these data and use the model to make short-term forecasts and high-resolution maps of current air pollution levels. We also show via an experiment that mobile networks can provide more nuanced information than an equally sized fixed-location network. This modeling framework has important real-world implications in understanding citizens’ personal environments, as data production and real-time availability continue to be driven by the ongoing development and improvement of mobile measurement technologies. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 532) |
| Halaman | : | 1675-1688 |
| Abstrak | : | Does receiving a medical education outside the United States impact a surgeon’s performance? We study this question by matching operations performed by internationally trained surgeons to those performed by US-trained surgeons in reanalysis of a large health outcomes study. An effective matched design must achieve several goals, including balancing covariate distributions marginally, ensuring units within individual pairs have similar values on key covariates, and using a sufficiently large sample from the raw data. Yet in our study, optimizing some of these goals forces less desirable results on others. We address such tradeoffs from a multi-objective optimization perspective by creating matched designs that are Pareto optimal with respect to two goals. We provide general tools for generating representative subsets of Pareto optimal solution sets and articulate how they can be used to improve decision-making in observational study design. In the motivating surgical outcomes study, formulating a multi-objective version of the problem helps us balance an important variable without sacrificing two other design goals, average closeness of matched pairs on a multivariate distance and size of the final matched sample. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | Kenichiro McAlinn |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 531) |
| Halaman | : | 1092-1110 |
| Abstrak | : | We present new methodology and a case study in use of a class of Bayesian predictive synthesis (BPS) models for multivariate time series forecasting. This extends the foundational BPS framework to the multivariate setting, with detailed application in the topical and challenging context of multistep macroeconomic forecasting in a monetary policy setting. BPS evaluates—sequentially and adaptively over time—varying forecast biases and facets of miscalibration of individual forecast densities for multiple time series, and—critically—their time-varying inter-dependencies. We define BPS methodology for a new class of dynamic multivariate latent factor models implied by BPS theory. Structured dynamic latent factor BPS is here motivated by the application context—sequential forecasting of multiple U.S. macroeconomic time series with forecasts generated from several traditional econometric time series models. The case study highlights the potential of BPS to improve of forecasts of multiple series at multiple forecast horizons, and its use in learning dynamic relationships among forecasting models or agents. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 532) |
| Halaman | : | 1664-1674 |
| Abstrak | : | Calibration refers to the estimation of unknown parameters which are present in computer experiments but not available in physical experiments. An accurate estimation of these parameters is important because it provides a scientific understanding of the underlying system which is not available in physical experiments. Most of the work in the literature is limited to the analysis of continuous responses. Motivated by a study of cell adhesion experiments, we propose a new calibration framework for binary responses. Its application to the T cell adhesion data provides insight into the unknown values of the kinetic parameters which are difficult to determine by physical experiments due to the limitation of the existing experimental techniques. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 532) |
| Halaman | : | 1645-1663 |
| Abstrak | : | Kidney obstruction, if untreated in a timely manner, can lead to irreversible loss of renal function. A widely used technology for evaluations of kidneys with suspected obstruction is diuresis renography. However, it is generally very challenging for radiologists who typically interpret renography data in practice to build high level of competency due to the low volume of renography studies and insufficient training. Another challenge is that there is currently no gold standard for detection of kidney obstruction. Seeking to develop a computer-aided diagnostic (CAD) tool that can assist practicing radiologists to reduce errors in the interpretation of kidney obstruction, a recent study collected data from diuresis renography, interpretations on the renography data from highly experienced nuclear medicine experts as well as clinical data. To achieve the objective, we develop a statistical model that can be used as a CAD tool for assisting radiologists in kidney interpretation. We use a Bayesian latent class modeling approach for predicting kidney obstruction through the integrative analysis of time-series renogram data, expert ratings, and clinical variables. A nonparametric Bayesian latent factor regression approach is adopted for modeling renogram curves in which the coefficients of the basis functions are parameterized via the factor loadings dependent on the latent disease status and the extended latent factors that can also adjust for clinical variables. A hierarchical probit model is used for expert ratings, allowing for training with rating data from multiple experts while predicting with at most one expert, which makes the proposed model operable in practice. An efficient MCMC algorithm is developed to train the model and predict kidney obstruction with associated uncertainty. We demonstrate the superiority of the proposed method over several existing methods through extensive simulations. Analysis of the renal study also lends support to the usefulness of our model as a CAD tool to assist less experienced radiologists in the field. Supplementary materials for this article are available online. |
| Pengarang | : | - |
| Nama Majalah/Jurnal | : | Journal of the American Statistical Association |
| Volume / Edisi | : | 115 (No. 531) |
| Halaman | : | 1079-1091 |
| Abstrak | : | Studying the effects of groups of single nucleotide polymorphisms (SNPs), as in a gene, genetic pathway, or network, can provide novel insight into complex diseases such as breast cancer, uncovering new genetic associations and augmenting the information that can be gleaned from studying SNPs individually. Common challenges in set-based genetic association testing include weak effect sizes, correlation between SNPs in a SNP-set, and scarcity of signals, with individual SNP effects often ranging from extremely sparse to moderately sparse in number. Motivated by these challenges, we propose the Generalized Berk–Jones (GBJ) test for the association between a SNP-set and outcome. The GBJ extends the Berk–Jones statistic by accounting for correlation among SNPs, and it provides advantages over the Generalized Higher Criticism test when signals in a SNP-set are moderately sparse. We also provide an analytic p-value calculation for SNP-sets of any finite size, and we develop an omnibus statistic that is robust to the degree of signal sparsity. An additional advantage of our work is the ability to conduct inference using individual SNP summary statistics from a genome-wide association study (GWAS). We evaluate the finite sample performance of the GBJ through simulation and apply the method to identify breast cancer risk genes in a GWAS conducted by the Cancer Genetic Markers of Susceptibility Consortium. Our results suggest evidence of association between FGFR2 and breast cancer and also identify other potential susceptibility genes, complementing conventional SNP-level analysis. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement. |