Hua Zhou 0001

dblp:58/329-1 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-1320-7118ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 An approximate-copula distribution for statistical modeling
abstract
Copulas, generalized estimating equations, and generalized linear mixed models promote the analysis of grouped data where non-normal responses are correlated. Unfortunately, parameter estimation remains challenging in these three frameworks. Based on prior work of Tonda, we derive a new class of probability density functions that allow explicit calculation of moments, marginal and conditional distributions, and the score and observed information needed in maximum likelihood estimation. We also illustrate how the new distribution flexibly models longitudinal data following a non-Gaussian distribution. Finally, we conduct a tri-variate genome-wide association analysis on dichotomized systolic and diastolic blood pressure and body mass index data from the UK-Biobank, showcasing the modeling potential and computational scalability of the new distributional family.
Sarah S. Ji, Benjamin B. Chu, Hua Zhou 0001, Kenneth Lange
PLoS Comput. Biol.3
2026 An MM estimation algorithm for independent component analysis
Xun-Jian Li, Hua Zhou 0001, Kenneth Lange
Signal Process.2
2023 Multivariate genome-wide association analysis by iterative hard thresholding
abstract
MOTIVATION: In a genome-wide association study, analyzing multiple correlated traits simultaneously is potentially superior to analyzing the traits one by one. Standard methods for multivariate genome-wide association study operate marker-by-marker and are computationally intensive. RESULTS: We present a sparsity constrained regression algorithm for multivariate genome-wide association study based on iterative hard thresholding and implement it in a convenient Julia package MendelIHT.jl. In simulation studies with up to 100 quantitative traits, iterative hard thresholding exhibits similar true positive rates, smaller false positive rates, and faster execution times than GEMMA's linear mixed models and mv-PLINK's canonical correlation analysis. On UK Biobank data with 470 228 variants, MendelIHT completed a three-trait joint analysis (n=185 656) in 20 h and an 18-trait joint analysis (n=104 264) in 53 h with an 80 GB memory footprint. In short, MendelIHT enables geneticists to fit a single regression model that simultaneously considers the effect of all SNPs and dozens of traits. AVAILABILITY AND IMPLEMENTATION: Software, documentation, and scripts to reproduce our results are available from https://github.com/OpenMendel/MendelIHT.jl.
Benjamin B. Chu, Seyoon Ko, Jin J. Zhou, Aubrey Jensen, Hua Zhou 0001, Janet S. Sinsheimer, Kenneth Lange
Bioinform.5
2022 Extensions to the Proximal Distance Method of Constrained Optimization
abstract
The current paper studies the problem of minimizing a loss $f(\boldsymbol{x})$ subject to constraints of the form $\boldsymbol{D}\boldsymbol{x} \in S$, where $S$ is a closed set, convex or not, and $\boldsymbol{D}$ is a matrix that fuses parameters. Fusion constraints can capture smoothness, sparsity, or more general constraint patterns. To tackle this generic class of problems, we combine the Beltrami-Courant penalty method of optimization with the proximal distance principle. The latter is driven by minimization of penalized objectives $f(\boldsymbol{x})+\frac{\rho}{2}\text{dist}(\boldsymbol{D}\boldsymbol{x},S)^2$ involving large tuning constants $\rho$ and the squared Euclidean distance of $\boldsymbol{D}\boldsymbol{x}$ from $S$. The next iterate $\boldsymbol{x}_{n+1}$ of the corresponding proximal distance algorithm is constructed from the current iterate $\boldsymbol{x}_n$ by minimizing the majorizing surrogate function $f(\boldsymbol{x})+\frac{\rho}{2}\|\boldsymbol{D}\boldsymbol{x}-\mathcal{P}_{S}(\boldsymbol{D}\boldsymbol{x}_n)\|^2$. For fixed $\rho$ and a subanalytic loss $f(\boldsymbol{x})$ and a subanalytic constraint set $S$, we prove convergence to a stationary point. Under stronger assumptions, we provide convergence rates and demonstrate linear local convergence. We also construct a steepest descent variant to avoid costly linear system solves. To benchmark our algorithms, we compare their results to those delivered by the alternating direction method of multipliers. Our extensive numerical tests include problems on metric projection, convex regression, convex clustering, total variation image denoising, and projection of a matrix to a good condition number. These experiments demonstrate the superior speed and acceptable accuracy of our steepest variant on high-dimensional problems.
Alfonso Landeros, Oscar Hernan Madrid Padilla, Hua Zhou 0001, Kenneth Lange
J. Mach. Learn. Res.3
2021 A fast data-driven method for genotype imputation, phasing and local ancestry inference: MendelImpute.jl
abstract
MOTIVATION: Current methods for genotype imputation and phasing exploit the volume of data in haplotype reference panels and rely on hidden Markov models (HMMs). Existing programs all have essentially the same imputation accuracy, are computationally intensive and generally require prephasing the typed markers. RESULTS: We introduce a novel data-mining method for genotype imputation and phasing that substitutes highly efficient linear algebra routines for HMM calculations. This strategy, embodied in our Julia program MendelImpute.jl, avoids explicit assumptions about recombination and population structure while delivering similar prediction accuracy, better memory usage and an order of magnitude or better run-times compared to the fastest competing method. MendelImpute operates on both dosage data and unphased genotype data and simultaneously imputes missing genotypes and phase at both the typed and untyped SNPs (single nucleotide polymorphisms). Finally, MendelImpute naturally extends to global and local ancestry estimation and lends itself to new strategies for data compression and hence faster data transport and sharing. AVAILABILITY AND IMPLEMENTATION: Software, documentation and scripts to reproduce our results are available from https://github.com/OpenMendel/MendelImpute.jl. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin B. Chu, Eric M. Sobel, Rory Wasiolek, Seyoon Ko, Janet S. Sinsheimer, Hua Zhou 0001, Kenneth Lange
Bioinform.6
2021 Modern simulation utilities for genetic analysis
abstract
BACKGROUND: Statistical geneticists employ simulation to estimate the power of proposed studies, test new analysis tools, and evaluate properties of causal models. Although there are existing trait simulators, there is ample room for modernization. For example, most phenotype simulators are limited to Gaussian traits or traits transformable to normality, while ignoring qualitative traits and realistic, non-normal trait distributions. Also, modern computer languages, such as Julia, that accommodate parallelization and cloud-based computing are now mainstream but rarely used in older applications. To meet the challenges of contemporary big studies, it is important for geneticists to adopt new computational tools. RESULTS: We present TraitSimulation, an open-source Julia package that makes it trivial to quickly simulate phenotypes under a variety of genetic architectures. This package is integrated into our OpenMendel suite for easy downstream analyses. Julia was purpose-built for scientific programming and provides tremendous speed and memory efficiency, easy access to multi-CPU and GPU hardware, and to distributed and cloud-based parallelization. TraitSimulation is designed to encourage flexible trait simulation, including via the standard devices of applied statistics, generalized linear models (GLMs) and generalized linear mixed models (GLMMs). TraitSimulation also accommodates many study designs: unrelateds, sibships, pedigrees, or a mixture of all three. (Of course, for data with pedigrees or cryptic relationships, the simulation process must include the genetic dependencies among the individuals.) We consider an assortment of trait models and study designs to illustrate integrated simulation and analysis pipelines. Step-by-step instructions for these analyses are available in our electronic Jupyter notebooks on Github. These interactive notebooks are ideal for reproducible research. CONCLUSION: The TraitSimulation package has three main advantages. (1) It leverages the computational efficiency and ease of use of Julia to provide extremely fast, straightforward simulation of even the most complex genetic models, including GLMs and GLMMs. (2) It can be operated entirely within, but is not limited to, the integrated analysis pipeline of OpenMendel. And finally (3), by allowing a wider range of more realistic phenotype models, TraitSimulation brings power calculations and diagnostic tools closer to what investigators might see in real-world analyses.
Sarah S. Ji, Christopher A. German, Kenneth Lange, Janet S. Sinsheimer, Hua Zhou 0001, Eric M. Sobel
BMC Bioinform.5
2021 A novel nonlinear dimension reduction approach to infer population structure for low-coverage sequencing data
abstract
BACKGROUND: Low-depth sequencing allows researchers to increase sample size at the expense of lower accuracy. To incorporate uncertainties while maintaining statistical power, we introduce MCPCA_PopGen to analyze population structure of low-depth sequencing data. RESULTS: The method optimizes the choice of nonlinear transformations of dosages to maximize the Ky Fan norm of the covariance matrix. The transformation incorporates the uncertainty in calling between heterozygotes and the common homozygotes for loci having a rare allele and is more linear when both variants are common. CONCLUSIONS: We apply MCPCA_PopGen to samples from two indigenous Siberian populations and reveal hidden population structure accurately using only a single chromosome. The MCPCA_PopGen package is available on https://github.com/yiwenstat/MCPCA_PopGen .
Hua Zhou 0001, Joseph Watkins
BMC Bioinform.3
2020 Domain Knowledge-Assisted Automatic Diagnosis of Idiopathic Pulmonary Fibrosis (IPF) Using High Resolution Computed Tomography (HRCT) (Student Abstract)
abstract
Domain knowledge acquired from pilot studies is important for medical diagnosis. This paper leverages the population-level domain knowledge based on the D-optimal design criterion to judiciously select CT slices that are meaningful for the disease diagnosis task. As an illustrative example, the diagnosis of idiopathic pulmonary fibrosis (IPF) among interstitial lung disease (ILD) patients is used for this work. IPF diagnosis is complicated and is subject to inter-observer variability. We aim to construct a time/memory-efficient IPF diagnosis model using high resolution computed tomography (HRCT) with domain knowledge-assisted data dimension reduction methods. Four two-dimensional convolutional neural network (2D-CNN) architectures (MobileNet, VGG16, ResNet, and DenseNet) are implemented for an automatic diagnosis of IPF among ILD patients. Axial lung CT images are acquired from five multi-center clinical trials, which sum up to 330 IPF patients and 650 non-IPF ILD patients. Model performance is evaluated using five-fold cross-validation. Depending on the model setup, MobileNet achieved satisfactory results with overall sensitivity, specificity, and accuracy greater than 90%. Further evaluation of independent datasets is underway. Based on our knowledge, this is the first work that (1) uses population-level domain knowledge with optimal design criterion in selecting CT slices and (2) focuses on patient-level IPF diagnosis.
Wenxi Yu, Hua Zhou 0001, Jonathan G. Goldin, Hyun J. Kim
AAAI2
2020 Provable Convex Co-clustering of Tensors
abstract
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between statistical guarantees and computational efficiency for existing tensor clustering solutions due to the nature of their non-convex formulations. In this work, we bridge this gap by developing a provable convex formulation of tensor co-clustering. Our convex co-clustering (CoCo) estimator enjoys stability guarantees and its computational and storage costs are polynomial in the size of the data. We further establish a non-asymptotic error bound for the CoCo estimator, which reveals a surprising “blessing of dimensionality” phenomenon that does not exist in vector or matrix-variate cluster analysis. Our theoretical findings are supported by extensive simulated studies. Finally, we apply the CoCo estimator to the cluster analysis of advertisement click tensor data from a major online company. Our clustering results provide meaningful business insights to improve advertising effectiveness.
Eric C. Chi, Brian J. Gaines, Will Wei Sun, Hua Zhou 0001, Jian Yang 0002
J. Mach. Learn. Res.4
2019 Proximal Distance Algorithms: Theory and Practice
abstract
Proximal distance algorithms combine the classical penalty method of constrained minimization with distance majorization. If $f(x)$ is the loss function, and $C$ is the constraint set in a constrained minimization problem, then the proximal distance principle mandates minimizing the penalized loss $f(x)+\frac{\rho}{2}dist(x,C)^2$ and following the solution $x_{\rho}$ to its limit as $\rho$ tends to $\infty$. At each iteration the squared Euclidean distance $dist(x,C)^2$ is majorized by the spherical quadratic $\|x-P_C(x_k)\|^2$, where $P_C(x_k)$ denotes the projection of the current iterate $x_k$ onto $C$. The minimum of the surrogate function $f(x)+\frac{\rho}{2}\|x-P_C(x_k)\|^2$ is given by the proximal map $prox_{\rho^{-1}f}[P_C(x_k)]$. The next iterate $x_{k+1}$ automatically decreases the original penalized loss for fixed $\rho$. Since many explicit projections and proximal maps are known, it is straightforward to derive and implement novel optimization algorithms in this setting. These algorithms can take hundreds if not thousands of iterations to converge, but the simple nature of each iteration makes proximal distance algorithms competitive with traditional algorithms. For convex problems, proximal distance algorithms reduce to proximal gradient algorithms and therefore enjoy well understood convergence properties. For nonconvex problems, one can attack convergence by invoking Zangwill's theorem. Our numerical examples demonstrate the utility of proximal distance algorithms in various high-dimensional settings, including a) linear programming, b) constrained least squares, c) projection to the closest kinship matrix, d) projection onto a second-order cone constraint, e) calculation of Horn's copositive matrix index, f) linear complementarity programming, and g) sparse principal components analysis. The proximal distance algorithm in each case is competitive or superior in speed to traditional methods such as the interior point method and the alternating direction method of multipliers (ADMM). Source code for the numerical examples can be found at https://github.com/klkeys/proxdist.
Kevin L. Keys, Hua Zhou 0001, Kenneth Lange
J. Mach. Learn. Res.2
2017 Classification based on neuroimaging data by tensor boosting
abstract
Recent advances in medical imaging technologies generate a high volume of imaging data. Classification of cognitive outcome and disease status based on brain images is one of the most important tasks in neuroimaging studies. However it poses great challenge to the current classification methods due to the extremely high dimensionality and low signal to noise ratio in brain image data. In this article we propose a tensor boosting algorithm for classification based on neuroimaging data. The method is off-the-shelf, computationally simple and amenable to various modalities of neuroimaging data. The method is applied to an EEG data set from an alcoholism study and an MRI data set from an ADHD Global Competition and shows significantly improved classification performance.
Hua Zhou 0001, Liwei Wang 0008, Chul Sung
IJCNN2
2013 Mendel: the Swiss army knife of genetic analysis programs
abstract
Abstract Summary: Mendel is one of the few statistical genetics packages that provide a full spectrum of gene mapping methods, ranging from parametric linkage in large pedigrees to genome-wide association with rare variants. Our latest additions to Mendel anticipate and respond to the needs of the genetics community. Compared with earlier versions, Mendel is faster and easier to use and has a wider range of applications. Supported platforms include Linux, MacOS and Windows. Availability: Free from www.genetics.ucla.edu/software/mendel Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Kenneth Lange, Jeanette C. Papp, Janet S. Sinsheimer, Ram Sripracha, Hua Zhou 0001, Eric M. Sobel
Bioinform.5
2012 Nonlinear dimension reduction with Wright-Fisher kernel for genotype aggregation and association mapping
abstract
MOTIVATION: Association tests based on next-generation sequencing data are often under-powered due to the presence of rare variants and large amount of neutral or protective variants. A successful strategy is to aggregate genetic information within meaningful single-nucleotide polymorphism (SNP) sets, e.g. genes or pathways, and test association on SNP sets. Many existing methods for group-wise tests require specific assumptions about the direction of individual SNP effects and/or perform poorly in the presence of interactions. RESULTS: We propose a joint association test strategy based on two key components: a nonlinear supervised dimension reduction approach for effective SNP information aggregation and a novel kernel specially designed for qualitative genotype data. The new test demonstrates superior performance in identifying causal genes over existing methods across a large variety of disease models simulated from sequence data of real genes. In general, the proposed method provides an association test strategy that can (i) detect both rare and common causal variants, (ii) deal with both additive and interaction effect, (iii) handle both quantitative traits and disease dichotomies and (iv) incorporate non-genetic covariates. In addition, the new kernel can potentially boost the power of the entire family of kernel-based methods for genetic data analysis. AVAILABILITY: The method is implemented in MATLAB. Source code is available upon request. CONTACT: [email protected].
Lexin Li, Hua Zhou 0001
Bioinform.3
2010 Association screening of common and rare genetic variants by penalized regression
abstract
Abstract Motivation: This article extends our recent research on penalized estimation methods in genome-wide association studies to the realm of rare variants. Results: The new strategy is tested on both simulated and real data. Our findings on breast cancer data replicate previous results and shed light on variant effects within genes. Availability: Rare variant discovery by group penalized regression is now implemented in the free program Mendel at http://www.genetics.ucla.edu/software/ Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Hua Zhou 0001, Mary E. Sehl, Janet S. Sinsheimer, Kenneth Lange
Bioinform.1