Weihao Gao

dblp:145/3338 · DBLP profile ↗
← Back
27ranked-venue papers
13as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Theory of computation · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Tumor Segmentation and Basal Diameter Prediction Network for Uveal Melanoma: TSBPNet-UM
abstract
The basal diameter of uveal melanoma is critical for its prognosis and therapy, and it can indicate the metastatic risk of the tumor. Tumor segmentation is also significant for guiding clinical diagnosis, as it provides morphological and other essential information for clinicians. However, precise uveal melanoma segmentation and basal diameter prediction remain challenging for current computer-aided methods. The scarcity of large, annotated datasets and appropriate multi-task models hinders further exploration to simultaneously and accurately segment the tumors and predict their basal diameter for uveal melanoma. To address this challenge, we collected a novel dataset and proposed the Tumor Segmentation and Basal Diameter Prediction Network for Uveal Melanoma (TSBPNet-UM), which utilizes readily accessible fundus images to accomplish both tasks concurrently. Its two sub-networks, the Uveal Melanoma Segmentation and Basal Diameter Prediction networks, are designed for mutual performance enhancement. The Uveal Melanoma Segmentation network transfers spatial segmentation information to assist the Basal Diameter Prediction network. Conversely, the Basal Diameter Prediction network updates all parameters using basal diameter-derived gradient information. This model effectively overcomes the limitation that conventional models often struggle to converge on the direct prediction task of basal diameter. The superior performance and efficacy of the TSBPNetUM have been demonstrated through extensive experiments.
Weihao Gao, Zhuo Deng 0001, Jingyan Yang, Haihan Zhang, Wenbin Wei 0007
BIBM3
2025 MRAFF: Multi-resolution Attention Feature Fusion Method for Multi-label Fundus Image Ocular Disease Screening
abstract
Retinal images play a crucial role in diagnosing various eye diseases. However, due to the shortage of ophthalmologists, many patients do not receive timely diagnosis and treatment. Though automatic eye disease diagnosis has been improved due to recent advancements in deep learning, current deep learning models often exhibit unsatisfactory performance when applied to multi-labeled ocular disease classification tasks with fundus images. This paper addresses the limitations of existing models by proposing a new retinopathy dataset, the Resk dataset, and the Multi-Resolution Attention Feature Fusion (MRAFF) architecture. The MRAFF modifies current backbones with multiple downsample steps and incorporates three types of attention mechanisms: channel attention, spatial attention, and linear attention. Extensive experiments demonstrate that MRAFF-modified models exhibit improved feature extraction and classification abilities in handling ocular disease screening. The MRAFF is expected to enhance ocular disease diagnosis and may also suit other multi-label disease screenings, thereby inspiring future computer-aided diagnostic methods.
Chucheng Chen, Weihao Gao, Zhuo Deng 0001, Wenbin Wei 0007
IJCNN3
2025 UMGen: Multi-scale and Multi-region Joint Prediction Pipeline for Uveal Melanoma Prognosis-related Gene Subtyping
abstract
Uveal Melanoma(UM) is a highly aggressive ocular malignancy. Once metastasis occurs, the survival period is very short. An effective prognosis for UM is necessary. The BAP1 gene and Somatic Copy Number Alterations (SCNA) are two of the important indicators for clinical UM prognosis. However, current methods for detecting prognosis-related genes have limitations, including high costs and the need for advanced technical expertise, which makes them challenging to apply in routine clinical practice. Low-cost and efficient prognosis-related gene detection is worth studying. In this paper, we investigate the prognosis-related gene subtyping prediction problem based on Whole Slide Images(WSI). We propose a novel method, UMGen, for WSI-based UM prognosis-related gene subtyping prediction. Specifically, UMGen consists of a multi-region sampling module, a classifier module, and a joint decision module. The classifier module, SAGNet, applies spatial attention to extract multi-scale features. To improve accuracy and generalization, the multi-region joint decision is introduced. Comprehensive quantitative ablation experiments demonstrate that both our UMGen and SAGNet can surpass other competitors and have excellent generalization and robustness. The code and models will be released to the public for further research.
Zhuo Deng 0001, Weihao Gao, Wenbin Wei 0007
IJCNN5
2025 Revisiting the numerical feature embeddings structure in neural network-based tabular modelling
abstract
Tabular data is one of the most common forms of data in real-life applications. In tabular data modelling, methods based on Gradient Boosting Decision Trees (GBDT) and neural networks have demonstrated unique advantages on different datasets. However, existing neural network methods still face bottlenecks in processing continuous numerical features. This study revisits the structural design of numerical embedding in tabular modelling based on neural networks to overcome these bottlenecks. We conceptually decouple the numerical embedding process into three core functional modules: numerical augmentation, normalization, and encoding. Based on this design, we have developed three novel models–TabMLPNet, TabKANet, and TabKANet-PLNL. These models apply batch normalization to numerical features and feed them into a single, shared Multi-Layer Perceptrons (MLP) or Kolmogorov-Arnold networks (KAN) encoder. We propose a Piecewise Learnable Noise Layer (PLNL), which enriches data representation by introducing piecewise noise into numerical features, further enhancing the model’s generalization ability. Extensive experiments on 18 public datasets demonstrate that our models achieve significant performance improvements, outperforming other neural network models. This study underscores the significant benefits of integrating Kolmogorov-Arnold Network (KAN) with batch normalization. This combination not only enables efficient normalization of numerical features, but also allows dynamic learning of numerical distributions within each batch. As a result, it effectively mitigates performance bottlenecks that are often caused by numerical skewness. Additionally, we have confirmed that introducing augmented noise during the training process can enhance the model’s feature learning capabilities. Our implementation is publicly available on GitHub and can be accessed at https://github.com/AI-thpremed/TabKANet .
Weihao Gao, Zhuo Deng 0001
Knowl. Based Syst.1
2024 ProTeM: Unifying Protein Function Prediction via Text Matching
Ming Qin, Yuhao Wang 0006, Hongbin Ye, Zongbing Wang, Weihao Gao, Shangsong Liang, Qiang Zhang 0026, Keyan Ding
ICANN (8)7
2024 Enhancing LM's Task Adaptability: Powerful Post-training Framework with Reinforcement Learning from Model Feedback
Fuju Rong, Weihao Gao, Zhuo Deng 0001, Chucheng Chen, Wenze Zhang, Zhiyuan Niu
ICANN (7)2
2024 Predicting City Origin-Destination Flow with Generative Pre-training
Lizhong Gao, Weihao Gao
ICANN (9)4
2024 OphGLM: An ophthalmology large language-and-vision assistant
abstract
Vision computer-aided diagnostic methods have been used in early ophthalmic disease screening and diagnosis. However, the limited output formats of these methods lead to poor human-computer interaction and low clinical applicability value. Thus, ophthalmic visual question answering is worth studying. Unfortunately, no practical solutions exist before Large Language Models(LLMs). In this paper, we investigate the ophthalmic visual diagnostic interaction problem. We construct an ophthalmology large language-and-vision assistant, OphGLM, consisting of an image encoder, a text encoder, a fusion module, and an LLM module. We establish a new Chinese ophthalmic fine-tuning dataset, FundusTuning-CN, including the fundus instruction and conversation sets. Based on FundusTuning-CN, we establish a novel LLM-tuning strategy to introduce visual model understanding and ophthalmic knowledge into LLMs at a low cost and high efficiency. Leveraging the pre-training of the image encoder, OphGLM demonstrates strong visual understanding and surpasses open-source visual language models in common fundus disease classification tasks. The FundusTuning-CN enables OphGLM to surpass open-source medical LLMs in both ophthalmic knowledge and interactive capabilities. Our proposed OphGLM has the potential to revolutionize clinical applications in ophthalmology. The dataset, code, and models will be publicly available at https://github.com/ML-AILab/OphGLM.
Zhuo Deng 0001, Weihao Gao, Chucheng Chen, Zhiyuan Niu, Zhenjie Cao, Zhaoyi Ma, Wenbin Wei 0007
Artif. Intell. Medicine2
2024 Exploiting local detail in single image super-resolution via hypergraph convolution
Bufan Wang, Yongjun Zhang 0007, Weihao Gao, He Yao, Ruzhong Chen
Multim. Syst.3
2024 A novel attention-based network for single image dehazing
Weihao Gao, Yongjun Zhang 0007, Huachun Jian
Vis. Comput.1
2023 Machine Learning Force Fields with Data Cost Aware Training
abstract
Machine learning force fields (MLFF) have been proposed to accelerate molecular dynamics (MD) simulation, which finds widespread applications in chemistry and biomedical research. Even for the most data-efficient MLFFs, reaching chemical accuracy can require hundreds of frames of force and energy labels generated by expensive quantum mechanical algorithms, which may scale as $O(n^3)$ to $O(n^7)$, with $n$ proportional to the number of basis functions. To address this issue, we propose a multi-stage computational framework -- ASTEROID, which lowers the data cost of MLFFs by leveraging a combination of cheap inaccurate data and expensive accurate data. The motivation behind ASTEROID is that inaccurate data, though incurring large bias, can help capture the sophisticated structures of the underlying force field. Therefore, we first train a MLFF model on a large amount of inaccurate training data, employing a bias-aware loss function to prevent the model from overfitting the potential bias of this data. We then fine-tune the obtained model using a small amount of accurate training data, which preserves the knowledge learned from the inaccurate training data while significantly improving the model's accuracy. Moreover, we propose a variant of ASTEROID based on score matching for the setting where the inaccurate training data are unlabeled. Extensive experiments on MD datasets and downstream tasks validate the efficacy of ASTEROID. Our code and data are available at https://github.com/abukharin3/asteroid.
Alexander Bukharin, Shengjie Wang 0001, Simiao Zuo, Weihao Gao, Tuo Zhao
ICML5
2023 A deraining with detail-recovery network via context aggregation
Weihao Gao, Yongjun Zhang 0007, Zhongwei Cui
Multim. Syst.1
2022 Label Leakage and Protection in Two-party Split Learning
Oscar Li, Jiankai Sun, Xin Yang 0017, Weihao Gao, Junyuan Xie, Virginia Smith, Chong Wang 0002
ICLR4
2022 PathFlow: A normalizing flow generator that finds transition paths
abstract
Sampling from a Boltzmann distribution to calculate important macro statistics is one of the central tasks in the study of large atomic and molecular systems. Recently, a one-shot configuration sampler, the Boltzmann generator [Noé et al., 2019], is introduced. Though a Boltzmann generator can directly generate independent metastable states, it lacks the ability to find transition pathways and describe the whole transition process. In this paper, we propose PathFlow that can function as a one-shot generator as well as a transition pathfinder. More specifically, a normalizing flow model is constructed to map the base distribution and linear interpolated path in the latent space to the Boltzmann distribution and a minimum (free) energy path in the configuration space simultaneously. PathFlow can be trained by standard gradient-based optimizers using the proposed gradient estimator with a theoretical guarantee. PathFlow, validated with the extensively studied examples including a synthetic Müller potential and Alanine dipeptide, shows a remarkable performance.
Weihao Gao, Chong Wang 0002
UAI2
2021 Learning An End-to-End Structure for Retrieval in Large-Scale Recommendations
abstract
One of the core problems in large-scale recommendations is to retrieve top relevant candidates accurately and efficiently, preferably in sub-linear time. Previous approaches are mostly based on a two-step procedure: first learn an inner-product model, and then use some approximate nearest neighbor (ANN) search algorithm to find top candidates. In this paper, we present Deep Retrieval (DR), to learn a retrievable structure directly with user-item interaction data (e.g. clicks) without resorting to the Euclidean space assumption in ANN algorithms. DR's structure encodes all candidate items into a discrete latent space. Those latent codes for the candidates are model parameters and learnt together with other neural network parameters to maximize the same objective function. With the model learnt, a beam search over the structure is performed to retrieve the top candidates for reranking. Empirically, we first demonstrate that DR, with sub-linear computational complexity, can achieve almost the same accuracy as the brute-force baseline on two public datasets. Moreover, we show that, in a live production recommendation system, a deployed DR approach significantly outperforms a well-tuned ANN baseline in terms of engagement metrics. To the best of our knowledge, DR is among the first non-ANN algorithms successfully deployed at the scale of hundreds of millions of items for industrial recommendation systems.
Weihao Gao, Xiangjun Fan, Chong Wang 0002, Jiankai Sun, Wenzhi Xiao, Ruofan Ding, Xingyan Bin
CIKM1
2020 Information-Theoretic Understanding of Population Risk Improvement with Model Compression
abstract
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalization error; this allows for an interpretation of model compression as a regularization technique to avoid overfitting. We then characterize the increase in empirical risk with model compression using rate distortion theory. These results imply that the population risk could be improved by model compression if the decrease in generalization error exceeds the increase in empirical risk. We show through a linear regression example that such a decrease in population risk due to model compression is indeed possible. Our theoretical results further suggest that the Hessian-weighted K-means clustering compression approach can be improved by regularizing the distance between the clustering centers. We provide experiments with neural networks to support our theoretical assertions.
Yuheng Bu, Weihao Gao, Shaofeng Zou, Venugopal V. Veeravalli
AAAI2
2019 Learning One-hidden-layer Neural Networks under General Input Distributions
abstract
Significant advances have been made recently on training neural networks, where the main challenge is in solving an optimization problem with abundant critical points. However, existing approaches to address this issue crucially rely on a restrictive assumption: the training data is drawn from a Gaussian distribution. In this paper, we provide a novel unified framework to design loss functions with desirable landscape properties for a wide range of general input distributions. On these loss functions, remarkably, stochastic gradient descent theoretically recovers the true parameters with \emph{global} initializations and empirically outperforms the existing approaches. Our loss function design bridges the notion of score functions with the topic of neural network optimization. Central to our approach is the task of estimating the score function from samples, which is of basic and independent interest to theoretical statistics. Traditional estimation methods (example: kernel based) fail right at the outset; we bring statistical methods of local likelihood to design a novel estimator of score functions, that provably adapts to the local geometry of the unknown density.
Weihao Gao, Ashok Vardhan Makkuva, Sewoong Oh, Pramod Viswanath
AISTATS1
2019 Rate Distortion For Model Compression: From Theory To Practice
abstract
The enormous size of modern deep neural net-works makes it challenging to deploy those models in memory and communication limited scenarios. Thus, compressing a trained model without a significant loss in performance has become an increasingly important task. Tremendous advances has been made recently, where the main technical building blocks are pruning, quantization, and low-rank factorization. In this paper, we propose principled approaches to improve upon the common heuristics used in those building blocks, by studying the fundamental limit for model compression via the rate distortion theory. We prove a lower bound for the rate distortion function for model compression and prove its achievability for linear models. Although this achievable compression scheme is intractable in practice, this analysis motivates a novel objective function for model compression, which can be used to improve classes of model compressor such as pruning or quantization. Theoretically, we prove that the proposed scheme is optimal for compressing one-hidden-layer ReLU neural networks. Empirically,we show that the proposed scheme improves upon the baseline in the compression-accuracy tradeoff.
Weihao Gao, Yu-Han Liu, Chong Wang 0002, Sewoong Oh
ICML1
2018 The Nearest Neighbor Information Estimator is Adaptively Near Minimax Rate-Optimal
abstract
We analyze the Kozachenko–Leonenko (KL) fixed k-nearest neighbor estimator for the differential entropy. We obtain the first uniform upper bound on its performance for any fixed k over H\"{o}lder balls on a torus without assuming any conditions on how close the density could be from zero. Accompanying a recent minimax lower bound over the H\"{o}lder ball, we show that the KL estimator for any fixed k is achieving the minimax rates up to logarithmic factors without cognizance of the smoothness parameter s of the H\"{o}lder ball for $s \in (0,2]$ and arbitrary dimension d, rendering it the first estimator that provably satisfies this property.
Jiantao Jiao, Weihao Gao, Yanjun Han
NeurIPS2
2018 Breaking the Bandwidth Barrier: Geometrical Adaptive Entropy Estimation
abstract
Estimators of information theoretic measures, such as entropy and mutual information, are a basic workhorse for many downstream applications in modern data science. State-of-the-art approaches have been either geometric [nearest neighbor (NN)-based] or kernel-based (with a globally chosen bandwidth). In this paper, we combine both these approaches to design new estimators of entropy and mutual information that outperform the state-of-the-art methods. Our estimator uses local bandwidth choices of k -NN distances with a finite k , independent of the sample size. Such a local and data dependent choice ameliorates boundary bias and improves performance in practice, but the bandwidth is vanishing at a fast rate, leading to a non-vanishing bias. We show that the asymptotic bias of the proposed estimator is universal; it is independent of the underlying distribution. Hence, it can be precomputed and subtracted from the estimate. As a byproduct, we obtain a unified way of obtaining both the kernel and NN estimators. The corresponding theoretical contribution relating the asymptotic geometry of nearest neighbors to order statistics is of independent mathematical interest.
Weihao Gao, Sewoong Oh, Pramod Viswanath
IEEE Trans. Inf. Theory1
2018 Demystifying Fixed k-Nearest Neighbor Information Estimators
abstract
Estimating mutual information from independent identically distributed samples drawn from an unknown joint density function is a basic statistical problem of broad interest with multitudinous applications. The most popular estimator is the one proposed by Kraskov, Stögbauer, and Grassberger (KSG) in 2004 and is nonparametric and based on the distances of each sample to its kthnearest neighboring sample, where k is a fixed small integer. Despite of its widespread use (part of scientific software packages), theoretical properties of this estimator have been largely unexplored. In this paper, we demonstrate that the estimator is consistent and also identify an upper bound on the rate of convergence of the ℓ2error as a function of a number of samples. We argue that the performance benefits of the KSG estimator stems from a curious “correlation boosting” effect and build on this intuition to modify the KSG estimator in novel ways to construct a superior estimator. As a by-product of our investigations, we obtain nearly tight rates of convergence of the ℓ2error of the well-known fixed k-nearest neighbor estimator of differential entropy by Kozachenko and Leonenko.
Weihao Gao, Sewoong Oh, Pramod Viswanath
IEEE Trans. Inf. Theory1
2017 Demystifying fixed k-nearest neighbor information estimators
abstract
Estimating mutual information from i.i.d. samples drawn from an unknown joint density function is a basic statistical problem of broad interest with multitudinous applications. The most popular estimator is one proposed by Kraskov and Stogbauer and Grassberger (KSG) in 2004, and is nonparametric and based on the distances of each sample to its kthnearest neighboring sample, where k is a fixed small integer. Despite its widespread use (part of scientific software packages), theoretical properties of this estimator have been largely unexplored. In this paper we demonstrate that the estimator is consistent and also identify an upper bound on the rate of convergence of the ℓ2error as a function of number of samples. We argue that the performance benefits of the KSG estimator stems from a curious “correlation boosting” effect and build on this intuition to modify the KSG estimator in novel ways to construct a superior estimator. As a byproduct of our investigations, we obtain nearly tight rates of convergence of the ℓ2error of the well known fixed k nearest neighbor estimator of differential entropy by Kozachenko and Leonenko.
Weihao Gao, Sewoong Oh, Pramod Viswanath
ISIT1
2017 Density functional estimators with k-nearest neighbor bandwidths
abstract
Estimating expected polynomials of density functions from samples is a basic problem with numerous applications in statistics and information theory. Although kernel density estimators are widely used in practice for such functional estimation problems, practitioners are left on their own to choose an appropriate bandwidth for each application in hand. Further, kernel density estimators suffer from boundary biases, which are prevalent in real world data with lower dimensional structures. We propose using the fixed-k nearest neighbor distances for the bandwidth, which adaptively adjusts to local geometry. Further, we propose a novel estimator based on local likelihood density estimators, that mitigates the boundary biases. Although such a choice of fixed-k nearest neighbor distances to bandwidths results in inconsistent estimators, we provide a simple debiasing scheme that precomputes the asymptotic bias and divides off this term. With this novel correction, we show consistency of this debiased estimator. We provide numerical experiments suggesting that it improves upon competing state-of-the-art methods.
Weihao Gao, Sewoong Oh, Pramod Viswanath
ISIT1
2017 Estimating Mutual Information for Discrete-Continuous Mixtures
abstract
Estimation of mutual information from observed samples is a basic primitive in machine learning, useful in several learning tasks including correlation mining, information bottleneck, Chow-Liu tree, and conditional independence testing in (causal) graphical models. While mutual information is a quantity well-defined for general probability spaces, estimators have been developed only in the special case of discrete or continuous pairs of random variables. Most of these estimators operate using the 3H -principle, i.e., by calculating the three (differential) entropies of X, Y and the pair (X,Y). However, in general mixture spaces, such individual entropies are not well defined, even though mutual information is. In this paper, we develop a novel estimator for estimating mutual information in discrete-continuous mixtures. We prove the consistency of this estimator theoretically as well as demonstrate its excellent empirical performance. This problem is relevant in a wide-array of applications, where some variables are discrete, some continuous, and others are a mixture between continuous and discrete components.
Weihao Gao, Sreeram Kannan, Sewoong Oh, Pramod Viswanath
NIPS1
2017 Discovering Potential Correlations via Hypercontractivity
abstract
Discovering a correlation from one variable to another variable is of fundamental scientific and practical interest. While existing correlation measures are suitable for discovering average correlation, they fail to discover hidden or potential correlations. To bridge this gap, (i) we postulate a set of natural axioms that we expect a measure of potential correlation to satisfy; (ii) we show that the rate of information bottleneck, i.e., the hypercontractivity coefficient, satisfies all the proposed axioms; (iii) we provide a novel estimator to estimate the hypercontractivity coefficient from samples; and (iv) we provide numerical experiments demonstrating that this proposed estimator discovers potential correlations among various indicators of WHO datasets, is robust in discovering gene interactions from gene expression time series data, and is statistically more powerful than the estimators for other correlation measures in binary hypothesis testing of canonical examples of potential correlations.
Hyeji Kim, Weihao Gao, Sreeram Kannan, Sewoong Oh, Pramod Viswanath
NIPS2
2016 Conditional Dependence via Shannon Capacity: Axioms, Estimators and Applications
abstract
We consider axiomatically the problem of estimating the strength of a conditional dependence relationship P_Y|X from a random variables X to a random variable Y. This has applications in determining the strength of a known causal relationship, where the strength depends only on the conditional distribution of the effect given the cause (and not on the driving distribution of the cause). Shannon capacity, appropriately regularized, emerges as a natural measure under these axioms. We examine the problem of calculating Shannon capacity from the observed samples and propose a novel fixed-k nearest neighbor estimator, and demonstrate its consistency. Finally, we demonstrate an application to single-cell flow-cytometry, where the proposed estimators significantly reduce sample complexity.
Weihao Gao, Sreeram Kannan, Sewoong Oh, Pramod Viswanath
ICML1
2016 Breaking the Bandwidth Barrier: Geometrical Adaptive Entropy Estimation
abstract
Estimators of information theoretic measures such as entropy and mutual information from samples are a basic workhorse for many downstream applications in modern data science. State of the art approaches have been either geometric (nearest neighbor (NN) based) or kernel based (with bandwidth chosen to be data independent and vanishing sub linearly in the sample size). In this paper we combine both these approaches to design new estimators of entropy and mutual information that strongly outperform all state of the art methods. Our estimator uses bandwidth choice of fixed $k$-NN distances; such a choice is both data dependent and linearly vanishing in the sample size and necessitates a bias cancellation term that is universal and independent of the underlying distribution. As a byproduct, we obtain a unified way of obtaining both kernel and NN estimators. The corresponding theoretical contribution relating the geometry of NN distances to asymptotic order statistics is of independent mathematical interest.
Weihao Gao, Sewoong Oh, Pramod Viswanath
NIPS1