VLDB 2026 Research / reviewers in the wild / expert
Guosheng Yin
dblp:185/3223
· DBLP profile ↗
34ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0003-3276-1392ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 23 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIB-MIL: Sparsity-Induced Bayesian Neural Network for Robust Multiple Instance Learning on Whole Slide Image AnalysisabstractMultiple instance learning (MIL) has shown prominent success in analyzing whole slide histopathology images (WSIs). However, existing MIL methods often suffer from overfitting due to weak supervision and the "needle-in-a-haystack" nature of WSIs. Additionally, most deterministic approaches lack a mechanism for uncertainty quantification. While Bayesian neural networks (BNNs) have emerged as a promising solution to mitigate overfitting and enable uncertainty estimation by imposing prior constraints, commonly used Gaussian BNNs exhibit unstable posterior predictive distributions under weak supervision and suffer from high prediction variance. To tackle these challenges, we propose a sparsity-induced Bayesian Neural Network to be adopted in the MIL scheme, named SIB-MIL, for robust WSI prediction. Instead of using Gaussian prior distributions, we place a sparsity-induced prior, the Horseshoe prior, on the BNN parameters to address the variance overflowing issue. Such sparsity also filters unimportant noise and highlights salient regions, which only occupy a small proportion in WSIs. Empirical evaluations on cancer classification and subtyping tasks corroborate that not only can our method improve the existing MIL networks, but it also performs well in uncertainty quantification. Codes are available at https://github.com/HKU-MedAI/SIB-MIL. Yihang Chen 0001, Tsai Hor Chan, Jianning Chen, Guosheng Yin, Lequan Yu |
IEEE Trans. Medical Imaging | 5 |
| 2025 | MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait AnimationabstractRecent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce videos from multiple viewpoints, resulting in less controllable and expressive animations. Moreover, text-guided portrait animation remains underexplored, despite its user-friendly nature. We present a novel two-stage text-guided framework, MVPortrait (Multi-view Vivid Portrait), to generate expressive multi-view portrait animations that faithfully capture the described motion and emotion. MVPortrait is the first to introduce FLAME as an intermediate representation, effectively embedding facial movements, expressions, and view transformations within its parameter space. In the first stage, we separately train the FLAME motion and emotion diffusion models based on text input. In the second stage, we train a multi-view video generation model conditioned on a reference portrait image and multi-view FLAME rendering sequences from the first stage. Experimental results exhibit that MVPortrait outperforms existing methods in terms of motion and emotion control, as well as view consistency. Furthermore, by leveraging FLAME as a bridge, MVPortrait becomes the first controllable portrait animation framework that is compatible with text, speech, and video as driving signals. Yukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren, Adela S. M. Lau, Guosheng Yin, Xiu Li 0001 |
CVPR | 6 |
| 2025 | Ensemble Manifold Learning via Model AveragingabstractManifold learning aims to extract information of high-dimensional data and provide low-dimensional representations while preserving nonlinear structures of the input data. Numerous manifold learning algorithms have been proposed in the literature. We develop a model averaging procedure to combine different manifold learning algorithms for enhancing the robustness of the result. Toward this goal, we propose a new quality metric that is tuning-free and scale-invariant by utilizing the Mahalanobis distance. The quality metric can also be used for selection of tuning parameters. By optimizing the weights and ensembling different approaches, our model averaging manifold learning method is shown to achieve more robust results, because no single method in the ensemble can dominate others in all scenarios. Through synthetic and real data examples, we show that the new metric outperforms existing ones and the model averaging outcome provides a unanimously superior outcome that is always competitive with respect to visualization or classification under different contexts. Ruoxu Tan, Guosheng Yin |
ECAI | 2 |
| 2025 | Effective Sample Size Estimation Based on the Concordance Between p-Value and Posterior Probability of the Null HypothesisabstractEstimating the effective sample size (ESS) of a prior distribution is an age-old yet pivotal challenge, with great implications for clinical trials and various biomedical applications. Although numerous endeavors have been dedicated to this pursuit, most of them neglect the likelihood context in which the prior is embedded, thereby considering all priors as “beneficial”. In the limited studies of addressing harmful priors, specifying a baseline prior remains an indispensable step. By means of the elegant bridge between the p-value and the posterior probability of the null hypothesis, we propose a new ESS estimation method based on p-value in the framework of hypothesis testing, expanding the scope of existing ESS estimation methods in three key aspects: (i) We address the specific likelihood context of the prior, enabling the possibility of negative ESS values in case of prior–likelihood disconcordance; (ii) By leveraging the well-established bridge between the frequentist and Bayesian configurations under noninformative priors, there is no need to specify a baseline prior which incurs another criticism of subjectivity; (iii) By incorporating ESS into the hypothesis testing framework, our hypothesis test-based ESS estimation method transcends the conventional one-ESS-one-prior paradigm and accommodates one-ESS-multiple-priors paradigm, where the sole ESS may reflect the collaborative impact of multiple priors in diverse contexts. Through comprehensive simulation analyses, we demonstrate the superior performance of the hypothesis test-based ESS estimation method in comparison with existing approaches. Furthermore, by applying this approach to an expression quantitative trait loci (eQTL) data analysis, we show the effectiveness of informative priors in uncovering gene eQTL loci. Han Wang 0061, Yan Dora Zhang, Guosheng Yin |
ECAI | 3 |
| 2025 | Cross-Modal Alignment via Variational Copula ModellingabstractVarious data modalities are common in real-world applications. (e.g., EHR, medical images and clinical notes in healthcare). Thus, it is essential to develop multimodal learning methods to aggregate information from multiple modalities. The main challenge is appropriately aligning and fusing the representations of different modalities into a joint distribution. Existing methods mainly rely on concatenation or the Kronecker product, oversimplifying interactions structure between modalities and indicating a need to model more complex interactions. Additionally, the joint distribution of latent representations with higher-order interactions is underexplored. Copula is a powerful statistical structure in modelling the interactions between variables, as it bridges the joint distribution and marginal distributions of multiple variables. In this paper, we propose a novel copula modelling-driven multimodal learning framework, which focuses on learning the joint distribution of various modalities to capture the complex interaction among them. The key idea is interpreting the copula model as a tool to align the marginal distributions of the modalities efficiently. By assuming a Gaussian mixture distribution for each modality and a copula model on the joint distribution, our model can also generate accurate representations for missing modalities. Extensive experiments on public MIMIC datasets demonstrate the superior performance of our model over other competitors. The code is anonymously available at https://github.com/HKU-MedAI/CMCM. Tsai Hor Chan, Fuying Wang, Guosheng Yin, Lequan Yu |
ICML | 4 |
| 2025 | Automatic Radiotherapy Treatment Planning with Deep Functional Reinforcement Learning
Bin Liu 0022, Yu Liu 0129, Zhiqian Li, Jianghong Xiao, Guosheng Yin, Huazhen Lin |
KDD (1) | 5 |
| 2025 | Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet ProcessabstractDeveloping effective multimodal fusion approaches has become increasingly essential in many real-world scenarios, such as health care and finance.
The key challenge is how to preserve the feature expressiveness in each modality while learning cross-modal interactions.
Previous approaches primarily focus on the cross-modal alignment,
while over-emphasis on the alignment of marginal distributions of modalities may impose excess regularization and obstruct meaningful representations within each modality.
The Dirichlet process (DP) mixture model is a powerful Bayesian non-parametric method that can amplify the most prominent features by its richer-gets-richer property, which allocates increasing weights to them.
Inspired by this unique characteristic of DP, we propose a new DP-driven multimodal learning framework that automatically achieves an optimal balance between prominent intra-modal representation learning and cross-modal alignment.
Specifically, we assume that each modality follows a mixture of multivariate Gaussian distributions and further adopt DP to calculate the mixture weights for all the components. This paradigm allows DP to dynamically allocate the contributions of features and select the most prominent ones, leveraging its richer-gets-richer property, thus facilitating multimodal feature fusion.
Extensive experiments on several multimodal datasets demonstrate the superior performance of our model over other competitors.
Ablation analysis further validates the effectiveness of DP in aligning modality distributions and its robustness to changes in key hyperparameters.
Code is anonymously available at https://github.com/HKU-MedAI/DPMM.git Tsai Hor Chan, Yihang Chen 0001, Guosheng Yin, Lequan Yu |
NeurIPS | 4 |
| 2025 | Variational Pólya TreeabstractDensity estimation is essential for generative modeling, particularly with the rise of modern neural networks. While existing methods capture complex data distributions, they often lack interpretability and uncertainty quantification. Bayesian nonparametric methods, especially the Pólya tree, offer a robust framework that addresses these issues by accurately capturing function behavior over small intervals. Traditional techniques like Markov chain Monte Carlo (MCMC) face high computational complexity and scalability limitations, hindering the use of Bayesian nonparametric methods in deep learning. To tackle this, we introduce the variational Pólya tree (VPT) model, which employs stochastic variational inference to compute posterior distributions. This model provides a flexible, nonparametric Bayesian prior that captures latent densities and works well with stochastic gradient optimization. We also leverage the joint distribution likelihood for a more precise variational posterior approximation than traditional mean-field methods. We evaluate the model performance on both real data and images, and demonstrate its competitiveness with other state-of-the-art deep density estimation methods. We also explore its ability in enhancing interpretability and uncertainty quantification. Code is available at https://github.com/howardchanth/var-polya-tree. Tsai Hor Chan, Lequan Yu, Kwok Fai Lam, Guosheng Yin |
NeurIPS | 5 |
| 2025 | Hierarchical and Stochastic Crystallization Learning: Geometrically Leveraged Nonparametric Regression with Delaunay TriangulationabstractHigh-dimensionality is known to be the bottleneck for both nonparametric regression and the Delaunay triangulation. To efficiently exploit the advantage of the Delaunay triangulation in utilizing geometry information for nonparametric regression without conducting the Delaunay triangulation for the entire feature space, we develop the crystallization search for the neighbor Delaunay simplices of the target point similar to crystal growth and estimate the conditional expectation function by fitting a local linear model to the data points of the constructed Delaunay simplices. Because the shapes and volumes of Delaunay simplices are adaptive to the density of feature data points, our method selects neighbor data points more uniformly in all directions in comparison with Euclidean distance based methods and thus it is more robust to the local geometric structure of the data. We further develop the stochastic approach to hyperparameter selection and the hierarchical crystallization learning under multimodal feature data densities, where an approximate global Delaunay triangulation is obtained by first triangulating the local centers and then constructing local Delaunay triangulations in parallel. We study the asymptotic properties of our method and conduct numerical experiments on both synthetic and real data to demonstrate the advantages of our method over the existing ones. Jiaqi Gu 0005, Guosheng Yin |
J. Mach. Learn. Res. | 2 |
| 2025 | Causal Effect of Functional TreatmentabstractWe study the causal effect with a functional treatment variable, where practical applications often arise in neuroscience, biomedical sciences, etc. Previous research concerning the effect of a functional variable on an outcome is typically restricted to exploring correlation rather than causality. The generalized propensity score, which is often used to calibrate the selection bias, is not directly applicable to a functional treatment variable due to a lack of definition of probability density function for functional data. We propose three estimators for the average dose-response functional based on the functional linear model, namely, the functional stabilized weight estimator, the outcome regression estimator and the doubly robust estimator, each of which has its own merits. We study their theoretical properties, which are corroborated through extensive numerical experiments. A real data application on electroencephalography data and disease severity demonstrates the practical value of our methods. Ruoxu Tan, Guosheng Yin |
J. Mach. Learn. Res. | 4 |
| 2025 | Democratizing large language model-based graph data augmentation via latent knowledge graphsabstractData augmentation is necessary for graph representation learning due to the scarcity and noise present in graph data. Most of the existing augmentation methods overlook the context information inherited from the dataset as they rely solely on the graph structure for augmentation. Despite the success of some large language model-based (LLM) graph learning methods, they are mostly white-box which require access to the weights or latent features from the open-access LLMs, making them difficult to be democratized for everyone as the most advanced LLMs are often closed-source for commercial considerations. To overcome these limitations, we propose a black-box context-driven graph data augmentation approach, with the guidance of LLMs - DemoGraph. Leveraging the text prompt as context-related information, we task the LLM with generating knowledge graphs (KGs), which allow us to capture the structural interactions from the text outputs. We then design a dynamic merging schema to stochastically integrate the LLM-generated KGs into the original graph during training. To control the sparsity of the augmented graph, we further devise a granularity-aware prompting strategy and an instruction fine-tuning module, which seamlessly generates text prompts according to different granularity levels of the dataset. Extensive experiments on various graph learning tasks validate the effectiveness of our method over existing graph data augmentation methods. Notably, our approach excels in scenarios involving electronic health records (EHRs), which validates its maximal utilization of contextual knowledge, leading to enhanced predictive performance and interpretability. Yushi Feng, Tsai Hor Chan, Guosheng Yin, Lequan Yu |
Neural Networks | 3 |
| 2025 | Feature Preserving Shrinkage on Bayesian Neural Networks Via the R2D2 PriorabstractBayesian neural networks (BNNs) treat neural network weights as random variables, which aim to provide posterior uncertainty estimates and avoid overfitting by performing inference on the posterior weights. However, selection of appropriate prior distributions remains a challenging task, and BNNs may suffer from catastrophic inflated variance or poor predictive performance when poor choices are made for the priors. Existing BNN designs apply different priors to weights, while the behaviours of these priors make it difficult to sufficiently shrink noisy signals or they are prone to overshrinking important signals in the weights. To alleviate this problem, we propose a novel R2D2-Net, which imposes the $R^{2}$R2-induced Dirichlet Decomposition (R2D2) prior to the BNN weights. The R2D2-Net can effectively shrink irrelevant coefficients towards zero, while preventing key features from over-shrinkage. To approximate the posterior distribution of weights more accurately, we further propose a variational Gibbs inference algorithm that combines the Gibbs updating procedure and gradient-based optimization. This strategy enhances stability and consistency in estimation when the variational objective involving the shrinkage parameters is non-convex. We also analyze the evidence lower bound (ELBO) and the posterior concentration rates from a theoretical perspective. Experiments on both natural and medical image classification and uncertainty estimation tasks demonstrate satisfactory performances of our method. Tsai Hor Chan, Dora Yan Zhang, Guosheng Yin, Lequan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Graph Portfolio: High-Frequency Factor Predictors via Heterogeneous Continual GNNsabstractThis study aims to address the challenges of financial price prediction in high-frequency trading (HFT) by introducing a novel continual learning framework based on factor predictors via graph neural networks. The model integrates multi-factor pricing theory with real-time market dynamics, effectively bypassing the limitations of conventional time series forecasting methods, which often lack financial theory guidance and ignore market correlations. We propose three heterogeneous tasks, including price gap regression, changepoint detection, and price moving average regression to trace the short-, intermediate-, and long-term trend factors present in the data. We also account for the cross-sectional correlations inherent in the financial market, where prices of different assets show strong dynamic correlations. To accurately capture these dynamic relationships, we resort to spatio-temporal graph neural network (STGNN) to enhance the predictive power of the model. Our model allows a continual learning strategy to simultaneously consider these tasks (factors). To tackle the catastrophic forgetting in continual learning while considering the heterogeneity of tasks, we propose to calculate parameter importance with mutual information between original observations and the extracted features. Empirical studies on the Chinese futures data and U.S. equity data demonstrate the superior performance of the proposed model compared to other state-of-the-art approaches. Zhi-zhong Tan, Bin Liu 0022, Guosheng Yin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Enhancing Semi-supervised Domain Adaptation via Effective Target LabelingabstractExisting semi-supervised domain adaptation (SSDA) models have exhibited impressive performance on the target domain by effectively utilizing few labeled target samples per class (e.g., 3 samples per class). To guarantee an equal number of labeled target samples for each class, however, they require domain experts to manually recognize a considerable amount of the unlabeled target data. Moreover, as the target samples are not equally informative for shaping the decision boundaries of the learning models, it is crucial to select the most informative target samples for labeling, which is, however, impossible for human selectors. As a remedy, we propose an EFfective Target Labeling (EFTL) framework that harnesses active learning and pseudo-labeling strategies to automatically select some informative target samples to annotate. Concretely, we introduce a novel sample query strategy, called non-maximal degree node suppression (NDNS), that iteratively performs maximal degree node query and non-maximal degree node removal to select representative and diverse target samples for labeling. To learn target-specific characteristics, we propose a novel pseudo-labeling strategy that attempts to label low-confidence target samples accurately via clustering consistency (CC), and then inject information of the model uncertainty into our query process. CC enhances the utilization of the annotation budget and increases the number of “labeled” target samples while requiring no additional manual effort. Our proposed EFTL framework can be easily coupled with existing SSDA models, showing significant improvements on three benchmarks Jiujun He, Guosheng Yin |
AAAI | 3 |
| 2024 | cDP-MIL: Robust Multiple Instance Learning via Cascaded Dirichlet Process
Yihang Chen 0001, Tsai Hor Chan, Guosheng Yin, Yuming Jiang 0005, Lequan Yu |
ECCV (54) | 3 |
| 2024 | Futures Quantitative Investment With Heterogeneous Continual Graph Neural NetworkabstractIt is challenging to predict futures prices with traditional econometric models as it necessitates a comprehensive consideration of both historical observations and correlations among various futures. Spatial-temporal graph neural networks (STGNNs) offer a promising approach for modeling such complex spatio-temporal data. Nevertheless, the direct application of STGNNs to high-frequency futures data remains a challenge, as traders must account for both short- and long-term characteristics when making decisions. To capture these distinct timeframes, we leverage additional label information by devising four heterogeneous tasks: price regression, price gap regression, price moving average regression, and change-point detection. To make full use of these labels, we train the model in a continual manner. Traditional continual GNNs define the gradient of losses as parameter importance to overcome the catastrophic forgetting (CF) issue, while this is unsuitable for our heterogeneous tasks with losses in distinct spaces. Therefore, we propose to calculate the parameter importance with mutual information between the original observations and the extracted features. The empirical results based on 49 varieties of commodity futures demonstrate that our model achieves superior performance compared to the leading approaches, highlighting the effectiveness of our approach in modeling the intricate dynamics of the futures market. Zhi-zhong Tan, Bin Liu 0022, Guosheng Yin |
ICDM | 4 |
| 2024 | Multi-task heterogeneous graph learning on electronic health records
Tsai Hor Chan, Guosheng Yin, Kyongtae Bae, Lequan Yu |
Neural Networks | 2 |
| 2023 | Histopathology Whole Slide Image Analysis with Heterogeneous Graph Representation LearningabstractGraph-based methods have been extensively applied to whole slide histopathology image (WSI) analysis due to the advantage of modeling the spatial relationships among different entities. However, most of the existing methods focus on modeling WSIs with homogeneous graphs (e.g., with homogeneous node type). Despite their successes, these works are incapable of mining the complex structural relations between biological entities (e.g., the diverse interaction among different cell types) in the WSI. We propose a novel heterogeneous graph-based framework to leverage the inter-relationships among different types of nuclei for WSI analysis. Specifically, we formulate the WSI as a heterogeneous graph with “nucleus-type” attribute to each node and a semantic similarity attribute to each edge. We then present a new heterogeneous-graph edge attribute transformer (HEAT) to take advantage of the edge and node heterogeneity during massage aggregating. Further, we design a new pseudo-label-based semantic-consistent pooling mechanism to obtain graph-level features, which can mitigate the over-parameterization issue of conventional cluster-based pooling. Additionally, observing the limitations of existing association-based localization methods, we propose a causal-driven approach attributing the contribution of each node to improve the interpretability of our framework. Extensive experiments on three public TCGA benchmark datasets demonstrate that our frame-work outperforms the state-of-the-art methods with considerable margins on various tasks. Our codes are available at https://github.com/HKU-MedAI/WSI-HGNN. Tsai Hor Chan, Fernando Julio Cendra, Guosheng Yin, Lequan Yu |
CVPR | 4 |
| 2023 | Interpret ESG Rating's Impact on the Industrial Chain Using Graph Neural NetworksabstractWe conduct a quantitative analysis of the development of the industry chain from the environmental, social, and governance (ESG) perspective, which is an overall measure of sustainability. Factors that may impact the performance of the industrial chain have been studied in the literature, such as government regulation, monetary policy, etc. Our interest lies in how the sustainability change (i.e., ESG shock) affects the performance of the industrial chain. To achieve this goal, we model the industrial chain with a graph neural network (GNN) and conduct node regression on two financial performance metrics, namely, the aggregated profitability ratios and operating margin. To quantify the effects of ESG, we propose to compute the interaction between ESG shocks and industrial chain features with a cross-attention module, and then filter the original node features in the graph regression. Experiments on two real datasets demonstrate that (i) there are significant effects of ESG shocks on the industrial chain, and (ii) model parameters including regression coefficients and the attention map can explain how ESG shocks affect the performance of the industrial chain. Jiujun He, Guosheng Yin |
IJCAI | 6 |
| 2023 | 3D-Polishing for Triangular Mesh Compression of Point Cloud DataabstractTriangular meshes are commonly used to reconstruct the surfaces of 3-dimensional (3D) objects based on the point cloud data. With an increasing demand for high-quality approximation, the sizes of point cloud data and the generated triangular meshes continue to increase, resulting in high computational cost in data processing, visualization, analysis, transmission and storage. Motivated by the process of sculpture polishing, we develop a novel progressive mesh compression approach, called the greedy 3D-polishing algorithm, to sequentially remove redundant points and triangles in a greedy manner while maintaining the approximation quality of the surface. Based on the polishing algorithm, we propose the approximate curvature radius to evaluate the scale of features polished at each iteration. By reformulating the compression rate selection as a change-point detection problem, a rank-based procedure is developed to select the optimal compression rate so that most of the global features of the 3D object surface can be preserved with statistical guarantee. Experiments on both moderate- and large-scale 3D shape datasets show that the proposed method can substantially reduce the size of the point cloud data and the corresponding triangular mesh, so that most of the surface information can be retained by a much smaller number of points and triangles. Jiaqi Gu 0005, Guosheng Yin |
KDD | 2 |
| 2023 | Adaptive Uncertainty Estimation via High-Dimensional Testing on Latent RepresentationsabstractUncertainty estimation aims to evaluate the confidence of a trained deep neural network. However, existing uncertainty estimation approaches rely on low-dimensional distributional assumptions and thus suffer from the high dimensionality of latent features. Existing approaches tend to focus on uncertainty on discrete classification probabilities, which leads to poor generalizability to uncertainty estimation for other tasks. Moreover, most of the literature requires seeing the out-of-distribution (OOD) data in the training for better estimation of uncertainty, which limits the uncertainty estimation performance in practice because the OOD data are typically unseen. To overcome these limitations, we propose a new framework using data-adaptive high-dimensional hypothesis testing for uncertainty estimation, which leverages the statistical properties of the feature representations. Our method directly operates on latent representations and thus does not require retraining the feature encoder under a modified objective. The test statistic relaxes the feature distribution assumptions to high dimensionality, and it is more discriminative to uncertainties in the latent representations. We demonstrate that encoding features with Bayesian neural networks can enhance testing performance and lead to more accurate uncertainty estimation. We further introduce a family-wise testing procedure to determine the optimal threshold of OOD detection, which minimizes the false discovery rate (FDR). Extensive experiments validate the satisfactory performance of our framework on uncertainty estimation and task-specific prediction over a variety of competitors. The experiments on the OOD detection task also show satisfactory performance of our method when the OOD data are unseen in the training. Codes are available at https://github.com/HKU-MedAI/bnn_uncertainty. Tsai Hor Chan, Kin Wai Lau, Guosheng Yin, Lequan Yu |
NeurIPS | 4 |
| 2023 | Graphmax for Text GenerationabstractIn text generation, a large language model (LM) makes a choice of each new word based only on the former selection of its context using the softmax function. Nevertheless, the link statistics information of concurrent words based on a scene-specific corpus is valuable in choosing the next word, which can help to ensure the topic of the generated text to be aligned with the current task. To fully explore the co-occurrence information, we propose a graphmax function for task-specific text generation. Using the graph-based regularization, graphmax enables the final word choice to be determined by both the global knowledge from the LM and the local knowledge from the scene-specific corpus. The traditional softmax function is regularized with a graph total variation (GTV) term, which incorporates the local knowledge into the LM and encourages the model to consider the statistical relationships between words in a scene-specific corpus. The proposed graphmax is versatile and can be readily plugged into any large pre-trained LM for text generation and machine translation. Through extensive experiments, we demonstrate that the new GTV-based regularization can improve performances in various natural language processing (NLP) tasks in comparison with existing methods. Moreover, through human experiments, we observe that participants can easily distinguish the text generated by graphmax or softmax. Guosheng Yin |
J. Artif. Intell. Res. | 2 |
| 2022 | Deep Reinforcement Learning for Bandit Arm LocalizationabstractIn the multi-armed bandit (MAB) framework, we investigate the problem of learning the means of distributions that are associated with a finite n umber o f a rms under a monotonic constraint. Different from the traditional MAB, our problem involves a parameter constraint and a limited trial budget (i.e., the number of arm pulls is small). However, the number of training samples can be as large as possible through (infinite) simulations, while each training sample is of limited size. This situation arises when some additional information is provided before the trial starts and each arm pull (or testing) could be of extraordinary cost. For example, in cancer dose-finding clinical trials, higher toxicity probabilities are typically associated with higher dose levels (i.e., the monotonic dose–toxicity constraint), and the loss due to the drug’s toxicity, side-effects or death of patients can be enormous. We formulate this problem in the reinforcement learning (RL) paradigm, which is referred to as a bandit arm localization problem. We propose a novel approach in a double deep Q-learning framework, which is integrated with a state-of-the-art statistical model to preserve the parameter constraint and develop a more effective learning strategy. The double deep Q-learning model can be trained with a large (can be as large as infinite) number of simulated trials, which is the first time to cast dose finding in the RL framework. We evaluate the performance of our approach through extensive simulation studies in realistic settings of phase I clinical trials. The proposed double deep Q-learning is shown to outperform the baseline methods in cancer dose-finding trials. Wenbin Du, Huaqing Jin, Chao Yu 0004, Guosheng Yin |
IEEE Big Data | 4 |
| 2022 | Asymmetric Self-Supervised Graph Neural NetworksabstractAlthough self-supervised learning (SSL) has been successfully applied to graph data using graph neural networks (GNNs), most of the existing methods only consider undirected graphs where relationships among connected nodes are two-way symmetric (i.e., information can be passed back and forth between two connected nodes). However, there is a vast amount of applications where the information flow is asymmetric, leading to directed graphs where information can only be passed along one direction. For example, a directed edge indicates that the information can only be conveyed forwardly from the start node to the end node, but not backwardly. To accommodate such an asymmetric structure of directed graphs, we propose a simple yet remarkably effective SSL framework for directed graph analysis to incorporate such one-way information passing. We define an incoming embedding and an outgoing embedding for each node to model its schemes of sending and receiving features respectively. We propose an auxiliary SSL task to predict the existence of the directed edges with the incoming and outgoing embeddings of nodes. The auxiliary SSL task is jointly trained with a downstream primary task that updates nodes’ incoming features and outgoing features in accordance with labels. Extensive experiments on multiple real-world directed graph datasets demonstrate outstanding performances of the proposed self-supervised GNNs in both node-level and graph-level tasks. Zhuo Tan, Bin Liu 0022, Guosheng Yin |
IEEE Big Data | 3 |
| 2022 | The GR2D2 estimator for the precision matricesabstractBiological networks are important for the analysis of human diseases, which summarize the regulatory interactions and other relationships between different molecules. Understanding and constructing networks for molecules, such as DNA, RNA and proteins, can help elucidate the mechanisms of complex biological systems. The Gaussian Graphical Models (GGMs) are popular tools for the estimation of biological networks. Nonetheless, reconstructing GGMs from high-dimensional datasets is still challenging. The current methods cannot handle the sparsity and high-dimensionality issues arising from datasets very well. Here, we developed a new GGM, called the GR2D2 (Graphical $R^2$-induced Dirichlet Decomposition) model, based on the R2D2 priors for linear models. Besides, we provided a data-augmented block Gibbs sampler algorithm. The R code is available at https://github.com/RavenGan/GR2D2. The GR2D2 estimator shows superior performance in estimating the precision matrices compared with the existing techniques in various simulation settings. When the true precision matrix is sparse and of high dimension, the GR2D2 provides the estimates with smallest information divergence from the underlying truth. We also compare the GR2D2 estimator with the graphical horseshoe estimator in five cancer RNA-seq gene expression datasets grouped by three cancer types. Our results show that GR2D2 successfully identifies common cancer pathways and cancer-specific pathways for each dataset. Dailin Gan, Guosheng Yin, Yan Dora Zhang |
Briefings Bioinform. | 2 |
| 2021 | Crystallization Learning with the Delaunay TriangulationabstractBased on the Delaunay triangulation, we propose the crystallization learning to estimate the conditional expectation function in the framework of nonparametric regression. By conducting the crystallization search for the Delaunay simplices closest to the target point in a hierarchical way, the crystallization learning estimates the conditional expectation of the response by fitting a local linear model to the data points of the constructed Delaunay simplices. Instead of conducting the Delaunay triangulation for the entire feature space which would encounter enormous computational difficulty, our approach focuses only on the neighborhood of the target point and thus greatly expedites the estimation for high-dimensional cases. Because the volumes of Delaunay simplices are adaptive to the density of feature data points, our method selects neighbor data points uniformly in all directions and thus is more robust to the local geometric structure of the data than existing nonparametric regression methods. We develop the asymptotic properties of the crystallization learning and conduct numerical experiments on both synthetic and real data to demonstrate the advantages of our method in estimation of the conditional expectation function and prediction of the response. Jiaqi Gu 0005, Guosheng Yin |
ICML | 2 |
| 2021 | Beyond COVID-19 Diagnosis: Prognosis with Hierarchical Graph Representation Learning
Jinze Cui, Dailin Gan, Guosheng Yin |
MICCAI (7) | 4 |
| 2020 | Learning distributed sentence vectors with bi-directional 3D convolutionsabstractWe propose to learn distributed sentence representation using the text's visual features as input.Different from the existing methods that render the words (or characters) of a sentence into images separately, we fold these images into a 3-dimensional sentence tensor.Then, multiple 3dimensional convolutions with different lengths (the third dimension) are applied to the sentence tensor, which would act as bi-gram, tri-gram, quad-gram, and even five-gram detectors jointly.Similar to the Bi-LSTMs, these n-gram detectors learn both forward and backward distributional semantic knowledge from the sentence tensor.The proposed model uses bi-directional convolutions to learn text embedding according to the semantic order of words.The feature maps from the two directions are concatenated for final sentence embedding learning.Our model involves only a single layer of convolution which makes it easy and fast to train.We evaluate the sentence embeddings on several downstream natural language processing (NLP) tasks, which demonstrate surprisingly excellent performance of the proposed model. Bin Liu 0022, Guosheng Yin |
COLING | 3 |
| 2020 | Chinese Document Classification with Bi-directional Convolutional Language ModelabstractBy setting a typeface, each character of the Chinese text can be converted to a glyph pixel matrix. We propose to conduct text classification with such glyph features using bi-directional convolution. Although the pixel embedding can be applied to all languages, it is much more convenient to be used to represent Chinese scripts due to the square shape of Chinese characters. We extract both the forward and backward n-gram features of the text via bi-directional convolutional operations and then concatenate them. A subsequent 1-dimensional max-over-time pooling is applied to the bi-directional feature maps, and then three fully connected layers are used for conducting text classification. The proposed model has a light-weight architecture that only contains a single-layer convolutional neural network. Experiments on several Chinese text classification datasets demonstrate surprisingly excellent results for the training speed and superior performance of the proposed model in comparison with traditional methods. Bin Liu 0022, Guosheng Yin |
SIGIR | 2 |
| 2020 | Functional Martingale Residual Process for High-Dimensional Cox Regression with Model AveragingabstractRegularization methods for the Cox proportional hazards regression with high-dimensional survival data have been studied extensively in the literature. However, if the model is misspecified, this would result in misleading statistical inference and prediction. To enhance the prediction accuracy for the relative risk and the survival probability, we propose three model averaging approaches for the high-dimensional Cox proportional hazards regression. Based on the martingale residual process, we define the delete-one cross-validation (CV) process, and further propose three novel CV functionals, including the end-time CV, integrated CV, and supremum CV, to achieve more accurate prediction for the risk quantities of clinical interest. The optimal weights for candidate models, without the constraint of summing up to one, can be obtained by minimizing these functionals, respectively. The proposed model averaging approach can attain the lowest possible prediction loss asymptotically. Furthermore, we develop a greedy model averaging algorithm to overcome the computational obstacle when the dimension is high. The performances of the proposed model averaging procedures are evaluated via extensive simulation studies, demonstrating that our methods achieve superior prediction accuracy over the existing regularization methods. As an illustration, we apply the proposed methods to the mantle cell lymphoma study. Baihua He, Yuanshan Wu, Guosheng Yin |
J. Mach. Learn. Res. | 4 |
| 2019 | Fast Algorithm for Generalized Multinomial Models with Ranking DataabstractWe develop a framework of generalized multinomial models, which includes both the popular Plackett–Luce model and Bradley–Terry model as special cases. From a theoretical perspective, we prove that the maximum likelihood estimator (MLE) under generalized multinomial models corresponds to the stationary distribution of an inhomogeneous Markov chain uniquely. Based on this property, we propose an iterative algorithm that is easy to implement and interpret, and is guaranteed to converge. Numerical experiments on synthetic data and real data demonstrate the advantages of our Markov chain based algorithm over existing ones. Our algorithm converges to the MLE with fewer iterations and at a faster convergence rate. The new algorithm is readily applicable to problems such as page ranking or sports ranking data. Jiaqi Gu 0005, Guosheng Yin |
ICML | 2 |
| 2019 | Fast and Stable Maximum Likelihood Estimation for Incomplete Multinomial ModelsabstractWe propose a fixed-point iteration approach to the maximum likelihood estimation for the incomplete multinomial model, which provides a unified framework for ranking data analysis. Incomplete observations typically fall in a subset of categories, and thus cannot be distinguished as belonging to a unique category. We develop a minorization–maximization (MM) type of algorithm, which requires relatively fewer iterations and shorter time to achieve convergence. Under such a general framework, incomplete multinomial models can be reformulated to include several well-known ranking models as special cases, such as the Bradley–Terry, Plackett–Luce models and their variants. The simple form of iteratively updating equations in our algorithm involves only basic matrix operations, which makes it efficient and easy to implement with large data. Experimental results show that our algorithm runs faster than existing methods on synthetic data and real data. Guosheng Yin |
ICML | 2 |
| 2019 | Bayesian Model Selection Approach to Multiple Change-Points Detection with Non-Local Prior DistributionsabstractWe propose a Bayesian model selection (BMS) boundary detection procedure using non-local prior distributions for a sequence of data with multiple systematic mean changes. By using the non-local priors in the BMS framework, the BMS method can effectively suppress the non-boundary spike points with large instantaneous changes. Further, we speedup the algorithm by reducing the multiple change points to a series of single change point detection problems. We establish the consistency of the estimated number and locations of the change points under various prior distributions. From both theoretical and numerical perspectives, we show that the non-local inverse moment prior leads to the fastest convergence rate in identifying the true change points on the boundaries. Extensive simulation studies are conducted to compare the BMS with existing methods, and our method is illustrated with application to the magnetic resonance imaging guided radiation therapy data. Guosheng Yin, Francesca Dominici |
ACM Trans. Knowl. Discov. Data | 2 |
| 2018 | Bayesian Model Selection Approach to Boundary Detection with Non-Local PriorsabstractBased on non-local prior distributions, we propose a Bayesian model selection (BMS) procedure for boundary detection in a sequence of data with multiple systematic mean changes. The BMS method can effectively suppress the non-boundary spike points with large instantaneous changes. We speed up the algorithm by reducing the multiple change points to a series of single change point detection problems. We establish the consistency of the estimated number and locations of the change points under various prior distributions. Extensive simulation studies are conducted to compare the BMS with existing methods, and our approach is illustrated with application to the magnetic resonance imaging guided radiation therapy data. Guosheng Yin, Francesca Dominici |
NeurIPS | 2 |