Yingheng Wang

dblp:265/6357 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-4828-4757ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MIRNet: Integrating Constrained Graph-Based Reasoning with Pre-training for Diagnostic Medical Imaging
abstract
Automated interpretation of medical images demands robust modeling of complex visual-semantic relationships while addressing annotation scarcity, label imbalance, and clinical plausibility constraints. We introduce MIRNet (Medical Image Reasoner Network), a novel framework that integrates self-supervised pre-training with constrained graph-based reasoning. Tongue image diagnosis is a particularly challenging domain that requires fine-grained visual and semantic understanding. Our approach leverages self-supervised masked autoencoder (MAE) to learn transferable visual representations from unlabeled data; employs graph attention networks (GAT) to model label correlations through expert-defined structured graphs; enforces clinical priors via constraint-aware optimization using KL divergence and regularization losses; and mitigates imbalance using asymmetric loss (ASL) and boosting ensembles. To address annotation scarcity, we also introduce TongueAtlas-4K, a comprehensive expert-curated benchmark comprising 4,000 images annotated with 22 diagnostic labels–representing the largest public dataset in tongue analysis. Validation shows our method achieves state-of-the-art performance. While optimized for tongue diagnosis, the framework readily generalizes to broader diagnostic medical imaging tasks.
Shufeng Kong, Nuan Cui, Yihan Meng, Yuanyuan Wei 0011, Feifan Chen, Yingheng Wang, Zhuo Cai 0004, Yuzheng Li, Zibin Zheng, Caihua Liu
AAAI8
2026 Unsupervised Combinatorial Probabilistic Reasoning: Probabilistic Coin Change Problem
abstract
We introduce the Probabilistic Coin Change Problem (PCCP), a novel variant of the classical Combination Coin Change Problem (CCCP), motivated by a real-world scientific inverse task. The goal of CCCP is to enumerate all unordered combinations of coin denominations that sum to a given target. In PCCP, each coin type’s value follows a discrete probability distribution, and the aggregate value of a combination of coins is thus stochastic. Given a set of such coin types and noisy observations of total sums, the task is to infer the most likely latent coin combination. To address the combinatorial and probabilistic complexity of PCCP, we propose DeepProReasoner (Deep Combinatorial Probabilistic Reasoning with Embedded Representations), an unsupervised, end-to-end, deep-learning framework that integrates combinatorial reasoning, latent-space modeling, and differentiable probabilistic reasoning. The model is trained using a reconstruction loss between the observed empirical distribution and a decoded probability mass function (PMF), enabling efficient gradient-based search over a continuous relaxation of the combinatorial space. We evaluate DeepProReasoner on two instances of PCCP: (1) a synthetic Candy Mix problem for ablation studies, and (2) a real-world task of molecular formula inference from ultrahigh resolution mass spectrometry (MS) data. Besides the two given instances, PCCP captures a wide range of inverse settings in biology, chemistry, environmental sciences, and medicine, where latent combinatorial structures give rise to noisy aggregate observations through stochastic processes. Our results show that DeepProReasoner achieves high accuracy and robustness, outperforming state-of-the-art methods.
Zhongdi Qu, Yingheng Wang, Utku Umur Acikalin, Aaron M. Ferber, Goncalo J. Gouveia, Brandon Bills, Joshua Kline, Sunandini Yedla, Frank C. Schroeder, Carla P. Gomes
AAAI2
2025 Denoising Diffusion Variational Inference: Diffusion Models as Expressive Variational Posteriors
abstract
We propose denoising diffusion variational inference (DDVI), a black-box variational inference algorithm for latent variable models which relies on diffusion models as flexible approximate posteriors. Specifically, our method introduces an expressive class of diffusion-based variational posteriors that perform iterative refinement in latent space; we train these posteriors with a novel regularized evidence lower bound (ELBO) on the marginal likelihood inspired by the wake-sleep algorithm. Our method is easy to implement (it fits a regularized extension of the ELBO), is compatible with black-box variational inference, and outperforms alternative classes of approximate posteriors based on normalizing flows or adversarial networks. We find that DDVI improves inference and learning in deep latent variable models across common benchmarks as well as on a motivating task in biology-inferring latent ancestry from human genomes-where it outperforms strong baselines on 1000 Genomes dataset.
Wasu Piriyakulkij, Yingheng Wang, Volodymyr Kuleshov
AAAI2
2024 Conformal Crystal Graph Transformer with Robust Encoding of Periodic Invariance
abstract
Machine learning techniques, especially in the realm of materials design, hold immense promise in predicting the properties of crystal materials and aiding in the discovery of novel crystals with desirable traits. However, crystals possess unique geometric constraints—namely, E(3) invariance for primitive cell and periodic invariance—which need to be accurately reflected in crystal representations. Though past research has explored various construction techniques to preserve periodic invariance in crystal representations, their robustness remains inadequate. Furthermore, effectively capturing angular information within 3D crystal structures continues to pose a significant challenge for graph-based approaches. This study introduces novel solutions to these challenges. We first present a graph construction method that robustly encodes periodic invariance and a strategy to capture angular information in neural networks without compromising efficiency. We further introduce CrystalFormer, a pioneering graph transformer architecture that emphasizes angle preservation and enhances long-range information. Through comprehensive evaluation, we verify our model's superior performance in 5 crystal prediction tasks, reaffirming the efficiency of our proposed methods.
Yingheng Wang, Shufeng Kong, John M. Gregoire, Carla P. Gomes
AAAI1
2024 GEM-RAG: Graphical Eigen Memories for Retrieval Augmented Generation
abstract
The ability to form, retrieve, and reason about memories in response to stimuli is central to general intelligence, enabling learning, adaptation, and insight. Large Language Models (LLMs), when given proper memories or context, can reason and respond effectively. However, they still struggle to optimally encode, store, and retrieve memories, a limitation that constrains their full potential as specialized AI agents. Retrieval Augmented Generation (RAG) seeks to address this by enriching LLMs with in-context examples. Inspired by human memory, we introduce Graphical Eigen Memories for Retrieval Augmented Generation (GEM-RAG), which tags information with LLMgenerated “utility” questions, links information together in a graph by similarity, and uses eigendecomposition to form higherlevel summary information. This approach not only enhances RAG tasks but also offers a way to explore text data sets. Using UnifiedQA, GPT-3.5 Turbo, SBERT, and OpenAI text encoders, we show that GEM-RAG outperforms state-of-the-art RAG methods on two standard QA tasks and discuss its implications for robust RAG systems.
Brendan Rappazzo, Yingheng Wang, Aaron M. Ferber, Carla P. Gomes
ICMLA2
2024 Growing Like a Tree: Finding Trunks From Graph Skeleton Trees
abstract
The message-passing paradigm has served as the foundation of graph neural networks (GNNs) for years, making them achieve great success in a wide range of applications. Despite its elegance, this paradigm presents several unexpected challenges for graph-level tasks, such as the long-range problem, information bottleneck, over-squashing phenomenon, and limited expressivity. In this study, we aim to overcome these major challenges and break the conventional "node- and edge-centric" mindset in graph-level tasks. To this end, we provide an in-depth theoretical analysis of the causes of the information bottleneck from the perspective of information influence. Building on the theoretical results, we offer unique insights to break this bottleneck and suggest extracting a skeleton tree from the original graph, followed by propagating information in a distinctive manner on this tree. Drawing inspiration from natural trees, we further propose to find trunks from graph skeleton trees to create powerful graph representations and develop the corresponding framework for graph-level tasks. Extensive experiments on multiple real-world datasets demonstrate the superiority of our model. Comprehensive experimental analyses further highlight its capability of capturing long-range dependencies and alleviating the over-squashing problem, thereby providing novel insights into graph-level tasks.
Zhongyu Huang, Yingheng Wang, Chaozhuo Li, Huiguang He
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Time Series Contrastive Learning with Information-Aware Augmentations
abstract
Various contrastive learning approaches have been proposed in recent years and achieve significant empirical success. While effective and prevalent, contrastive learning has been less explored for time series data. A key component of contrastive learning is to select appropriate augmentations imposing some priors to construct feasible positive samples, such that an encoder can be trained to learn robust and discriminative representations. Unlike image and language domains where "desired'' augmented samples can be generated with the rule of thumb guided by prefabricated human priors, the ad-hoc manual selection of time series augmentations is hindered by their diverse and human-unrecognizable temporal structures. How to find the desired augmentations of time series data that are meaningful for given contrastive learning tasks and datasets remains an open question. In this work, we address the problem by encouraging both high fidelity and variety based on information theory. A theoretical analysis leads to the criteria for selecting feasible data augmentations. On top of that, we propose a new contrastive learning approach with information-aware augmentations, InfoTS, that adaptively selects optimal augmentations for time series representation learning. Experiments on various datasets show highly competitive performance with up to a 12.0% reduction in MSE on forecasting tasks and up to 3.7% relative improvement in accuracy on classification tasks over the leading baselines.
Wei Cheng 0002, Yingheng Wang, Dongkuan Xu, Jingchao Ni, Wenchao Yu, Xuchao Zhang, Yanchi Liu, Yuncong Chen, Xiang Zhang 0001
AAAI3
2023 InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models
abstract
While diffusion models excel at generating high-quality samples, their latent variables typically lack semantic meaning and are not suitable for representation learning. Here, we propose InfoDiffusion, an algorithm that augments diffusion models with low-dimensional latent variables that capture high-level factors of variation in the data. InfoDiffusion relies on a learning objective regularized with the mutual information between observed and hidden variables, which improves latent space quality and prevents the latents from being ignored by expressive diffusion-based decoders. Empirically, we find that InfoDiffusion learns disentangled and human-interpretable latent representations that are competitive with state-of-the-art generative and contrastive methods, while retaining the high sample quality of diffusion models. Our method enables manipulating the attributes of generated images and has the potential to assist tasks that require exploring a learned latent space to generate quality samples, e.g., generative design.
Yingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan, Fei Wang 0001, Christopher De Sa, Volodymyr Kuleshov
ICML1
2023 M2Hub: Unlocking the Potential of Machine Learning for Materials Discovery
abstract
We introduce M$^2$Hub, a toolkit for advancing machine learning in materials discovery. Machine learning has achieved remarkable progress in modeling molecular structures, especially biomolecules for drug discovery. However, the development of machine learning approaches for modeling materials structures lag behind, which is partly due to the lack of an integrated platform that enables access to diverse tasks for materials discovery. To bridge this gap, M$^2$Hub will enable easy access to materials discovery tasks, datasets, machine learning methods, evaluations, and benchmark results that cover the entire workflow. Specifically, the first release of M$^2$Hub focuses on three key stages in materials discovery: virtual screening, inverse design, and molecular simulation, including 9 datasets that covers 6 types of materials with 56 tasks across 8 types of material properties. We further provide 2 synthetic datasets for the purpose of generative tasks on materials. In addition to random data splits, we also provide 3 additional data partitions to reflect the real-world materials discovery scenarios. State-of-the-art machine learning methods (including those are suitable for materials structures but never compared in the literature) are benchmarked on representative tasks. Our codes and library are publicly available at \url{https://github.com/yuanqidu/M2Hub}.
Yuanqi Du, Yingheng Wang, Yining Huang, Jianan Canal Li, Yanqiao Zhu 0001, Chenru Duan, John M. Gregoire, Carla P. Gomes
NeurIPS2
2023 Graph-Enhanced Emotion Neural Decoding
abstract
Brain signal-based emotion recognition has recently attracted considerable attention since it has powerful potential to be applied in human-computer interaction. To realize the emotional interaction of intelligent systems with humans, researchers have made efforts to decode human emotions from brain imaging data. The majority of current efforts use emotion similarities (e.g., emotion graphs) or brain region similarities (e.g., brain networks) to learn emotion and brain representations. However, the relationships between emotions and brain regions are not explicitly incorporated into the representation learning process. As a result, the learned representations may not be informative enough to benefit specific tasks, e.g., emotion decoding. In this work, we propose a novel idea of graph-enhanced emotion neural decoding, which takes advantage of a bipartite graph structure to integrate the relationships between emotions and brain regions into the neural decoding process, thus helping learn better representations. Theoretical analyses conclude that the suggested emotion-brain bipartite graph inherits and generalizes the conventional emotion graphs and brain networks. Comprehensive experiments on visually evoked emotion datasets demonstrate the effectiveness and superiority of our approach.
Zhongyu Huang, Changde Du, Yingheng Wang, Kaicheng Fu, Huiguang He
IEEE Trans. Medical Imaging3
2022 Going Deeper into Permutation-Sensitive Graph Neural Networks
abstract
The invariance to permutations of the adjacency matrix, i.e., graph isomorphism, is an overarching requirement for Graph Neural Networks (GNNs). Conventionally, this prerequisite can be satisfied by the invariant operations over node permutations when aggregating messages. However, such an invariant manner may ignore the relationships among neighboring nodes, thereby hindering the expressivity of GNNs. In this work, we devise an efficient permutation-sensitive aggregation mechanism via permutation groups, capturing pairwise correlations between neighboring nodes. We prove that our approach is strictly more powerful than the 2-dimensional Weisfeiler-Lehman (2-WL) graph isomorphism test and not less powerful than the 3-WL test. Moreover, we prove that our approach achieves the linear sampling complexity. Comprehensive experiments on multiple synthetic and real-world datasets demonstrate the superiority of our model.
Zhongyu Huang, Yingheng Wang, Chaozhuo Li, Huiguang He
ICML2
2022 Graph Emotion Decoding from Visually Evoked Neural Responses
Zhongyu Huang, Changde Du, Yingheng Wang, Huiguang He
MICCAI (8)3
2021 MolCloze: A Unified Cloze-style Self-supervised Molecular Structure Learning Model for Chemical Property Prediction
abstract
Machine Learning approaches are required to predict accurately on test samples that are distributionally different from training ones in the fields of drug discovery, computational biology, and cheminformatics. However, (i) labeled task-specific molecule data are often scarce, and (ii) poor generalization due to test molecules that are structurally different from those seen during training. To alleviate the problems, we propose a cloze-style self-supervised learning model (MolCloze) to obtain universal informative representations for molecular property prediction tasks. With carefully designed self-supervised tasks unifying generative- and discriminative-paradigm, MolCloze can learn rich structural and semantic information of molecules from enormous unlabelled molecular data. To capture such complex information, we design two novel strategies - Structural Fingerprint Tokenization (SFT) for better tokenizing molecule graphs, and Normalized Graph Raw Shortcut-connection (NGRS) for better latent representations by training a deeper model. We pretrain the MolCloze model via three tasks, which are Unordered Masked Language Modeling (UMLM), Replaced Masked Token Detection (RMTD), and Contrastive Energy-based Unmasked Token Clozing (CE-UTC). Then, we transfer the pre-trained model to a broad range of downstream molecular property prediction tasks via minor architecture modification. Extensive experiments demonstrate the generalizability of MolCloze by predicting a broad range of chemical properties which are related to drug discovery. We also observe significant performance boost on different downstream molecular property prediction datasets, achieving higher performance than the state-of-the-art baseline approaches and previous pre-training techniques developed for molecule data.
Yingheng Wang, Yaosen Min, Ji Wu 0002
BIBM1
2021 Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations
abstract
Learning generalizable, transferable, and robust representations for molecule data has always been a challenge. The recent success of contrastive learning (CL) for self-supervised graph representation learning provides a novel perspective to learn molecule representations. However, existing graph CL frameworks usually adopt stochastic augmentations or schemes according to pre-defined rules ont he input graph to obtain different graph views in various scales, which may destroy topological semantemes and domain prior in molecule data, leading to suboptimal performance. Therefore, a well-designed parameterized augmentation scheme that preserves chemically meaningful structural information and intrinsically essential attributes is crucial for molecular graph contrastive learning, helping to learn representations that are insensitive to perturbation on unimportant atoms and bonds. In this paper, we propose a novel method, Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations, that adaptively incorporates chemically significative information from both topological and semantic aspects of molecular graphs. Specifically, we apply deep neural networks to parameterize the augmentation process for both the molecular graph topology and atom attributes, to highlight contributive molecular substructures and recognize underlying chemical semantemes. Comprehensive experiments demonstrate that our method consistently outperforms compared baselines, verifying the effectiveness of the proposed framework. Our self-supervised model only uses one percent of the parameters to achieve comparative results against the state-of-the-art baseline, which has hundreds of millions of parameters. We also provide detailed case studies to validate the explainability of augmented views.
Yingheng Wang, Yaosen Min, Erzhuo Shao, Ji Wu 0002
BIBM1
2021 One-shot Transfer Learning for Population Mapping
abstract
Fine-grained population distribution data is of great importance for many applications, e.g., urban planning, traffic scheduling, epidemic modeling, and risk control. However, due to the limitations of data collection, including infrastructure density, user privacy, and business security, such fine-grained data is hard to collect and usually, only coarse-grained data is available. Thus, obtaining fine-grained population distribution from coarse-grained distribution becomes an important problem. To tackle this problem, existing methods mainly rely on sufficient fine-grained ground truth for training, which is not often available for the majority of cities. That limits the applications of these methods and brings the necessity to transfer knowledge between data-sufficient source cities to data-scarce target cities.
Erzhuo Shao, Jie Feng 0002, Yingheng Wang, Tong Xia, Yong Li 0008
CIKM3
2021 Multi-view Graph Contrastive Representation Learning for Drug-Drug Interaction Prediction
abstract
Potential Drug-Drug Interactions (DDI) occur while treating complex or co-existing diseases with drug combinations, which may cause changes in drugs’ pharmacological activity. Therefore, DDI prediction has been an important task in the medical health machine learning community. Graph-based learning methods have recently aroused widespread interest and are proved to be a priority for this task. However, these methods are often limited to exploiting the inter-view drug molecular structure and ignoring the drug’s intra-view interaction relationship, vital to capturing the complex DDI patterns. This study presents a new method, multi-view graph contrastive representation learning for drug-drug interaction prediction, MIRACLE for brevity, to capture inter-view molecule structure and intra-view interactions between molecules simultaneously. MIRACLE treats a DDI network as a multi-view graph where each node in the interaction graph itself is a drug molecular graph instance. We use GCN to encode DDI relationships and a bond-aware attentive message propagating method to capture drug molecular structure information in the MIRACLE learning stage. Also, we propose a novel unsupervised contrastive learning component to balance and integrate the multi-view information. Comprehensive experiments on multiple real datasets show that MIRACLE outperforms the state-of-the-art DDI prediction models consistently.
Yingheng Wang, Yaosen Min, Ji Wu 0002
WWW1