Yan Han 0001

dblp:79/4311-1 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-7164-2295ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Graph learning · 29% Language models and text generation · 18% Image recognition and object detection · 16%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 56% Knowledge graphs · 44%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
1.322023
Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity Modeling · NeurIPS 2023
PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor · AAAI 2023
Machine learning › Graph learning
graph representation learning
1.322023
Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity Modeling · NeurIPS 2023
PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor · AAAI 2023
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation
1.012026
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data · ACL (1) 2026
Natural language and speech › Language models and text generation
behavior simulation
1.012026
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data · ACL (1) 2026
Machine learning › Graph learning › hypergraph learning
hypergraph neural network
0.712023
Vision HGNN: An Image is More than a Graph of Nodes · ICCV 2023
Computer vision › Image recognition and object detection
image classification
0.712023
Vision HGNN: An Image is More than a Graph of Nodes · ICCV 2023
Machine learning › Deep learning architectures and training
mixture of experts
0.712023
Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity Modeling · NeurIPS 2023
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.712023
PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor · AAAI 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Vision HGNN: An Image is More than a Graph of Nodes · ICCV 2023
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
performance prediction
0.712023
PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor · AAAI 2023
Knowledge graphs
link prediction
0.712023
Search Behavior Prediction: A Hypergraph Perspective · WSDM 2023
Information retrieval
query suggestion
0.712023
Search Behavior Prediction: A Hypergraph Perspective · WSDM 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.512021
SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata · ICDM 2021
Computer vision › Image recognition and object detection › medical image analysis
medical image classification
0.512021
SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata · ICDM 2021
Machine learning › Representation and self-supervised learning › contrastive learning
supervised contrastive learning
0.512021
SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata · ICDM 2021
Machine learning › Deep learning architectures and training
transformer
0.212023
PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor · AAAI 2023
Information retrieval
personalized search
0.212023
Search Behavior Prediction: A Hypergraph Perspective · WSDM 2023
Medical and health informatics
computer-aided diagnosis
0.112021
SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata · ICDM 2021

Methods — techniques the papers use, named apart from their topics

triplet attention · 1.0contrastive learning · 1.0Grad-CAM++ · 1.0transformer · 0.7permutation invariance · 0.7mixture of experts · 0.7laplacian positional encoding · 0.7hypergraph structure learning · 0.7hypergraph neural network · 0.7graph representation learning · 0.7graph augmentation · 0.7fuzzy c-means · 0.7
YearPublicationVenuePosition
2026 Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
abstract
Yuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao, Sisong Bei, Yaochen Xie, Yisi Sang, Qi He, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxuan Lu 0003, Yan Han 0001, Bingsheng Yao, Sisong Bei, Yaochen Xie, Yisi Sang, Qi He 0002, Dakuo Wang
ACL (1)3
2023 PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor
abstract
Time-consuming performance evaluation is the bottleneck of traditional Neural Architecture Search (NAS) methods. Predictor-based NAS can speed up performance evaluation by directly predicting performance, rather than training a large number of sub-models and then validating their performance. Most predictor-based NAS approaches use a proxy dataset to train model-based predictors efficiently but suffer from performance degradation and generalization problems. We attribute these problems to the poor abilities of existing predictors to character the sub-models' structure, specifically the topology information extraction and the node feature representation of the input graph data. To address these problems, we propose a Transformer-like NAS predictor PINAT, consisting of a Permutation INvariance Augmentation module serving as both token embedding layer and self-attention head, as well as a Laplacian matrix to be the positional encoding. Our design produces more representative features of the encoded architecture and outperforms state-of-the-art NAS predictors on six search spaces: NAS-Bench-101, NAS-Bench-201, DARTS, ProxylessNAS, PPI, and ModelNet. The code is available at https://github.com/ShunLu91/PINAT.
Shun Lu 0001, Yu Hu 0001, Peihao Wang, Yan Han 0001, Jianchao Tan, Sen Yang 0004, Ji Liu 0002
AAAI4
2023 Vision HGNN: An Image is More than a Graph of Nodes
abstract
The realm of graph-based modeling has proven its adaptability across diverse real-world data types. However, its applicability to general computer vision tasks had been limited until the introduction of the Vision Graph Neural Network (ViG). ViG divides input images into patches, conceptualized as nodes, constructing a graph through connections to nearest neighbors. Nonetheless, this method of graph construction confines itself to simple pairwise relationships, leading to surplus edges and unwarranted memory and computation expenses. In this paper, we enhance ViG by transcending conventional "pairwise" linkages and harnessing the power of the hypergraph to encapsulate image information. Our objective is to encompass more intricate inter-patch associations. In both training and inference phases, we adeptly establish and update the hypergraph structure using the Fuzzy C-Means method, ensuring minimal computational burden. This augmentation yields the Vision HyperGraph Neural Network (ViHGNN). The model’s efficacy is empirically substantiated through its state-of-the-art performance on both image classification and object detection tasks, courtesy of the hypergraph structure learning module that uncovers higher-order relationships. Our code is available at: https://github.com/VITA-Group/ViHGNN.
Yan Han 0001, Peihao Wang, Souvik Kundu 0009, Ying Ding 0001, Zhangyang Wang
ICCV1
2023 Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity Modeling
abstract
Graph neural networks (GNNs) have found extensive applications in learning from graph data. However, real-world graphs often possess diverse structures and comprise nodes and edges of varying types. To bolster the generalization capacity of GNNs, it has become customary to augment training graph structures through techniques like graph augmentations and large-scale pre-training on a wider array of graphs. Balancing this diversity while avoiding increased computational costs and the notorious trainability issues of GNNs is crucial. This study introduces the concept of Mixture-of-Experts (MoE) to GNNs, with the aim of augmenting their capacity to adapt to a diverse range of training graph structures, without incurring explosive computational overhead. The proposed Graph Mixture of Experts (GMoE) model empowers individual nodes in the graph to dynamically and adaptively select more general information aggregation experts. These experts are trained to capture distinct subgroups of graph structures and to incorporate information with varying hop sizes, where those with larger hop sizes specialize in gathering information over longer distances. The effectiveness of GMoE is validated through a series of experiments on a diverse set of tasks, including graph, node, and link prediction, using the OGB benchmark. Notably, it enhances ROC-AUC by $1.81\%$ in ogbg-molhiv and by $1.40\%$ in ogbg-molbbbp, when compared to the non-MoE baselines. Our code is publicly available at https://github.com/VITA-Group/Graph-Mixture-of-Experts.
Haotao Wang, Ziyu Jiang, Yuning You, Yan Han 0001, Gaowen Liu, Jayanth Srinivasa, Ramana Rao Kompella, Zhangyang Wang
NeurIPS4
2023 Search Behavior Prediction: A Hypergraph Perspective
abstract
At E-Commerce stores such as Amazon, eBay, and Taobao, the shopping items and the query words that customers use to search for the items form a bipartite graph that captures search behavior. Such a query-item graph can be used to forecast search trends or improve search results. For example, generating query-item associations, which is equivalent to predicting links in the bipartite graph, can yield recommendations that can customize and improve the user search experience. Although the bipartite shopping graphs are straightforward to model search behavior, they suffer from two challenges: 1) The majority of items are sporadically searched and hence have noisy/sparse query associations, leading to a long-tail distribution. 2) Infrequent queries are more likely to link to popular items, leading to another hurdle known as disassortative mixing.
Yan Han 0001, Edward W. Huang, Wenqing Zheng, Nikhil Rao 0001, Zhangyang Wang, Karthik Subbian
WSDM1
2023 Radiomics-Guided Global-Local Transformer for Weakly Supervised Pathology Localization in Chest X-Rays
abstract
Before the recent success of deep learning methods for automated medical image analysis, practitioners used handcrafted radiomic features to quantitatively describe local patches of medical images. However, extracting discriminative radiomic features relies on accurate pathology localization, which is difficult to acquire in real-world settings. Despite advances in disease classification and localization from chest X-rays, many approaches fail to incorporate clinically-informed domainspecific radiomic features. For these reasons, we propose a Radiomics-Guided Transformer (RGT) that fuses global image information with local radiomics-guided auxiliary information to provide accurate cardiopulmonary pathology localization and classification without any bounding box annotations. RGT consists of an image Transformer branch, a radiomics Transformer branch, and fusion layers that aggregate image and radiomics information. Using the learned self-attention of its image branch, RGT extracts a bounding box for which to compute radiomic features, which are further processed by the radiomics branch; learned image and radiomic features are then fused and mutually interact via cross-attention layers. Thus, RGT utilizes a novel end-to-end feedback loop that can bootstrap accurate pathology localization only using image-level disease labels. Experiments on the NIH ChestXRay dataset demonstrate that RGT outperforms prior works in weakly supervised disease localization (by an average margin of 3.6% over various intersection-over-union thresholds) and classification (by 1.1% in average area under the receiver operating characteristic curve). We publicly release our codes and pre-trained models at https://github.com/VITAGroup/chext.
Yan Han 0001, Gregory Holste, Ying Ding 0001, Ahmed H. Tewfik, Yifan Peng 0002, Zhangyang Wang
IEEE Trans. Medical Imaging1
2022 Knowledge-Augmented Contrastive Learning for Abnormality Classification and Localization in Chest X-rays with Radiomics using a Feedback Loop
abstract
Accurate classification and localization of abnormalities in chest X-rays play an important role in clinical diagnosis and treatment planning. Building a highly accurate predictive model for these tasks usually requires a large number of manually annotated labels and pixel regions (bounding boxes) of abnormalities. However, it is expensive to acquire such annotations, especially the bounding boxes. Recently, contrastive learning has shown strong promise in leveraging unlabeled natural images to produce highly generalizable and discriminative features. However, extending its power to the medical image domain is under-explored and highly non-trivial, since medical images are much less amendable to data augmentations. In contrast, their prior knowledge, as well as radiomic features, is often crucial. To bridge this gap, we propose an end-to-end semi-supervised knowledge-augmented contrastive learning framework, that simultaneously performs disease classification and localization tasks. The key knob of our framework is a unique positive sampling approach tailored for the medical images, by seamlessly integrating radiomic features as a knowledge augmentation. Specifically, we first apply an image encoder to classify the chest X-rays and to generate the image features. We next leverage Grad-CAM to highlight the crucial (abnormal) regions for chest X-rays (even when unannotated), from which we extract radiomic features. The radiomic features are then passed through another dedicated encoder to act as the positive sample for the image features generated from the same chest X-ray. In this way, our framework constitutes a feedback loop for image and radiomic features to mutually reinforce each other. Their contrasting yields knowledge-augmented representations that are both robust and interpretable. Extensive experiments on the NIH Chest X-ray dataset demonstrate that our approach outperforms existing baselines in both classification and localization tasks.
Yan Han 0001, Chongyan Chen, Ahmed H. Tewfik, Benjamin S. Glicksberg, Ying Ding 0001, Yifan Peng 0002, Zhangyang Wang
WACV1
2021 Using Radiomics as Prior Knowledge for Thorax Disease Classification and Localization in Chest X-rays
Yan Han 0001, Chongyan Chen, Liyan Tang, Mingquan Lin, Ajay Jaiswal, Song Wang 0026, Ahmed H. Tewfik, George Shih, Ying Ding 0001, Yifan Peng 0002
AMIA1
2021 SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
abstract
Computer-aided diagnosis plays a salient role in more accessible and accurate cardiopulmonary diseases classification and localization on chest radiography. Millions of people get affected and die due to these diseases without an accurate and timely diagnosis. Recently proposed contrastive learning heavily relies on data augmentation, especially positive data augmentation. However, generating clinically-accurate data augmentations for medical images is extremely difficult because the common data augmentation methods in computer vision, such as sharp, blur, and crop operations, can severely alter the clinical settings of medical images. In this paper, we proposed a novel and simple data augmentation method based on patient metadata and supervised knowledge to create clinically accurate positive and negative augmentations for chest X-rays. We introduce an end-to-end framework, SCALP, which extends the self-supervised contrastive approach to a supervised setting. Specifically, SCALP pulls together chest X-rays from the same patient (positive keys) and pushes apart chest X-rays from different patients (negative keys). In addition, it uses ResNet-50 along with the triplet-attention mechanism to identify cardiopulmonary diseases, and Grad-CAM++ to highlight the abnormal regions. Our extensive experiments demonstrate that SCALP outperforms existing baselines with significant margins in both classification and localization tasks. Specifically, the average classification AUCs improve from 82.8% (SOTA using DenseNet-121) to 83.9% (SCALP using ResNet-50), while the localization results improve on average by 3.7% over different IoU thresholds.
Ajay Jaiswal, Cyprian Zander, Yan Han 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001
ICDM4
2020 Speech Synthesis Using EEG
abstract
In this paper we demonstrate speech synthesis using different electroencephalography (EEG) feature sets recently introduced in [1]. We make use of a recurrent neural network (RNN) regression model to predict acoustic features directly from EEG features. We demonstrate our results using EEG features recorded in parallel with spoken speech as well as using EEG recorded in parallel with listening utterances. We provide EEG based speech synthesis results for four subjects in this paper and our results demonstrate the feasibility of synthesizing speech directly from EEG features.
Gautam Krishna, Co Tran, Yan Han 0001, Mason Carnahan, Ahmed H. Tewfik
ICASSP3