Hongjiao Guan

dblp:198/9484 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-4372-7416ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI Prediction
abstract
Enming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan, Yinglong Wang, Sibo Wei, Jianbin Guo, Ruifeng Xu, Wenpeng Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Enming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan, Yinglong Wang 0001, Sibo Wei, Ruifeng Xu 0001, Wenpeng Lu
ACL (1)4
2025 AiCoder: Exploring Automated ICD Coding on Chinese EMRs with a Multi-Agent Framework
abstract
The International Classification of Diseases (ICD) coding is a crucial component for standardizing medical information, and automated coding represents an important direction for improving the efficiency and accuracy of this process. However, current automated ICD coding research still faces significant challenges. First, existing studies are predominantly based on English contexts, making their models difficult to directly apply to China's localized ICD versions, while the immense pressure on its healthcare system makes this research particularly urgent in China. In addition, mainstream approaches oversimplify coding as multi-label classification, thus ignoring the coder workflow and failing to generalize to unseen codes or uncover implicit diagnoses. To address these challenges, we propose AiCoder, a novel multi-agent framework that simulates real-world coding workflows through a two-stage design. The first stage focuses on accuracy by parsing discharge diagnoses, correlating medical record evidence, and utilizing knowledge graphs for standardization; the second stage emphasizes completeness by leveraging ICD hierarchical structures and integrating clinical reasoning patterns to identify potentially missed codes. Furthermore, we construct a publicly available high-quality Chinese ICD annotation dataset, providing a valuable resource for the research community. Comprehensive experiments including ablation study and case study validate our framework's effectiveness. We also systematically analyze common coding errors of LLMs in Chinese ICD coding tasks, hoping to provide reference for future research.
Zhenpeng Liang, Hongjiao Guan, Weiyu Zhang 0001, Ying Lian, Wenpeng Lu
BIBM2
2025 HiRes: Hierarchical Feature Optimization and Rescorer for Automatic ICD Coding
abstract
The International Classification of Diseases (ICD) coding assigns standardized codes to diseases. Automating this process enhances the efficiency and accuracy of clinical records processing. However, current methods struggle with noisy and lengthy clinical texts, making it difficult to ensure the reliability of feature extraction. Furthermore, they typically make separate binary predictions for each code, overlooking the dependencies between them. To address these issues, we propose a novel model called Hierarchical Feature Optimization and Rescorer (HiRes). We employ a cascaded convolution architecture to mitigate noise and enhance feature representation. Additionally, a masked autoencoder-based rescorer is introduced to capture ICD code interdependencies, refining initial predictions for improved accuracy. The model also considers the inconsistencies in code representations. Experiments on the MIMIC datasets demonstrate the effectiveness of our model.
Zhenpeng Liang, Hongjiao Guan, Wenpeng Lu, Xueping Peng, Muyun Yang
ICASSP2
2025 From Feature Alignment to Multimodal Fusion: A Two-Stage Primary Modality-Guided Approach for MSA
abstract
Multimodal Sentiment Analysis (MSA) aims to leverage heterogeneous data—typically language, vision, and acoustic modalities—to accurately interpret human emotional states. Despite recent advances, challenges persist due to the feature distribution difference caused by intrinsic modality heterogeneity. Prior works either neglect the contribution disparity among modalities, especially the dominant role of language in sentiment reasoning, or emphasize language dominance in fusion-space alignment, ignoring coordination in the early feature space. To address these limitations, we propose a novel Two-Stage Primary Modality-Guided (TSPMG) framework, which introduces primary-modality supervision into both feature-space distribution alignment and fusion-space attention modulation. This dual-level cooperative mechanism progressively amplifies the dominant modality’s influence throughout the entire representation learning pipeline. Extensive experiments on two benchmark datasets demonstrate that TSPMG achieves superior or comparable results to state-of-the-art baselines, with ablation studies further validating the effectiveness of primary-modality-guided strategies for robust and interpretable multimodal sentiment analysis. The code is available at https://github.com/Kaisa777/TSPMG.
Xiaoqiang Ren, Hongjiao Guan
MMAsia4
2025 Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought
Zhe Tao, Muyun Yang, Hongjiao Guan, Wenpeng Lu, Hailong Cao, Conghui Zhu, Tiejun Zhao
NLPCC (4)4
2025 MedConMA: A Confidence-Driven Multi-agent Framework for Medical Q&A
Rui Wang 0043, Yonghe Chen, Weiyu Zhang 0001, Jiasheng Si, Hongjiao Guan, Xueping Peng, Wenpeng Lu
PAKDD (3)5
2024 Medical Entity Disambiguation with Medical Mention Relation and Fine-grained Entity Knowledge
abstract
Medical entity disambiguation (MED) plays a crucial role in natural language processing and biomedical domains, which is the task of mapping ambiguous medical mentions to structured candidate medical entities from knowledge bases (KBs). However, existing methods for MED often fail to fully utilize the knowledge within medical KBs and overlook essential interactions between medical mentions and candidate entities, resulting in knowledge- and interaction-inefficient modeling and suboptimal disambiguation performance. To address these limitations, this paper proposes a novel approach, MED with Medical Mention Relation and Fine-grained Entity Knowledge (MMR-FEK). Specifically, MMR-FEK incorporates a mention relation fusion module and an entity knowledge fusion module, followed by an interaction module. The former employs a relation graph convolutional network to fuse mention relation information between medical mentions to enhance mention representations, while the latter leverages an attention mechanism to fuse synonym and type information of candidate entities to enhance entity representations. Afterwards, an interaction module is designed to employ a bidirectional attention mechanism to capture interactions between mentions and entities to generate the matching representation. Extensive experiments on two publicly available real-world datasets demonstrate MMR-FEK’s superiority over state-of-the-art(SOTA) MED baselines across all metrics. Our source code is publicly available.
Wenpeng Lu, Guobiao Zhang, Xueping Peng, Hongjiao Guan, Shoujin Wang
LREC/COLING4
2024 Multi-view Contrastive Learning for Medical Question Summarization
abstract
Most Seq2Seq neural model-based medical question summarization (MQS) systems have a severe mismatch between training and inference, i.e., exposure bias. However, this problem remains unexplored in the MQS task. To bridge this research gap and alleviate the problem of exposure bias, we propose a novel re-ranking training framework for MQS called Multi-view Contrastive Learning (MvCL). MvCL simultaneously considers the similarity scores between medical questions and candidate summaries as well as the average similarity scores between candidate summaries and other candidates within the same group, and utilizes contrastive learning to optimize the model’s ranking ability. Additionally, we propose a new multilevel inference approach to adapt to this training strategy. The approach first filters out candidate summaries that are dissimilar to the original medical question, and then selects the summary with the highest average similarity to other candidate summaries from the remaining candidates as the final output. We conducted extensive experiments, and the results demonstrate that our proposed MvCL framework achieves state-of-the-art results on the majority of evaluation metrics across four datasets.1
Sibo Wei, Xueping Peng, Hongjiao Guan, Lina Geng, Ping Jian, Hao Wu 0066, Wenpeng Lu
CSCWD3
2024 Filter-Enhanced Hypergraph Transformer for Multi-Behavior Sequential Recommendation
abstract
Sequential recommendation has been developed to predict the next item in which users are most interested by capturing user behavior patterns embedded in their historical interaction sequences. However, most existing methods appear to exhibit limitations in modeling fine-grained dependencies embedded in users’ various periodic behavior patterns and heterogeneous dependencies across multi-behaviors. Towards this end, we propose a Filter-enhanced Hypergraph Transformer framework for Multi-Behavior Sequential Recommendation (FHT-MB) to address the above challenges. Specifically, a multi-scale filter layer equipped with multi-learnable filters is devised to encode behavior-aware sequential patterns emerging from different periodic trends (e.g., daily or weekly routines), and then a hypergraph structure is devised to extract heterogeneous dependencies across users’ multiple types of behaviors. Extensive experiments on two real-world e-commerce datasets show the superiority of our proposed FHT-MB over various state-of-the-art methods.1
Zhufeng Shao, Shoujin Wang, Wenpeng Lu, Weiyu Zhang 0001, Hongjiao Guan, Long Zhao 0002
ICASSP5
2024 SN-RNSP: Mining self-adaptive nonoverlapping repetitive negative sequential patterns in transaction sequences
Chuanhou Sun, Yongshun Gong, Ying Guo 0030, Long Zhao 0002, Hongjiao Guan, Xinwang Liu 0002, Xiangjun Dong 0001
Knowl. Based Syst.5
2023 Ga-RFR: Recurrent Feature Reasoning with Gated Convolution for Chinese Inscriptions Image Inpainting
Long Zhao 0002, Yuhao Lou, Zonglong Yuan, Xiangjun Dong 0001, Xiaoqiang Ren, Hongjiao Guan
ICANN (2)6
2023 Fusion of Dynamic Hypergraph and Clinical Event for Sequential Diagnosis Prediction
abstract
Sequential diagnosis prediction (SDP) is a challenging task, aiming to predict patients’ future diagnoses based on their historical medical records. While methods based on graph neural networks (GNNs) have proven successful for this task, they typically focus on modeling pairwise diseases using a global disease combination graph. However, these approaches neglect the fine-grained higher-order relations among persistent and emerging diseases within a single visit, which may contain crucial clues to predict the next diagnosis. Additionally, they fail to fully leverage patient-related clinical information present in electronic health records (EHRs). To address these challenges, this paper proposes a novel approach called the fusion of Dynamic Hypergraph and Clinical Event (DHCE) for sequential diagnosis prediction. The proposed method aims to exploit the fine-grained higher-order relations among diagnoses within a visit and leverage clinical event information from EHRs to improve the accuracy of predicting the next diagnosis. Specifically, DHCE categorizes diagnoses within a single visit in a fine-grained granularity into persistent and emerging categories based on a patient’s historical diagnoses. It then constructs dynamic hypergraphs to capture higher-order disease relations within each visit. Next, we design a transition function to extract the transitional context from previous visits in order to generate the visit representation. Furthermore, to fully leverage patient-related clinical events in a visit, we utilize Bio-Clinical BERT to encode them and generate the clinical event representation for each visit. Finally, we combine the visit representation and event representation to generate a comprehensive patient representation, which is then used to predict the patient’s next diagnosis. Experimental results on two benchmark datasets consistently demonstrate that DHCE outperforms state-of-the-art methods1.
Xueping Peng, Hongjiao Guan, Long Zhao 0002, Xinxiao Qiao, Wenpeng Lu
ICPADS3
2023 Extended natural neighborhood for SMOTE and its variants in imbalanced classification
Hongjiao Guan, Long Zhao 0002, Xiangjun Dong 0001
Eng. Appl. Artif. Intell.1
2023 Dual objective bounded abstaining model to control performance for safety-critical applications
Hongjiao Guan, Xiangjun Dong 0001, Long Zhao 0002, Xiaoqiang Ren
Eng. Appl. Artif. Intell.1
2022 Dual Self-Paced SMOTE for Imbalanced Data
abstract
Imbalanced classification has always been a challenging issue. The minority class usually has degraded recognition rate. The key factors are sample scarcity of the minority class and intrinsic complex distribution characteristics in imbalanced data. SMOTE and its extensions are commonly used data-level methods. These SMOTE-related methods oversample all or unsafe minority class examples, without considering sample difference or only considering sample difference in spatial distribution. In this paper, we propose a dual self-paced SMOTE (DSP-SMOTE) method, which considers temporal-spatial distribution of samples. DSP-SMOTE adopts a two-way self-paced mechanism at different stages of sampling. The majority class is undersampled according to easy example first principle, whereas the minority class is oversampled following difficult sample first principle. In DSP-SMOTE, a hybrid (including static and dynamic) difficulty is introduced to estimate easy or difficult examples. Furthermore, the trade-off between static and dynamic difficulties is adjusted adaptively according to sample distribution and training status. Experimental results demonstrate that DSP-SMOTE outperforms previous SMOTE-related methods significantly in terms of AUC and sensitivity.
Yangguang Shao, Hongjiao Guan
ICPR3
2021 ExNN-SMOTE: Extended Natural Neighbors Based SMOTE to Deal with Imbalanced Data
abstract
Many practical applications suffer from the problem of imbalanced classification. The minority class has poor classification performance; on the other hand, its misclassification cost is high. One reason for classification difficulty is the intrinsic complicated distribution characteristics (CDCs) in imbalanced data itself. Classical oversampling method SMOTE generates synthetic minority class examples between neighbors, which is parameter dependent. Furthermore, due to blindness of neighbor selection, SMOTE suffers from overgeneralization in the minority class. To solve such problems, we propose an oversampling method, called extended natural neighbors based SMOTE (ExNN-SMOTE). In ExNN-SMOTE, neighbors are determined adaptively by capturing data distribution characteristics. Extensive experiments over synthetic and real datasets demonstrate the effectiveness of ExNN-SMOTE dealing with CDCs and the superiority of ExNN-SMOTE over other SMOTE-related methods.
Hongjiao Guan, Bin Ma 0003, Yingtao Zhang, Xianglong Tang
ACML1
2021 A Generalized Optimization Embedded Framework of Undersampling Ensembles for Imbalanced Classification
abstract
Imbalanced classification exists commonly in practical applications, and it has always been a challenging issue. Traditional classification methods have poor performance on imbalanced data, especially, on the minority class. However, the minority class is usually of our interest, and its misclassification cost is higher. The critical factor is the intrinsic complicated distribution characteristics in imbalanced data itself. Resampling ensemble learning achieves promising results and is a research focus recently. However, some resampling ensembles do not consider complicated distribution characteristics, thus limiting the performance improvement. In this paper, a generalized optimization embedded framework (GOEF) is proposed based on undersampling bagging. The GOEF aims to pay more attention to the learning of local regions to handle the complicated distribution characteristics. Specifically, the GOEF utilizes out-of-bag data to explore heterogeneous local areas and chooses misclassified examples to optimize base classifiers. The optimization can focus on a single class or both classes. Extensive experiments over synthetic and real datasets demonstrate that GOEF with the minority class optimization performs the best in terms of AUC, G-mean, and sensitivity, compared with five resampling ensemble methods.
Hongjiao Guan, Yingtao Zhang, Bin Ma 0003, Jian Li 0034, Chunpeng Wang 0001
DSAA1
2021 SMOTE-WENN: Solving class imbalance and small sample problems by oversampling and distance scaling
Hongjiao Guan, Yingtao Zhang, Min Xian, Heng-Da Cheng, Xianglong Tang
Appl. Intell.1
2019 BA2Cs: Bounded abstaining with two constraints of reject rates in binary classification
Hongjiao Guan, Yingtao Zhang, Heng-Da Cheng, Min Xian, Xianglong Tang
Neurocomputing1
2016 WENN for individualized cleaning in imbalanced data
abstract
This paper proposes individualized cleaning for diverse imbalanced data sets. Existing techniques for data cleaning have difficulties with rare cases and outliers in minority class, especially, in highly unbalanced data. The drawback leads incomplete and imprecise examples to removal. In order to enhance the robustness and perform thorough data cleaning, we propose a weighted edited nearest neighbor (WENN), which detects and removes noisy examples from both classes intelligently. It considers individual characteristics of each imbalanced data, involving global class imbalance and local distribution. The main idea of the proposed method is to carefully put more focus on the majority class than the minority class during data cleaning. Extensive experiments over synthetic and real data clearly validate the superiority of our approach against other data cleaning methods.
Hongjiao Guan, Yingtao Zhang, Min Xian, Heng-Da Cheng, Xianglong Tang
ICPR1