EDBT 2026 Demo / reviewers in the wild / expert
Xinpan Yuan
dblp:70/10857
· DBLP profile ↗
34ranked-venue papers
18as first author
33since 2021 · last 2026
0000-0001-9509-0755ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, this assumption may not be able to hold for the cross-modal matching due to weak or even corrupted correlations between textual descriptions and visual content, referred to as noisy correspondence (NC). Such NC largely disrupts the correspondence learning between visual and semantic modalities. Though prior works have improved single-modal robustness against noisy labels, systematic modeling of both cross-modal and intra-modal geometric structures in TBPR remains limited attention. In this paper, we propose Geometric Structure Consistency Alignment (GSCA) to TBPR, which leverages cross-modal cosine similarity and intra-modal nearest-neighbor affinity to learn visual-semantic consistency under noisy correspondence. To mitigate the structural corruption caused by noisy pairs, we introduce the Structure Refinement and Mining (SRAM) module. By partitioning training data into clean, ambiguous, and noisy subsets, SRAM enables the model to strategically refine the cross-modal correspondence by mining reliable pairs, thus enhancing the reliability of positive or negative samples discrimination and preserving structural consistency across modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets. On CUHK-PEDES, it boosts Rank-1 by 1.42% in noise-free conditions, sustaining a robust 74.25% Rank-1 under a 50% noise ratio. Xinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang 0001, Guihu Zhao, Lin Wu 0001 |
AAAI | 1 |
| 2026 | Syntax-Aware Dependency Parsing for Dual-Origin Noisy Correspondence in Text-Based Person Search
Xinpan Yuan, Wenguang Gan, Shaomin Xie, Chengyuan Zhang 0001, Liujie Hua |
DASFAA (1) | 1 |
| 2026 | Set-to-One Structured Captioning for Heterogeneous Nasal Image Collections via Spatio-Temporal Modeling and Clinical Knowledge Integration
Xinpan Yuan, Jianuo Ju, Liujie Hua, Mingzhu Huang |
DASFAA (3) | 1 |
| 2026 | Bidirectional Conditional Diffusion of Cross-modal Feature Enhancement for Text-based Person Search
Xinpan Yuan, Xingyu Jin |
ICIC (8) | 4 |
| 2026 | Graph-Enhanced Dynamic Interactive Cross-Modal Learning for Text-Based Person Search
Xinpan Yuan |
ICIC (21) | 4 |
| 2026 | MSCAF: Multi-scale convolutional and adaptive fusion cloud workload forecasting model based on iTransformer
Qiang Liu 0032, Buqing Cao, Xinpan Yuan |
J. Syst. Softw. | 5 |
| 2026 | Multi-scene topic-aware for novel single continuous shot multiple scenes endoscopy report generation
Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao, Siming Jin |
Knowl. Based Syst. | 1 |
| 2026 | A Rolling Bearing Fault Diagnosis Model Integrating Adaptive Distribution-Aware Discriminative Loss FunctionabstractIn industrial scenarios, noise interference and feature overlap often result in blurred classification boundaries, compromising the reliability of rolling bearing fault diagnosis. An adaptive distribution-aware discriminative loss (ADADL) is introduced, through which intraclass thresholds are dynamically adjusted and interclass boundaries are optimized, thereby enhancing compactness and separability in the feature space. By integrating it with the cross-entropy loss, ADADL yields marked gains in diagnostic accuracy on both the Case Western Reserve University benchmark and real-world datasets, particularly under conditions of class imbalance and high noise. Visualization analyses further confirm its ability to sharpen clustering boundaries, suppress feature overlap, and effectively mitigate blurred decision regions. Cheng Peng 0015, Xin Liu 0172, Weihua Gui 0001, Zhaohui Tang 0004, Longxin Zhang, Xinpan Yuan |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | HCSeer2: A Deep Learning-Based Multi-Scale Modeling Framework for Predicting Cold and Hot Spots of Variants in the Human ExomeabstractAccurately identifying variant hot and cold spots in human exonic regions is a crucial step in applying the PM1 and benign interpretation criteria of the ACMG-AMP guidelines. Addressing the issue that the existing tool HCSeeker relies heavily on the ClinVar database and has limited predictive performance in low-frequency variant regions, this study proposes HCSeer2-an innovative deep learning framework. By integrating the local feature extraction capability of CNN with the global dependency modeling of the Self-Attention mechanism, HCSeer2 achieves high-precision prediction of variant cold and hot spots across the entire human exome. HCSeer2 is trained on existing cold and hot spot data computed by tools such as HCSeeker. Through a multimodal feature fusion architecture, it simultaneously integrates sequence information and four types of genomic prior knowledge, effectively capturing the local aggregation patterns and global distribution rules of genomic variants. It enhances the understanding of variant clustering mechanisms in a self-supervised manner to predict variant cold and hot spots. Using this model, we identified 28,907 variant hot spots and 159,970 variant cold spots in the entire human exome, covering all human exonic regions. We also verified that the pathogenic potential of sites in hot spots is significantly higher than that in cold spots. This study not only provides scalable variant annotation resources for the clinical application of the ACMG-AMP guidelines but also offers new insights into the identification of genomic functional elements through the proposed multi-scale modeling approach. The code and data of HCSeer2 are freely available at https://github.com/xq-xia/HCSeer2. Xingquan Xia, Guihu Zhao, Jinchen Li, Xinpan Yuan |
BIBM | 4 |
| 2025 | VRSegNet: Visual-Relation-Guided Segmentation of Clear Nasal Discharge Under Anatomical Constraints for Rhinitis AssessmentabstractAllergic rhinitis (AR) affects hundreds of millions globally and presents diverse, recurring symptoms that impair quality of life. Despite the widespread clinical use of nasal endoscopy, current evaluations remain largely subjective, lacking structured image-based assessment. Clear nasal discharge (CND), as a key visual indicator of AR, reflects inflammation and provides critical cues for grading and subtyping through its anatomical distribution and morphology. However, CND often presents with high visual heterogeneity and irregular anatomical distribution in nasal endoscopic images, making its segmentation particularly challenging—especially in the presence of optical artifacts, blurred anatomical boundaries, and structural diversity. To address these challenges, we propose VRSegNet, a segmentation framework that integrates anatomical structure information with inter-regional visual relations, aiming to enhance segmentation robustness and accuracy in complex scenarios. To accommodate structural and morphological variability, VRSegNet decomposes the nasal cavity into three functional zones: the Central Lumen Zone (CLZ), Nasal Groove Zone (NGZ), and Turbinate Sidewall Zone (TSZ). Based on this anatomical prior, the framework further incorporates a Regional Dynamic Perception (RDP) module and Visual-Relation Contrastive Learning (VR-CL) to improve the model's ability to identify spatially fragmented fluid anomalies within the same region, particularly under challenging conditions such as boundary ambiguity and partial occlusion. Evaluated per subset on RhinoMucus, VRSegNet attains mIoU of$\mathbf{7 3. 3 8 \%}$on MucusC (CLZ),$\mathbf{5 8. 8 7 \%}$on MucusG (NGZ), and 50.57 % on MucusW (TSZ), confirming robust performance across heterogeneous anatomical zones. Xinpan Yuan, Mingzhu Huang |
BIBM | 1 |
| 2025 | PBD: A Manually Curated Full-Chain Benchmark Dataset for Evaluating LLMs on ACMG PS3/BS3 Functional Evidence AcquisitionabstractThe PS3 (Pathogenic Functional Evidence) and BS3 (Benign Functional Evidence) criteria in the American College of Medical Genetics and Genomics (ACMG) guidelines are critical for genetic variant classification. However, manual evaluation is time-consuming and prone to inter-laboratory inconsistencies, limiting the clinical interpretation of Variants of Uncertain Significance (VUS). Large Language Models (LLMs) offer potential for automated assessment, but their performance validation is hindered by the lack of standardized, high-quality datasets. This study introduces the PS3/BS3 Full-Chain Evidence Benchmark Dataset (PBD), the first manually curated dataset comprising 77 peer-reviewed publications, covering 266 cDNA variants (including duplicates) and their PS3/BS3 rating results. Adhering to ClinGen Sequence Variant Interpretation (SVI) standards, PBD includes structured, comprehensive evidence chains spanning genes, diseases, variants, experiments, and evidence ratings, designed to evaluate LLM capabilities in functional evidence extraction. We detail the dataset construction process, including literature screening, data extraction, standardization, and quality control, and developed a Python-based automated evaluation pipeline for reproducible, standardized analysis. Experiments using DeepSeek models$(1.5 \mathrm{b} / 7 \mathrm{b} / 14 \mathrm{b})$demonstrate PBD's potential in supporting automated variant interpretation. PBD provides a vital resource for bioinformatics and precision medicine, facilitating the development and standardization of variant classification tools. Data examples are available at https://github.com/User8588/PBD, with full data and code to be released upon paper acceptance. Xinpan Yuan, Bozhao Li, Chenbin Liu, Xinxue Li, Liujie Hua, Jinchen Li, Lin Wu 0001, Guihu Zhao |
BIBM | 1 |
| 2025 | MMAG: Multimodal Learning for Mucus Anomaly Grading in Nasal Endoscopy via Semantic Attribute PromptingabstractAccurate grading of rhinitis severity in nasal endoscopy relies heavily on the characterization of key secretion types, notably clear nasal discharge (CND) and purulent nasal secretion (PUS).However, both exhibit ambiguous appearance and high structural variability, posing challenges to automated grading under weak supervision.To address this, we propose Multimodal Learning for Mucus Anomaly Grading (MMAG), which integrates structured prompts with rank-aware vision-language modeling for joint detection and grading.Attribute prompts are constructed from clinical descriptors (e.g., secretion type, severity, location) and aligned with multi-level visual features via a dual-branch encoder.During inference, the model localizes mucus anomalies and maps the input image to severity-specific prompts (e.g., "moderate pus"), projecting them into a rankaware feature space for progressive similarity scoring.Extensive evaluations on CND and PUS datasets show that our method achieves consistent gains over Baseline, improving AUC by 6.31% and 4.79%, and F1 score by 12.85% and 6.03%, respectively.This framework enables interpretable, annotation-efficient, and semantically grounded assessment of rhinitis severity based on mucus anomalies. Xinpan Yuan, Mingzhu Huang, Liujie Hua, Jianuo Ju |
EMNLP | 1 |
| 2025 | OF-AR Relation Aware Representation Learning for Lesion Image Segmentation and GradingabstractMedical image segmentation provides important supplementary information for lesion grading, but existing segmentation models usually only focus on the lesion area, which is susceptible to the influence of shooting distance and angle, leading to feature extraction errors. We found an "as one falls, another rises"(OF-AR) relationship between the lesion and the surrounding non lesion areas, and introduced OF-AR relationship aware learning representation to jointly extract features of lesions and non lesions. By dividing the image into black and white dual zones, expanding the feature extraction area, and using a dual zone contrastive learning module to increase the inter-class distance, accurate grading is ensured by utilizing the relative information of two complete targets. In addition, using weighted fusion methods to enhance the comprehensiveness and objectivity of grading. The experiment used adenoids as an example to verify the high accuracy of this method. Xinpan Yuan, Siming Jin, Liujie Hua, Guihu Zhao |
ICASSP | 1 |
| 2025 | A Novel Single Continuous Shot Multiple Lesions Endoscopy Report GenerationabstractAutomatic Report Generation(ARG), which aims to automatically provide observations on images, is challenged by the lack of coherence between multiple scenes and precise description of multiple lesions. In order to explore the task of multi-scene multi-lesion report generation(MSMLRG) in one shot, in this paper, we introduce a multi-scene multi-lesion report generation framework to extract the scene-report alignment relation and scene-topic relation. Specifically, we design a scene-report feature aligner to achieve fine-grained alignment of different lesion in different scenes in images and reports, and incorporate a topic-aware module to help generate a topic text vocabulary for different scenes. Our framework has been successfully experimented on several automatic report generation models, and performs well on automatic evaluation metrics. The framework for one-shot report generation during multi-scene not only fills the gap of multi-scene image report generation, but also effectively improves the accuracy and consistency of diagnostic reports. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
ICASSP | 1 |
| 2025 | TWRLR: Composed Image Retrieval Based via Two-Way Reciprocity Learning and Reasoning
Qiang Liu 0032, Xinpan Yuan, Mengxi Ying, Changyuan Zhang, Wenguang Gan |
ICIC (14) | 3 |
| 2025 | Prompting Large Models for Knowledge and Reasoning Augmentation in KB-VQA
Qiang Liu 0032, Mengxi Ying, Gan Li, Xinpan Yuan |
ICIC (23) | 5 |
| 2025 | VAMP: Visual Attribute-Guided Multi-dimensional Perception for Nasal Endoscopy Report Generation
Xinpan Yuan, Jianuo Ju, Liujie Hua, Mingzhu Huang, Shaomin Xie, Wenguang Gan |
ICIC (5) | 1 |
| 2025 | MAS-ZSAS: A Zero-Shot Anomaly Segmentation Framework with Multi-attribute Guided Text Prompts
Xinpan Yuan, Guorong Liang, Liujie Hua, Shaomin Xie, Wenguang Gan |
ICIC (6) | 1 |
| 2025 | Advanced Font-Aware Document Hierarchy Reconstruction for Enhanced Structured Parsing
Xinpan Yuan, Gan Li, Liujie Hua, Guihu Zhao, Shaomin Xie |
ICIC (16) | 1 |
| 2025 | TPS-SG: A Semantic-Generalization Model with Lightweight Mix-Augmentation for Text-Based Person Search
Xinpan Yuan, Yuxiang Luo, Qiang Liu 0032, Gan Li |
ICIC (12) | 1 |
| 2025 | CLIO: A Unified Framework for Consistency-Aware Learning and Intra-Modal Optimization in Text-Based Person Re-identification
Xinpan Yuan, Shaomin Xie, Guihu Zhao, Liujie Hua, Wenguang Gan |
ICIC (5) | 1 |
| 2025 | GCA-Net: Global Contextual Attention Network with Lightweight Hierarchical Alignment for Text-Guided Fashion Image Retrieval
Liujie Hua, Ruihui Yi, Pinjie He, Xinpan Yuan |
ICIC (9) | 6 |
| 2025 | Adapting Cross-Modal Semantic Discrepancy in Text-based Person SearchabstractText-Based Person Search (TBPS) aims to retrieve target pedestrian images through language descriptions. However, the visual attributes and textual descriptions of different identities (pedestrians) tend to exhibit considerable similarity, leading to Similar Semantic Interference (SSI). To mitigate this issue, we propose the Adapting Cross-Modal Semantic Discrepancy (ACMSD) method, employing a cross-modal constraint approach to alleviate interference in model training. Specifically, we introduce the Consistent Constraint Alignment (CCA) strategy, which establishes both inter-modal alignment and intra-modal alignment, along with an Identity-Balanced Distribution (IBD) loss. This paradigm utilizes Cyclic Image-Text Contrastive to regularize the spatial distribution of the modalities, while the IBD loss implicitly clusters strong positive samples by using identity as a key index. Additionally, we incorporate an Attention-based Implicit Alignment (AIA) module to enforce modality-specific embeddings, thereby strengthening the interaction between cross-modal information. Extensive experiments are conducted on three public benchmark datasets to evaluate the performance of the ACMSD method. Xinpan Yuan, Wenguang Gan, Mengxi Ying, Liujie Hua |
ICME | 1 |
| 2025 | RSTA: A Recurrent Scene Topic-Aware Model for Multi-Scene Endoscopic Report GenerationabstractAutomated Report Generation (ARG) aims to automatically provide observations on images based on images, but faces difficulties due to the lack of graphic correspondence between multi-scenes and precise descriptions of multi-scene medical knowledge topics. To explore the task of multi-scene multi-lesion report generation (MSMLRG) and to address the above difficulties, this study proposes a recurrent scene topic-aware (RSTA) report generation model. We simulate the process of report writing by ear, nose, and throat (ENT) specialists with an innovative combination of the scene topic-aware module and recurrent generation module architectures. The scene-report feature aligner is used to realize the fine-grained alignment of images and different lesions in different scenes in the report, and incorporates a topic-aware module to help generate the topic text vocabulary for different scenes. The recurrent module includes an initialization statement generation module and a recurrent paragraph generation module to ensure coherent multi-sentence report generation for multiple complex images of endoscopy. Our approach has been successfully experimented on existing single-scene public datasets and multi-scene datasets using nasal endoscopy as an example, with excellent performance on natural language generation (NLG) and clinical efficacy (CE) metrics. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
IJCNN | 1 |
| 2025 | HADFF: Composed Image Retrieval Based on Hybrid-Attention and Dynamic Feature FusionabstractComposed Image Retrieval (CIR), which aims at retrieving a target image according to a reference image along with complement text, has drawn considerable attention. The critical challenge lies in accurately integrating the semantics from the two heterogeneous modalities. However, most existing approaches fail to fully exploit dynamic connection between the reference image and the complementary text. This causes the model to overlook the potential correlations between the image and text, resulting in poor fine-grained retrieval performance. To overcome the limitation, we propose a composed image retrieval method based on Hybrid-Attention and Dynamic Feature Fusion(HADFF). Specifically, Hybrid-Attention Module(HAM) is incorporated to extract rich image features from different levels. And then Q-former is used to fuse image and text features by dynamically allocating weights. Experiments on the public datasets CIRR and FashionIQ suggest that HADFF performs well in terms of retrieval accuracy and stability. Ruihui Yi, Liujie Hua, Pinjie He, Xinpan Yuan |
IJCNN | 6 |
| 2025 | HCSeer: A Classification Tool for Human Genetic Variant Hot and Cold Spots Designed for PM1 and Benign Criteria in the ACMG Guideline
Xingquan Xia, Guihu Zhao, Xinpan Yuan |
ISBRA (1) | 3 |
| 2025 | Configurable Platform for Biomedical Literature Mining via Multimodal-Driven Extraction
Xinpan Yuan, Bozhao Li, Guihu Zhao, Liujie Hua, Junhua Kuang, Shaomin Xie, Gan Li |
MICCAI (5) | 1 |
| 2025 | Diffusion with Awareness: An Adaptive Framework for Multi-class Anomaly Detection
Pinjie He, Zhihai Wu, Xinpan Yuan, Hongyong Duan |
PRCV (1) | 3 |
| 2025 | TimeSAST: Cloud Workload Sequence Prediction Method Based on CLEC-SSAabstractCloud platform elastic resource management relies on accurate workload sequence prediction to optimize dynamic resource allocation. However, cloud workloads exhibit highly nonlinear characteristics with stochastic fluctuations and noise interference, which severely degrade the feature extraction capability and prediction accuracy of traditional models. To address these challenges, this paper proposes Time Sparrow Search Algorithm with Attention and Soft Thresholding (TimeSAST), a novel prediction framework based on the Chaotic Levy Elite Cauchy Sparrow Search Algorithm (CLEC-SSA). The CLEC-SSA framework systematically overcomes three limitations of the traditional Sparrow Search Algorithm (SSA) through Tent chaotic mapping for uniform population initialization, Levy flight-enhanced global search to avoid local optima, and Cauchy mutation for adaptive parameter updates. Additionally, we design a neural network integrating self-attention mechanisms and soft thresholding, which effectively suppresses noise while capturing long-term dependencies. The experiments show that on the Google Cluster dataset, its MAE and RMSE are reduced by 32.6% and 54.2% respectively, while on the Alibaba Cloud dataset, the average reduction rates of the two metrics reach 38.2% and 52.6%. These results validate the robustness and practicality of TimeSAST in noisy and non-stationary cloud workload prediction tasks. Qiang Liu 0032, Xinpan Yuan, Xianchao Zhou |
QRS | 3 |
| 2025 | Bearing fault diagnosis based on multimodal knowledge graphs under few-shot samples
Cheng Peng 0015, Yanyan Sheng, Weihua Gui 0001, Zhaohui Tang 0004, Longxin Zhang, Xinpan Yuan |
Knowl. Based Syst. | 6 |
| 2024 | PaSeMix: A Multi-modal Partitional Semantic Data Augmentation Method for Text-Based Person Search
Xinpan Yuan, Wenguang Gan, Yanbin Weng |
ICIC (3) | 1 |
| 2024 | HeteroPP: A directive-based heterogeneous cooperative parallel programming frameworkabstractAbstract Heterogeneous platforms composed of multiple different types of computing devices (such as CPUs, GPUs, and Intel MICs) have been widely used recently. However, most of parallel applications developed in such a heterogeneous platform usually only utilize a certain kind of computing device due to the lack of easy‐to‐use heterogeneous cooperative parallel programming models. To reduce the difficulty of heterogeneous cooperative parallel programming, a directive‐based heterogeneous cooperative parallel programming framework called HeteroPP is proposed. HeteroPP provides an easier way for programmers to fully exploit multiple different types of computing devices to concurrently and cooperatively perform data‐parallel applications on heterogeneous platforms. An extension to OpenMP directives and clauses is proposed to make it possible for programmers to easily offload a data‐parallel compute kernel to multiple different types of computing devices. A source‐to‐source compiler is designed to help programmers to automatically generate multiple device‐specific compute kernels that can be concurrently and cooperatively performed on heterogeneous platforms. Many experiments are conducted with 12 typical data‐parallel applications implemented with HeteroPP on a heterogeneous CPU‐GPU‐MIC platform. The results show that HeteroPP not only greatly simplifies the heterogeneous cooperative parallel programming, but also can fully utilize the CPUs, GPU, and MIC to efficiently perform these applications. Lanjun Wan, Xueyan Cui, Xinpan Yuan |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | Research of image recognition method based on enhanced inception-ResNet-V2
Xinpan Yuan |
Multim. Tools Appl. | 3 |
| 2019 | Hierarchical one permutation hashing: efficient multimedia near duplicate detection
Chengyuan Zhang 0001, Yunwu Lin, Lei Zhu 0005, Xinpan Yuan, Fang Huang 0004 |
Multim. Tools Appl. | 4 |