EDBT 2026 Demo / reviewers in the wild / expert
Guihu Zhao
dblp:157/0188
· DBLP profile ↗
15ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0003-4033-1843ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, this assumption may not be able to hold for the cross-modal matching due to weak or even corrupted correlations between textual descriptions and visual content, referred to as noisy correspondence (NC). Such NC largely disrupts the correspondence learning between visual and semantic modalities. Though prior works have improved single-modal robustness against noisy labels, systematic modeling of both cross-modal and intra-modal geometric structures in TBPR remains limited attention. In this paper, we propose Geometric Structure Consistency Alignment (GSCA) to TBPR, which leverages cross-modal cosine similarity and intra-modal nearest-neighbor affinity to learn visual-semantic consistency under noisy correspondence. To mitigate the structural corruption caused by noisy pairs, we introduce the Structure Refinement and Mining (SRAM) module. By partitioning training data into clean, ambiguous, and noisy subsets, SRAM enables the model to strategically refine the cross-modal correspondence by mining reliable pairs, thus enhancing the reliability of positive or negative samples discrimination and preserving structural consistency across modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets. On CUHK-PEDES, it boosts Rank-1 by 1.42% in noise-free conditions, sustaining a robust 74.25% Rank-1 under a 50% noise ratio. Xinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang 0001, Guihu Zhao, Lin Wu 0001 |
AAAI | 5 |
| 2026 | Enhancing interpretation of clinical disease-associated copy number variations from multiple sequencing strategies with CNVSeekerabstractMOTIVATION: DNA copy number variations (CNVs) exert a profound impact on major genetic disorders in humans. Although multiple sequencing technologies have become the first line of molecular diagnosis for CNVs, existing tools are unable to resolve the pathogenicity of CNVs directly from raw sequencing data. RESULTS: We developed CNVSeeker, a one-stop and easy-to-use pipeline that provides comprehensive analysis from raw sequencing data to variant interpretation reports, and supports multiple types of sequencing data including short-read data such as whole genome sequencing data and whole exome sequencing data, and long-read sequencing data from Pacific Biosciences HiFi platform or Oxford Nanopore Technologies platform. Through extensive benchmarking, CNVSeeker demonstrated comparable enhancement over the state-of-the-art methods for CNV calling. Moreover, CNVSeeker enables significantly precise variant classification with an accuracy of ∼87%. By applying CNVSeeker to 1946 individuals with autism spectrum disorder (ASD), a total of 133 ASD-associated CNVs in 122 patients were identified, yielding a diagnostic yield of ∼6.3%. Additionally, we have also provided a user-friendly webserver for intuitive visualization of results. This study highlights the potential of CNVSeeker to benefit clinicians and geneticists with limited bioinformatic skill by aiding them interpret CNVs directly from various types of raw sequencing data for auxiliary disease diagnosis. AVAILABILITY AND IMPLEMENTATION: The web server is freely available at https://genemed.tech/cnvseeker and the open-source code can be found at https://github.com/lovelycatZ/CNVSeeker. Xudong Xiang, Xinxin Mao, Tengfei Luo, Chenbin Liu, Bozhao Li, Pei Yu, Dai Wu, Yixiao Zhu, Guihu Zhao, Jinchen Li |
Bioinform. | 14 |
| 2026 | Multi-scene topic-aware for novel single continuous shot multiple scenes endoscopy report generation
Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao, Siming Jin |
Knowl. Based Syst. | 5 |
| 2025 | HCSeer2: A Deep Learning-Based Multi-Scale Modeling Framework for Predicting Cold and Hot Spots of Variants in the Human ExomeabstractAccurately identifying variant hot and cold spots in human exonic regions is a crucial step in applying the PM1 and benign interpretation criteria of the ACMG-AMP guidelines. Addressing the issue that the existing tool HCSeeker relies heavily on the ClinVar database and has limited predictive performance in low-frequency variant regions, this study proposes HCSeer2-an innovative deep learning framework. By integrating the local feature extraction capability of CNN with the global dependency modeling of the Self-Attention mechanism, HCSeer2 achieves high-precision prediction of variant cold and hot spots across the entire human exome. HCSeer2 is trained on existing cold and hot spot data computed by tools such as HCSeeker. Through a multimodal feature fusion architecture, it simultaneously integrates sequence information and four types of genomic prior knowledge, effectively capturing the local aggregation patterns and global distribution rules of genomic variants. It enhances the understanding of variant clustering mechanisms in a self-supervised manner to predict variant cold and hot spots. Using this model, we identified 28,907 variant hot spots and 159,970 variant cold spots in the entire human exome, covering all human exonic regions. We also verified that the pathogenic potential of sites in hot spots is significantly higher than that in cold spots. This study not only provides scalable variant annotation resources for the clinical application of the ACMG-AMP guidelines but also offers new insights into the identification of genomic functional elements through the proposed multi-scale modeling approach. The code and data of HCSeer2 are freely available at https://github.com/xq-xia/HCSeer2. Xingquan Xia, Guihu Zhao, Jinchen Li, Xinpan Yuan |
BIBM | 2 |
| 2025 | PBD: A Manually Curated Full-Chain Benchmark Dataset for Evaluating LLMs on ACMG PS3/BS3 Functional Evidence AcquisitionabstractThe PS3 (Pathogenic Functional Evidence) and BS3 (Benign Functional Evidence) criteria in the American College of Medical Genetics and Genomics (ACMG) guidelines are critical for genetic variant classification. However, manual evaluation is time-consuming and prone to inter-laboratory inconsistencies, limiting the clinical interpretation of Variants of Uncertain Significance (VUS). Large Language Models (LLMs) offer potential for automated assessment, but their performance validation is hindered by the lack of standardized, high-quality datasets. This study introduces the PS3/BS3 Full-Chain Evidence Benchmark Dataset (PBD), the first manually curated dataset comprising 77 peer-reviewed publications, covering 266 cDNA variants (including duplicates) and their PS3/BS3 rating results. Adhering to ClinGen Sequence Variant Interpretation (SVI) standards, PBD includes structured, comprehensive evidence chains spanning genes, diseases, variants, experiments, and evidence ratings, designed to evaluate LLM capabilities in functional evidence extraction. We detail the dataset construction process, including literature screening, data extraction, standardization, and quality control, and developed a Python-based automated evaluation pipeline for reproducible, standardized analysis. Experiments using DeepSeek models$(1.5 \mathrm{b} / 7 \mathrm{b} / 14 \mathrm{b})$demonstrate PBD's potential in supporting automated variant interpretation. PBD provides a vital resource for bioinformatics and precision medicine, facilitating the development and standardization of variant classification tools. Data examples are available at https://github.com/User8588/PBD, with full data and code to be released upon paper acceptance. Xinpan Yuan, Bozhao Li, Chenbin Liu, Xinxue Li, Liujie Hua, Jinchen Li, Lin Wu 0001, Guihu Zhao |
BIBM | 8 |
| 2025 | OF-AR Relation Aware Representation Learning for Lesion Image Segmentation and GradingabstractMedical image segmentation provides important supplementary information for lesion grading, but existing segmentation models usually only focus on the lesion area, which is susceptible to the influence of shooting distance and angle, leading to feature extraction errors. We found an "as one falls, another rises"(OF-AR) relationship between the lesion and the surrounding non lesion areas, and introduced OF-AR relationship aware learning representation to jointly extract features of lesions and non lesions. By dividing the image into black and white dual zones, expanding the feature extraction area, and using a dual zone contrastive learning module to increase the inter-class distance, accurate grading is ensured by utilizing the relative information of two complete targets. In addition, using weighted fusion methods to enhance the comprehensiveness and objectivity of grading. The experiment used adenoids as an example to verify the high accuracy of this method. Xinpan Yuan, Siming Jin, Liujie Hua, Guihu Zhao |
ICASSP | 4 |
| 2025 | A Novel Single Continuous Shot Multiple Lesions Endoscopy Report GenerationabstractAutomatic Report Generation(ARG), which aims to automatically provide observations on images, is challenged by the lack of coherence between multiple scenes and precise description of multiple lesions. In order to explore the task of multi-scene multi-lesion report generation(MSMLRG) in one shot, in this paper, we introduce a multi-scene multi-lesion report generation framework to extract the scene-report alignment relation and scene-topic relation. Specifically, we design a scene-report feature aligner to achieve fine-grained alignment of different lesion in different scenes in images and reports, and incorporate a topic-aware module to help generate a topic text vocabulary for different scenes. Our framework has been successfully experimented on several automatic report generation models, and performs well on automatic evaluation metrics. The framework for one-shot report generation during multi-scene not only fills the gap of multi-scene image report generation, but also effectively improves the accuracy and consistency of diagnostic reports. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
ICASSP | 4 |
| 2025 | Advanced Font-Aware Document Hierarchy Reconstruction for Enhanced Structured Parsing
Xinpan Yuan, Gan Li, Liujie Hua, Guihu Zhao, Shaomin Xie |
ICIC (16) | 4 |
| 2025 | CLIO: A Unified Framework for Consistency-Aware Learning and Intra-Modal Optimization in Text-Based Person Re-identification
Xinpan Yuan, Shaomin Xie, Guihu Zhao, Liujie Hua, Wenguang Gan |
ICIC (5) | 3 |
| 2025 | RSTA: A Recurrent Scene Topic-Aware Model for Multi-Scene Endoscopic Report GenerationabstractAutomated Report Generation (ARG) aims to automatically provide observations on images based on images, but faces difficulties due to the lack of graphic correspondence between multi-scenes and precise descriptions of multi-scene medical knowledge topics. To explore the task of multi-scene multi-lesion report generation (MSMLRG) and to address the above difficulties, this study proposes a recurrent scene topic-aware (RSTA) report generation model. We simulate the process of report writing by ear, nose, and throat (ENT) specialists with an innovative combination of the scene topic-aware module and recurrent generation module architectures. The scene-report feature aligner is used to realize the fine-grained alignment of images and different lesions in different scenes in the report, and incorporates a topic-aware module to help generate the topic text vocabulary for different scenes. The recurrent module includes an initialization statement generation module and a recurrent paragraph generation module to ensure coherent multi-sentence report generation for multiple complex images of endoscopy. Our approach has been successfully experimented on existing single-scene public datasets and multi-scene datasets using nasal endoscopy as an example, with excellent performance on natural language generation (NLG) and clinical efficacy (CE) metrics. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
IJCNN | 4 |
| 2025 | HCSeer: A Classification Tool for Human Genetic Variant Hot and Cold Spots Designed for PM1 and Benign Criteria in the ACMG Guideline
Xingquan Xia, Guihu Zhao, Xinpan Yuan |
ISBRA (1) | 2 |
| 2025 | Configurable Platform for Biomedical Literature Mining via Multimodal-Driven Extraction
Xinpan Yuan, Bozhao Li, Guihu Zhao, Liujie Hua, Junhua Kuang, Shaomin Xie, Gan Li |
MICCAI (5) | 3 |
| 2018 | Exploring a smart pathological brain detection method on pseudo Zernike moment
Yudong Zhang 0001, Yongyan Jiang, Weiguo Zhu, Siyuan Lu 0001, Guihu Zhao |
Multim. Tools Appl. | 5 |
| 2018 | Cat Swarm Optimization applied to alcohol use disorder identification
Yudong Zhang 0001, Yuxiu Sui, Junding Sun, Guihu Zhao, Pengjiang Qian |
Multim. Tools Appl. | 4 |
| 2018 | Smart pathological brain detection by synthetic minority oversampling technique, extreme learning machine, and Jaya algorithm
Yudong Zhang 0001, Guihu Zhao, Junding Sun, Xiaosheng Wu, Zhiheng Wang 0001, Hongmin Liu 0001, Vishnuvarthanan Govindaraj, Tianming Zhan, Jianwu Li |
Multim. Tools Appl. | 2 |