VLDB 2026 Research / reviewers in the wild / expert
Liujie Hua
dblp:313/8326
· DBLP profile ↗
25ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-2756-1540ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, this assumption may not be able to hold for the cross-modal matching due to weak or even corrupted correlations between textual descriptions and visual content, referred to as noisy correspondence (NC). Such NC largely disrupts the correspondence learning between visual and semantic modalities. Though prior works have improved single-modal robustness against noisy labels, systematic modeling of both cross-modal and intra-modal geometric structures in TBPR remains limited attention. In this paper, we propose Geometric Structure Consistency Alignment (GSCA) to TBPR, which leverages cross-modal cosine similarity and intra-modal nearest-neighbor affinity to learn visual-semantic consistency under noisy correspondence. To mitigate the structural corruption caused by noisy pairs, we introduce the Structure Refinement and Mining (SRAM) module. By partitioning training data into clean, ambiguous, and noisy subsets, SRAM enables the model to strategically refine the cross-modal correspondence by mining reliable pairs, thus enhancing the reliability of positive or negative samples discrimination and preserving structural consistency across modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets. On CUHK-PEDES, it boosts Rank-1 by 1.42% in noise-free conditions, sustaining a robust 74.25% Rank-1 under a 50% noise ratio. Xinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang 0001, Guihu Zhao, Lin Wu 0001 |
AAAI | 3 |
| 2026 | Syntax-Aware Dependency Parsing for Dual-Origin Noisy Correspondence in Text-Based Person Search
Xinpan Yuan, Wenguang Gan, Shaomin Xie, Chengyuan Zhang 0001, Liujie Hua |
DASFAA (1) | 6 |
| 2026 | Set-to-One Structured Captioning for Heterogeneous Nasal Image Collections via Spatio-Temporal Modeling and Clinical Knowledge Integration
Xinpan Yuan, Jianuo Ju, Liujie Hua, Mingzhu Huang |
DASFAA (3) | 3 |
| 2026 | FeatDeNoiseNet: Multi-scale Structure-Consistent Manifold Modeling for Unsupervised Anomaly Localization
Hongxiao Fei, Zhubang Qu, Liujie Hua, Qianqian Qi 0009, Yueyi Luo |
ICIC (8) | 4 |
| 2026 | RiskGate: Risk-Calibrated Task-Aware Evidence Gating for Reliable Industrial VQA
Yueyi Luo, Liujie Hua, Qianqian Qi 0009 |
ICIC (21) | 3 |
| 2026 | Multi-scene topic-aware for novel single continuous shot multiple scenes endoscopy report generation
Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao, Siming Jin |
Knowl. Based Syst. | 4 |
| 2025 | PBD: A Manually Curated Full-Chain Benchmark Dataset for Evaluating LLMs on ACMG PS3/BS3 Functional Evidence AcquisitionabstractThe PS3 (Pathogenic Functional Evidence) and BS3 (Benign Functional Evidence) criteria in the American College of Medical Genetics and Genomics (ACMG) guidelines are critical for genetic variant classification. However, manual evaluation is time-consuming and prone to inter-laboratory inconsistencies, limiting the clinical interpretation of Variants of Uncertain Significance (VUS). Large Language Models (LLMs) offer potential for automated assessment, but their performance validation is hindered by the lack of standardized, high-quality datasets. This study introduces the PS3/BS3 Full-Chain Evidence Benchmark Dataset (PBD), the first manually curated dataset comprising 77 peer-reviewed publications, covering 266 cDNA variants (including duplicates) and their PS3/BS3 rating results. Adhering to ClinGen Sequence Variant Interpretation (SVI) standards, PBD includes structured, comprehensive evidence chains spanning genes, diseases, variants, experiments, and evidence ratings, designed to evaluate LLM capabilities in functional evidence extraction. We detail the dataset construction process, including literature screening, data extraction, standardization, and quality control, and developed a Python-based automated evaluation pipeline for reproducible, standardized analysis. Experiments using DeepSeek models$(1.5 \mathrm{b} / 7 \mathrm{b} / 14 \mathrm{b})$demonstrate PBD's potential in supporting automated variant interpretation. PBD provides a vital resource for bioinformatics and precision medicine, facilitating the development and standardization of variant classification tools. Data examples are available at https://github.com/User8588/PBD, with full data and code to be released upon paper acceptance. Xinpan Yuan, Bozhao Li, Chenbin Liu, Xinxue Li, Liujie Hua, Jinchen Li, Lin Wu 0001, Guihu Zhao |
BIBM | 5 |
| 2025 | MMAG: Multimodal Learning for Mucus Anomaly Grading in Nasal Endoscopy via Semantic Attribute PromptingabstractAccurate grading of rhinitis severity in nasal endoscopy relies heavily on the characterization of key secretion types, notably clear nasal discharge (CND) and purulent nasal secretion (PUS).However, both exhibit ambiguous appearance and high structural variability, posing challenges to automated grading under weak supervision.To address this, we propose Multimodal Learning for Mucus Anomaly Grading (MMAG), which integrates structured prompts with rank-aware vision-language modeling for joint detection and grading.Attribute prompts are constructed from clinical descriptors (e.g., secretion type, severity, location) and aligned with multi-level visual features via a dual-branch encoder.During inference, the model localizes mucus anomalies and maps the input image to severity-specific prompts (e.g., "moderate pus"), projecting them into a rankaware feature space for progressive similarity scoring.Extensive evaluations on CND and PUS datasets show that our method achieves consistent gains over Baseline, improving AUC by 6.31% and 4.79%, and F1 score by 12.85% and 6.03%, respectively.This framework enables interpretable, annotation-efficient, and semantically grounded assessment of rhinitis severity based on mucus anomalies. Xinpan Yuan, Mingzhu Huang, Liujie Hua, Jianuo Ju |
EMNLP | 3 |
| 2025 | A Reinforcement Learning Agent Controlled Multi-branch Small Object Detection FrameworkabstractThe past few years have witnessed the immense development of small object detection, which is aimed at detecting size-limited targets in high-resolution images. The prevailing methods focus on extracting fine-grained information by expanding the receptive fields and then generating the potential small object region. However, these solutions inevitably add sophisticated detectors and extra learning components, which incur time-consuming and computation-costing. Meanwhile, we observe that it’s suboptimal to extract fine features in such a generic way. To alleviate the issues, we propose a multi-branch small object detection framework with a regular-scale detection branch and a small-scale detection branch. Specifically, we design and pre-train a reinforcement learning agent to control feature extractors in both branches according to the results of small object areas. Moreover, we present a region clipping algorithm to rebuild the small object to regular size, which can be input into a mature detector directly. The extensive experiments on COCO, VisDrone, SODA-D, and our collecting TVDS datasets demonstrate our method outperforms the state-of-the-art methods in several metrics. Junkun Hong, Yitian Long, Yueyi Luo, Liujie Hua, Qianqian Qi 0009 |
ICASSP | 4 |
| 2025 | HieClip: Hierarchical CLIP with Explicit Alignment for Zero-Shot Anomaly DetectionabstractLarge image-language models(LLM) have made significant progress in zero-shot anomaly detection(ZSAD), however, the semantic gap between images and text limits their performance in hierarchical learning. In this paper, we propose the hierarchical alignment clip(HieClip) framework, to achieve hierarchical alignment between images and text. Specifically, we introduce learnable hierarchical textual(LHT) to reduce the representation differences between various levels of images and text, while performing multi-level comprehensive discrimination. Additionally, the dynamically adjusting the weights of features at different levels, improving the model’s ability to capture both global and local information. Experiments on public industrial datasets demonstrate HieClip’s effectiveness, showing significant accuracy improvement, and its strong generalization capabilities were further validated on medical datasets. Compared to existing methods, HieClip excels in anomaly detection tasks, particularly in industrial inspection and medical diagnosis scenarios. Liujie Hua, Xiu Su, Yueyi Luo, Shan You |
ICASSP | 1 |
| 2025 | OF-AR Relation Aware Representation Learning for Lesion Image Segmentation and GradingabstractMedical image segmentation provides important supplementary information for lesion grading, but existing segmentation models usually only focus on the lesion area, which is susceptible to the influence of shooting distance and angle, leading to feature extraction errors. We found an "as one falls, another rises"(OF-AR) relationship between the lesion and the surrounding non lesion areas, and introduced OF-AR relationship aware learning representation to jointly extract features of lesions and non lesions. By dividing the image into black and white dual zones, expanding the feature extraction area, and using a dual zone contrastive learning module to increase the inter-class distance, accurate grading is ensured by utilizing the relative information of two complete targets. In addition, using weighted fusion methods to enhance the comprehensiveness and objectivity of grading. The experiment used adenoids as an example to verify the high accuracy of this method. Xinpan Yuan, Siming Jin, Liujie Hua, Guihu Zhao |
ICASSP | 3 |
| 2025 | A Novel Single Continuous Shot Multiple Lesions Endoscopy Report GenerationabstractAutomatic Report Generation(ARG), which aims to automatically provide observations on images, is challenged by the lack of coherence between multiple scenes and precise description of multiple lesions. In order to explore the task of multi-scene multi-lesion report generation(MSMLRG) in one shot, in this paper, we introduce a multi-scene multi-lesion report generation framework to extract the scene-report alignment relation and scene-topic relation. Specifically, we design a scene-report feature aligner to achieve fine-grained alignment of different lesion in different scenes in images and reports, and incorporate a topic-aware module to help generate a topic text vocabulary for different scenes. Our framework has been successfully experimented on several automatic report generation models, and performs well on automatic evaluation metrics. The framework for one-shot report generation during multi-scene not only fills the gap of multi-scene image report generation, but also effectively improves the accuracy and consistency of diagnostic reports. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
ICASSP | 3 |
| 2025 | VAMP: Visual Attribute-Guided Multi-dimensional Perception for Nasal Endoscopy Report Generation
Xinpan Yuan, Jianuo Ju, Liujie Hua, Mingzhu Huang, Shaomin Xie, Wenguang Gan |
ICIC (5) | 3 |
| 2025 | MAS-ZSAS: A Zero-Shot Anomaly Segmentation Framework with Multi-attribute Guided Text Prompts
Xinpan Yuan, Guorong Liang, Liujie Hua, Shaomin Xie, Wenguang Gan |
ICIC (6) | 3 |
| 2025 | Advanced Font-Aware Document Hierarchy Reconstruction for Enhanced Structured Parsing
Xinpan Yuan, Gan Li, Liujie Hua, Guihu Zhao, Shaomin Xie |
ICIC (16) | 3 |
| 2025 | CLIO: A Unified Framework for Consistency-Aware Learning and Intra-Modal Optimization in Text-Based Person Re-identification
Xinpan Yuan, Shaomin Xie, Guihu Zhao, Liujie Hua, Wenguang Gan |
ICIC (5) | 4 |
| 2025 | GCA-Net: Global Contextual Attention Network with Lightweight Hierarchical Alignment for Text-Guided Fashion Image Retrieval
Liujie Hua, Ruihui Yi, Pinjie He, Xinpan Yuan |
ICIC (9) | 3 |
| 2025 | Adapting Cross-Modal Semantic Discrepancy in Text-based Person SearchabstractText-Based Person Search (TBPS) aims to retrieve target pedestrian images through language descriptions. However, the visual attributes and textual descriptions of different identities (pedestrians) tend to exhibit considerable similarity, leading to Similar Semantic Interference (SSI). To mitigate this issue, we propose the Adapting Cross-Modal Semantic Discrepancy (ACMSD) method, employing a cross-modal constraint approach to alleviate interference in model training. Specifically, we introduce the Consistent Constraint Alignment (CCA) strategy, which establishes both inter-modal alignment and intra-modal alignment, along with an Identity-Balanced Distribution (IBD) loss. This paradigm utilizes Cyclic Image-Text Contrastive to regularize the spatial distribution of the modalities, while the IBD loss implicitly clusters strong positive samples by using identity as a key index. Additionally, we incorporate an Attention-based Implicit Alignment (AIA) module to enforce modality-specific embeddings, thereby strengthening the interaction between cross-modal information. Extensive experiments are conducted on three public benchmark datasets to evaluate the performance of the ACMSD method. Xinpan Yuan, Wenguang Gan, Mengxi Ying, Liujie Hua |
ICME | 6 |
| 2025 | RSTA: A Recurrent Scene Topic-Aware Model for Multi-Scene Endoscopic Report GenerationabstractAutomated Report Generation (ARG) aims to automatically provide observations on images based on images, but faces difficulties due to the lack of graphic correspondence between multi-scenes and precise descriptions of multi-scene medical knowledge topics. To explore the task of multi-scene multi-lesion report generation (MSMLRG) and to address the above difficulties, this study proposes a recurrent scene topic-aware (RSTA) report generation model. We simulate the process of report writing by ear, nose, and throat (ENT) specialists with an innovative combination of the scene topic-aware module and recurrent generation module architectures. The scene-report feature aligner is used to realize the fine-grained alignment of images and different lesions in different scenes in the report, and incorporates a topic-aware module to help generate the topic text vocabulary for different scenes. The recurrent module includes an initialization statement generation module and a recurrent paragraph generation module to ensure coherent multi-sentence report generation for multiple complex images of endoscopy. Our approach has been successfully experimented on existing single-scene public datasets and multi-scene datasets using nasal endoscopy as an example, with excellent performance on natural language generation (NLG) and clinical efficacy (CE) metrics. Xinpan Yuan, Junhua Kuang, Liujie Hua, Guihu Zhao |
IJCNN | 3 |
| 2025 | HADFF: Composed Image Retrieval Based on Hybrid-Attention and Dynamic Feature FusionabstractComposed Image Retrieval (CIR), which aims at retrieving a target image according to a reference image along with complement text, has drawn considerable attention. The critical challenge lies in accurately integrating the semantics from the two heterogeneous modalities. However, most existing approaches fail to fully exploit dynamic connection between the reference image and the complementary text. This causes the model to overlook the potential correlations between the image and text, resulting in poor fine-grained retrieval performance. To overcome the limitation, we propose a composed image retrieval method based on Hybrid-Attention and Dynamic Feature Fusion(HADFF). Specifically, Hybrid-Attention Module(HAM) is incorporated to extract rich image features from different levels. And then Q-former is used to fuse image and text features by dynamically allocating weights. Experiments on the public datasets CIRR and FashionIQ suggest that HADFF performs well in terms of retrieval accuracy and stability. Ruihui Yi, Liujie Hua, Pinjie He, Xinpan Yuan |
IJCNN | 3 |
| 2025 | Configurable Platform for Biomedical Literature Mining via Multimodal-Driven Extraction
Xinpan Yuan, Bozhao Li, Guihu Zhao, Liujie Hua, Junhua Kuang, Shaomin Xie, Gan Li |
MICCAI (5) | 5 |
| 2024 | Estate: Expert-Guided State Text Enhancement for Zero-Shot Industrial Anomaly DetectionabstractThe Expert-Guided State Text Enhancement Anomaly Detection (ESTATE) framework addresses the challenges in industrial anomaly detection arising from diverse product categories and limited defective samples. This framework, integrating expert insights through comparative state prompts, leverages two innovative text-guided networks, CLS-Refiner and SEG-Refiner, enhancing model training. These networks, connected to residual textual features of standard vision-language pre-trained models, focus on amplifying adjectives’ significance in text for improved image block and pixel-level alignment. ESTATE’s effectiveness is demonstrated through evaluations on MVTecAD and VisA datasets, achieving AUROC scores of 89.6%/89.6% for classification and 95.1%/85.0% for segmentation tasks, alongside setting new benchmarks in F1Max and PRO metrics. The AUC-cls on MVTecAD and VisA demonstrated an enhancement of 5.06% and 8.97%, respectively, compared to the APRIL-GAN approach. Bingke Zhu, Hao Li 0115, Changlin Chen, Liujie Hua, Jinqiao Wang |
ICIP | 4 |
| 2024 | Image Anomaly Detection Based on Controllable Self-AugmentationabstractBased on data synthesis, anomaly detection (AD) methods often rely on external data for data synthesis. However, most external abnormal data exhibits strong randomness, which may lead to a reduced range of diversity among the synthesized data. In order to achieve a broader diversity in data synthesis, it is necessary to not only have highly diverse data but also to incorporate low-diversity noise data. To enhance the diversity range of the synthesized data, this study proposes a diversity measurement assisted by image self-representation: measuring the distance between noise data and normal data and quantitatively synthesizing diversified data by selecting diverse noise data for synthesis, namely, Diversified Synthesis (DS). Diversified Synthesis introduces patch measurement and a controllable enhancement module to establish controllable diversified enhanced data. The contribution of this study lies in proposing a novel diversified synthesis method, which achieves a broader diversity synthesis through the introduction of image self-representation-assisted diversity measurement and quantitative synthesis. Furthermore, through the self-enhancement data augmentation method, the use of image intrinsic features for enhancement achieves diversity and multi-scale characteristics in the synthesized data, thereby improving the training performance of the discriminative model. This provides an effective optimization solution for comprehensive anomaly detection methods. Liujie Hua, Yichao Cao, Yitian Long, Shan You, Xiu Su, Yueyi Luo, Chang Xu 0002 |
IJCNN | 1 |
| 2022 | Self-supervised Augmented Patches Segmentation for Anomaly Detection
Yuxi Yang, Liujie Hua, Yiqi Ou |
ACCV (2) | 3 |
| 2022 | Label embedding semantic-guided hashing
Longzhi Sun, Lin Guo 0014, Liujie Hua, Zhan Yang 0001 |
Neurocomputing | 4 |