VLDB 2026 Research / reviewers in the wild / expert
Zailong Chen
dblp:310/8119
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0003-8431-5471ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience StrategiesabstractLarge Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of token positions, it induces progressive attention decay as token distance increases, especially with progressive attention decay over distant token pairs, which severely impairs the model's ability to remember global context. To alleviate this issue, we propose inference-only Three-step Decay Resilience Strategies (T-DRS), comprising (1) Semantic-Driven DRS (SD-DRS), amplifying semantically meaningful but distant signals via content-aware residuals, (2) Distance-aware Control DRS (DC-DRS), which can purify attention by smoothly modulating weights based on positional distances, suppressing noise while preserving locality, and (3) re-Reinforce Distant DRS (reRD-DRS), consolidating the remaining informative remote dependencies to maintain global coherence. Together, the T-DRS recover suppressed long-range token pairs without harming local inductive biases. Extensive experiments on Vision Question Answering (VQA) benchmarks demonstrate that T-DRS can consistently improve performance in an inference-only manner. Yujian Lee, Zailong Chen, Hui Zhang 0062 |
AAAI | 4 |
| 2025 | NCL-CIR: Noise-aware Contrastive Learning for Composed Image RetrievalabstractComposed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through data augmentation or model design. These methods often assume perfect alignment between queries and target images, an idealized scenario rarely encountered in practice. In reality, pairs are often partially or completely mismatched due to issues like inaccurate modification texts, low-quality target images, and annotation errors. Ignoring these mismatches leads to numerous False Positive Pair (FFPs) denoted as noise pairs in the dataset, causing the model to overfit and ultimately reducing its performance. To address this problem, we propose the Noise-aware Contrastive Learning for CIR (NCL-CIR), comprising two key components: the Weight Compensation Block (WCB) and the Noise-pair Filter Block (NFB). The WCB coupled with diverse weight maps can ensure more stable token representations of multi-modal queries and target images. Meanwhile, the NFB, in conjunction with the Gaussian Mixture Model (GMM) predicts noise pairs by evaluating loss distributions, and generates soft labels correspondingly, allowing for the design of the soft-label based Noise Contrastive Estimation (NCE) loss function. Consequently, the overall architecture helps to mitigate the influence of mismatched and partially matched samples, with experimental results demonstrating that NCL-CIR achieves exceptional performance on the benchmark datasets. Yujian Lee, Zailong Chen, Yiyang Hu, Guquan Jing |
ICASSP | 3 |
| 2025 | Optimizing Efficiency and Visual-Textual Alignment for LLM-Based Radiology Report GenerationabstractLLM-based radiology report generation (R2Gen) systems have demonstrated promising performance but face significant challenges in bridging the gap between the visual encoder and the LLM. Specifically, two issues hinder progress: (1) parameter-heavy visual projector that increases complexity and degrades performance, and (2) insufficient alignment between visual and textual modalities, limiting system efficacy. To address these, we propose R2Gen-EVA, a novel framework emphasizing Efficiency and Visual-Textual Alignment (VTA), which introduces two key innovations: (1) a parameter-free visual projector that enhances model efficiency while improving performance, and (2) an LLM-adapted VTA module that enhances the alignment of visual features with LLM’s textual embeddings. Our design significantly improves model efficacy without adding extra parameters, achieving both streamlined complexity and higher computational efficiency during inference. Extensive experiments demonstrate that R2Gen-EVA enhances the fluency and clinical accuracy of generated reports, establishing it as a more effective and efficient solution for LLM-based R2Gen. The code is available at https://github.com/zailongchen/R2Gen-EVA. Zailong Chen, Yujian Lee, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
ICME | 1 |
| 2025 | Boosting Audio-Visual Segmentation via Triple-Modalities AlignmentabstractThe Audio-Visual Segmentation (AVS) task aims to identify sound-producing objects in the visual domain using auditory cues. Enhancing segmentation efficiency by incorporating prior knowledge, such as object locations and textual prompts, has proven to be crucial. However, existing methods suffer from feature misalignment during model training, leading to ineffective integration and reduced performance. To address this, we propose Triple-modalities alignment (TM-align), which combines audio signals, visual images, and textual prompts. By leveraging prompts from a frozen multi-modal large language model (MLLM), we extract two types of semantic information: contextual semantic description (C.S.D) and prompt specific summary (P.S.S). TM-align yields three pairs of aligned features: visual and C.S.D, visual and P.S.S, visual and audio, within two of our proposed cross-modalities alignment (CMA) models. To further enhance the alignment, we employ Jensen-Shannon Divergence (JSD) to regulate the domain distribution of the latter two features. By effectively aligning the three modalities, TM-align reduces redundancy and improves the overall AVS performance. Experimental results demonstrate that TM-align outperforms the mainstream AVS models.1 Yujian Lee, Zailong Chen, Wentao Fan 0001, Guquan Jing, Yiyang Hu |
ICME | 3 |
| 2025 | Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language ModelsabstractComposed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git. Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai |
ICMR | 5 |
| 2025 | Enhancing Radiology Report Generation via Multi-Phased SupervisionabstractRadiology report generation using large language models has recently produced reports with more realistic styles and better language fluency. However, their clinical accuracy remains inadequate. Considering the significant imbalance between clinical phrases and general descriptions in a report, we argue that using an entire report for supervision is problematic as it fails to emphasize the crucial clinical phrases, which require focused learning. To address this issue, we propose a multi-phased supervision method, inspired by the spirit of curriculum learning where models are trained by gradually increasing task complexity. Our approach organizes the learning process into structured phases at different levels of semantical granularity, each building on the previous one to enhance the model. During the first phase, disease labels are used to supervise the model, equipping it with the ability to identify underlying diseases. The second phase progresses to use entity-relation triples to guide the model to describe associated clinical findings. Finally, in the third phase, we introduce conventional whole-report-based supervision to quickly adapt the model for report generation. Throughout the phased training, the model remains the same and consistently operates in the generation mode. As experimentally demonstrated, this proposed change in the way of supervision enhances report generation, achieving state-of-the-art performance in both language fluency and clinical accuracy. Our work underscores the importance of training process design in radiology report generation. Our code is available on https://github.com/zailongchen/MultiP-R2Gen. Zailong Chen, Yingshu Li 0002, Zhanyu Wang, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Question-Aware Global-Local Video Understanding Network for Audio-Visual Question AnsweringabstractAs a newly emerging task, audio-visual question answering (AVQA) has attracted research attention. Compared with traditional single-modality (e.g., audio or visual) QA tasks, it poses new challenges due to the higher complexity of feature extraction and fusion brought by the multimodal inputs. First, AVQA requires more comprehensive understanding of the scene which involves both audio and visual information; Second, in the presence of more information, feature extraction has to be better connected with a given question; Third, features from different modalities need to be sufficiently correlated and fused. To address this situation, this work proposes a novel framework for multimodal question answering task. It characterises an audiovisual scene at both global and local levels, and within each level, the features from different modalities are well fused. Furthermore, the given question is utilised to guide not only the feature extraction at the local level but also the final fusion of global and local features to predict the answer. Our framework provides a new perspective for audio-visual scene understanding through focusing on both general and specific representations as well as aggregating multimodalities by prioritizing question-related information. As experimentally demonstrated, our method significantly improves the existing audio-visual question answering performance, with the averaged absolute gain of 3.3% and 3.1% on MUSIC-AVQA and AVQA datasets, respectively. Moreover, the ablation study verifies the necessity and effectiveness of our design. Our code will be publicly released. Zailong Chen, Lei Wang 0001, Peng Wang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | EPMC: efficient parallel memory compression in deep neural network training
Zailong Chen, Shenghong Yang, Chubo Liu, Yikun Hu 0001, Kenli Li 0001, Keqin Li 0001 |
Neural Comput. Appl. | 1 |
| 2022 | LAP: Latency-aware automated pruning with dynamic-based filter selection
Zailong Chen, Chubo Liu, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
Neural Networks | 1 |