EDBT 2026 Demo / reviewers in the wild / expert
Yujian Lee
dblp:375/1204
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0003-2514-3913ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience StrategiesabstractLarge Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of token positions, it induces progressive attention decay as token distance increases, especially with progressive attention decay over distant token pairs, which severely impairs the model's ability to remember global context. To alleviate this issue, we propose inference-only Three-step Decay Resilience Strategies (T-DRS), comprising (1) Semantic-Driven DRS (SD-DRS), amplifying semantically meaningful but distant signals via content-aware residuals, (2) Distance-aware Control DRS (DC-DRS), which can purify attention by smoothly modulating weights based on positional distances, suppressing noise while preserving locality, and (3) re-Reinforce Distant DRS (reRD-DRS), consolidating the remaining informative remote dependencies to maintain global coherence. Together, the T-DRS recover suppressed long-range token pairs without harming local inductive biases. Extensive experiments on Vision Question Answering (VQA) benchmarks demonstrate that T-DRS can consistently improve performance in an inference-only manner. Yujian Lee, Zailong Chen, Hui Zhang 0062 |
AAAI | 2 |
| 2026 | Describing-Verifying-Scoring: A Hierarchical Reasoning Framework for Zero-Shot Composed Image RetrievalabstractZero-Shot Composed Image Retrieval (ZS-CIR) aims to identify target images using a composed query of a reference image and modification text without labeled triplets. While recent advances leverage Multimodal Large Language Models (MLLMs) for intent reasoning, they often suffer from hallucination-induced inaccuracies where misaligned descriptions degrade retrieval reliability, and insufficient reasoning due to shallow prompting strategies. To address these challenges, we propose DVSCIR, a novel training-free framework featuring a hierarchical Describing-Verifying-Scoring pipeline with MLLM. Specifically, the Describing stage generates an initial candidate caption, followed by a Verifying stage that rectifies potential hallucinations to ensure description accuracy. The Scoring stage performs a fine-grained re-ranking to identify the optimal match. Within each stage, a hierarchical Chain-of-Thought (CoT) process tailored for ZS-CIR guides the MLLM from low-level perception to deep intentional reasoning via sequential steps within structured sections. This progression ensures robust cross-modal correspondence through a hierarchical refinement of the retrieval process. Extensive experiments across four benchmarks demonstrate that DVSCIR achieves state-of-the-art performance, validating its effectiveness in ZS-CIR. Guquan Jing, Yujian Lee, Hui Zhang 0062 |
ICMR | 3 |
| 2026 | Bayesian Hyperspherical Graph Mixture-of-Experts Deciphers Cell-Cell Interaction in Spatial TranscriptomicsabstractSpatial transcriptomics (ST) technologies have transformed our understanding of tissue biology by capturing gene expression with spatial context, enabling systematic analysis of cell-cell interactions (CCIs) and spatial domains in complex tissues. However, existing computational approaches often rely on fixed proximity graphs, curated ligand-receptor (LR) databases, or deep graph neural networks that are prone to over-smoothing and lack principled uncertainty quantification. These limitations hinder the discovery of heterogeneous, directional, and long-range CCIs essential for interpreting tissue organization and disease mechanisms. Here, we present B-HGME (Bayesian Hyperspherical Graph Mixture of Experts), a scalable, unsupervised framework that jointly delineates spatial domains and infers CCI networks from ST data with principled uncertainty estimates. B-HGME integrates spatial and gene-regulatory graphs into a dual-scale structure, encodes cell representations on a unit hypersphere via coupled message passing, and decodes edges using a Bayesian mixture-of-experts governed by a Dirichlet-regularized gating network. This design enables the model to capture multi-scale, directional, and biologically coherent interactions while avoiding the over-smoothing and posterior collapse of conventional models. The hyperspherical embedding geometry ensures angular similarity is preserved in high dimensions, and edge-level credibility is derived from the Bayesian posterior, facilitating interpretable and confident CCI inference. Across multiple datasets from six major ST platforms, B-HGME consistently achieves state-of-the-art spatial clustering accuracy and uncovers biologically coherent and diverse CCIs, including novel interactions beyond curated ligand-receptor pairs. B-HGME's hyperspherical embeddings accurately localize canonical astrocytic and laminar markers (e.g.,Gfap,Pcp4,Calb1, andCamk2a) to their expected spatial niches, confirming biochemical fidelity at single-gene resolution. Inferred ligand-receptor circuits not only recover known pathways but also reveal previously uncharacterized interactions (e.g., Astro1-eL2/3, VIP-Oligo), furnishing mechanistic hypotheses for cortical layer formation and tumor-stroma crosstalk. Together, these results demonstrate that B-HGME offers a powerful tool for spatial systems biology and hypothesis generation in development, immunity, and cancer. The source code of our model is available athttps://github.com/zxj8806/B-HGME. Wenchuan Zhang, Yujian Lee, Ricky Yuen-Tan Hou, Weifeng Su, Hong Yan 0006, Wentao Fan 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | NCL-CIR: Noise-aware Contrastive Learning for Composed Image RetrievalabstractComposed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through data augmentation or model design. These methods often assume perfect alignment between queries and target images, an idealized scenario rarely encountered in practice. In reality, pairs are often partially or completely mismatched due to issues like inaccurate modification texts, low-quality target images, and annotation errors. Ignoring these mismatches leads to numerous False Positive Pair (FFPs) denoted as noise pairs in the dataset, causing the model to overfit and ultimately reducing its performance. To address this problem, we propose the Noise-aware Contrastive Learning for CIR (NCL-CIR), comprising two key components: the Weight Compensation Block (WCB) and the Noise-pair Filter Block (NFB). The WCB coupled with diverse weight maps can ensure more stable token representations of multi-modal queries and target images. Meanwhile, the NFB, in conjunction with the Gaussian Mixture Model (GMM) predicts noise pairs by evaluating loss distributions, and generates soft labels correspondingly, allowing for the design of the soft-label based Noise Contrastive Estimation (NCE) loss function. Consequently, the overall architecture helps to mitigate the influence of mismatched and partially matched samples, with experimental results demonstrating that NCL-CIR achieves exceptional performance on the benchmark datasets. Yujian Lee, Zailong Chen, Yiyang Hu, Guquan Jing |
ICASSP | 2 |
| 2025 | How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
Yujian Lee, Yongqi Xu, Wentao Fan 0001 |
ICCV | 1 |
| 2025 | Optimizing Efficiency and Visual-Textual Alignment for LLM-Based Radiology Report GenerationabstractLLM-based radiology report generation (R2Gen) systems have demonstrated promising performance but face significant challenges in bridging the gap between the visual encoder and the LLM. Specifically, two issues hinder progress: (1) parameter-heavy visual projector that increases complexity and degrades performance, and (2) insufficient alignment between visual and textual modalities, limiting system efficacy. To address these, we propose R2Gen-EVA, a novel framework emphasizing Efficiency and Visual-Textual Alignment (VTA), which introduces two key innovations: (1) a parameter-free visual projector that enhances model efficiency while improving performance, and (2) an LLM-adapted VTA module that enhances the alignment of visual features with LLM’s textual embeddings. Our design significantly improves model efficacy without adding extra parameters, achieving both streamlined complexity and higher computational efficiency during inference. Extensive experiments demonstrate that R2Gen-EVA enhances the fluency and clinical accuracy of generated reports, establishing it as a more effective and efficient solution for LLM-based R2Gen. The code is available at https://github.com/zailongchen/R2Gen-EVA. Zailong Chen, Yujian Lee, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
ICME | 3 |
| 2025 | ESTI: An Efficient Spatial-Temporal Interaction Network For Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to identify the target pedestrian from video sequences. However, redundant information exist in input frames. Extracting spatial-temporal features in whole adjacent frames can introduce additional computational overhead. Furthermore, this process leads to the loss of critical spatial and temporal details, causing suboptimal representations. To mitigate these issues, we propose an Efficient Spatial-Temporal Interaction (ESTI) network, which processes half of the input sequence separately through spatial and temporal branches, extracting high-level discriminative features across multiple layers and avoiding redundancy computations. In particular, we propose a Feature Enhancement Module (FEM) for the spatial branch to focus on enhancing spatial dependencies adaptively, and a Temporal Interaction Module (TIM) for temporal branch to capture temporal correlations effectively. Spatial-temporal interaction is performed at the final layer to generate distinctive representations. Extensive experiments on three challenging video Re-ID datasets show that our ESTI achieves competitive results while maintaining low computational complexity. Guquan Jing, Yiyang Hu, Yujian Lee, Hui Zhang 0062 |
ICME | 4 |
| 2025 | Boosting Audio-Visual Segmentation via Triple-Modalities AlignmentabstractThe Audio-Visual Segmentation (AVS) task aims to identify sound-producing objects in the visual domain using auditory cues. Enhancing segmentation efficiency by incorporating prior knowledge, such as object locations and textual prompts, has proven to be crucial. However, existing methods suffer from feature misalignment during model training, leading to ineffective integration and reduced performance. To address this, we propose Triple-modalities alignment (TM-align), which combines audio signals, visual images, and textual prompts. By leveraging prompts from a frozen multi-modal large language model (MLLM), we extract two types of semantic information: contextual semantic description (C.S.D) and prompt specific summary (P.S.S). TM-align yields three pairs of aligned features: visual and C.S.D, visual and P.S.S, visual and audio, within two of our proposed cross-modalities alignment (CMA) models. To further enhance the alignment, we employ Jensen-Shannon Divergence (JSD) to regulate the domain distribution of the latter two features. By effectively aligning the three modalities, TM-align reduces redundancy and improves the overall AVS performance. Experimental results demonstrate that TM-align outperforms the mainstream AVS models.1 Yujian Lee, Zailong Chen, Wentao Fan 0001, Guquan Jing, Yiyang Hu |
ICME | 1 |
| 2025 | Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language ModelsabstractComposed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git. Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai |
ICMR | 2 |
| 2025 | MEGA-GO: functions prediction of diverse protein sequence length using Multi-scalE Graph Adaptive neural networkabstractMOTIVATION: The increasing accessibility of large-scale protein sequences through advanced sequencing technologies has necessitated the development of efficient and accurate methods for predicting protein function. Computational prediction models have emerged as a promising solution to expedite the annotation process. However, despite making significant progress in protein research, graph neural networks face challenges in capturing long-range structural correlations and identifying critical residues in protein graphs. Furthermore, existing models have limitations in effectively predicting the function of newly sequenced proteins that are not included in protein interaction networks. This highlights the need for novel approaches integrating protein structure and sequence data. RESULTS: We introduce Multi-scalE Graph Adaptive neural network (MEGA-GO), highlighting the capability of capturing diverse protein sequence length features from multiple scales. The unique graph adaptive neural network architecture of MEGA-GO enables a more nuanced extraction of graph structure features, effectively capturing intricate relationships within biological data. Experimental results demonstrate that MEGA-GO outperforms mainstream protein function prediction models in the accuracy of Gene Ontology term classification, yielding 33.4%, 68.9%, and 44.6% of area under the precision-recall curve on biological process, molecular function, and cellular component domains, respectively. The rest of the experimental results reveal that our model consistently surpasses the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code and data of MEGA-GO are available at https://github.com/Cheliosoops/MEGA-GO. Yujian Lee, Yongqi Xu, Shuaicheng Li 0001 |
Bioinform. | 1 |
| 2025 | 3D-Aided Pedestrian Representation Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match the target pedestrian from video sequences. Recent methods perform frame-level feature extraction followed by temporal aggregation to obtain video representations. However, they pay insufficient attention to the quality of frame-level features, which suffer from issues including multi-frame misalignment, partial occlusion and appearance confusion. People live in a 3D space. 3D pedestrian representations can provide rich geometric information and shape cues that offer promising solutions to these challenges in video-based Re-ID. To mitigate these issues, this paper proposes a 3D-Aid Pedestrian Representation Learning (3DAPRL) network, which introduces 3D modality to video-based Re-ID. Specifically, two novel modules are designed,i.e., the Cross-Modal Fusion (CMF) module and the Shape-aware Spatial-Temporal Interaction (SSTI) module, to enhance pedestrian representation learning. The CMF module generates discriminative fusion representations by utilizing 3D pedestrian data, while the SSTI module learns spatial-temporal 3D shape representation which are distinguishable for finding the target pedestrian in video scenarios. Both features generated from the CMF and SSTI modules contribute to the final video representation. Extensive experiments on four challenging video-based Re-ID datasets demonstrate that our 3DAPRL network reaches better performance than state-of-the-arts methods. Guquan Jing, Yujian Lee, Yiyang Hu, Hui Zhang 0062 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |