VLDB 2026 Research / reviewers in the wild / expert
Lin Li 0065
dblp:73/2252-65
· DBLP profile ↗
19ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-5678-4487ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation ComprehensionabstractRecent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic roles. The core reason is the lack of modeling for structural semantic dependencies among multi-entities, leading to over-reliance on language priors (e.g., defaulting to "person drinks a milk" if a person is merely holding it). To this end, we propose Relation-R1, the first unified relation comprehension framework that explicitly integrates cognitive chain-of-thought (CoT)-guided supervised fine-tuning (SFT) and group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we first establish foundational reasoning capabilities via SFT, enforcing structured outputs with thinking processes. Then, GRPO is utilized to refine these outputs via multi-rewards optimization, prioritizing visual-semantic grounding over language-induced biases, thereby improving generalization capability. Furthermore, we investigate the impact of various CoT strategies within this framework, demonstrating that a specific-to-general progressive approach in CoT guidance further improves generalization, especially in capturing synonymous N-ary relations. Extensive experiments on widely-used PSG and SWiG datasets demonstrate that Relation-R1 achieves state-of-the-art performance in both binary and N-ary relation understanding. Lin Li 0065, Wei Chen 0070, Jiahui Li 0003, Kwang-Ting Cheng, Long Chen 0016 |
AAAI | 1 |
| 2026 | Multi-level Compositional Feature Augmentation for Unbiased Scene Graph GenerationabstractAbstract Scene Graph Generation (SGG) aims to detect all the visual relation triplets <, , > in a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recent years. However, due to the ubiquitous long-tailed predicate distributions, today’s SGG models are still easily biased to the head predicates. Currently, the most prevalent debiasing solutions for SGG are re-balancing methods, e . g ., changing the distributions of original training samples. In this paper, we argue that all existing re-balancing strategies fail to increase the diversity of the relation triplet features of each predicate, which is critical for robust SGG. To this end, we propose a novel M ulti-level C ompositional F eature A ugmentation ( MCFA ) strategy, which aims to mitigate the bias issue from the perspective of increasing the diversity of triplet features. Specifically, we enhance relationship diversity on not only feature-level , i . e ., replacing the intrinsic or extrinsic visual features of triplets with other correlated samples to create novel feature compositions for tail predicates, but also image-level , i . e ., manipulating the image to generate brand new visual appearance for triplets. Due to its model-agnostic nature, MCFA can be seamlessly incorporated into various SGG frameworks. Extensive ablations have shown that MCFA achieves a new state-of-the-art performance on the trade-off between different metrics. Lin Li 0065, Long Chen 0016 |
Int. J. Comput. Vis. | 1 |
| 2025 | CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationabstractInterleaved image-text generation has emerged as a vital multimodal task aimed at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit narrative coherence and entity and style consistency remains challenging due to poor training data quality. To this end, we introduce CoMM, a high-quality Coherent interleaved image-text MultiModal dataset designed to enhance the coherence, consistency, and alignment of generated multimodal content. Initially, CoMM harnesses raw data from diverse sources, focusing on instructional content and visual storytelling, establishing a foundation for coherent and consistent content. To further refine the data quality, we devise a multi-perspective filter strategy that leverages advanced pre-trained models to ensure the development of sentences, consistency of inserted images, and semantic alignment between them. Various quality evaluation metrics are designed to prove the high quality of the filtered dataset. Meanwhile, extensive few-shot experiments on various downstream tasks demonstrate CoMM’s effectiveness in significantly enhancing the in-context learning capabilities of MLLMs. Moreover, we propose four new tasks to evaluate MLLMs’ interleaved generation abilities, supported by a comprehensive evaluation framework. We believe CoMM opens a new avenue for advanced MLLMs with superior multimodal in-context learning and understanding ability. Wei Chen 0070, Lin Li 0065, Yongqi Yang, Fan Yang 0094, Tingting Gao, Yu Wu 0011, Long Chen 0016 |
CVPR | 2 |
| 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward RedistributionabstractReinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences.Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase.However, current reward models operate as sequence-to-one models, allocating a single, sparse, and delayed reward to an entire output sequence.This approach may overlook the significant contributions of individual tokens toward the desired outcome.To this end, we propose a more finegrained, token-level guidance approach for RL training.Specifically, we introduce RED, a novel REward reDistribition method that evaluates and assigns specific credit to each token using an off-the-shelf reward model.Utilizing these fine-grained rewards enhances the model's understanding of language nuances, leading to more precise performance improvements.Notably, our method does not require modifying the reward model or introducing additional training steps, thereby incurring minimal computational costs.Experimental results across diverse datasets and tasks demonstrate the superiority of our approach. Jiahui Li 0003, Lin Li 0065, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011, Cheng Yang 0002 |
EMNLP | 2 |
| 2025 | Compositional Zero-shot Learning via Progressive Language-based ObservationsabstractCompositional zero-shot learning aims to recognize unseen stateobject compositions by leveraging known primitives (state and object) during training. However, effectively modeling interactions between primitives and generalizing knowledge to novel compositions remains a perennial challenge. There are two crucial factors: large object-conditioned and state-conditioned variance, i.e., the appearance of states (or objects) can vary significantly when combined with different objects (or states). For instance, the state "old" can signify vintage design for a "car" or advanced age for a "cat". In this paper, we argue that these variances can be mitigated by predicting composition categories based on salient observation cues. Therefore, we propose Progressive Language-based Observations (PLO), which can automatically determine the order of observation cues. These "observation cues" comprise a series of primitive concepts or graduated descriptions that allow the model to understand image content in a step-by-step manner. Specifically, PLO adopts pre-trained vision-language models (VLMs) to empower the model with observation capabilities.We further devise two variants: a twostep method (PLO-VLM) with a pre-observing classifier dynamically selecting the order of primitive concept-based cues, and a multistep approach (PLO-LLM) using large language models (LLMs) to craft graduated description-based cues. Extensive tests on three datasets show PLO's effectiveness in compositional recognition. Lin Li 0065, Guikun Chen, Zhen Wang 0004, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 1 |
| 2025 | Zero-shot Compositional Action Recognition with Neural Logic ConstraintsabstractZero-shot compositional action recognition (ZS-CAR) aims to identify unseen verb-object compositions in the videos by exploiting the learned knowledge of verb and object primitives during training. Despite compositional learning's progress in ZS-CAR, two critical challenges persist: 1) Missing compositional structure constraint, leading to spurious correlations between primitives; 2) Neglecting semantic hierarchy constraint, leading to semantic ambiguity and impairing the training process. In this paper, we argue that human-like symbolic reasoning offers a principled solution to these challenges by explicitly modeling compositional and hierarchical structured abstraction. To this end, we propose a logic-driven ZS-CAR framework LogicCAR that integrates dual symbolic constraints: Explicit Compositional Logic and Hierarchical Primitive Logic. Specifically, the former models the restrictions within the compositions, enhancing the compositional reasoning ability of our model. The latter investigates the semantical dependencies among different primitives, empowering the models with fine-to-coarse reasoning capacity. By formalizing these constraints in first-order logic and embedding them into neural network architectures, LogicCAR systematically bridges the gap between symbolic abstraction and existing models. Extensive experiments on the Sth-com dataset demonstrate that our LogicCAR outperforms existing baseline methods, proving the effectiveness of our logic-driven constraints. Gefan Ye, Lin Li 0065, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 2 |
| 2025 | Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph GenerationabstractOpen-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Existing OVSGG methods always adopt a two-stage pipeline: 1) Infusing knowledge into large-scale models via pre-training on large datasets; 2) Transferring knowledge from pre-trained models with fully annotated scene graphs during supervised fine-tuning. However, due to a lack of explicit interaction modeling, these methods struggle to distinguish between interacting and non-interacting instances of the same object category. This limitation induces critical issues in both stages of OVSGG: it generates noisy pseudo-supervision from mismatched objects during knowledge infusion, and causes ambiguous query matching during knowledge transfer. To this end, in this paper, we propose an interACtion-Centric end-to-end OVSGG framework (ACC) in an interaction-driven paradigm to minimize these mismatches. For interaction-centric knowledge infusion, ACC employs a bidirectional interaction prompt for robust pseudo-supervision generation to enhance the model's interaction knowledge. For interaction-centric knowledge transfer, ACC first adopts interaction-guided query selection that prioritizes pairing interacting objects to reduce interference from non-interacting ones. Then, it integrates interaction-consistent knowledge distillation to bolster robustness by pushing relational foreground away from the background while retaining general knowledge. Extensive experimental results on three benchmarks show that ACC achieves state-of-the-art performance, demonstrating the potential of interaction-centric paradigms for real-world applications. Lin Li 0065, Long Chen 0016 |
NeurIPS | 1 |
| 2025 | From Easy to Hard: Learning Curricular Shape-Aware Features for Robust Panoptic Scene Graph Generation
Hanrong Shi, Lin Li 0065, Jun Xiao 0001, Yueting Zhuang, Long Chen 0016 |
Int. J. Comput. Vis. | 2 |
| 2025 | Knowledge Integration for Grounded Situation Recognition
Jiaming Lei, Sijing Wu, Lin Li 0065, Lei Chen 0082, Jun Xiao 0001, Yi Yang 0001, Long Chen 0016 |
Pattern Recognit. | 3 |
| 2024 | Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language ExplainerabstractBenefiting from strong generalization ability, pre-trained vision-language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple recognition tasks, grounded situation recognition (GSR) requires the model not only to classify salient activity (verb) in the image, but also to detect all semantic roles that participate in the action. This complex task usually involves three steps: verb recognition, semantic role grounding, and noun recognition. Directly employing class-based prompts with VLMs and grounding models for this task suffers from several limitations, e.g., it struggles to distinguish ambiguous verb concepts, accurately localize roles with fixed verb-centric template input, and achieve context-aware noun predictions. In this paper, we argue that these limitations stem from the model's poor understanding of verb/noun classes. To this end, we introduce a new approach for zero-shot GSR via Language EXplainer (LEX), which significantly boosts the model's comprehensive capabilities through three explainers: 1) verb explainer, which generates general verb-centric descriptions to enhance the discriminability of different verb classes; 2) grounding explainer, which rephrases verb-centric templates for clearer understanding, thereby enhancing precise semantic role localization; and 3) noun explainer, which creates scene-specific noun descriptions to ensure context-aware noun recognition. By equipping each step of the GSR process with an auxiliary explainer, LEX facilitates complex scene understanding in real-world scenarios. Our extensive validations on the SWiG dataset demonstrate LEX's effectiveness and interoperability in zero-shot GSR. Jiaming Lei, Lin Li 0065, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 2 |
| 2024 | NICEST: Noisy Label Correction and Training for Robust Scene Graph GenerationabstractNearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that neither of the assumptions applies to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions and harm the training of unbiased SGG models. To this end, we propose a novelNoIsy labelCorrEction andSampleTrainingstrategy for SGG:NICEST, which rules out these noisy label issues by generating high-quality samples and designing an effective training strategy. Specifically, it consists of: 1)NICE: it detects noisy samples and then reassigns higher-quality soft predicate labels to them. To achieve this goal, NICE contains three main steps: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, it is treated as an out-of-distribution detection problem, and the pseudo labels are assigned to all detected noisy negative samples. Then, in Pos-NSD, we use a density-based clustering algorithm to detect noisy positive samples. Lastly, in NSC, we use weighted KNN to reassign more robust soft predicate labels rather than hard labels to all noisy positive samples. 2)NIST: it is a multi-teacher knowledge distillation based training strategy, which enables the model to learn unbiased fusion knowledge. A dynamic trade-off weighting strategy in NIST is designed to penalize the bias of different teachers. Due to the model-agnostic nature of both NICE and NIST, NICEST can be seamlessly incorporated into any SGG architecture to boost its performance on different predicate categories. In addition, to better assess the generalization ability of SGG models, we propose a new benchmark,VG-OOD, by reorganizing the prevalent VG dataset. This reorganization deliberately makes the predicate distributions between the training and test sets as different as possible for each subject-object category pair. This new benchmark helps disentangle the influence of subject-object category biases. Extensive ablations and results on different backbones and tasks have attested to the effectiveness and generalization ability of each component of NICEST. Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Hanwang Zhang, Yi Yang 0001, Wei Liu 0005, Long Chen 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Label Semantic Knowledge Distillation for Unbiased Scene Graph GenerationabstractThe Scene Graph Generation (SGG) task aims to detect all the objects and their pairwise visual relationships in a given image. Although SGG has achieved remarkable progress over the last few years, almost all existing SGG models follow the same training paradigm: they treat both object and predicate classification in SGG as a single-label classification problem, and the ground-truths are one-hot target labels. However, this prevalent training paradigm has overlooked two characteristics of current SGG datasets: 1) For positive samples, some specific subject-object instances may have multiple reasonable predicates. 2) For negative samples, there are numerous missing annotations. Regardless of the two characteristics, SGG models are easy to be confused and make wrong predictions. To this end, we propose a novel model-agnostic Label Semantic Knowledge Distillation (LS-KD) for unbiased SGG. Specifically, LS-KD dynamically generates a “soft” label for each subject-object instance by fusing a predicted Label Semantic Distribution (LSD) with its original one-hot target label. LSD reflects the correlations between this instance and multiple predicate categories. Meanwhile, we propose two different strategies to predict LSD: iterative self-KD and synchronous self-KD. Extensive ablations and results on three SGG tasks have attested to the superiority and generality of our proposed LS-KD, which can consistently achieve decent trade-off performance between different predicate categories. Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Wenxiao Wang 0001, Jian Shao 0001, Anan Liu, Yi Yang 0001, Long Chen 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Compositional Feature Augmentation for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to detect all the visual relation tripletsin a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recent years. However, due to the ubiquitous long-tailed predicate distributions, today’s SGG models are still easily biased to the head predicates. Currently, the most prevalent debiasing solutions for SGG are re-balancing methods, e.g., changing the distributions of original training samples. In this paper, we argue that all existing re-balancing strategies fail to increase the diversity of the relation triplet features of each predicate, which is critical for robust SGG. To this end, we propose a novel Compositional Feature Augmentation (CFA) strategy, which is the first unbiased SGG work to mitigate the bias issue from the perspective of increasing the diversity of triplet features. Specifically, we first decompose each relation triplet feature into two components: intrinsic feature and extrinsic feature, which correspond to the intrinsic characteristics and extrinsic contexts of a relation triplet, respectively. Then, we design two different feature augmentation modules to enrich the feature diversity of original relation triplets by replacing or mixing up either their intrinsic or extrinsic features from other samples. Due to its model-agnostic nature, CFA can be seamlessly incorporated into various SGG frameworks. Extensive ablations have shown that CFA achieves a new state-of-the-art performance on the trade-off between different metrics. Lin Li 0065, Guikun Chen, Jun Xiao 0001, Yi Yang 0001, Chunping Wang 0001, Long Chen 0016 |
ICCV | 1 |
| 2023 | Addressing Predicate Overlap in Scene Graph Generation with Semantic Granularity ControllerabstractSemantic overlap between predicates (e.g., riding versus on) occurs inevitably when describing a scene. However, most existing Scene Graph Generation (SGG) works sidestep it by modeling the semantic overlap at category-level and assigning merely one-hot target to each sample, which hurt the performance on other reasonable predicates. In this paper, we argue that semantic overlap between predicates tends to vary in different abstract patterns, and a subject-object pair should retain multiple reasonable predicates. To this end, we make an early attempt to reformulate SGG as a partial multi-label learning problem and accordingly propose a model-agnostic Semantic Granularity Controller (SGC). SGC consists of a pattern-specific controller, partial multi-label learning, and controllable inference. The former two solve semantic confusion during training, while the latter makes the semantic granularity of prediction controllable. Extensive experiments demonstrate that SGC can improve the performance of SGG and guide the model to predict coarse/fine-grained predicates. Guikun Chen, Lin Li 0065, Yawei Luo, Jun Xiao 0001 |
ICME | 2 |
| 2023 | Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsabstractPretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs within an image. However, naively utilizing CLIP with prevalent class-based prompts for zero-shot VRD has several weaknesses, e.g., it struggles to distinguish between different fine-grained relation types and it neglects essential spatial information of two objects. To this end, we propose a novel method for zero-shot VRD: RECODE, which solves RElation detection via COmposite DEscription prompts. Specifically, RECODE first decomposes each predicate category into subject, object, and spatial components. Then, it leverages large language models (LLMs) to generate description-based prompts (or visual cues) for each component. Different visual cues enhance the discriminability of similar relation categories from different perspectives, which significantly boosts performance in VRD. To dynamically fuse different cues, we further introduce a chain-of-thought method that prompts LLMs to generate reasonable weights for different visual cues. Extensive experiments on four VRD benchmarks have demonstrated the effectiveness and interpretability of RECODE. Lin Li 0065, Jun Xiao 0001, Guikun Chen, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
NeurIPS | 1 |
| 2023 | Question-guided feature pyramid network for medical visual question answering
Yonglin Yu, Hanrong Shi, Lin Li 0065, Jun Xiao 0001 |
Expert Syst. Appl. | 4 |
| 2022 | The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationabstractUnbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that both assumptions are inapplicable to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions, and these noisy samples actually harm the training of unbiased SGG models. To this end, we propose a novel model-agnostic NoIsy label CorrEction strategy for SGG: NICE. NICE can not only detect noisy samples but also reassign more high-quality predicate labels to them. After the NICE training, we can obtain a cleaner version of SGG dataset for model training. Specifically, NICE consists of three components: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, we formulate this task as an out-of-distribution detection problem, and assign pseudo labels to all detected noisy negative samples. Then, in Pos-NSD, we use a clustering-based algorithm to divide all positive samples into multiple sets, and treat the samples in the noisiest set as noisy positive samples. Lastly, in NSC, we use a simple but effective weighted KNN to reassign new predicate labels to noisy positive samples. Extensive results on different backbones and tasks have attested to the effectiveness and generalization abilities of each component of NICE. Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jun Xiao 0001 |
CVPR | 1 |
| 2021 | Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based ExplanationabstractInterpreting model knowledge is an essential topic to improve human understanding of deep black-box models. Traditional methods contribute to providing intuitive instance-wise explanations which allocating importance scores for low-level features (e.g, pixels for images). To adapt to the human way of thinking, one strand of recent researches has shifted its spotlight to mining important concepts. However, these concept-based interpretation methods focus on computing the contribution of each discovered concept on the class level and can not precisely give instance-wise explanations. Besides, they consider each concept as an independent unit, and ignore the interactions among concepts. To this end, in this paper, we propose a novel COncept-based NEighbor Shapley approach (dubbed as CONE-SHAP) to evaluate the importance of each concept by considering its physical and semantic neighbors, and interpret model knowledge with both instance-wise and class-wise explanations. Thanks to this design, the interactions among concepts in the same image are fully considered. Meanwhile, the computational complexity of Shapley Value is reduced from exponential to polynomial. Moreover, for a more comprehensive evaluation, we further propose three criteria to quantify the rationality of the allocated contributions for the concepts, including coherency, complexity, and faithfulness. Extensive experiments and ablations have demonstrated that our CONE-SHAP algorithm outperforms existing concept-based methods and simultaneously provides precise explanations for each instance and class. Jiahui Li 0003, Kun Kuang 0001, Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 3 |
| 2021 | Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video ApplicationsabstractWe present a self-supervised spatiotemporal learning approach by exploring the temporal coherence of videos. The chronological order of shuffled clips from the video is used as the supervisory signal to guide the 3D Convolutional Neural Networks (CNNs) to learn meaningful visual knowledge. Unlike the existing approaches which use frames, we utilize dynamic video clips to reduce the uncertainty of order. We test three types of representative 3D CNNs, all of which benefit from the proposed approach. The learned 3D CNNs can be used either as a feature extractor or a pre-trained model for further fine-tuning on downstream tasks. We also propose two curriculum learning strategies to make the 3D CNNs easier to train and get the state-of-the-art results in nearest neighbor retrieval and action recognition tasks compared with other self-supervised learning methods. Meanwhile, it is further extended to the field of visual question answering application and has achieved promising results. Besides, comprehensive and extensive experimental results and analyses are provided for readers to better understand the video clip order we explore with self-supervised and curriculum learning for video application. Jun Xiao 0001, Lin Li 0065, Dejing Xu, Chengjiang Long, Jian Shao 0001, Shiliang Pu, Yueting Zhuang |
IEEE Trans. Multim. | 2 |