EDBT 2026 Demo / reviewers in the wild / expert
Xiaoye Qu
dblp:229/8206
· DBLP profile ↗
80ranked-venue papers
7as first author
74since 2021 · last 2026
0000-0002-4907-3978ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 3 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 36 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Video-Language Model from the Language Input PerspectiveabstractDriven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefined by the specific template. In real-world applications, such a strict assumption is impossible to satisfy since 1) predefining all the texts is extremely time-consuming and labor-intensive. 2) these predefined text inputs are too restrictive and user-unfriendly, limiting their applications. It is observed that given a video input, texts with similar semantics but different templates lead to various performances. To this end, in this paper, we propose a novel plug-and-play framework for various VLM-based methods to fully bridge videos and texts. Specifically, we first generate positive and negative texts from the original ones to target specific text components. Then, we propose an attribute-based text reasoning strategy to mine fine-grained textual semantics of generated texts. Finally, we utilize videos as guidance to conduct cross-modal bridging by designing a self-weighted loss. Extensive experiments show that the proposed method can serve as the plug-and-play module to effectively improve the performance of state-of-the-art VLMs. Wanlong Fang, Changshuo Wang 0001, Xiaoye Qu, Daizong Liu |
AAAI | 4 |
| 2026 | Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect such realistic conflict scenarios. Most focus solely on intra-memory conflicts, while context-memory and inter-context conflicts remain largely unaddressed. Furthermore, commonly used factual knowledge-based evaluations are often overlooked, and existing datasets lack a thorough investigation into conflict detection capabilities.To bridge this gap, we propose MMKC-Bench, a benchmark designed to evaluate factual knowledge conflicts in both context-memory and inter-context scenarios. MMKC-Bench encompasses four types of multimodal knowledge conflicts and includes 1,881 knowledge instances and 3,997 images across 32 broad types, collected through automated pipelines with human verification. We evaluate four representative series of LMMs on both model behavior analysis and conflict detection tasks. Our findings show that while current LMMs are capable of recognizing knowledge conflicts, they tend to favor internal parametric knowledge over external evidence. We hope MMKC-Bench will foster further research in multimodal knowledge conflict and enhance the development of multimodal RAG systems. Yuntao Du 0001, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin 0003, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang 0002, Xiaoye Qu, Qian Li 0043, Dongrui Liu |
AAAI | 12 |
| 2026 | Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning ModelsabstractInstruction-following is essential for aligning large language models (LLMs) with user intent.Yet recent reasoning-oriented models, despite their strong performance on complex mathematical problems, often fail to comply with simple natural language directives.In this work, we analyze the interaction between reasoning ability and instruction adherence in large reasoning models (LRMs).Using a controlled evaluation framework (MathIF), we uncover a persistent trade-off: as models scale reasoning capacity through long chains-of-thought or reinforcement learning on reasoning traces, their obedience to instructions degrades, particularly when generation length grows.We further show that interventions such as constraining or repeating instructions can partially restore compliance, but typically at the expense of reasoning performance.Taken together, our findings expose a dilemma between intelligence and obedience in current training paradigms and underscore the need for instruction-aware approaches to developing controllable reasoning models. Tingchen Fu, Yafu Li, Jiawei Gu, Xiaoye Qu, Yu Cheng 0001 |
ACL (1) | 4 |
| 2026 | Generating transferable attacks across large vision-language models using adversarial deformation learning
Daizong Liu, Wangqin Liu, Xiaowen Cai 0001, Pan Zhou 0001, Runwei Guan, Xiaoye Qu, Bo Du 0001 |
Pattern Recognit. | 6 |
| 2026 | Are Large Vision-Language Models Robust to Adversarial Visual Transformations?abstractLarge vision-language models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding and reasoning tasks. However, recent research shows that LVLMs are susceptible to adversarial examples. Existing attackers either optimize the perturbations on the visual input or manipulate prompts to fool the LVLM models, requiring extensive design and engineering on these adversarial manipulations. While straightforward visual transformation can boast training generalization-ability, its potential risks to LVLMs in terms of safety and trustworthiness have been largely neglected. In this paper, we ask an intriguing question:can simple yet easy-to-implement adversarial visual transformations be utilized to attack the LVLM models?Motivated by this research gap and new attack setting, we propose the first comprehensive assessment of LVLMs’ adversarial robustness to visual transformations by testing LVLMs’ resilience to all possible transformation operations. Our empirical observations suggest that with the appropriate combination of the most harmful transformations, we can build transformation-based attacks more adversarial to the LVLM models. Moreover, adversarial learning of visual transformations is further introduced to adaptively apply the malicious impacts of all potentially harmful transformations to the raw images via gradient approximation for improving the attack effectiveness and imperceptibility. We hope that this study can provide deeper insights into the potential vulnerability of LVLMs to adversarial visual transformations. Daizong Liu, Xiaowen Cai 0001, Pan Zhou 0001, Xiaoye Qu, Lichao Sun 0001, Wei Hu 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Mitigating Multilingual Hallucination in Large Vision-Language ModelsabstractWhile Large Vision-Language Models (LVLMs) have exhibited remarkable capabilities across a wide range of tasks, they suffer from hallucination problems, where models generate plausible yet incorrect answers given the input image-query pair. This hallucination phenomenon is even more severe when querying the image in non-English languages, while existing methods for mitigating hallucinations in LVLMs only consider the English scenarios. In this article, we make the first attempt to mitigate this important multilingual hallucination in LVLMs. With thorough experimental analysis, we found that multilingual hallucination in LVLMs is a systemic problem that could arise from deficiencies in multilingual capabilities or inadequate multimodal abilities. To this end, we propose a two-stage Multilingual Hallucination Removal (MHR) framework for LVLMs, aiming to improve resistance to hallucination for both high-resource and low-resource languages. Specifically, in the first stage, considering that most non-English languages cannot follow instructions well and output non-sense answers given the input image, we boost multilingual instruction-following ability with a multilingual supervised fine-tuning. The second phase is aimed at enhancing the LVLM’s ability to diminish multilingual hallucinations. Instead of relying on the intricate manual annotations of multilingual resources, we fully leverage the inherent capabilities of the LVLM and propose a novel cross-lingual alignment method, which generates multiple responses for each image-query input and then identifies the hallucination-aware pairs for each language. These data pairs are finally used for direct preference optimization to prompt the LVLMs to favor non-hallucinating responses. Experimental results show that our MHR achieves a substantial reduction in hallucination generation for LVLMs. Our code and model weights are available at https://github.com/ssmisya/MHR . Xiaoye Qu, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory PerspectiveabstractDespite the remarkable success of attentionbased large language models (LLMs), the precise interaction mechanisms between attention heads remain poorly understood.In contrast to prevalent methods that focus on individual head contributions, we rigorously analyze the intricate interplay among attention heads through a novel framework based on the Harsanyi dividend, a concept from cooperative game theory.Our analysis reveals that significant positive Harsanyi dividends are sparsely distributed across head combinations, indicating that most heads do not contribute cooperatively.Moreover, certain head combinations exhibit negative dividends, indicating implicit competitive relationships.To further optimize the interactions among attention heads, we propose a training-free Game-theoretic Attention Calibration (GAC) method.Specifically, GAC selectively retains heads demonstrating significant cooperative gains and applies fine-grained distributional adjustments to the remaining heads.Comprehensive experiments across 17 benchmarks demonstrate the effectiveness of our proposed GAC and its superior generalization capabilities across diverse model families, scales, and modalities.The source code is available at Xiaoye Qu, Zengqi Yu, Dongrui Liu, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
ACL (1) | 1 |
| 2025 | PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsabstractProcess-level Reward Models (PRMs) are crucial for complex reasoning and decisionmaking tasks, where each intermediate step plays an important role in the reasoning process.Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios.However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically.To address this gap, we introduce PRMBENCH, a processlevel benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs.PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity.In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs.These findings reveal the challenges inherent in processlevel evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development. Zhaochen Su, Xiaoye Qu, Yu Cheng 0001 |
ACL (1) | 3 |
| 2025 | Multi-level Association Refinement Network for Dialogue Aspect-based Sentiment Quadruple AnalysisabstractDialogue Aspect-based Sentiment Quadruple (DiaASQ) analysis aims to identify all quadruples (i.e., ) from the dialogue.This task is challenging as different elements within a quadruple may manifest in different utterances, requiring precise handling of associations at both the utterance and word levels.However, most existing methods tackling it predominantly leverage predefined dialogue structure (e.g., reply) and word semantics, resulting in a surficial understanding of the deep sentiment association between utterances and words.In this paper, we propose a novel Multi-level Association Refinement Network (MARN) designed to achieve more accurate and comprehensive sentiment associations between utterances and words.Specifically, for utterances, we dynamically capture their associations with enriched semantic features through a holistic understanding of the dialogue, aligning them more closely with sentiment associations within elements in quadruples.For words, we develop a novel crossutterance syntax parser (CU-Parser) that fully exploits syntactic information to enhance the association between word pairs within and across utterances.Moreover, to address the scarcity of labeled data in DiaASQ, we further introduce a multi-view data augmentation strategy to enhance the performance of MARN under low-resource conditions.Experimental results demonstrate that MARN achieves state-of-theart performance and maintains robustness even under low-resource conditions. Self-Attention Zeliang Tong, Wei Wei 0002, Xiaoye Qu, Rikui Huang, Zhixin Chen |
ACL (1) | 3 |
| 2025 | Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path ReasoningabstractRecently, Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multi-modal context comprehension. However, they still suffer from hallucination problems referring to generating inconsistent outputs with the image content. To mitigate hallucinations, previous studies mainly focus on retraining LVLMs with custom datasets. Although effective, they inherently come with additional computational costs. In this paper, we propose a training-free framework, MVP, that aims to reduce hallucinations by making the most of the innate capabilities of the LVLMs via Multi-View Multi-Path Reasoning. Specifically, we first devise a multi-view information-seeking strategy to thoroughly perceive the comprehensive information in the image, which enriches the general global information captured by the original vision encoder in LVLMs. Furthermore, during the answer decoding, we propose multi-path reasoning for each information view to quantify and aggregate the certainty scores for each potential answer among multiple decoding paths and finally decide the output answer. By fully grasping the information in the image and carefully considering the certainty of the potential answers when decoding, our MVP can effectively reduce hallucinations in LVLMs. The extensive experiments verify that our proposed MVP significantly mitigates the hallucination problem across four well-known LVLMs. Xiaoye Qu, Jiashuo Sun, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
COLING | 1 |
| 2025 | From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data CalibrationabstractLarge Vision-Language Models (LVLMs) have achieved significant progress in combining visual comprehension with language generation. Despite this success, the training data of LVLMs still suffers from Long-Tail (LT) problems, where the data distribution is highly imbalanced. Previous works have mainly focused on traditional VLM architectures, i.e., CLIP or ViT, and specific tasks such as recognition and classification. Nevertheless, the exploration of LVLM (e.g. LLaVA) and more general tasks (e.g. Visual Question Answering and Visual Reasoning) remains under-explored. In this paper, we first conduct an in-depth analysis of the LT issues in LVLMs and identify two core causes: the overrepresentation of head concepts and the underrepresentation of tail concepts. Based on the above observation, we propose an Adaptive Data Refinement Framework (ADR), which consists of two stages: Data Rebalancing (DR) and Data Synthesis (DS). In the DR stage, we adaptively rebalance the redundant data based on entity distributions, while in the DS stage, we leverage Denoising Diffusion Probabilistic Models (DDPMs) and scarce images to supplement under-represented portions. Through comprehensive evaluations across eleven benchmarks, our proposed ADR effectively mitigates the long-tail problem in the training data, improving the average performance of LLaVA 1.5 relatively by 4.36%, without increasing the training data volume. Xiaoye Qu, Yu Cheng 0001 |
CVPR | 2 |
| 2025 | Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You ThinkabstractImage-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance of the images. However, current I2V diffusion models (I2V-DMs) often produce videos with limited motion degrees or exhibit uncontrollable motion that conflicts with the textual condition. To address these limitations, we propose a novel Extrapolating and Decoupling framework, which introduces model merging techniques to the I2V domain for the first time. Specifically, our framework consists of three separate stages: (1) Starting with a base I2V-DM, we explicitly inject the textual condition into the temporal module using a lightweight, learnable adapter and fine-tune the integrated model to improve motion controllability. (2) We introduce a training-free extrapolation strategy to amplify the dynamic range of the motion, effectively reversing the fine-tuning process to enhance the motion degree significantly. (3) With the above two-stage models excelling in motion controllability and degree, we decouple the relevant parameters associated with each type of motion ability and inject them into the base I2V-DM. Since the I2V-DM handles different levels of motion controllability and dynamics at various denoising time steps, we adjust the motion-aware parameters accordingly over time. Extensive qualitative and quantitative experiments have been conducted to demonstrate the superiority of our framework over existing methods. Code is available at https://github.com/Chuge0335/EDG Xiaoye Qu, Zhenyi Lu, Wei Wei 0002, Yu Cheng 0001 |
CVPR | 2 |
| 2025 | CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingabstractContrastive Language-Image Pre-training (CLIP) has become a cornerstone in multimodal intelligence.However, recent studies discovered that CLIP can only encode one aspect of the feature space, leading to substantial information loss and indistinctive features.To mitigate this issue, this paper introduces a novel strategy that fine-tunes a series of complementary CLIP models and transforms them into a CLIP-MoE.Specifically, we propose a model-agnostic Diversified Multiplet Upcycling (DMU) framework for CLIP.Instead of training multiple CLIP models from scratch, DMU leverages a pre-trained CLIP and fine-tunes it into a diverse set with highly cost-effective multistage contrastive learning, thus capturing distinct feature subspaces efficiently.To fully exploit these fine-tuned models while minimizing computational overhead, we transform them into a CLIP-MoE, which dynamically activates a subset of CLIP experts, achieving an effective balance between model capacity and computational cost.Comprehensive experiments demonstrate the superior performance of CLIP-MoE across various zero-shot retrieval, zero-shot image classification tasks, and downstream Multimodal Large Language Model (MLLM) benchmarks when used as a vision encoder.Code is available at https: //github.com/OpenSparseLLMs/CLIP-MoE. Jihai Zhang 0002, Xiaoye Qu, Tong Zhu 0002, Yu Cheng 0001 |
EMNLP | 2 |
| 2025 | Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral Constraints
Guanjie Chen, Yucheng Zhou 0001, Xiaoye Qu, Tianlong Chen 0001, Yu Cheng 0001 |
ICCV | 4 |
| 2025 | LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation Learning
Jianfeng Dong, Daizong Liu, Jie Sun 0034, Xiaoye Qu, Xun Yang 0001, Dongsheng Liu 0003, Xun Wang 0007 |
ICCV | 5 |
| 2025 | Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentabstractWhile Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current methods optimize LoRA by initializing with static singular value decomposition (SVD) subsets, leading to suboptimal leveraging of pre-trained knowledge. Another path for improving LoRA is incorporating a Mixture-of-Experts (MoE) architecture. However, weight misalignment and complex gradient dynamics make it challenging to adopt SVD prior to the LoRA MoE architecture. To mitigate these issues, we propose Great LoRA Mixture-of-Expert (GOAT), a framework that (1) adaptively integrates relevant priors using an SVD-structured MoE, and (2) aligns optimization with full fine-tuned MoE by deriving a theoretical scaling factor. We demonstrate that proper scaling, without modifying the architecture or training algorithms, boosts LoRA MoE’s efficiency and performance. Experiments across 25 datasets, including natural language understanding, commonsense reasoning, image classification, and natural language generation, demonstrate GOAT’s state-of-the-art performance, closing the gap with Full FT. Our code is available at: https://github.com/Facico/GOAT-PEFT Chenghao Fan, Zhenyi Lu, Chengfeng Gu, Xiaoye Qu, Wei Wei 0002, Yu Cheng 0001 |
ICML | 5 |
| 2025 | Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningabstractWhile showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in sparse-reward scenarios. Inspired by the divide-and-conquer principle, we propose an innovative framework GLIDER (Grounding Language Models as EffIcient Decision-Making Agents via Offline HiErarchical Reinforcement Learning) that introduces a parameter-efficient and generally applicable hierarchy to LLM policies. We develop a scheme where the low-level controller is supervised with abstract, step-by-step plans that are learned and instructed by the high-level policy. This design decomposes complicated problems into a series of coherent chain-of-thought reasoning sub-tasks, providing flexible temporal abstraction to significantly enhance exploration and learning for long-horizon tasks. Furthermore, GLIDER facilitates fast online adaptation to non-stationary environments owing to the strong transferability of its task-agnostic low-level skills. Experiments on ScienceWorld and ALFWorld benchmarks show that GLIDER achieves consistent performance gains, along with enhanced generalization capabilities. Zican Hu, Wei Liu 0131, Xiaoye Qu, Xiangyu Yue 0001, Chunlin Chen 0001, Zhi Wang 0001, Yu Cheng 0001 |
ICML | 3 |
| 2025 | Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackabstractLarge language models (LLMs) have presented impressive performance but often lack the flexibility to adapt to human preferences quickly without retraining. Inspired by the recent efforts on test-time scaling, we make the first attempt to propose Test-time Preference Optimization (TPO), a framework that aligns LLM outputs with human preferences during inference, eliminating the need to update model parameters. Instead of relying on purely numerical rewards, TPO translates reward signals into \emph{textual} critiques and uses them as textual rewards to iteratively refine its response. Evaluations on benchmarks covering instruction following, preference alignment, safety, and mathematics reveal that TPO progressively improves alignment with human preferences. Notably, after only a few TPO steps, the initially unaligned Llama-3.1-70B-SFT model can surpass the aligned counterpart, Llama-3.1-70B-Instruct. Furthermore, TPO scales efficiently with both the search width and depth of the inference process. Through case studies, we illustrate how TPO exploits the innate capacity of LLM to interpret and act upon reward signals. Our findings establish TPO as a practical, lightweight alternative for test-time preference optimization, achieving alignment on the fly. Yafu Li, Xuyang Hu, Xiaoye Qu, Yu Cheng 0001 |
ICML | 3 |
| 2025 | Exploring Disentangled Appearance-Motion Contexts for Temporal Activity LocalizationabstractTemporal Activity Localization (TAL) is crucial and fundamental for multimedia understanding. Although many works have made great efforts and achieved significant progress on this task, most of them directly utilize the mixed visual features extracted by the 3D backbone network to match with the complicated query semantic, thus failing to capture the subtly distinct visual features associated with the interested entities or events for better activity modeling. To overcome this challenge, in this paper, we present a novel Disentangled Appearance-Motion Learning (DAML) framework that is able to learn the disentangled representations and capture finer levels of granularity across different modalities, such as nouns-related visual appearance or verbs-related visual motion for more interpretable cross-modal alignment. Specifically, without introducing any large feature extraction model, we disentangle the mixed video feature extracted by 3D backbone into separate appearance and motion contexts with the help of vector quantization. In this way, we can achieve more fine-grained correspondence between the visual appearance and textual nouns, visual motion and textual verbs for better modeling the object entities, events of the target activity. Extensive experiments on three challenging datasets (Charades-STA, TACoS and ActivityNet) show the effectiveness of DAML. Huashuo Lei, Xiaowen Cai 0001, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Jixiang Yu, Keyan Jin |
IJCNN | 5 |
| 2025 | MonoAttack: A Strong Attack Framework with Depth-Migration and Attribute-Tampering for Monocular 3D Object DetectionabstractAlthough many efforts have been made into attacks on deep neural networks (DNNs) in recent years, no research explores the vulnerability of monocular 3D object detection (M3D) models. This M3D task is fundamental but essential in safety-critical 3D applications, potentially bringing hazards to autonomous driving. In this paper, we thoroughly investigate the sensitivity of current M3D models to adversarial noise and propose a novel M3D adversarial attack method called MonoAttack. The key insight of our method is exploring both depth-migration and attribute-tampering for generating M3D adversarial samples. Specifically, in addition to the general misleading of the detection model, we deceive the M3D model by changing the potential object depth into its opposite position. We also guide the M3D model to mis-recognize the class attribute of its detected object for generating low-confidence bounding boxes. Moreover, we further disentangle the depth knowledge from the geometric and semantic perspectives to auxiliary correlate the detection and attribute information for jointly generating the latent perturbation. In this manner, our attack framework is strong and can effectively attack M3D models with trivial perturbations. Experimental results on the KITTI dataset demonstrate that our attack achieves high adversarial ability against current monocular 3D detection models. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 4 |
| 2025 | Manipulating the Bounding Box: Multimodal Controlled Backdoor Attacks on 3D Visual Grounding Modelsabstract3D visual grounding models, pivotal in interpreting and aligning objects within 3D spaces with textual descriptions, have become integral to the advancement of the multimedia community. As these models are widely used in daily life as real-world applications, they become more susceptible to be attacked. Backdoor attacks are designed to corrupt a model in such a way that it responds with adversary-wanted outputs when specific trigger patterns are introduced, while responding normally to clean inputs. Unlike traditional backdoor attack methods that focus on attacking simple classification models, attacking 3D visual grounding models presents unique challenges due to their multi-modal inputs and the nature of their output, which is the localization box of objects described by the text within the 3D scene. This necessitates distinct attack strategies and trigger designs, adding complexity to executing successful attacks. To this end, in this paper, we present a novel multimodal controlled backdoor attack aimed at manipulating the positioning and size of bounding boxes in the challenging multi-modal 3D visual grounding models. Specifically, we design triggers for both point cloud and textual modalities, along with specialized placement strategies for each, to enhance the stealth and precision of the attack. Furthermore, we develop optimization strategies to enhance the efficacy of the point cloud trigger. Experimental results across various standard models confirm the effectiveness of our backdoor attack method, with negligible impact on performance in clean datasets. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 4 |
| 2025 | Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
Junan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu, Xun Yang 0001, Jixiang Zhu, Sanyuan Zhang, Jianfeng Dong |
ACM Multimedia | 4 |
| 2025 | Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-ExpertsabstractTong Zhu, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tong Zhu 0002, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsabstractAlthough Large Vision-Language Models (LVLMs) exhibit impressive multimodal capabilities, their vulnerability to adversarial examples has raised serious security concerns.
Existing LVLM attackers simply optimize adversarial images that easily overfit a certain model/prompt, making them ineffective once they are transferred to attack a different model/prompt.
Motivated by this research gap, this paper aims to develop a more powerful attack that is transferable to black-box LVLM models of different structures and task-aware prompts of different semantics.
Specifically, we introduce a new perspective of information theory to investigate LVLMs' transferable characteristics by exploring the relative dependence between outputs of the LVLM model and input adversarial samples. Our empirical observations suggest that enlarging/decreasing the mutual information between outputs and the disentangled adversarial/benign patterns of input images helps to generate more agnostic perturbations for misleading LVLMs' perception with better transferability.
In particular, we formulate the complicated calculation of information gain as an estimation problem and incorporate such informative constraints into the adversarial learning process.
Extensive experiments on various LVLM models/prompts demonstrate our significant transfer-attack performance. Xiaowen Cai 0001, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Keke Tang, Pan Zhou 0001, Lichao Sun 0001, Wei Hu 0003 |
NeurIPS | 3 |
| 2025 | Learning to Reason under Off-Policy GuidanceabstractRecent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(RLVR).
However, existing RLVR approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities.
To address this issue, we introduce LUFFY (Learning to reason Under oFF-policY guidance), a framework that augments RLVR with off-policy reasoning traces.
LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training.
Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training.
Compared with previous RLVR methods, LUFFY achieves an over +6.4 average gain across six math benchmarks and an advantage of over +6.2 points in out-of-distribution tasks.
Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR. Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang 0001, Ganqu Cui, Xiaoye Qu, Yu Cheng 0001, Yue Zhang 0004 |
NeurIPS | 6 |
| 2025 | Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsabstractAlthough Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt pair or fixed training dataset, which often leads to overfitting. Consequently, these perturbations fail to remain malicious once transferred to attack unseen image-prompt pairs, suffering from significant resource costs to cover the diverse multimodal inputs in complicated real-world scenarios. To alleviate this issue, this paper proposes a novel adversarial attack on MLLMs based on distribution approximation theory, which models the potential image-prompt input distribution and adds the same distribution-fitting adversarial perturbation on multimodal input pairs to achieve effective cross-image/prompt transfer attacks. Specifically, we exploit the Laplace approximation to model the Gaussian distribution of the image and prompt inputs for the MLLM, deriving an estimate of the mean and covariance parameters. By sampling from this approximated distribution with Monte Carlo mechanism, we efficiently optimize and fit a single input‑agnostic perturbation over diverse image‑prompt pairs, yielding strong universality and transferability. Extensive experiments are conducted to verify the strong adversarial capabilities of our proposed attack against prevalent MLLMs spanning a spectrum of images/prompts. Hai Yan, Haijian Ma, Xiaowen Cai 0001, Daizong Liu, Zenghui Yuan, Xiaoye Qu, Jianfeng Dong, Runwei Guan, Hongyang He, Yulai Xie 0002, Pan Zhou 0001 |
NeurIPS | 6 |
| 2025 | Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge InfusionabstractAttribute-Specific Fashion Retrieval (ASFR) focuses on retrieving images based on fine-grained, attribute-specific criteria rather than naive global visual similarity, enabling more precise and interpretable search results. Existing ASFR methods ideally assume that all attribute semantics are in-domain distributions of the training datasets. However, realistic scenarios are generally more complex and naturally contain unseen attribute information, often resulting in ungeneralizable retrieval outcomes. In this paper, we take the first step to address the new and challenging open-world ASFR setting, which involves handling diverse and practical attributes instead of relying solely on predefined attribute sets in closed-world scenarios. Specifically, to comprehend unseen attributes, we propose a novel LLM-based Commonsense Knowledge Infusion (CoKi) framework that integrates commonsense knowledge as complementary context into attribute representations using a Large Language Model (LLM). By infusing such LLM-based commonsense knowledge through descriptive contexts, our method enables robust semantic enrichment and effective generalization to unseen attributes. Additionally, we introduce a modality-switchable prompt and an imputation mechanism to ensure model robustness across diverse input configurations by dynamically adapting to missing modalities. Extensive experiments demonstrate that our approach not only achieves state-of-the-art in-domain retrieval performance but also significantly enhances adaptability to unseen attributes and cross-domain generalization, establishing a new benchmark for fine-grained fashion retrieval in open-world scenarios. Our source code is publicly available at https://github.com/HuiGuanLab/CoKi. Jianfeng Dong, Daizong Liu, Xiaoye Qu, Cuizhu Bao, Zhike Han, Jixiang Zhu, Xun Wang 0007 |
SIGIR | 4 |
| 2025 | Path-Aware Reasoning Network for document-level relation extraction with co-regularization loss
Xuhua Ai, Yiting Yu, Xiaoye Qu, Wei Wei 0002 |
Knowl. Based Syst. | 4 |
| 2025 | A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future TrendsabstractWith the significant development of large models in recent years, large vision-language models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding and reasoning tasks. Compared with traditional large language models (LLMs), LVLMs present great potential and challenges due to their closer proximity to the multiresource real-world applications and the complexity of multimodal processing. However, the vulnerability of LVLMs is relatively underexplored, posing potential security risks in the daily use of LVLM applications. In this article, we provide a comprehensive review of the various forms of existing LVLM attacks. Specifically, we first introduce the background of attacks targeting LVLMs, including the attack preliminary, attack challenges, and attack resources. Then, we systematically review the development of LVLM attack methods, such as adversarial attacks that manipulate model outputs, jailbreak attacks that exploit model vulnerabilities for unauthorized actions, prompt injection attacks that engineer the prompt type and pattern, and data poisoning that affects model training. Finally, we discuss promising future research directions in LVLM attacks. We believe that our survey provides insights into the current landscape of LVLM vulnerabilities, inspiring more researchers to explore and mitigate potential safety issues in LVLM developments. Daizong Liu, Xiaoye Qu, Pan Zhou 0001, Yu Cheng 0001, Wei Hu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Alleviating Hallucination in Large Vision-Language Models with Active Retrieval AugmentationabstractDespite the remarkable ability of Large Vision-Language Models (LVLMs) in image comprehension, these models frequently generate plausible yet factually incorrect responses, a phenomenon known as hallucination. Recently, in Large Language Models (LLMs), augmenting LLMs by retrieving information from external knowledge resources has been proven as a promising solution to mitigate hallucinations. However, the retrieval augmentation in LVLM significantly lags behind the widespread applications of LVLM. Moreover, when transferred to augmenting LVLMs, sometimes the hallucination degree of the model is even exacerbated. Motivated by the research gap and counter-intuitive phenomenon, we introduce a novel framework, the Active Retrieval-Augmented (ARA) LVLM, specifically designed to address hallucinations by incorporating three critical dimensions: (i) dissecting the retrieval targets based on the inherent hierarchical structures of images; (ii) pinpointing the most effective retrieval methods and filtering out the reliable retrieval results; and (iii) timing the retrieval process to coincide with episodes of low certainty, while circumventing unnecessary retrieval during periods of high certainty. To assess the capability of our proposed ARA model in reducing hallucination, we employ three widely used LVLM models (LLaVA-1.5, Qwen-VL, and mPLUG-Owl2) across four benchmarks. Our empirical observations suggest that by utilizing fitting retrieval mechanisms and timing the retrieval judiciously, we can effectively mitigate the hallucination problem. We hope that this study can provide deeper insights into how to adapt the retrieval augmentation to LVLMs for reducing hallucinations with more effective retrieval and minimal retrieval occurrences. Xiaoye Qu, Qiyuan Chen 0003, Wei Wei 0002, Jiashuo Sun, Daizong Liu, Jianfeng Dong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Enhancing Low-Resource Relation Representations through Multi-View DecouplingabstractRecently, prompt-tuning with pre-trained language models (PLMs) has demonstrated the significantly enhancing ability of relation extraction (RE) tasks. However, in low-resource scenarios, where the available training data is scarce, previous prompt-based methods may still perform poorly for prompt-based representation learning due to a superficial understanding of the relation. To this end, we highlight the importance of learning high-quality relation representation in low-resource scenarios for RE, and propose a novel prompt-based relation representation method, named MVRE (Multi-View Relation Extraction), to better leverage the capacity of PLMs to improve the performance of RE within the low-resource prompt-tuning paradigm. Specifically, MVRE decouples each relation into different perspectives to encompass multi-view relation representations for maximizing the likelihood during relation inference. Furthermore, we also design a Global-Local loss and a Dynamic-Initialization method for better alignment of the multi-view relation-representing virtual words, containing the semantics of relation labels during the optimization learning process and initialization. Extensive experiments on three benchmark datasets show that our method can achieve state-of-the-art in low-resource settings. Chenghao Fan, Wei Wei 0002, Xiaoye Qu, Zhenyi Lu, Wenfeng Xie, Yu Cheng 0001, Dangyang Chen |
AAAI | 3 |
| 2024 | Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationabstractTemporal sentence localization (TSL) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant yet expensive manual annotations for training. Moreover, these trained data-dependent models usually can not generalize well to unseen scenarios because of the inherent domain shift. To facilitate this issue, in this paper, we target another more practical but challenging setting: unsupervised domain adaptative temporal sentence localization (UDA-TSL), which explores whether the localization knowledge can be transferred from a fully-annotated data domain (source domain) to a new unannotated data domain (target domain). Particularly, we propose an effective and novel baseline for UDA-TSL to bridge the multi-modal gap across different domains and learn the potential correspondence between the video-query pairs in target domain. We first develop separate modality-specific domain adaptation modules to smoothly balance the minimization of the domain shifts in cross-dataset video and query domains. Then, to fully exploit the semantic correspondence of both modalities in target domain for unsupervised localization, we devise a mutual information learning module to adaptively align the video-query pairs which are more likely to be relevant in target domain, leading to more truly aligned target pairs and ensuring the discriminability of target features. In this way, our model can learn domain-invariant and semantic-aligned cross-modal representations. Three sets of migration experiments show that our model achieves competitive performance compared to existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Yang Yang 0002, Pan Zhou 0001, Yu Cheng 0001 |
AAAI | 3 |
| 2024 | Confidence is not Timeless: Modeling Temporal Validity for Rule-based Temporal Knowledge Graph ForecastingabstractRecently, Temporal Knowledge Graph Forecasting (TKGF) has emerged as a pivotal domain for forecasting future events.Unlike black-box neural network methods, rule-based approaches are lauded for their efficiency and interpretability.For this line of work, it is crucial to correctly estimate the predictive effectiveness of the rules, i.e., the confidence.However, the existing literature lacks in-depth investigation into how confidence evolves with time.Moreover, inaccurate and heuristic confidence estimation limits the performance of rule-based methods.To alleviate such issues, we propose a framework named TempValid to explicitly model the temporal validity of rules for TKGF.Specifically, we design a time function to model the interaction between temporal information with confidence.TempValid conceptualizes confidence and other coefficients as learnable parameters to avoid inaccurate estimation and combinatorial explosion.Furthermore, we introduce a rule-adversarial negative sampling and a time-aware negative sampling strategies to facilitate TempValid learning.Extensive experiments show that TempValid significantly outperforms previous state-of-theart (SOTA) rule-based methods on six TKGF datasets.Moreover, it exhibits substantial advancements in cross-domain and resourceconstrained rule learning scenarios. Rikui Huang, Wei Wei 0002, Xiaoye Qu, Shengzhe Zhang, Dangyang Chen, Yu Cheng 0001 |
ACL (1) | 3 |
| 2024 | Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?abstractZhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaochen Su, Juntao Li 0005, Jun Zhang 0069, Tong Zhu 0002, Xiaoye Qu, Pan Zhou 0001, Yan Bowen, Yu Cheng 0001, Min Zhang 0005 |
ACL (1) | 5 |
| 2024 | Towards Robust Temporal Activity Localization Learning with Noisy LabelsabstractThis paper addresses the task of temporal activity localization (TAL). Although recent works have made significant progress in TAL research, almost all of them implicitly assume that the dense frame-level correspondences in each video-query pair are correctly annotated. However, in reality, such an assumption is extremely expensive and even impossible to satisfy due to subjective labeling. To alleviate this issue, in this paper, we explore a new TAL setting termed Noisy Temporal activity localization (NTAL), where a TAL model should be robust to the mixed training data with noisy moment boundaries. Inspired by the memorization effect of neural networks, we propose a novel method called Co-Teaching Regularizer (CTR) for NTAL. Specifically, we first learn a Gaussian Mixture Model to divide the mixed training data into preliminary clean and noisy subsets. Subsequently, we refine the labels of the two subsets by an adaptive prediction function so that their true positive and false positive samples could be identified. To avoid single model being prone to its mistakes learned by the mixed data, we adopt a co-teaching paradigm, which utilizes two models sharing the same framework to teach each other for robust learning. A curriculum strategy is further introduced to gradually learn the moment confidence from easy to hard. Experiments on three datasets demonstrate that our CTR is significantly more robust to the noisy training data compared to the existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Guoshun Nan, Keke Tang, Wanlong Fang, Yu Cheng 0001 |
LREC/COLING | 2 |
| 2024 | Rethinking Weakly-Supervised Video Temporal Grounding From a Game Perspective
Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen 0006, Jianfeng Dong, Keke Tang, Pan Zhou 0001, Yu Cheng 0001, Daizong Liu |
ECCV (45) | 4 |
| 2024 | Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
Jihai Zhang 0002, Xiang Lan 0004, Xiaoye Qu, Yu Cheng 0001, Mengling Feng, Bryan Hooi |
ECCV (83) | 3 |
| 2024 | LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-TrainingabstractMixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs).However, training MoE from scratch in a largescale setting still suffers from data-hungry and instability problems.Motivated by this limit, we investigate building MoE models from existing dense large language models.Specifically, based on the well-known LLaMA-2 7B model, we obtain an MoE model by: (1) Expert Construction, which partitions the parameters of original Feed-Forward Networks (FFNs) into multiple experts; (2) Continual pretraining, which further trains the transformed MoE model and additional gate networks.In this paper, we comprehensively explore different methods for expert construction and various data sampling strategies for continual pretraining.After these stages, our LLaMA-MoE models could maintain language abilities and route the input tokens to specific experts with part of the parameters activated.Empirically, by training 200B tokens, LLaMA-MoE-3.5Bmodels significantly outperform dense models that contain similar activation parameters. Tong Zhu 0002, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, Yu Cheng 0001 |
EMNLP | 2 |
| 2024 | SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved InformationabstractLarge Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing.However, the full potential of LVLMs' Retrieval-Augmented Generation (RAG) capabilities remains underutilized.Existing works either focus solely on the text modality or are limited to specific tasks.Moreover, most LVLMs struggle to selectively utilize retrieved information and are sensitive to irrelevant or misleading references.To address these challenges, we propose a self-refinement framework designed to teach LVLMs to Selectively Utilize Retrieved Information (SURf).Specifically, when given questions that are incorrectly answered by the LVLM backbone, we obtain references that help correct the answers (positive references) and those that do not (negative references).We then fine-tune the LVLM backbone using a combination of these positive and negative references.Our experiments across three tasks and seven datasets demonstrate that our framework significantly enhances LVLMs' ability to effectively utilize retrieved multimodal references and improves their robustness against irrelevant or misleading information.The source code is available at https://github.com/GasolSun36/SURf. * Work done during internship at Shanghai AI Laboratory.† Both are corresponding authors.How many apples in the images? VQAVanilla: There are three apples.The image depicting five apples on a tree...The picture shows 7 apples .... leaves...Ours: There are four apples.Describe this image in details. Captioning Vanilla: A person walking in snow.The image depicting a...the skier is in a crouched position... The image captures a dynamic scene ..a skier dressed in a ... Jiashuo Sun, Jihai Zhang 0002, Yucheng Zhou 0001, Zhaochen Su, Xiaoye Qu, Yu Cheng 0001 |
EMNLP | 5 |
| 2024 | Joint Multi-Facts Reasoning Network for Complex Temporal Question Answering Over Knowledge GraphabstractTemporal Knowledge Graph (TKG) is an extension of regular knowledge graph by attaching the time scope. Existing temporal knowledge graph question answering (TKGQA) models solely approach simple questions, owing to the prior assumption that each question only contains a single temporal fact with explicit/implicit temporal constraints. Hence, they perform poorly on questions which own multiple temporal facts. In this paper, we propose Joint Multi Facts Reasoning Network (JMFRN), to jointly reasoning multiple temporal facts for accurately answering complex temporal questions. Specifically, JMFRN first retrieves question-related temporal facts from TKG for each entity of the given complex question. For joint reasoning, we design two different attention (i.e., entity-aware and time-aware) modules, which are suitable for universal settings, to aggregate entities and timestamps information of retrieved facts. Moreover, to filter incorrect type answers, we introduce an additional answer type discrimination task. Extensive experiments demonstrate our proposed method significantly outperforms the state-of-art on the wellknown complex temporal question benchmark TimeQuestions. Rikui Huang, Wei Wei 0002, Xiaoye Qu, Wenfeng Xie, Xianling Mao, Dangyang Chen |
ICASSP | 3 |
| 2024 | Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
Zhuojun Ding, Wei Wei 0002, Xiaoye Qu, Dangyang Chen |
IJCAI | 3 |
| 2024 | Frequency-Aware GAN for Imperceptible Transfer Attack on 3D Point CloudsabstractWith the development of depth sensors and 3D vision, the vulnerability of 3D point cloud models has garnered heightened concern. Almost all existing 3D attackers are deployed in the white-box setting, where they access the model details and directly optimize coordinate-wise noises to perturb 3D objects. However, realistic 3D applications would not share any model information (model parameters, gradients, etc.) with users. Although a few recent works try to explore the black-box attack, they still achieve limited attack success rates (ASR) and fail to generate high-quality adversarial samples. In this paper, we focus on designing a transfer-based black-box attack method, called Transferable Frequency-aware 3D GAN, to delve into achieving a high black-box ASR by improving the adversarial transferability while making the adversarial samples more imperceptible. Considering that the 3D imperceptibility depends on whether the shape of the object is distorted, we utilize the spectral tool with the GAN design to explicitly perceive and preserve the 3D geometric structures. Specifically, we design the Graph Fourier Transform (GFT) encoding layer in the GAN generator to extract the geometries as guidance, and develop a corresponding Inverse-GFT decoding layer to decode latent features with this guidance to reconstruct high-quality adversarial samples. To further improve the transferability, we develop a dual learning scheme of discriminator from both frequency and feature perspectives to constrain the generator via adversarial learning. Finally, imperceptible and transferable perturbations are rapidly generated by our proposed attack. Experimental results demonstrate that our attack method achieves the highest transfer ASR while exhibiting stronger imperceptibility. Xiaowen Cai 0001, Yunbo Tao, Daizong Liu, Pan Zhou 0001, Xiaoye Qu, Jianfeng Dong, Keke Tang, Lichao Sun 0001 |
ACM Multimedia | 5 |
| 2024 | Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageabstractVideo Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent respectable works have made remarkable progress in this task, they implicitly are rooted in the closed-set assumption that all the given queries as video-relevant. Given an OOD query in open-set scenarios, they still utilize it for wrong retrieval, which might lead to irrecoverable losses in high-risk scenarios, e.g., criminal activity detection. To this end, we creatively explore a brand-new VMR setting termed Open-Set Video Moment Retrieval (OS-VMR), where we should not only retrieve the precise moments based on ID query, but also reject OOD queries. In this paper, we make the first attempt to step toward OS-VMR and propose a novel model OpenVMR, which first distinguishes ID and OOD queries based on the normalizing flow technology, and then conducts moment retrieval based on ID queries. Specifically, we first learn the ID distribution by constructing a normalizing flow, and assume the ID query distribution obeys the multi-variate Gaussian distribution. Then, we introduce an uncertainty score to search the ID-OOD separating boundary. After that, we refine the ID-OOD boundary by pulling together ID query features. Besides, video-query matching and frame-query matching are designed for coarse-grained and fine-grained cross-modal interaction, respectively. Finally, a positive-unlabeled learning module is introduced for moment retrieval. Experimental results on three VMR datasets show the effectiveness of our OpenVMR. Wanlong Fang, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Renfu Li, Zichuan Xu, Lixing Chen, Panpan Zheng, Yu Cheng 0001 |
ACM Multimedia | 4 |
| 2024 | GIST: Improving Parameter Efficient Fine-Tuning via Knowledge InteractionabstractRecently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009 of ViT-B/16). The code is available at https://github.com/JCruan519/GIST. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu |
ACM Multimedia | 8 |
| 2024 | Temporal Sentence Grounding with Relevance Feedback in VideosabstractAs a widely explored multi-modal task, Temporal Sentence Grounding in videos (TSG) endeavors to retrieve a specific video segment matched with a given query text from a video. The traditional paradigm for TSG generally assumes that relevant segments always exist within a given video. However, this assumption is restrictive and unrealistic in real-world applications where the existence of a query-related segment is uncertain, easily resulting in erroneous grounding. Motivated by the research gap and practical application, this paper introduces a new task, named Temporal Sentence Grounding with Relevance Feedback (TSG-RF) in videos, which accommodates the possibility that a video may or may not include a segment related to the query. This task entails localizing precise video segments that semantically align with the query text when such content is present, while delivering definitive feedback on the non-existence of related segments when absent. Moreover, we propose a novel Relation-aware Temporal Sentence Grounding (RaTSG) network for addressing this challenging task. This network first reformulates the TSG-RF task as a foreground-background detection problem by investigating whether the query-related semantics exist in both frame and video levels. Then, a multi-granularity relevance discriminator is exploited to produce precise video-query relevance feedback and a relation-aware segment grounding module is employed to selectively conduct the grounding process, dynamically adapting to the presence or absence of query-related segments in videos. To validate our RaTSG network, we reconstruct two popular TSG datasets, establishing a rigorous benchmark for TSG-RF. Experimental results demonstrate the effectiveness of our proposed RaTSG for the TSG-RF task. Our source code is available at https://github.com/HuiGuanLab/RaTSG. Jianfeng Dong, Xiaoman Peng, Daizong Liu, Xiaoye Qu, Xun Yang 0001, Cuizhu Bao, Meng Wang 0001 |
NeurIPS | 4 |
| 2024 | On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits FusionabstractEfficient fine-tuning of large language models for task-specific applications is imperative, yet the vast number of parameters in these models makes their training increasingly challenging.
Despite numerous proposals for effective methods, a substantial memory overhead remains for gradient computations during updates. \thm{Can we fine-tune a series of task-specific small models and transfer their knowledge directly to a much larger model without additional training?}
In this paper, we explore weak-to-strong specialization using logit arithmetic, facilitating a direct answer to this question.
Existing weak-to-strong methods often employ a static knowledge transfer ratio and a single small model for transferring complex knowledge, which leads to suboptimal performance.
To surmount these limitations,
we propose a dynamic logit fusion approach that works with a series of task-specific small models, each specialized in a different task.
This method adaptively allocates weights among these models at each decoding step,
learning the weights through Kullback-Leibler divergence constrained optimization problems.
We conduct extensive experiments across various benchmarks in both single-task and multi-task settings, achieving leading results.
By transferring expertise from the 7B model to the 13B model, our method closes the performance gap by 96.4\% in single-task scenarios and by 86.3\% in multi-task scenarios compared to full fine-tuning of the 13B model. Notably, we achieve surpassing performance on unseen tasks. Moreover, we further demonstrate that our method can effortlessly integrate in-context learning for single tasks and task arithmetic for multi-task scenarios. Chenghao Fan, Zhenyi Lu, Wei Wei 0002, Xiaoye Qu, Dangyang Chen, Yu Cheng 0001 |
NeurIPS | 5 |
| 2024 | Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding tasks. Nevertheless, these models are susceptible to adversarial examples. In real-world applications, existing LVLM attackers generally rely on the detailed prior knowledge of the model to generate effective perturbations. Moreover, these attacks are task-specific, leading to significant costs for designing perturbation. Motivated by the research gap and practical demands, in this paper, we make the first attempt to build a universal attacker against real-world LVLMs, focusing on two critical aspects: (i) restricting access to only the LVLM inputs and outputs. (ii) devising a universal adversarial patch, which is task-agnostic and can deceive any LVLM-driven task when applied to various inputs. Specifically, we start by initializing the location and the pattern of the adversarial patch through random sampling, guided by the semantic distance between their output and the target label. Subsequently, we maintain a consistent patch location while refining the pattern to enhance semantic resemblance to the target. In particular, our approach incorporates a diverse set of LVLM task inputs as query samples to approximate the patch gradient, capitalizing on the importance of distinct inputs. In this way, the optimized patch is universally adversarial against different tasks and prompts, leveraging solely gradient estimates queried from the model. Extensive experiments are conducted to verify the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs including LLaVA, MiniGPT-4, Flamingo, and BLIP-2, spanning a spectrum of tasks, all achieved without delving into the details of the model structures. Daizong Liu, Xiaoye Qu, Pan Zhou 0001, Keke Tang, Yao Wan 0001, Lichao Sun 0001 |
NeurIPS | 3 |
| 2024 | Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingabstractIn the era of large language models, model merging is a promising way to combine multiple task-specific models into a single multitask model without extra training.
However, two challenges remain: (a) interference between different models and (b) heterogeneous data during testing. Traditional model merging methods often show significant performance gaps compared to fine-tuned models due to these issues.
Additionally, a one-size-fits-all model lacks flexibility for diverse test data, leading to performance degradation.
We show that both shared and exclusive task-specific knowledge are crucial for merging performance, but directly merging exclusive knowledge hinders overall performance.
In view of this, we propose Twin-Merging, a method that encompasses two principal stages:
(1) modularizing knowledge into shared and exclusive components, with compression to reduce redundancy and enhance efficiency;
(2) dynamically merging shared and task-specific knowledge based on the input.
This approach narrows the performance gap between merged and fine-tuned models and improves adaptability to heterogeneous data.
Extensive experiments on $20$ datasets for both language and vision tasks demonstrate the effectiveness of our method, showing an average improvement of $28.34\%$ in absolute normalized score for discriminative tasks and even surpassing the fine-tuned upper bound on the generative tasks. Zhenyi Lu, Chenghao Fan, Wei Wei 0002, Xiaoye Qu, Dangyang Chen, Yu Cheng 0001 |
NeurIPS | 4 |
| 2024 | ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001 |
NeurIPS | 3 |
| 2024 | A Survey on Arabic Named Entity Recognition: Past, Recent Advances, and Future TrendsabstractAs more and more Arabic texts emerged on the Internet, extracting important information from these Arabic texts is especially useful. As a fundamental technology, Named entity recognition (NER) serves as the core component in information extraction technology, while also playing a critical role in many other Natural Language Processing (NLP) systems, such as question answering and knowledge graph building. In this paper, we provide a comprehensive review of the development of Arabic NER, especially the recent advances in deep learning and pre-trained language model. Specifically, we first introduce the background of Arabic NER, including the characteristics of Arabic and existing resources for Arabic NER. Then, we systematically review the development of Arabic NER methods. Traditional Arabic NER systems focus on feature engineering and designing domain-specific rules. In recent years, deep learning methods achieve significant progress by representing texts via continuous vector representations. With the growth of pre-trained language model, Arabic NER yields better performance. Finally, we conclude the method gap between Arabic NER and NER methods from other languages, which helps outline future directions for Arabic NER. Xiaoye Qu, Yingjie Gu, Qingrong Xia, Zechang Li, Zhefeng Wang 0001, Baoxing Huai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Rethinking Video Sentence Grounding From a Tracking Perspective With Memory Network and Masked AttentionabstractVideo sentence grounding (VSG) is the task of identifying the segment of an untrimmed video that semantically corresponds to a given natural language query. While many existing methods extract frame-grained features using pre-trained 2D or 3D convolution networks, often fail to capture subtle differences between ambiguous adjacent frames. Although some recent approaches incorporate object-grained features using Faster R-CNN to capture more fine-grained details, they are still primarily based on feature enhancement and lack spatio-temporal modeling to explore the semantics of the core persons/objects. To solve the problem of modeling the core target's behavior, in this paper, we propose a new perspective for addressing the VSG task by tracking pivotal objects and activities to learn more fine-grained spatio-temporal features. Specifically, we introduce the Video Sentence Tracker with Memory Network and Masked Attention (VSTMM), which comprises a cross-modal targets generator for producing multi-modal templates and search space, a memory-based tracker for dynamically tracking multi-modal targets using a memory network to record targets' behaviors, a masked attention localizer which learns local shared features between frames and eliminates interference from long-term dependencies, resulting in improved accuracy when localizing the moment. To evaluate the performance of our VSTMM, we conducted extensive experiments and comparisons with state-of-the-art methods on three challenging benchmarks, including Charades-STA, ActivityNet Captions, and TACoS. Without bells and whistles, our VSTMM achieves leading performance with a considerable real-time speed. Zeyu Xiong, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Jiahao Zhu 0003, Keke Tang, Pan Zhou 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Transform-Equivariant Consistency Learning for Temporal Sentence GroundingabstractThis paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-query paired data for training, but also easily fail into the dataset distribution bias. To alleviate these limitations, we introduce a novel Equivariant Consistency Regulation Learning (ECRL) framework to learn more discriminative query-related frame-wise representations for each video, in a self-supervised manner. Our motivation comes from that the temporal boundary of the query-guided activity should be consistently predicted under various video-level transformations. Concretely, we first design a series of spatio-temporal augmentations on both foreground and background video segments to generate a set of synthetic video samples. In particular, we devise a self-refine module to enhance the completeness and smoothness of the augmented video. Then, we present a novel self-supervised consistency loss (SSCL) applied on the original and augmented videos to capture their invariant query-related semantic by minimizing the KL-divergence between the sequence similarity of two videos and a prior Gaussian distribution of timestamp distance. At last, a shared grounding head is introduced to predict the transform-equivariant query-guided segment boundaries for both the original and augmented videos. Extensive experiments on three challenging datasets (ActivityNet, TACoS, and Charades-STA) demonstrate both effectiveness and efficiency of our proposed ECRL framework. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu, Haozhao Wang, Xing Di, Weining Lu, Yu Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Distantly-Supervised Named Entity Recognition with Adaptive Teacher Learning and Fine-Grained Student EnsembleabstractDistantly-Supervised Named Entity Recognition (DS-NER) effectively alleviates the data scarcity problem in NER by automatically generating training samples. Unfortunately, the distant supervision may induce noisy labels, thus undermining the robustness of the learned models and restricting the practical application. To relieve this problem, recent works adopt self-training teacher-student frameworks to gradually refine the training labels and improve the generalization ability of NER models. However, we argue that the performance of the current self-training frameworks for DS-NER is severely underestimated by their plain designs, including both inadequate student learning and coarse-grained teacher updating. Therefore, in this paper, we make the first attempt to alleviate these issues by proposing: (1) adaptive teacher learning comprised of joint training of two teacher-student networks and considering both consistent and inconsistent predictions between two teachers, thus promoting comprehensive student learning. (2) fine-grained student ensemble that updates each fragment of the teacher model with a temporal moving average of the corresponding fragment of the student, which enhances consistent predictions on each model fragment against noise. To verify the effectiveness of our proposed method, we conduct experiments on four DS-NER datasets. The experimental results demonstrate that our method significantly surpasses previous SOTA methods. The code is available at https://github.com/zenhjunpro/ATSEN. Xiaoye Qu, Daizong Liu, Zhefeng Wang 0001, Baoxing Huai, Pan Zhou 0001 |
AAAI | 1 |
| 2023 | TREA: Tree-Structure Reasoning Schema for Conversational RecommendationabstractConversational recommender systems (CRS) aim to timely trace the dynamic interests of users through dialogues and generate relevant responses for item recommendations.Recently, various external knowledge bases (especially knowledge graphs) are incorporated into CRS to enhance the understanding of conversation contexts.However, recent reasoning-based models heavily rely on simplified structures such as linear structures or fixed-hierarchical structures for causality reasoning, hence they cannot fully figure out sophisticated relationships among utterances with external knowledge.To address this, we propose a novel Treestructure Reasoning schEmA named TREA.TREA constructs a multi-hierarchical scalable tree as the reasoning structure to clarify the causal relationships between mentioned entities, and fully utilizes historical conversations to generate more reasonable and suitable responses for recommended results.Extensive experiments on two public CRS datasets have demonstrated the effectiveness of our approach.Our Wendi Li, Wei Wei 0002, Xiaoye Qu, Xianling Mao, Wenfeng Xie, Dangyang Chen |
ACL (1) | 3 |
| 2023 | Mirror: A Universal Framework for Various Information Extraction TasksabstractTong Zhu, Junfei Ren, Zijian Yu, Mengsong Wu, Guoliang Zhang, Xiaoye Qu, Wenliang Chen, Zhefeng Wang, Baoxing Huai, Min Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Tong Zhu 0002, Junfei Ren, Zijian Yu, Mengsong Wu, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 6 |
| 2023 | Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalabstractAlmost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR) task, which aims to retrieve partially relevant untrimmed videos with the query input. Particularly, we propose to address PRVR from a new perspective, i.e., distilling the generalization knowledge from the large-scale vision-language pre-trained model and transferring it to a task-specific PRVR network. To be specific, we introduce a Dual Learning framework with Dynamic Knowledge Distillation (DL-DKD), which exploits the knowledge of a large vision-language model as the teacher to guide a student model. During the knowledge distillation, an inheritance student branch is devised to absorb the knowledge from the teacher model. Considering that the large model may be of mediocre performance due to the domain gaps, we further develop an exploration student branch to take the benefits of task-specific information. In addition, a dynamical knowledge distillation strategy is further devised to adjust the effect of each student branch learning during the training. Experiment results demonstrate that our proposed model achieves state-of-the-art performance on ActivityNet and TVR datasets for PRVR. Jianfeng Dong, Minsong Zhang, Xianke Chen, Daizong Liu, Xiaoye Qu, Xun Wang 0007 |
ICCV | 6 |
| 2023 | Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalabstractThis paper addresses the challenging task of language-driven moment retrieval. Previous methods are typically trained to localize the target moment corresponding to a single sentence query in a complicated video. However, this specific moment generally delivers richer contents than the query, i.e., the semantics of one query may miss certain object details or actions in the complex foreground-background visual contents. Such information imbalance between two modalities makes it difficult to finely align their representations. To this end, instead of training with a single query, we propose to utilize the diversity and complementarity among different queries corresponding to the same video moment for enriching the textual semantics. Specifically, we develop a Teacher-Student Moment Retrieval (TSMR) framework to fill this cross-modal information gap. A teacher model is trained to not only encode a certain query but also capture extra complementary queries to aggregate contextual semantics for obtaining more comprehensive moment-related query representations. Since the additional queries are inaccessible during inference, we further introduce an adaptive knowledge distillation mechanism to train a student model with a single query input by selectively absorbing the knowledge from the teacher model. In this manner, the student model is more robust to the cross-modal information gap during the moment retrieval guided by a single query. Experimental results on two benchmarks demonstrate the effectiveness of our proposed method. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan, Pan Zhou 0001, Zichuan Xu, Lixing Chen, Yu Cheng 0001 |
ACM Multimedia | 2 |
| 2023 | Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionabstractExisting few-shot action recognition methods have placed primary focus on improving the recognition accuracy while neglecting another important indicator in practical scenarios, i.e., model efficiency. In this paper, we make the first attempt and propose a Lightweight Multi-modal Knowledge Distillation framework (Lite-MKD) for few-shot action recognition. In this framework, the teacher model conducts multi-modal learning to achieve a comprehensive fusion of the optical flow, depth, and appearance features of human movements, thus achieving a more robust representation of actions. The student model is utilized to learn to recognize actions from the single RGB modality at a lower computational cost under the guidance of the teacher. To fully explore and integrate multi-modal information, a hierarchical Multi-modal Fusion Module (MFM) is introduced in the teacher model. Besides, a multi-level Distinguish-to-Mimic (D2M) knowledge distillation component is proposed for the student model. D2M improves the ability of the student model to mimic the action classification probabilities of the teacher model by enhancing the distinguishability of the student model for different video categories in the support set. Extensive experiments on three action recognition datasets Kinetics, HMDB51, and UCF101 demonstrate our framework's effectiveness and stable generalization ability. With a much more lightweight network for inference, we achieve comparable performance to previous state-of-the-art methods. Our source code is available at https://github.com/HuiGuanLab/Lite-MKD Daizong Liu, Xiaoye Qu, Junyu Gao 0002, Jianfeng Dong, Xun Wang 0007 |
ACM Multimedia | 5 |
| 2023 | Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingabstractUnsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models (i.e., joint, bone, and motion), then integrate the multi-modal information for action understanding by a late-fusion strategy. Although these approaches have achieved significant performance, they suffer from the complex yet redundant multi-stream model designs, each of which is also limited to the fixed input skeleton modality. To alleviate these issues, in this paper, we propose a Unified Multimodal Unsupervised Representation Learning framework, called UmURL, which exploits an efficient early-fusion strategy to jointly encode the multi-modal features in a single-stream manner. Specifically, instead of designing separate modality-specific optimization processes for uni-modal unsupervised learning, we feed different modality inputs into the same stream with an early-fusion strategy to learn their multi-modal features for reducing model complexity. To ensure that the fused multi-modal features do not exhibit modality bias, i.e., being dominated by a certain modality input, we further propose both intra- and inter-modal consistency learning to guarantee that the multi-modal features contain the complete semantics of each modal via feature decomposition and distinct alignment. In this manner, our framework is able to learn the unified representations of uni-modal or multi-modal skeleton input, which is flexible to different kinds of modality input for robust action understanding in practical cases. Extensive experiments conducted on three large-scale datasets, i.e., NTU-60, NTU-120, and PKU-MMD II, demonstrate that UmURL is highly efficient, possessing the approximate complexity with the uni-modal methods, while achieving new state-of-the-art performance across various downstream task scenarios in skeleton-based action representation learning. Our source code is available at https://github.com/HuiGuanLab/UmURL. Shengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu, Junyu Gao 0002, Xun Yang 0001, Xun Wang 0007, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2023 | From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalabstractAttribute-specific fashion retrieval (ASFR) is a challenging information retrieval task, which has attracted increasing attention in recent years. Different from traditional fashion retrieval which mainly focuses on optimizing holistic similarity, the ASFR task concentrates on attribute-specific similarity, resulting in more fine-grained and interpretable retrieval results. As the attribute-specific similarity typically corresponds to the specific subtle regions of images, we propose a Region-to-Patch Framework (RPF) that consists of a region-aware branch and a patch-aware branch to extract fine-grained attribute-related visual features for precise retrieval in a coarse-to-fine manner. In particular, the region-aware branch is first to be utilized to locate the potential regions related to the semantic of the given attribute. Then, considering that the located region is coarse and still contains the background visual contents, the patch-aware branch is proposed to capture patch-wise attribute-related details from the previous amplified region. Such a hybrid architecture strikes a proper balance between region localization and feature extraction. Besides, different from previous works that solely focus on discriminating the attribute-relevant foreground visual features, we argue that the attribute-irrelevant background features are also crucial for distinguishing the detailed visual contexts in a contrastive manner. Therefore, a novel E-InfoNCE loss based on the foreground and background representations is further proposed to improve the discrimination of attribute-specific representation. Extensive experiments on three datasets demonstrate the effectiveness of our proposed framework, and also show a decent generalization of our RPF on out-of-domain fashion images. Our source code is available at https://github.com/HuiGuanLab/RPF. Jianfeng Dong, Xiaoman Peng, Zhe Ma 0002, Daizong Liu, Xiaoye Qu, Xun Yang 0001, Jixiang Zhu |
SIGIR | 5 |
| 2023 | Multi-level feature disentanglement network for cross-dataset face forgery detection
Zhixiao Fu, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Xuhong Zhang 0002, Shouling Ji |
Image Vis. Comput. | 4 |
| 2023 | Progressive Localization Networks for Language-Based Moment LocalizationabstractThis article targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large variation of the temporal lengths of video moments. Most existing methods prefer to first sample sufficient candidate moments with various temporal lengths, then match them with the given query to determine the target moment. However, candidate moments generated with a fixed temporal granularity may be suboptimal to handle the large variation in moment lengths. To this end, we propose a novel multi-stage Progressive Localization Network (PLN) that progressively localizes the target moment in a coarse-to-fine manner. Specifically, each stage of PLN has a localization branch and focuses on candidate moments that are generated with a specific temporal granularity. The temporal granularities of candidate moments are different across the stages. Moreover, we devise a conditional feature manipulation module and an upsampling connection to bridge the multiple localization branches. In this fashion, the later stages are able to absorb the previously learned information, thus facilitating the more fine-grained localization. Extensive experiments on three public datasets demonstrate the effectiveness of our proposed PLN for language-based moment localization, especially for localizing short moments in long videos. Jianfeng Dong, Xiaoye Qu, Xun Yang 0001, Pan Zhou 0001, Xun Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Memory-Guided Semantic Learning Network for Temporal Sentence GroundingabstractTemporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although existing methods train well-designed deep networks with large amount of data, we find that they can easily forget the rarely appeared cases during training due to the off-balance data distribution, which influences the model generalization and leads to unsatisfactory performance. To tackle this issue, we propose a memory-augmented network, called Memory-Guided Semantic Learning Network (MGSL-Net), that learns and memorizes the rarely appeared content in TSG task. Specifically, our proposed model consists of three main parts: cross-modal interaction module, memory augmentation module, and heterogeneous attention module. We first align the given video-query pair by a cross-modal graph convolutional network, and then utilize memory module to record the cross-modal shared semantic features in the domain-specific persistent memory. During training, the memory slots are dynamically associated with both common and rare cases, alleviating the forgetting issue. In testing, the rare cases can thus be enhanced by retrieving the stored memories, leading to better generalization. At last, the heterogeneous attention module is utilized to integrate the enhanced multi-modal features in both video and query domains. Experimental results on three benchmarks show the superiority of our method on both effectiveness and efficiency, which substantially improves the accuracy not only on the entire dataset but also on the rare cases. Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 2 |
| 2022 | Unsupervised Temporal Video Grounding with Deep Semantic ClusteringabstractTemporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant video-query paired data, which is expensive to collect in real-world scenarios. In this paper, we explore whether a video grounding model can be learned without any paired annotations. To the best of our knowledge, this paper is the first work trying to address TVG in an unsupervised setting. Considering there is no paired supervision, we propose a novel Deep Semantic Clustering Network (DSCNet) to leverage all semantic information from the whole query set to compose the possible activity in each video for grounding. Specifically, we first develop a language semantic mining module, which extracts implicit semantic features from the whole query set. Then, these language semantic features serve as the guidance to compose the activity in video via a video-based semantic aggregation module. Finally, we utilize a foreground attention branch to filter out the redundant background activities and refine the grounding results. To validate the effectiveness of our DSCNet, we conduct experiments on both ActivityNet Captions and Charades-STA datasets. The results demonstrate that our DSCNet achieves competitive performance, and even outperforms most weakly-supervised approaches. Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 2 |
| 2022 | Exploring Motion and Appearance Information for Temporal Sentence GroundingabstractThis paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A major limitation of these works is that they fail to distinguish ambiguous video frames with subtle appearance differences due to frame-level feature extraction. Recently, a few methods adopt Faster R-CNN to extract detailed object features in each frame to differentiate the fine-grained appearance similarities. However, the object-level features extracted by Faster R-CNN suffer from missing motion analysis since the object detection model lacks temporal modeling. To solve this issue, we propose a novel Motion-Appearance Reasoning Network (MARN), which incorporates both motion-aware and appearance-aware object features to better reason object relations for modeling the activity among successive frames. Specifically, we first introduce two individual video encoders to embed the video into corresponding motion-oriented and appearance-aspect object representations. Then, we develop separate motion and appearance branches to learn motion-guided and appearance-guided object relations, respectively. At last, both motion and appearance information from two branches are associated to generate more representative features for final grounding. Extensive experiments on two challenging datasets (Charades-STA and TACoS) show that our proposed MARN significantly outperforms previous state-of-the-art methods by a large margin. Daizong Liu, Xiaoye Qu, Pan Zhou 0001 |
AAAI | 2 |
| 2022 | Efficient Document-level Event Extraction via Pseudo-Trigger-aware Pruned Complete GraphabstractMost previous studies of document-level event extraction mainly focus on building argument chains in an autoregressive way, which achieves a certain success but is inefficient in both training and inference. In contrast to the previous studies, we propose a fast and lightweight model named as PTPCG. In our model, we design a novel strategy for event argument combination together with a non-autoregressive decoding algorithm via pruned complete graphs, which are constructed under the guidance of the automatically selected pseudo triggers. Compared to the previous systems, our system achieves competitive results with 19.8% of parameters and much lower resource consumption, taking only 3.8% GPU hours for training and up to 8.5 times faster for inference. Besides, our model shows superior compatibility for the datasets with (or without) triggers and the pseudo triggers can be the supplements for annotated triggers to make further improvements. Codes are available at https://github.com/Spico197/DocEE . Tong Zhu 0002, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Min Zhang 0005 |
IJCAI | 2 |
| 2022 | Reducing the Vision and Language Bias for Temporal Sentence GroundingabstractTemporal sentence grounding (TSG) is an important yet challenging task in multimedia information retrieval. Although previous TSG methods have achieved decent performance, they tend to capture the selection biases of frequently appeared video-query pairs in the dataset rather than present robust multimodal reasoning abilities, especially for the rarely appeared pairs. In this paper, we study the above issue of selection biases and accordingly propose a Debiasing-TSG (D-TSG) model to filter and remove the negative biases in both vision and language modalities for enhancing the model generalization ability. Specifically, we propose to alleviate the issue from two perspectives: 1) Feature distillation. We built a multi-modal debiasing branch to firstly capture the vision and language biases, and then apply a bias identification module to explicitly recognize the true negative biases and remove them from the benign multi-modal representations. 2) Contrastive sample generation. We construct two types of negative samples to enforce the model to accurately learn the aligned multi-modal semantics and make complete semantic reasoning. We apply the proposed model to both commonly and rarely appeared TSG cases, and demonstrate its effectiveness by achieving the state-of-the-art performance on three benchmark datasets (ActivityNet Caption, TACoS, and Charades-STA). Daizong Liu, Xiaoye Qu, Wei Hu 0003 |
ACM Multimedia | 2 |
| 2022 | Reading-Strategy Inspired Visual Representation Learning for Text-to-Video RetrievalabstractThis paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relevant to the given query, from a great number of unlabeled videos. The success of this task depends on cross-modal representation learning that projects both videos and sentences into common spaces for semantic similarity computation. In this work, we concentrate on video representation learning, an essential component for text-to-video retrieval. Inspired by the reading strategy of humans, we propose a Reading-strategy Inspired Visual Representation Learning (RIVRL) to represent videos, which consists of two branches: a previewing branch and an intensive-reading branch. The previewing branch is designed to briefly capture the overview information of videos, while the intensive-reading branch is designed to obtain more in-depth information. Moreover, the intensive-reading branch is aware of the video overview captured by the previewing branch. Such holistic information is found to be useful for the intensive-reading branch to extract more fine-grained features. Extensive experiments on three datasets are conducted, where our model RIVRL achieves a new state-of-the-art on TGIF and VATEX. Moreover, on MSR-VTT, our model using two video features shows comparable performance to the state-of-the-art using seven video features and even outperforms models pre-trained on the large-scale HowTo100M dataset. Code is available athttps://github.com/LiJiaBei-7/rivrl. Jianfeng Dong, Xianke Chen, Xiaoye Qu, Xirong Li 0001, Yuan He 0011, Xun Wang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Read, Retrospect, Select: An MRC Framework to Short Text Entity LinkingabstractEntity linking (EL) for the rapidly growing short text (e.g. search queries and news titles) is critical to industrial applications. Most existing approaches relying on adequate context for long text EL are not effective for the concise and sparse short text. In this paper, we propose a novel framework called Multi-turn Multiple-choice Machine reading comprehension (M3) to solve the short text EL from a new perspective: a query is generated for each ambiguous mention exploiting its surrounding context, and an option selection module is employed to identify the golden entity from candidates using the query. In this way, M3 framework sufficiently interacts limited context with candidate entities during the encoding process, as well as implicitly considers the dissimilarities inside the candidate bunch in the selection stage. In addition, we design a two-stage verifier incorporated into M3 to address the commonly existed unlinkable problem in short text. To further consider the topical coherence and interdependence among referred entities, M3 leverages a multi-turn fashion to deal with mentions in a sequence manner by retrospecting historical cues. Evaluation shows that our M3 framework achieves the state-of-the-art performance on five Chinese and English datasets for the real-world short text EL. Yingjie Gu, Xiaoye Qu, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Xiaolin Gui |
AAAI | 2 |
| 2021 | Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingabstractThis paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or directly regress the boundary timestamps of the target segment. In this paper, we propose a novel localization framework that scores all pairs of start and end indices within the video simultaneously with a biaffine mechanism. In particular, we present a Context-aware Biaffine Localizing Network (CBLN) which incorporates both local and global contexts into features of each start/end position for biaffine-based localization. The local contexts from the adjacent frames help distinguish the visually similar appearance, and the global contexts from the entire video contribute to reasoning the temporal relation. Besides, we also develop a multi-modal self-attention module to provide fine-grained query-guided video representation for this biaffine strategy. Extensive experiments show that our CBLN significantly outperforms state-of-thearts on three public datasets (ActivityNet Captions, TACoS, and Charades-STA), demonstrating the effectiveness of the proposed localization framework. The code is available at https://github.com/liudaizong/CBLN. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Yu Cheng 0001, Wei Wei 0002, Zichuan Xu, Yulai Xie 0002 |
CVPR | 2 |
| 2021 | Adaptive Proposal Generation Network for Temporal Sentence Localization in VideosabstractWe address the problem of temporal sentence localization in videos (TSLV).Traditional methods follow a top-down framework which localizes the target segment with predefined segment proposals.Although they have achieved decent performance, the proposals are handcrafted and redundant.Recently, bottom-up framework attracts increasing attention due to its superior efficiency.It directly predicts the probabilities for each frame as a boundary.However, the performance of bottom-up model is inferior to the top-down counterpart as it fails to exploit the segmentlevel interaction.In this paper, we propose an Adaptive Proposal Generation Network (APGN) to maintain the segment-level interaction while speeding up the efficiency.Specifically, we first perform a foregroundbackground classification upon the video and regress on the foreground frames to adaptively generate proposals.In this way, the handcrafted proposal design is discarded and the redundant proposals are decreased.Then, a proposal consolidation module is further developed to enhance the semantic of the generated proposals.Finally, we locate the target moments with these generated proposals following the top-down framework.Extensive experiments on three challenging benchmarks show that our proposed APGN significantly outperforms previous state-of-the-art methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001 |
EMNLP (1) | 2 |
| 2021 | Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence GroundingabstractA key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence description.Existing methods mainly leverage vanilla soft attention to perform the alignment in a single-step process.However, such single-step attention is insufficient in practice, since complicated relations between inter-and intra-modality are usually obtained through multi-step reasoning.In this paper, we propose an Iterative Alignment Network (IA-Net) for TSG task, which iteratively interacts inter-and intra-modal features within multiple steps for more accurate grounding.Specifically, during the iterative reasoning process, we pad multi-modal features with learnable parameters to alleviate the nowhere-toattend problem of non-matched frame-word pairs, and enhance the basic co-attention mechanism in a parallel manner.To further calibrate the misaligned attention caused by each reasoning step, we also devise a calibration module following each attention module to refine the alignment knowledge.With such iterative alignment scheme, our IA-Net can robustly capture the fine-grained relations between vision and language domains step-bystep for progressively reasoning the temporal boundaries.Extensive experiments conducted on three challenging benchmarks demonstrate that our proposed model performs better than the state-of-the-arts. Daizong Liu, Xiaoye Qu, Pan Zhou 0001 |
EMNLP (1) | 2 |
| 2021 | Hierarchical Similarity Learning for Language-Based Product Image RetrievalabstractThis paper aims for the language-based product image retrieval task. The majority of previous works have made significant progress by designing network structure, similarity measurement, and loss function. However, they typically perform vision-text matching at certain granularity regardless of the intrinsic multiple granularities of images. In this paper, we focus on the cross-modal similarity measurement, and propose a novel Hierarchical Similarity Learning (HSL) network. HSL first learns multi-level representations of input data by stacked encoders, and object-granularity similarity and image-granularity similarity are computed at each level. All the similarities are combined as the final hierarchical cross-modal similarity. Experiments on a large-scale product retrieval dataset demonstrate the effectiveness of our proposed method. Code and data are available at https://github.com/liufh1/hsl. Zhe Ma 0002, Fenghao Liu, Jianfeng Dong, Xiaoye Qu, Yuan He 0011, Shouling Ji |
ICASSP | 4 |
| 2021 | Coarse to Fine: Domain Adaptive Crowd Counting via Adversarial Scoring NetworkabstractRecent deep networks have convincingly demonstrated high capability in crowd counting, which is a critical task attracting widespread attention due to its various industrial applications. Despite such progress, trained data-dependent models usually can not generalize well to unseen scenarios because of the inherent domain shift. To facilitate this issue, this paper proposes a novel adversarial scoring network (ASNet) to gradually bridge the gap across domains from coarse to fine granularity. In specific, at the coarse-grained stage, we design a dual-discriminator strategy to adapt source domain to be close to the targets from the perspectives of both global and local feature space via adversarial learning. The distributions between two domains can thus be aligned roughly. At the fine-grained stage, we explore the transferability of source characteristics by scoring how similar the source samples are to target ones from multiple levels based on generative probability derived from coarse stage. Guided by these hierarchical scores, the transferable source features are properly selected to enhance the knowledge transfer during the adaptation process. With the coarse-to-fine design, the generalization bottleneck induced from the domain discrepancy can be effectively alleviated. Three sets of migration experiments show that the proposed methods achieve state-of-the-art counting performance compared with major unsupervised methods. Zhikang Zou, Xiaoye Qu, Pan Zhou 0001, Shuangjie Xu, Xiaoqing Ye, Jin Ye 0006 |
ACM Multimedia | 2 |
| 2020 | Reasoning Step-by-Step: Temporal Sentence Localization in Videos via Deep Rectification-Modulation NetworkabstractTemporal sentence localization in videos aims to ground the best matched segment in an untrimmed video according to a given sentence query.Previous works in this field mainly rely on single-step attentional frameworks to align the temporal boundaries by a soft selection.Although they focus on the visual content relevant to the query, these attention strategies are insufficient to model complex video contents and restrict the higher-level reasoning demand for temporal relation.In this paper, we propose a novel deep rectification-modulation network (RMN), transforming this task into a multi-step reasoning process by repeating rectification and modulation.In each rectification-modulation layer, unlike existing methods directly conducting the cross-modal interaction, we first devise a rectification module to correct implicit attention misalignment which focuses on wrong position during the interaction process.Then, a modulation module is developed to model the frame-to-frame relation with the help of specific sentence information for better correlating and composing the video contents over time.With multiple such layers cascaded in depth, our RMN progressively refines video and query interactions, thus enabling a further precise localization.Experimental evaluations on three public datasets show that the proposed method achieves state-of-the-art performance. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001 |
COLING | 2 |
| 2020 | Fine-Grained Text Sentiment Transfer via Dependency Parsing
Lulu Xiao, Xiaoye Qu, Ruixuan Li 0001, Jun Wang 0018, Pan Zhou 0001, Yuhua Li 0003 |
ECAI | 2 |
| 2020 | Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment LocalizationabstractQuery-based moment localization is a new task that localizes the best matched segment in an untrimmed video according to a given sentence query. In this localization task, one should pay more attention to thoroughly mine visual and linguistic information. To this end, we propose a novel Cross- and Self-Modal Graph Attention Network (CSMGAN) that recasts this task as a process of iterative messages passing over a joint graph. Specifically, the joint graph consists of Cross-Modal interaction Graph (CMG) and Self-Modal relation Graph (SMG), where frames and words are represented as nodes, and the relations between cross- and self-modal node pairs are described by an attention mechanism. Through parametric message passing, CMG highlights relevant instances across video and sentence, and then SMG models the pairwise relation inside each modality for frame (word) correlating. With multiple layers of such a joint graph, our CSMGAN is able to effectively capture high-order interactions between two modalities, thus enabling a further precise localization. Besides, to better comprehend the contextual details in the query, we develop a hierarchical sentence encoder to enhance the query understanding. Extensive experiments on four public datasets demonstrate the effectiveness of our proposed model, and GCSMAN significantly outperforms the state-of-the-arts. Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu |
ACM Multimedia | 2 |
| 2020 | Fine-grained Iterative Attention Network for Temporal Language Localization in VideosabstractTemporal language localization in videos aims to ground one video segment in an untrimmed video based on a given sentence query. To tackle this task, designing an effective model to extract ground-ing information from both visual and textual modalities is crucial. However, most previous attempts in this field only focus on unidirectional interactions from video to query, which emphasizes which words to listen and attends to sentence information via vanilla soft attention, but clues from query-by-video interactions implying where to look are not taken into consideration. In this paper, we propose a Fine-grained Iterative Attention Network (FIAN) that consists of an iterative attention module for bilateral query-video in-formation extraction. Specifically, in the iterative attention module, each word in the query is first enhanced by attending to each frame in the video through fine-grained attention, then video iteratively attends to the integrated query. Finally, both video and query information is utilized to provide robust cross-modal representation for further moment localization. In addition, to better predict the target segment, we propose a content-oriented localization strategy instead of applying recent anchor-based localization. We evaluate the proposed method on three challenging public benchmarks: ActivityNet Captions, TACoS, and Charades-STA. FIAN significantly outperforms the state-of-the-art approaches. Xiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 0001, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu |
ACM Multimedia | 1 |
| 2019 | Enhanced 3D convolutional networks for crowd counting
Zhikang Zou, Huiliang Shao, Xiaoye Qu, Wei Wei 0002, Pan Zhou 0001 |
BMVC | 3 |
| 2019 | Attend to count: Crowd counting with adaptive capacity multi-scale CNNs
Zhikang Zou, Yu Cheng 0001, Xiaoye Qu, Shouling Ji, Pan Zhou 0001 |
Neurocomputing | 3 |