VLDB 2026 Research / reviewers in the wild / expert
Long Chen 0016
dblp:64/5725-16
· DBLP profile ↗
113ranked-venue papers
11as first author
103since 2021 · last 2026
0000-0001-6148-9709ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 11 first-author · 74 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 9 first-author · 58 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation ComprehensionabstractRecent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic roles. The core reason is the lack of modeling for structural semantic dependencies among multi-entities, leading to over-reliance on language priors (e.g., defaulting to "person drinks a milk" if a person is merely holding it). To this end, we propose Relation-R1, the first unified relation comprehension framework that explicitly integrates cognitive chain-of-thought (CoT)-guided supervised fine-tuning (SFT) and group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we first establish foundational reasoning capabilities via SFT, enforcing structured outputs with thinking processes. Then, GRPO is utilized to refine these outputs via multi-rewards optimization, prioritizing visual-semantic grounding over language-induced biases, thereby improving generalization capability. Furthermore, we investigate the impact of various CoT strategies within this framework, demonstrating that a specific-to-general progressive approach in CoT guidance further improves generalization, especially in capturing synonymous N-ary relations. Extensive experiments on widely-used PSG and SWiG datasets demonstrate that Relation-R1 achieves state-of-the-art performance in both binary and N-ary relation understanding. Lin Li 0065, Wei Chen 0070, Jiahui Li 0003, Kwang-Ting Cheng, Long Chen 0016 |
AAAI | 5 |
| 2026 | Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningabstractComposed Image Retrieval (CIR) enables image search by combining a reference image with modification text. Intrinsic noise in CIR triplets incurs intrinsic uncertainty and threatens model's robustness. Probabilistic learning approaches have shown promise in addressing such issues; however, they fall short for CIR due to their instance-level holistic modeling and homogeneous treatments for queries and targets. This paper introduces a Heterogeneous Uncertainty-Guided (HUG) paradigm to overcome these limitations. HUG utilizes a fine-grained probabilistic learning framework, where queries and targets are represented by Gaussian embeddings capturing detailed concepts and uncertainties. We customize heterogeneous uncertainty estimations for multi-modal queries and uni-modal targets. Given a query, we capture uncertainties not only regarding uni-modal content quality but also multi-modal coordination, followed by a provable dynamic weighting mechanism to derive the comprehensive query uncertainty. We further design uncertainty-guided objectives, including query-target holistic contrast and fine-grained contrasts with comprehensive negative sampling strategies, which effectively enhance discriminative learning. Experiments on benchmarks demonstrate HUG's effectiveness beyond state-of-the-art baselines, with faithful analysis justifying the technical contributions. Haomiao Tang, Jinpeng Wang 0002, Minyi Zhao, Guanghao Meng, Ruisheng Luo, Long Chen 0016, Shutao Xia |
AAAI | 6 |
| 2026 | Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single ImageabstractPersonalizing 3D scenes from a single reference image enables intuitive user-guided editing, which requires achieving both multi-view consistency across perspectives and referential consistency with the input image. However, these goals are particularly challenging due to the viewpoint bias caused by the limited perspective provided in a single image. Lacking the mechanisms to effectively expand reference information beyond the original view, existing methods of image-conditioned 3DGS personalization often suffer from this viewpoint bias and struggle to produce consistent results. Therefore, in this paper, we present Consistent Personalization for 3D Gaussian Splatting (CP-GS), a framework that progressively propagates the single-view reference appearance to novel perspectives. In particular, CP-GS integrates pre-trained image-to-3D generation and iterative LoRA fine-tuning to extract and extend the reference appearance, and finally produces faithful multi-view guidance images and the personalized 3DGS outputs through a view-consistent generation process guided by geometric cues. Extensive experiments on real-world scenes show that our CP-GS effectively mitigates the viewpoint bias, achieving high-quality image-conditioned 3DGS personalization that significantly outperforms existing methods. Xuanyu Yi, Qingshan Xu 0001, Yuan Zhou 0016, Long Chen 0016, Hanwang Zhang |
AAAI | 5 |
| 2026 | Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting
Guangxing Han, Long Chen 0016, Jiawei Ma, Shiyuan Huang 0001, Rama Chellappa, Shih-Fu Chang |
Int. J. Comput. Vis. | 2 |
| 2026 | Multi-level Compositional Feature Augmentation for Unbiased Scene Graph GenerationabstractAbstract Scene Graph Generation (SGG) aims to detect all the visual relation triplets <, , > in a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recent years. However, due to the ubiquitous long-tailed predicate distributions, today’s SGG models are still easily biased to the head predicates. Currently, the most prevalent debiasing solutions for SGG are re-balancing methods, e . g ., changing the distributions of original training samples. In this paper, we argue that all existing re-balancing strategies fail to increase the diversity of the relation triplet features of each predicate, which is critical for robust SGG. To this end, we propose a novel M ulti-level C ompositional F eature A ugmentation ( MCFA ) strategy, which aims to mitigate the bias issue from the perspective of increasing the diversity of triplet features. Specifically, we enhance relationship diversity on not only feature-level , i . e ., replacing the intrinsic or extrinsic visual features of triplets with other correlated samples to create novel feature compositions for tail predicates, but also image-level , i . e ., manipulating the image to generate brand new visual appearance for triplets. Due to its model-agnostic nature, MCFA can be seamlessly incorporated into various SGG frameworks. Extensive ablations have shown that MCFA achieves a new state-of-the-art performance on the trade-off between different metrics. Lin Li 0065, Long Chen 0016 |
Int. J. Comput. Vis. | 5 |
| 2026 | Physically Plausible Human-Object Rendering From Sparse Views via 3D Gaussian SplattingabstractRendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applications. Existing methods often struggle to simultaneously achieve high rendering quality, physical plausibility, and computational efficiency. To address these limitations, we propose HOGS (Human-Object Rendering via 3D Gaussian Splatting), a novel framework for efficient HOI rendering with physically plausible geometric constraints from sparse views. HOGS represents both humans and objects as dynamic 3D Gaussians. Central to HOGS is a novel optimization process that operates directly on these Gaussians to enforce geometric consistency (i.e., preventing inter-penetration or floating contacts) to achieve physical plausibility. To support this core optimization under sparse-view ambiguity, our framework incorporates two pre-trained modules: an optimization-guided Human Pose Refiner for robust estimation under sparse-view occlusions, and a Human-Object Contact Predictor that efficiently identifies interaction regions to guide our novel contact and separation losses. Extensive experiments on both human-object and hand-object interaction datasets demonstrate that HOGS achieves state-of-the-art rendering quality and maintains high computational efficiency. Jun Xiao 0001, Yi Yang 0001, Yueting Zhuang, Long Chen 0016 |
IEEE Trans. Image Process. | 5 |
| 2026 | Multi-Way Cascade-Attention Network for Multi-Modal Sequential RecommendationabstractSequential recommendation has become a hot topic, which aims to predict the desired items for each user based on his/her historical actions. The mainstream advancements for this task focus on modeling user behaviors in a pure item ID-based manner, which often fail to provide satisfactory results due to data sparsity and cold-start issues. Recently, several studies that leverage multi-modal information (i.e., multi-modal sequential recommendation) have shed light on alleviating such issues. However, we argue that three limitations are still not well addressed: 1) they usually extract ID modality features of an item with an one-hot encoding, which does not include any semantic information; 2) they fail to effectively mitigate the semantic gap issue and explicitly explore the asynchronous interplay between any two modalities; and 3) during the model prediction stage, they neglect the significance of adaptively fusing multi-modal embeddings for each user. To address such defects, we propose a novel framework for multi-modal sequential recommendation, namely, a multi-way cascade-attention network (MCN). Specifically, we apply a lightweight graph propagation network to derive informative representations of the ID-oriented modality, explicitly encoding collaborative signals in the user-item interaction graph. Next, we develop a multi-way cascade-attention module (CAM) to accomplish user behavior sequence alignment across different modality spaces. Each CAM consists of a cross-attention block followed by a series of self-attention blocks. The former encodes the asynchronous interplay between two modalities, while the latter captures intra-modal temporal dependencies. Finally, we design a modality-aware attentive strategy to dynamically fuse the user's dynamic interests across different modality spaces. Our extensive experiments on four public datasets demonstrate the superiority of MCN over recent state-of-the-art recommenders. Bin Wu 0019, Long Chen 0016, Yuheng Fu, Yunshan Ma 0002, Mingliang Xu 0001, Tat-Seng Chua |
IEEE Trans. Multim. | 2 |
| 2026 | FreeTuner: Any Subject in Any Style with Training-free DiffusionabstractWith the advance of diffusion models, various personalized image generation methods have been proposed. However, almost all existing work only focuses on either subject-driven or style-driven personalization. Meanwhile, state-of-the-art methods face several challenges in realizing compositional personalization , that is, composing different subject and style concepts, such as concept disentanglement, unified reconstruction paradigm, and insufficient training data. To address these issues, we introduce FreeTuner , a flexible and training-free method for compositional personalization that can generate any user-provided subject in any user-provided style . Our approach employs a disentanglement strategy that separates the generation process into two stages to effectively mitigate concept entanglement. FreeTuner leverages the intermediate features within the diffusion model for subject concept representation and introduces style guidance to align the synthesized images with the style concept, ensuring the preservation of both the subject’s structure and the style’s aesthetic features. Extensive experiments have demonstrated the generation ability of FreeTuner across various personalization settings. Youcan Xu, Zhen Wang 0004, Jun Xiao 0001, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Learning Causal Transition Matrix for Instance-dependent Label NoiseabstractNoisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. However, the transition matrix is often considered unidentifiable. One strand of methods typically addresses this problem by assuming that the transition matrix is instance-independent; that is, the probability of mislabeling a particular instance is not influenced by its characteristics or attributes. This assumption is clearly invalid in complex real-world scenarios. To better understand the transition relationship and relax this assumption, we propose to study the data generation process of noisy labels from a causal perspective. We discover that an unobservable latent variable can affect either the instance itself, the label annotation procedure, or both, which complicates the identification of the transition matrix. To address various scenarios, we have unified these observations within a new causal graph. In this graph, the input instance is divided into a noise-resistant component and a noise-sensitive component based on whether they are affected by the latent variable. These two components contribute to identifying the “causal transition matrix”, which approximates the true transition matrix with theoretical guarantee. In line with this, we have designed a novel training framework that explicitly models this causal relationship and, as a result, achieves a more accurate model for inferring the clean label. Jiahui Li 0003, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011 |
AAAI | 5 |
| 2025 | Open-World Multimodal Understanding and Generation with Efficiently Finetuned Foundation ModelsabstractWith the astonishing ability of different pretrained foundation models (e.g., large language models (LLMs), vision-language models, diffusion models), today’s AI research and development tendency has been revolutionized. In this talk, I will answer two questions: Q1: How can we efficiently train or fine-tune foundation models? Q2: How can we build strong open-world multimodal understanding and generation models with these pretrained foundation models? Long Chen 0016 |
AAAI | 1 |
| 2025 | Modeling Uncertainty in Composed Image Retrieval via Probabilistic EmbeddingsabstractHaomiao Tang, Jinpeng Wang, Yuang Peng, GuangHao Meng, Ruisheng Luo, Bin Chen, Long Chen, Yaowei Wang, Shu-Tao Xia. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haomiao Tang, Jinpeng Wang 0002, Yuang Peng, Guanghao Meng, Ruisheng Luo, Bin Chen 0011, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
ACL (1) | 7 |
| 2025 | Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningabstractVisual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which in practice may not hold true. Multiple suitable prompts may exist, but individually they often fall short, leading to difficulties in selection and the exclusion of useful context. To address this, we propose a new perspective: prompt condensation. Rather than relying on a single prompt, candidate prompts collaborate to efficiently integrate informative contexts without sacrificing resolution. We devise Condenser, a lightweight external plugin that compresses relevant fine-grained context across multiple prompts. Optimized end-to-end with the backbone, Condenser ensures accurate integration of contextual cues. Experiments demonstrate Condenser outperforms state-of-the-arts across benchmark tasks, showing superior context compression, scalability with more prompts, and enhanced computational efficiency compared to ensemble methods, positioning it as a highly competitive solution for VICL. Code is open-sourced at https://github.com/gimpong/CVPR25-Condenser. Jinpeng Wang 0002, Tianci Luo, Yaohua Zha, Ruisheng Luo, Bin Chen 0011, Tao Dai 0001, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
CVPR | 8 |
| 2025 | IterIS: Iterative Inference-Solving Alignment for LoRA MergingabstractLow-rank adaptations (LoRA) are widely used to fine-tune large models across various domains for specific downstream tasks. While task-specific LoRAs are often available, concerns about data privacy and intellectual property can restrict access to training data, limiting the acquisition of a multi-task model through gradient-based training. In response, LoRA merging presents an effective solution by combining multiple LoRAs into a unified adapter while maintaining data privacy. Prior works on LoRA merging primarily frame it as an optimization problem, yet these approaches face several limitations, including the rough assumption about input features utilized in optimization, massive sample requirements, and the unbalanced optimization objective. These limitations can significantly degrade performance. To address these, we propose a novel optimization-based method, named IterIS: 1) We formulate LoRA merging as an advanced optimization problem to mitigate the rough assumption. Additionally, we employ an iterative inference-solving framework in our algorithm. It can progressively refine the optimization objective for improved performance. 2) We introduce an efficient regularization term to reduce the need for massive sample requirements (requiring only 1-5% of the unlabeled samples compared to prior methods). 3) We utilize adaptive weights in the optimization objective to mitigate potential unbalances in LoRA merging process. Our method demonstrates significant improvements over multiple baselines and state-of-the-art methods in composing tasks for text-to-image diffusion, vision-language models, and large language models. Furthermore, our layer-wise algorithm can achieve convergence with minimal steps, ensuring efficiency in both memory and computation. Hongxu Chen 0002, Zhen Wang 0004, Runshi Li, Bowei Zhu, Long Chen 0016 |
CVPR | 5 |
| 2025 | CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationabstractInterleaved image-text generation has emerged as a vital multimodal task aimed at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit narrative coherence and entity and style consistency remains challenging due to poor training data quality. To this end, we introduce CoMM, a high-quality Coherent interleaved image-text MultiModal dataset designed to enhance the coherence, consistency, and alignment of generated multimodal content. Initially, CoMM harnesses raw data from diverse sources, focusing on instructional content and visual storytelling, establishing a foundation for coherent and consistent content. To further refine the data quality, we devise a multi-perspective filter strategy that leverages advanced pre-trained models to ensure the development of sentences, consistency of inserted images, and semantic alignment between them. Various quality evaluation metrics are designed to prove the high quality of the filtered dataset. Meanwhile, extensive few-shot experiments on various downstream tasks demonstrate CoMM’s effectiveness in significantly enhancing the in-context learning capabilities of MLLMs. Moreover, we propose four new tasks to evaluate MLLMs’ interleaved generation abilities, supported by a comprehensive evaluation framework. We believe CoMM opens a new avenue for advanced MLLMs with superior multimodal in-context learning and understanding ability. Wei Chen 0070, Lin Li 0065, Yongqi Yang, Fan Yang 0094, Tingting Gao, Yu Wu 0011, Long Chen 0016 |
CVPR | 8 |
| 2025 | Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse RewardsabstractDiffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1. Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001 |
CVPR | 3 |
| 2025 | Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce ClassificationabstractData Augmentation (DA), i.e., synthesizing faithful and diverse samples to expand the original training set, is an effective strategy to improve the performance of various data-scarce tasks. With the powerful image generation ability, diffusion-based DA has shown strong performance gains on different image classification benchmarks. In this paper, we analyze today’s diffusion-based DA methods, and argue that they cannot take account of both faithfulness and diversity, which are two critical keys for generating high-quality samples and boosting classification performance. To this end, we propose a novel Diffusion-based DA method: Diff-II. Specifically, it consists of three steps: 1) Category concepts learning: Learning concept embeddings for each category. 2) Inversion interpolation: Calculating the inversion for each image, and conducting circle interpolation for two randomly sampled inversions from the same category. 3) Two-stage denoising: Using different prompts to generate synthesized images in a coarse-to-fine manner. Extensive experiments on various data-scarce image classification tasks (e.g., few-shot, long-tailed, and out-of-distribution classification) have demonstrated their effectiveness over state-of-the-art diffusion-based DA methods. Yanghao Wang, Long Chen 0016 |
CVPR | 2 |
| 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward RedistributionabstractReinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences.Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase.However, current reward models operate as sequence-to-one models, allocating a single, sparse, and delayed reward to an entire output sequence.This approach may overlook the significant contributions of individual tokens toward the desired outcome.To this end, we propose a more finegrained, token-level guidance approach for RL training.Specifically, we introduce RED, a novel REward reDistribition method that evaluates and assigns specific credit to each token using an off-the-shelf reward model.Utilizing these fine-grained rewards enhances the model's understanding of language nuances, leading to more precise performance improvements.Notably, our method does not require modifying the reward model or introducing additional training steps, thereby incurring minimal computational costs.Experimental results across diverse datasets and tasks demonstrate the superiority of our approach. Jiahui Li 0003, Lin Li 0065, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011, Cheng Yang 0002 |
EMNLP | 5 |
| 2025 | KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual ClassificationabstractFine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive and resource-intensive as models grow larger. Recently, parameter-efficient fine-tuning (PEFT) methods have emerged as a popular solution to improve training efficiency and reduce storage needs by tuning additional low-rank modules within pre-trained backbones. Despite their advantages, they struggle with limited representation capabilities and misalignment with pre-trained intermediate features. To address these issues, we introduce an innovative Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission (KARST) for various recognition tasks. Specifically, KARST’s multi-kernel design extends Kronecker projections horizontally and separates adaptation matrices into multiple complementary spaces, reducing parameter dependency and creating more compact subspaces. Besides, it incorporates extra learnable re-scaling factors to better align with pre-trained feature distributions, allowing for more flexible and balanced feature aggregation. Extensive experiments on diverse downstream datasets validate that our KARST not only outperforms other PEFT counterparts across model types and data domains, but also surpasses full fine-tuning with a negligible inference cost due to its re-parameterization characteristics. Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Long Chen 0016, Huchuan Lu |
ICASSP | 4 |
| 2025 | Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
Jun Li 0131, Jinpeng Wang 0002, Chaolei Tan, Niu Lian, Long Chen 0016, Yaowei Wang 0001, Min Zhang 0005, Shutao Xia, Bin Chen 0011 |
ICCV | 5 |
| 2025 | Nautilus: Locality-Aware Autoencoder for Scalable Mesh GenerationabstractTriangle meshes are fundamental to 3D applications, enabling efficient modification and rasterization while maintaining compatibility with standard rendering pipelines. However, current automatic mesh generation methods typically rely on intermediate representations that lack the continuous surface quality inherent to meshes. Converting these representations into meshes produces dense, suboptimal outputs. Although recent autoregressive approaches demonstrate promise in directly modeling mesh vertices and faces, they are constrained by the limitation in face count, scalability, and structural fidelity. To address these challenges, we propose Nautilus, a locality-aware autoencoder for artist-like mesh generation that leverages the local properties of manifold meshes to achieve structural fidelity and efficient representation. Our approach introduces a novel tokenization algorithm that preserves face proximity relationships and compresses sequence length through locally shared vertices and edges, enabling the generation of meshes with an unprecedented scale of up to 5,000 faces. Furthermore, we develop a Dual-stream Point Conditioner that provides multi-scale geometric guidance, ensuring global consistency and local structural fidelity by capturing fine-grained geometric features. Extensive experiments demonstrate that Nautilus significantly outperforms state-of-the-art methods in both fidelity and scalability. The project page is at https://nautilusmeshgen.github.io. Xuanyu Yi, Haohan Weng, Qingshan Xu 0001, Xiaokang Wei, Xianghui Yang, Chunchao Guo, Long Chen 0016, Hanwang Zhang |
ICCV | 8 |
| 2025 | CLIPDrag: Combining Text-based and Drag-based Instructions for Image EditingabstractPrecise and flexible image editing remains a fundamental challenge in computer vision. Based on the modified areas, most editing methods can be divided into two main types: global editing and local editing. In this paper, we choose the two most common editing approaches (\ie text-based editing and drag-based editing) and analyze their drawbacks. Specifically, text-based methods often fail to describe the desired modifications precisely, while drag-based methods suffer from ambiguity. To address these issues, we proposed \textbf{CLIPDrag}, a novel image editing method that is the first to combine text and drag signals for precise and ambiguity-free manipulations on diffusion models. To fully leverage these two signals, we treat text signals as global guidance and drag points as local information. Then we introduce a novel global-local motion supervision method to integrate text signals into existing drag-based methods by adapting a pre-trained language-vision model like CLIP. Furthermore, we also address the problem of slow convergence in CLIPDrag by presenting a fast point-tracking method that enforces drag points moving toward correct directions. Extensive experiments demonstrate that CLIPDrag outperforms existing single drag-based methods or text-based methods. Ziqi Jiang, Zhen Wang 0004, Long Chen 0016 |
ICLR | 3 |
| 2025 | DisPose: Disentangling Pose Guidance for Controllable Human Image AnimationabstractControllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignment. However, such strict dense guidance impairs the quality of the generated video when the body shape of the reference character differs significantly from that of the driving video. In this paper, we present DisPose to mine more generalizable and effective control signals without additional dense input, which disentangles the sparse skeleton pose in human image animation into motion field guidance and keypoint correspondence. Specifically, we generate a dense motion field from a sparse motion field and the reference image, which provides region-level dense guidance while maintaining the generalization of the sparse pose control. We also extract diffusion features corresponding to pose keypoints from the reference image, and then these point features are transferred to the target pose to provide distinct identity information. To seamlessly integrate into existing models, we propose a plug-and-play hybrid ControlNet that improves the quality and consistency of generated videos while freezing the existing model parameters. Extensive qualitative and quantitative experiments demonstrate the superiority of DisPose compared to current methods. Project page: https://github.com/lihxxx/DisPose. Hongxiang Li 0004, Yaowei Li 0001, Zhihong Zhu 0001, Xuxin Cheng, Long Chen 0016 |
ICLR | 7 |
| 2025 | Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object DetectionabstractIn pursuit of detecting unstinted objects that extend beyond predefined categories, prior arts of open-vocabulary object detection (OVD) typically resort to pretrained vision-language models (VLMs) for base-to-novel category generalization. However, to mitigate the misalignment between upstream image-text pretraining and downstream region-level perception, additional supervisions are indispensable, e.g., image-text pairs or pseudo annotations generated via self-training strategies. In this work, we propose CCKT-Det trained without any extra supervision. The proposed framework constructs a cyclic and dynamic knowledge transfer from language queries and visual region features extracted from VLMs, which forces the detector to closely align with the visual-semantic space of VLMs. Specifically, 1) we prefilter and inject semantic priors to guide the learning of queries, and 2) introduce a regional contrastive loss to improve the awareness of queries on novel objects. CCKT-Det can consistently improve performance as the scale of VLMs increases, all while requiring the detector at a moderate level of computation overhead. Comprehensive experimental results demonstrate that our method achieves performance gain of +2.9% and +10.2% AP_{50} over previous state-of-the-arts on the challenging COCO benchmark, both without and with a stronger teacher model. Pingcheng Dong, Long Chen 0016 |
ICLR | 4 |
| 2025 | Multi-Resolution Decomposable Diffusion Model for Non-Stationary Time Series Anomaly DetectionabstractRecently, generative models have shown considerable promise in unsupervised time series anomaly detection. Nonetheless, the task of effectively capturing complex temporal patterns and minimizing false alarms becomes increasingly challenging when dealing with non-stationary time series, characterized by continuously fluctuating statistical attributes and joint distributions. To confront these challenges, we underscore the benefits of multi-resolution modeling, which improves the ability to distinguish between anomalies and non-stationary behaviors by leveraging correlations across various resolution scales. In response, we introduce a **M**ulti-Res**o**lution **De**composable Diffusion **M**odel (MODEM), which integrates a coarse-to-fine diffusion paradigm with a frequency-enhanced decomposable network to adeptly navigate the intricacies of non-stationarity. Technically, the coarse-to-fine diffusion model embeds cross-resolution correlations into the forward process to optimize diffusion transitions mathematically. It then innovatively employs low-resolution recovery to guide the reverse trajectories of high-resolution series in a coarse-to-fine manner, enhancing the model's ability to learn and elucidate underlying temporal patterns. Furthermore, the frequency-enhanced decomposable network operates in the frequency domain to extract globally shared time-invariant information and time-variant temporal dynamics for accurate series reconstruction. Extensive experiments conducted across five real-world datasets demonstrate that our proposed MODEM achieves state-of-the-art performance and can be generalized to other time series tasks. Guojin Zhong, Pan Wang 0011, Jin Yuan 0002, Zhiyong Li 0001, Long Chen 0016 |
ICLR | 5 |
| 2025 | Event-Customized Image GenerationabstractCustomized Image Generation, generating customized images with user-specified concepts, has raised significant attention due to its creativity and novelty. With impressive progress achieved in subject customization, some pioneer works further explored the customization of action and interaction beyond entity (i.e., human, animal, and object) appearance. However, these approaches only focus on basic actions and interactions between two entities, and their effects are limited by insufficient ”exactly same” reference images. To extend customized image generation to more complex scenes for general real-world applications, we propose a new task: event-customized image generation. Given a single reference image, we define the ”event” as all specific actions, poses, relations, or interactions between different entities in the scene. This task aims at accurately capturing the complex event and generating customized images with various target entities. To solve this task, we proposed a novel training-free event customization method: FreeEvent. Specifically, FreeEvent introduces two extra paths alongside the general diffusion denoising process: 1) Entity switching path: it applies cross-attention guidance and regulation for target entity generation. 2) Event transferring path: it injects the spatial feature and self-attention maps from the reference image to the target image for event generation. To further facilitate this new task, we collected two evaluation benchmarks: SWiG-Event and Real-Event. Extensive experiments and ablations have demonstrated the effectiveness of FreeEvent. Zhen Wang 0004, Yilei Jiang, Jun Xiao 0001, Long Chen 0016 |
ICML | 5 |
| 2025 | Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache SharingabstractWith the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an autoregressive manner, i.e., generating subsequent clips conditioned on the last frame(s) of the previous clip. However, existing autoregressive VDMs are highly inefficient and redundant: The model must re-compute all the conditional frames that are overlapped between adjacent clips. This issue is exacerbated when the conditional frames are extended autoregressively to provide the model with long-term context. In such cases, the computational demands increase significantly (i.e., with a quadratic complexity w.r.t. the autoregression step). In this paper, we propose **Ca2-VDM**, an efficient autoregressive VDM with **Ca**usal generation and **Ca**che sharing. For **causal generation**, it introduces unidirectional feature computation, which ensures that the cache of conditional frames can be precomputed in previous autoregression steps and reused in every subsequent step, eliminating redundant computations. For **cache sharing**, it shares the cache across all denoising steps to avoid the huge cache storage cost. Extensive experiments demonstrated that our Ca2-VDM achieves state-of-the-art quantitative and qualitative video generation results and significantly improves the generation speed. Code is available: https://github.com/Dawn-LX/CausalCache-VDM Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016 |
ICML | 6 |
| 2025 | SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationabstractAudio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring semantic information, such as the classes of sounding sources present in the audio, limiting their ability to generate videos with accurate content and spatial composition. In contrast, we humans can not only naturally identify the semantic categories of sounding sources but also determine their deeply encoded spatial attributes, including locations and movement directions. This useful information can be elucidated by considering specific spatial indicators derived from the inherent physical properties of sound, such as loudness or frequency. As prior methods largely ignore this factor, we present SpA2V, the first framework explicitly exploits these spatial auditory cues from audios to generate videos with high semantic and spatial correspondence. SpA2V decomposes the generation process into two stages: 1) Audio-guided Video Planning: We meticulously adapt a state-of-the-art MLLM for a novel task of harnessing spatial and semantic cues from input audio to construct Video Scene Layouts (VSLs). This serves as an intermediate representation to bridge the gap between the audio and video modalities. 2) Layout-grounded Video Generation: We develop an efficient and effective approach to seamlessly integrate VSLs as conditional guidance into pre-trained diffusion models, enabling VSL-grounded video generation in a training-free manner. Extensive experiments demonstrate that SpA2V excels in generating realistic videos with semantic and spatial alignment to the input audios. Kien T. Pham 0001, Yingqing He, Yazhou Xing, Qifeng Chen 0001, Long Chen 0016 |
ACM Multimedia | 5 |
| 2025 | Compositional Zero-shot Learning via Progressive Language-based ObservationsabstractCompositional zero-shot learning aims to recognize unseen stateobject compositions by leveraging known primitives (state and object) during training. However, effectively modeling interactions between primitives and generalizing knowledge to novel compositions remains a perennial challenge. There are two crucial factors: large object-conditioned and state-conditioned variance, i.e., the appearance of states (or objects) can vary significantly when combined with different objects (or states). For instance, the state "old" can signify vintage design for a "car" or advanced age for a "cat". In this paper, we argue that these variances can be mitigated by predicting composition categories based on salient observation cues. Therefore, we propose Progressive Language-based Observations (PLO), which can automatically determine the order of observation cues. These "observation cues" comprise a series of primitive concepts or graduated descriptions that allow the model to understand image content in a step-by-step manner. Specifically, PLO adopts pre-trained vision-language models (VLMs) to empower the model with observation capabilities.We further devise two variants: a twostep method (PLO-VLM) with a pre-observing classifier dynamically selecting the order of primitive concept-based cues, and a multistep approach (PLO-LLM) using large language models (LLMs) to craft graduated description-based cues. Extensive tests on three datasets show PLO's effectiveness in compositional recognition. Lin Li 0065, Guikun Chen, Zhen Wang 0004, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 5 |
| 2025 | Zero-shot Compositional Action Recognition with Neural Logic ConstraintsabstractZero-shot compositional action recognition (ZS-CAR) aims to identify unseen verb-object compositions in the videos by exploiting the learned knowledge of verb and object primitives during training. Despite compositional learning's progress in ZS-CAR, two critical challenges persist: 1) Missing compositional structure constraint, leading to spurious correlations between primitives; 2) Neglecting semantic hierarchy constraint, leading to semantic ambiguity and impairing the training process. In this paper, we argue that human-like symbolic reasoning offers a principled solution to these challenges by explicitly modeling compositional and hierarchical structured abstraction. To this end, we propose a logic-driven ZS-CAR framework LogicCAR that integrates dual symbolic constraints: Explicit Compositional Logic and Hierarchical Primitive Logic. Specifically, the former models the restrictions within the compositions, enhancing the compositional reasoning ability of our model. The latter investigates the semantical dependencies among different primitives, empowering the models with fine-to-coarse reasoning capacity. By formalizing these constraints in first-order logic and embedding them into neural network architectures, LogicCAR systematically bridges the gap between symbolic abstraction and existing models. Extensive experiments on the Sth-com dataset demonstrate that our LogicCAR outperforms existing baseline methods, proving the effectiveness of our logic-driven constraints. Gefan Ye, Lin Li 0065, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 5 |
| 2025 | Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language ModelsabstractAlthough multimodal large language models (MLLMs) exhibit remarkable reasoning capabilities on complex multimodal understanding tasks, they still suffer from the notorious 'hallucination' issue: generating outputs misaligned with obvious visual or factual evidence. Currently, training-based solutions, like direct preference optimization (DPO), leverage paired preference data to suppress hallucinations. However, they risk sacrificing general reasoning capabilities due to the likelihood displacement. Meanwhile, training-free solutions, like contrastive decoding, achieve this goal by subtracting the estimated hallucination pattern from a distorted input. Yet, these handcrafted perturbations (e.g., add noise to images) may poorly capture authentic hallucination patterns. To avoid these weaknesses of existing methods, and realize ``robust'' hallucination mitigation (\ie, maintaining general reasoning performance), we propose a novel framework: Decoupling Contrastive Decoding (DCD). Specifically, DCD decouples the learning of positive and negative samples in preference datasets, and trains separate positive and negative image projections within the MLLM. The negative projection implicitly models real hallucination patterns, which enables vision-aware negative images in the contrastive decoding inference stage. Our DCD alleviates likelihood displacement by avoiding pairwise optimization and generalizes robustly without handcrafted degradation. Extensive ablations across hallucination benchmarks and general reasoning tasks demonstrate the effectiveness of DCD, \ie, it matches DPO’s hallucination suppression while preserving general capabilities and outperforms the handcrafted contrastive decoding methods. Wei Chen 0070, Fan Yang 0094, Tingting Gao, Di Zhang 0026, Long Chen 0016 |
NeurIPS | 7 |
| 2025 | Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph GenerationabstractOpen-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Existing OVSGG methods always adopt a two-stage pipeline: 1) Infusing knowledge into large-scale models via pre-training on large datasets; 2) Transferring knowledge from pre-trained models with fully annotated scene graphs during supervised fine-tuning. However, due to a lack of explicit interaction modeling, these methods struggle to distinguish between interacting and non-interacting instances of the same object category. This limitation induces critical issues in both stages of OVSGG: it generates noisy pseudo-supervision from mismatched objects during knowledge infusion, and causes ambiguous query matching during knowledge transfer. To this end, in this paper, we propose an interACtion-Centric end-to-end OVSGG framework (ACC) in an interaction-driven paradigm to minimize these mismatches. For interaction-centric knowledge infusion, ACC employs a bidirectional interaction prompt for robust pseudo-supervision generation to enhance the model's interaction knowledge. For interaction-centric knowledge transfer, ACC first adopts interaction-guided query selection that prioritizes pairing interacting objects to reduce interference from non-interacting ones. Then, it integrates interaction-consistent knowledge distillation to bolster robustness by pushing relational foreground away from the background while retaining general knowledge. Extensive experimental results on three benchmarks show that ACC achieves state-of-the-art performance, demonstrating the potential of interaction-centric paradigms for real-world applications. Lin Li 0065, Long Chen 0016 |
NeurIPS | 6 |
| 2025 | Noise Matters: Optimizing Matching Noise for Diffusion ClassifiersabstractAlthough today's pretrained discriminative vision-language models (e.g., CLIP) have demonstrated strong perception abilities, such as zero-shot image classification, they also suffer from the bag-of-words problem and spurious bias. To mitigate these problems, some pioneering studies leverage powerful generative models (e.g., pretrained diffusion models) to realize generalizable image classification, dubbed Diffusion Classifier (DC). Specifically, by randomly sampling a Gaussian noise, DC utilizes the differences of denoising effects with different category conditions to classify categories. Unfortunately, an inherent and notorious weakness of existing DCs is noise instability: different random sampled noises lead to significant performance changes. To achieve stable classification performance, existing DCs always ensemble the results of hundreds of sampled noises, which significantly reduces the classification speed. To this end, we firstly explore the role of noise in DC, and conclude that: there are some ``good noises'' that can relieve the instability. Meanwhile, we argue that these good noises should meet two principles: 1) Frequency Matching: noise should destroy the specific frequency signals; 2) Spatial Matching: noise should destroy the specific spatial areas. Regarding both principles, we propose a novel Noise Optimization method to learn matching (i.e., good) noise for DCs: NoOp. For frequency matching, NoOp first optimizes a dataset-specific noise: Given a dataset and a timestep $t$, optimize one randomly initialized parameterized noise. For Spatial Matching, NoOp trains a Meta-Network that adopts an image as input and outputs image-specific noise offset. The sum of optimized noise and noise offset will be used in DC to replace random noise. Extensive ablations on various datasets demonstrated the effectiveness of NoOp. It is worth noting that our noise optimization is orthogonal to existing optimization methods (e.g., prompt tuning), our NoOP can even benefit from these methods to further boost performance. Yanghao Wang, Long Chen 0016 |
NeurIPS | 2 |
| 2025 | From Easy to Hard: Learning Curricular Shape-Aware Features for Robust Panoptic Scene Graph Generation
Hanrong Shi, Lin Li 0065, Jun Xiao 0001, Yueting Zhuang, Long Chen 0016 |
Int. J. Comput. Vis. | 5 |
| 2025 | Learning Combinatorial Prompts for Universal Controllable Image CaptioningabstractAbstract Controllable Image Captioning (CIC)—generating natural language descriptions about images under the guidance of given control signals—is one of the most promising directions toward next-generation captioning systems. Till now, various kinds of control signals for CIC have been proposed, ranging from content-related control to structure-related control. However, due to the format and target gaps of different control signals, all existing CIC works (or architectures) only focus on one certain control signal, and overlook the human-like combinatorial ability. By “combinatorial", we mean that our humans can easily meet multiple needs (or constraints) simultaneously when generating descriptions. To this end, we propose a novel prompt-based framework for CIC by learning Com binatorial Pro mpts, dubbed as ComPro . Specifically, we directly utilize a pretrained language model GPT-2 Radford et al. (OpenAI blog 1:9, 2019) as our language model, which can help to bridge the gap between different signal-specific CIC architectures. Then, we reformulate the CIC as a prompt-guide sentence generation problem, and propose a new lightweight prompt generation network to generate the combinatorial prompts for different kinds of control signals. For different control signals, we further design a new mask attention mechanism to realize the prompt-based CIC. Due to its simplicity, our ComPro can be further extended to more kinds of combined control signals by concatenating these prompts. Extensive experiments on two prevalent CIC benchmarks have verified the effectiveness and efficiency of our ComPro on both single and combined control signals. Zhen Wang 0004, Jun Xiao 0001, Yueting Zhuang, Fei Gao 0014, Jian Shao 0001, Long Chen 0016 |
Int. J. Comput. Vis. | 6 |
| 2025 | Knowledge Integration for Grounded Situation Recognition
Jiaming Lei, Sijing Wu, Lin Li 0065, Lei Chen 0082, Jun Xiao 0001, Yi Yang 0001, Long Chen 0016 |
Pattern Recognit. | 7 |
| 2025 | Disentangled Graph Contrastive Learning for Socially-Aware Next-Item RecommendationabstractNext-item recommendation has received considerable attention in academia and industry, which aims to predict the next desired item for each user based on his historical behaviors. It has been validated that user behaviors are driven by two key factors:social influence, which leverages social relationships to better infer user preference, andsequential influence, which captures item transition patterns to model the dynamic evolution of user interest. While previous next-item recommenders have made great progress, we argue that there still exist three critical limitations: (1) they always follow the paradigm of entangling social and sequential influences, resulting in poor interpretability; (2) they fail to explicitly capture high-order influences at social-level and sequential-level, resulting in the suboptimal performance; (3) they fail to dynamically distinguish the importance of each influence factor when predicting user preference. To settle these three defects, we contribute a novel solution for socially-aware next-item recommendation, namely Disentangled Graph Contrastive Learning (DGCL) method, which explicitly disentangles social and sequential influences on user behavior data. Specifically, we first reorganize social relationships and all users' behavior sequences as two separate graphs. Then, a disentangled graph propagation module is developed on these two graphs to independently capture high-order influences at social-level and sequential-level. Furthermore, we formalize a dual contrastive learning paradigm as an auxiliary task to supervise the thorough disentanglement. Finally, we devise a user-specific attention mechanism to adaptively differentiate the importance of each influence factor for model prediction. Empirical results on four benchmark datasets demonstrate the superiority of DGCL over recent state-of-the-art recommenders. Further analysis verifies the rationality and necessity of each part in our solution. Our implemented codes and used datasets are available athttps://github.com/wubinzzu/DGCL. Bin Wu 0019, Xun Su, Long Chen 0016, Jing J. Liang, Yangdong Ye |
IEEE Trans. Big Data | 3 |
| 2025 | Eliminating Semantic Ambiguity in Human Pose Estimation via Stable Feature UpsamplingabstractHuman pose estimation is a challenging research task in the computer vision community due to the semantic ambiguity problem caused by inevitable occlusions, varying body shapes, and complex articulations. Although deep learning-based methods have significantly improved the performance of this task, existing feature upsampling operations,e.g., bilinear interpolation and transposed convolution, within current convolutional neural networks and Transformer frameworks suffer from a multitude of limitations, including the inability to adapt to specific tasks and the loss of fine-grained semantic details. In this work, we propose a simple yet effective two-step stable feature upsampling (SIU) strategy that addresses these limitations by leveraging a learnable and efficient upsampling operation. Specifically, wefirstapply periodic shuffling to increase the resolution of the feature maps.Secondly, we utilize convolution layers to adjust the size of feature channels to match those of the input feature maps. The proposed SIU enables the entire network to adapt to the specific feature requirements of the human pose estimation task, making it more effective in preserving spatial information. Quantitatively, extensive experimental results on the challenging COCO-WholeBody dataset validate that our approach outperforms state-of-the-art methods accurately and efficiently, and possesses strong transferability, making it applicable to a wide range of baselines. Moreover, the qualitative results validate that SIU can effectively eliminate the semantic ambiguity problem in challenging pose scenarios, such as occlusions and overlapping1. Rui Yan 0010, Xiangbo Shu, Pingcheng Dong, Long Chen 0016, Xiaoyu Du 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior ModelingabstractLong-term user behavior sequences are a goldmine for businesses to explore users’ interests to improve Click-Through Rate (CTR). However, it is very challenging to accurately capture users’ long-term interests from their long-term behavior sequences and give quick responses from the online serving systems. To meet such requirements, existing methods “inadvertently” destroy two basic requirements in long-term sequence modeling:R1) make full use of the entire sequence to keep the information as much as possible;R2) extract information from the most relevant behaviors to keep high relevance between learned interests and current target items. The performance of online serving systems is significantly affected by incomplete and inaccurate user interest information obtained by existing methods. To this end, we propose an efficient two-stage long-term sequence modeling approach, named asEfficieNtClustering based twO-stage interest moDEling (ENCODE), consisting of offline extraction stage and online inference stage. It not only meets the aforementioned two basic requirements but also achieves a desirable balance between online service efficiency and precision. Specifically, in the offline extraction stage, ENCODE clusters the entire behavior sequence and extracts accurate interests. To reduce the overhead of the clustering process, we design a metric learning-based dimension reduction algorithm that preserves the relative pairwise distances of behaviors in the new feature space. While in the online inference stage, ENCODE takes the off-the-shelf user interests to predict the associations with target items. Besides, to further ensure the relevance between user interests and target items, we adopt the same relevance metric throughout the whole pipeline of ENCODE. The extensive experiment and comparison with SOTA on both industrial and public datasets have demonstrated the effectiveness and efficiency of our proposed ENCODE. Yuhang Zheng 0003, Yinfu Feng, Yunan Ye, Rong Xiao 0005, Long Chen 0016, Xiaosong Yang, Jun Xiao 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Cross-Modal Conditioned Reconstruction for Language-Guided Medical Image SegmentationabstractRecent developments underscore the potential of textual information in enhancing learning models for a deeper understanding of medical visual semantics. However, language-guided medical image segmentation still faces a challenging issue. Previous works employ implicit architectures to embed textual information. This leads to segmentation results that are inconsistent with the semantics represented by the language, sometimes even diverging significantly. To this end, we propose a novel cross-modal conditioned Reconstruction for Language-guided Medical Image Segmentation (RecLMIS) to explicitly capture cross-modal interactions, which assumes that well-aligned medical visual features and medical notes can effectively reconstruct each other. We introduce conditioned interaction to adaptively predict patches and words of interest. Subsequently, they are utilized as conditioning factors for mutual reconstruction to align with regions described in the medical notes. Extensive experiments demonstrate the superiority of our RecLMIS, surpassing LViT by 3.74% mIoU on the MosMedData+ dataset and 1.89% mIoU on the QATA-CoV19 dataset. More importantly, we achieve a relative reduction of 20.2% in parameter count and a 55.5% decrease in computational load. The code will be available at https://github.com/ShawnHuang497/RecLMIS. Xiaoshuang Huang, Hongxiang Li 0004, Meng Cao 0002, Long Chen 0016, Chenyu You, Dong An 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Decomposed Prototype Learning for Few-Shot Scene Graph GenerationabstractToday's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world applications with massive uncommon predicate categories whose annotations are hard to collect. In this article, we focus on Few-Shot SGG (FSSGG) , which encourages SGG models to be able to quickly transfer previous knowledge and recognize unseen predicates well with only a few examples. However, current methods for FSSGG are hindered by the high intra-class variance of predicate categories in SGG: On one hand, each predicate category commonly has multiple semantic meanings under different contexts. On the other hand, the visual appearance of relation triplets with the same predicate differs greatly under different subject–object compositions. Such great variance of inputs makes it hard to learn generalizable representation for each predicate category with current few-shot learning (FSL) methods. However, we found that this intra-class variance of predicates is highly related to the composed subjects and objects. To model the intra-class variance of predicates with subject–object context, we propose a novel Decomposed Prototype Learning (DPL) model for FSSGG. Specifically, we first construct a decomposable prototype space to capture diverse semantics and visual patterns of subjects and objects for predicates by decomposing them into multiple prototypes. Afterwards, we integrate these prototypes with different weights to generate query-adaptive predicate representation with more reliable semantics for each query sample. We conduct extensive experiments and compare with various baseline methods to show the effectiveness of our method. Jun Xiao 0001, Guikun Chen, Yinfu Feng, Yi Yang 0001, Anan Liu, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Beyond Grounding: Extracting Fine-Grained Event Hierarchies across ModalitiesabstractEvents describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical (via grounding) and thus, on the same semantic level. However, grounding fails to capture the intricate cross-event relations that exist due to the same events being referred to on many semantic levels. For example, the abstract event of "war'' manifests at a lower semantic level through subevents "tanks firing'' (in video) and airplane "shot'' (in text), leading to a hierarchical, multimodal relationship between the events. In this paper, we propose the task of extracting event hierarchies from multimodal (video and text) data to capture how the same event manifests itself in different modalities at different semantic levels. This reveals the structure of events and is critical to understanding them. To support research on this task, we introduce the Multimodal Hierarchical Events (MultiHiEve) dataset. Unlike prior video-language datasets, MultiHiEve is composed of news video-article pairs, which makes it rich in event hierarchies. We densely annotate a part of the dataset to construct the test benchmark. We show the limitations of state-of-the-art unimodal and multimodal baselines on this task. Further, we address these limitations via a new weakly supervised model, leveraging only unannotated video-article pairs from MultiHiEve. We perform a thorough evaluation of our proposed method which demonstrates improved performance on this task and highlight opportunities for future research. Data: https://github.com/hayyubi/multihieve Hammad A. Ayyubi, Christopher Thomas 0004, Lovish Chum, Rahul Lokesh, Long Chen 0016, Yulei Niu, Xudong Lin 0003, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang |
AAAI | 5 |
| 2024 | UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and MemoryabstractParameter-efficient transfer learning (PETL), i.e., finetuning a small portion of parameters, is an effective strategy for adapting pre-trained models to downstream domains. To further reduce the memory demand, recent PETL works focus on the more valuable memory-efficient characteristic. In this paper, we argue that the scalability, adaptability, and generalizability of state-of-the-art methods are hindered by structural dependency and pertinency on specific pretrained backbones. To this end, we propose a new memoryefficient PETL strategy, Universal Parallel Tuning (UniPT), to mitigate these weaknesses. Specifically, we facilitate the transfer process via a lightweight and learnable parallel network, which consists of: 1) A parallel interaction module that decouples the sequential connections and processes the intermediate activations detachedly from the pre-trained network. 2) A confidence aggregation module that learns optimal strategies adaptively for integrating cross-layer features. We evaluate UniPT with different backbones (e.g., T5 [69], VSE∞[12], CLIP4Clip [58], Clip-ViL [73], and MDETR [42]) on various vision-and-language and pure NLP tasks. Extensive ablations on 18 datasets have validated that UniPT can not only dramatically reduce memory consumption and outperform the best competitor, but also achieve competitive performance over other plain PETL methods with lower training memory overhead. Our code is publicly available at: https://github.com/Paranioar/UniPT. Haiwen Diao, Ying Zhang 0021, Xu Jia 0012, Huchuan Lu, Long Chen 0016 |
CVPR | 6 |
| 2024 | Distributionally Generative Augmentation for Fair Facial Attribute ClassificationabstractFacial Attribute Classification (FAC) holds substantial promise in widespread applications. However, FAC models trained by traditional methodologies can be unfair by exhibiting accuracy inconsistencies across varied data sub-populations. This unfairness is largely attributed to bias in data, where some spurious attributes (e.g., Male) statistically correlate with the target attribute (e.g., Smiling). Most of existing fairness-aware methods rely on the labels of spurious attributes, which may be unavailable in practice. This work proposes a novel, generation-based two-stage framework to train a fair FAC model on biased data without additional annotation. Initially, we identify the potential spurious attributes based on generative models. Notably, it enhances interpretability by explicitly showing the spurious attributes in image space. Following this, for each image, we first edit the spurious attributes with a random degree sampled from a uniform distribution, while keeping target attribute unchanged. Then we train a fair FAC model by fostering model invariance to these augmentation. Extensive experiments on three common datasets demonstrate the effectiveness of our method in promoting fairness in FAC without compromising accuracy. Codes are in https://github.com/heqianpei/DiGA. Fengda Zhang, Qianpei He, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001, Hanwang Zhang |
CVPR | 5 |
| 2024 | An Efficient and Effective Transformer Decoder-Based Framework for Multi-task Visual Grounding
Wei Chen 0005, Long Chen 0016, Yu Wu 0011 |
ECCV (45) | 2 |
| 2024 | SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
Haiwen Diao, Xu Jia 0012, Yunzhi Zhuge, Ying Zhang 0021, Huchuan Lu, Long Chen 0016 |
ECCV (44) | 7 |
| 2024 | DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
Zhen Wang 0004, Xinyun Jiang, Jun Xiao 0001, Long Chen 0016 |
ECCV (43) | 5 |
| 2024 | View-Consistent 3D Editing with Gaussian Splatting
Xuanyu Yi, Zike Wu, Na Zhao 0004, Long Chen 0016, Hanwang Zhang |
ECCV (35) | 5 |
| 2024 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement LearningabstractReinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs).These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape reward models that steer LM optimization.Nonetheless, the effective integration of rewards from diverse sources presents a significant challenge due to their disparate characteristics.To address this, recent research has developed algorithms incorporating strategies such as weighting, ranking, and constraining to handle this complexity.Despite these innovations, a bias toward disproportionately high rewards can still skew the reinforcement learning process and negatively impact LM performance.This paper explores a methodology for reward composition that enables simultaneous improvements in LMs across multiple dimensions.Inspired by fairness theory, we introduce a training algorithm that aims to reduce Disparity and enhance Stability among various rewards.Our method treats the aggregate reward as a dynamic weighted sum of individual rewards, with alternating updates to the weights and model parameters.For efficient and straightforward implementation, we employ an estimation technique rooted in the mirror descent method for weight updates, eliminating the need for gradient computations.The empirical results under various types of rewards across a wide range of scenarios demonstrate the effectiveness of our method. Jiahui Li 0003, Fengda Zhang, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011 |
EMNLP | 6 |
| 2024 | MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase UnderstandingabstractBaixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Long Chen, Yangqiu Song. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Baixuan Xu, Weiqi Wang 0001, Wenxuan Ding 0001, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu 0039, Changlong Yu, Zheng Li 0018, Chen Luo 0003, Qingyu Yin, Long Chen 0016, Yangqiu Song |
EMNLP | 14 |
| 2024 | Mrtnet: Multi-Resolution Temporal Network for Video Sentence GroundingabstractVideo sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolution Temporal (MRT) module, a Query-aware Attention (QAM) module, and a predictor. The MRT module uses an encoder-decoder network and Transformers to predict start and end times. The QAM module fuses visual and text features. Both MRT and QAM modules are easily integrated into existing VSG models. We also employ a loss function for cross-modal feature supervision at multiple scales. Extensive experiments on two prevalent datasets have shown the effectiveness of MRTNet. Wei Ji 0008, You Qin, Long Chen 0016, Yinwei Wei, Yiming Wu 0005, Roger Zimmermann |
ICASSP | 3 |
| 2024 | SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosabstractWe study the problem of procedure planning in instructional videos, which aims to make a goal-oriented sequence of action steps given partial visual state observations. The motivation of this problem is to learn a structured and plannable state and action space. Recent works succeeded in sequence modeling of steps with only sequence-level annotations accessible during training, which overlooked the roles of states in the procedures. In this work, we point out that State CHangEs MAtter (SCHEMA) for procedure planning in instructional videos. We aim to establish a more structured state space by investigating the causal relations between steps and states in procedures. Specifically, we explicitly represent each step as state changes and track the state changes in procedures. For step representation, we leveraged the commonsense knowledge in large language models (LLMs) to describe the state changes of steps via our designed chain-of-thought prompting. For state changes tracking, we align visual state observations with language state descriptions via cross-modal contrastive learning, and explicitly model the intermediate states of the procedure using LLM-generated state descriptions. Experiments on CrossTask, COIN, and NIV benchmark datasets demonstrate that our proposed SCHEMA model achieves state-of-the-art performance and obtains explainable visualizations. Yulei Niu, Wenliang Guo, Long Chen 0016, Xudong Lin 0003, Shih-Fu Chang |
ICLR | 3 |
| 2024 | ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces
Libing Yang, Long Chen 0016 |
IJCAI | 3 |
| 2024 | Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum LearningabstractBeing widely used in learning unbiased visual question answering (VQA) models, Data Augmentation (DA) helps mitigate language biases by generating extra training samples beyond the original samples. While today's DA methods can generate robust samples, the augmented training set, significantly larger than the original dataset, often exhibits redundancy in terms of difficulty or content repetition, leading to inefficient model training and even compromising the model performance. To this end, we design an Effective Curriculum Learning strategy ECL to enhance DA-based VQA methods. Intuitively, ECL trains VQA models on relatively "easy'' samples first, and then gradually changes to "harder'' samples, and less-valuable samples are dynamically removed. Compared to training on the entire augmented dataset, ECL strategy can further enhance VQA models' performance with fewer training samples. Extensive ablations have demonstrated the effectiveness of ECL on various methods. Yuhang Zheng 0003, Zhen Wang 0004, Long Chen 0016 |
ICMR | 3 |
| 2024 | Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language ExplainerabstractBenefiting from strong generalization ability, pre-trained vision-language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple recognition tasks, grounded situation recognition (GSR) requires the model not only to classify salient activity (verb) in the image, but also to detect all semantic roles that participate in the action. This complex task usually involves three steps: verb recognition, semantic role grounding, and noun recognition. Directly employing class-based prompts with VLMs and grounding models for this task suffers from several limitations, e.g., it struggles to distinguish ambiguous verb concepts, accurately localize roles with fixed verb-centric template input, and achieve context-aware noun predictions. In this paper, we argue that these limitations stem from the model's poor understanding of verb/noun classes. To this end, we introduce a new approach for zero-shot GSR via Language EXplainer (LEX), which significantly boosts the model's comprehensive capabilities through three explainers: 1) verb explainer, which generates general verb-centric descriptions to enhance the discriminability of different verb classes; 2) grounding explainer, which rephrases verb-centric templates for clearer understanding, thereby enhancing precise semantic role localization; and 3) noun explainer, which creates scene-specific noun descriptions to ensure context-aware noun recognition. By equipping each step of the GSR process with an auxiliary explainer, LEX facilitates complex scene understanding in real-world scenarios. Our extensive validations on the SWiG dataset demonstrate LEX's effectiveness and interoperability in zero-shot GSR. Jiaming Lei, Lin Li 0065, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016 |
ACM Multimedia | 5 |
| 2024 | PROMOTE: Prior-Guided Diffusion Model with Global-Local Contrastive Learning for Exemplar-Based Image TranslationabstractExemplar-based image translation has garnered significant interest from researchers due to its broad applications in multimedia/multimodal processing. Existing methods primarily employ Euclidean-based losses to implicitly establish cross-domain correspondences between exemplar and conditional images, aiming to produce high-fidelity images. However, these methods often suffer from two challenges: 1) Insufficient excavation of domain-invariant features leads to low-quality cross-domain correspondences, and 2) Inaccurate correspondences result in errors propagated during the translation process due to a lack of reliable prior guidance. To tackle these issues, we propose a novel prior-guided diffusion model with global-local contrastive learning (PROMOTE), which is trained in a self-supervised manner. Technically, global-local contrastive learning is designed to align two cross-domain images within hyperbolic space and reduce the gap between their semantic correlation distributions using the Fisher-Rao metric, allowing the visual encoders to extract domain-invariant features more effectively. Moreover, a prior-guided diffusion model is developed that propagates the structural prior to all timesteps in the diffusion process. It is optimized by a novel prior denoising loss, mathematically derived from the transitions modified by prior information in a self-supervised manner, successfully alleviating the impact of inaccurate correspondences on image translation. Extensive experiments conducted across seven datasets demonstrate that our proposed PROMOTE significantly exceeds state-of-the-art performance in diverse exemplar-based image translation tasks. The source code is publicly available at http://github.com/zgj77/PROMOTE. Guojin Zhong, Yihu Guo, Jin Yuan 0002, Qianjun Zhang, Weili Guan, Long Chen 0016 |
ACM Multimedia | 6 |
| 2024 | $\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose EstimationabstractDiffusion models have demonstrated their effectiveness in addressing the inherent uncertainty and indeterminacy in monocular 3D human pose estimation (HPE).
Despite their strengths, the need for large search spaces and the corresponding demand for substantial training data make these models prone to generating biomechanically unrealistic poses.
This challenge is particularly noticeable in occlusion scenarios, where the complexity of inferring 3D structures from 2D images intensifies.
In response to these limitations, we introduce the **Di**screte **Di**ffusion **Pose** (**$\text{Di}^2\text{Pose}$**), a novel framework designed for occluded 3D HPE that capitalizes on the benefits of a discrete diffusion model.
Specifically, **$\text{Di}^2\text{Pose}$** employs a two-stage process: it first converts 3D poses into a discrete representation through a pose quantization step, which is subsequently modeled in latent space through a discrete diffusion process.
This methodological innovation restrictively confines the search space towards physically viable configurations and enhances the model’s capability to comprehend how occlusions affect human pose within the latent space.
Extensive evaluations conducted on various benchmarks (e.g., Human3.6M, 3DPW, and 3DPW-Occ) have demonstrated its effectiveness. Jun Xiao 0001, Chunping Wang 0001, Wei Liu 0005, Long Chen 0016 |
NeurIPS | 6 |
| 2024 | LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016 |
NeurIPS | 10 |
| 2024 | NICEST: Noisy Label Correction and Training for Robust Scene Graph GenerationabstractNearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that neither of the assumptions applies to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions and harm the training of unbiased SGG models. To this end, we propose a novelNoIsy labelCorrEction andSampleTrainingstrategy for SGG:NICEST, which rules out these noisy label issues by generating high-quality samples and designing an effective training strategy. Specifically, it consists of: 1)NICE: it detects noisy samples and then reassigns higher-quality soft predicate labels to them. To achieve this goal, NICE contains three main steps: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, it is treated as an out-of-distribution detection problem, and the pseudo labels are assigned to all detected noisy negative samples. Then, in Pos-NSD, we use a density-based clustering algorithm to detect noisy positive samples. Lastly, in NSC, we use weighted KNN to reassign more robust soft predicate labels rather than hard labels to all noisy positive samples. 2)NIST: it is a multi-teacher knowledge distillation based training strategy, which enables the model to learn unbiased fusion knowledge. A dynamic trade-off weighting strategy in NIST is designed to penalize the bias of different teachers. Due to the model-agnostic nature of both NICE and NIST, NICEST can be seamlessly incorporated into any SGG architecture to boost its performance on different predicate categories. In addition, to better assess the generalization ability of SGG models, we propose a new benchmark,VG-OOD, by reorganizing the prevalent VG dataset. This reorganization deliberately makes the predicate distributions between the training and test sets as different as possible for each subject-object category pair. This new benchmark helps disentangle the influence of subject-object category biases. Extensive ablations and results on different backbones and tasks have attested to the effectiveness and generalization ability of each component of NICEST. Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Hanwang Zhang, Yi Yang 0001, Wei Liu 0005, Long Chen 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | CrossFormer++: A Versatile Vision Transformer Hinging on Cross-Scale AttentionabstractWhile features of different scales are perceptually important to visual inputs, existing vision transformers do not yet take advantage of them explicitly. To this end, we first propose a cross-scale vision transformer, CrossFormer. It introduces a cross-scale embedding layer (CEL) and a long-short distance attention (LSDA). On the one hand, CEL blends each token with multiple patches of different scales, providing the self-attention module itself with cross-scale features. On the other hand, LSDA splits the self-attention module into a short-distance one and a long-distance counterpart, which not only reduces the computational burden but also keeps both small-scale and large-scale features in the tokens. Moreover, through experiments on CrossFormer, we observe another two issues that affect vision transformers' performance, i.e., the enlarging self-attention maps and amplitude explosion. Thus, we further propose a progressive group size (PGS) paradigm and an amplitude cooling layer (ACL) to alleviate the two issues, respectively. The CrossFormer incorporating with PGS and ACL is called CrossFormer++. Extensive experiments show that CrossFormer++ outperforms the other vision transformers on image classification, object detection, instance segmentation, and semantic segmentation tasks. Wenxiao Wang 0001, Wei Chen 0005, Qibo Qiu, Long Chen 0016, Boxi Wu 0001, Binbin Lin 0001, Xiaofei He 0001, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and FutureabstractAs the most fundamental scene understanding tasks, object detection and segmentation have made tremendous progress in deep learning era. Due to the expensive manual labeling cost, the annotated categories in existing datasets are often small-scale and pre-defined, i.e., state-of-the-art fully-supervised detectors and segmentors fail to generalize beyond the closed vocabulary. To resolve this limitation, in the last few years, the community has witnessed an increasing attention toward Open-Vocabulary Detection (OVD) and Segmentation (OVS). By "open-vocabulary", we mean that the models can classify objects beyond pre-defined categories. In this survey, we provide a comprehensive review on recent developments of OVD and OVS. A taxonomy is first developed to organize different tasks and methodologies. We find that the permission and usage of weak supervision signals can well discriminate different methodologies, including: visual-semantic space mapping, novel visual feature synthesis, region-aware training, pseudo-labeling, knowledge distillation, and transfer learning. The proposed taxonomy is universal across different tasks, covering object detection, semantic/instance/panoptic segmentation, 3D and video understanding. The main design principles, key challenges, development routes, methodology strengths, and weaknesses are thoroughly analyzed. Long Chen 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Label Semantic Knowledge Distillation for Unbiased Scene Graph GenerationabstractThe Scene Graph Generation (SGG) task aims to detect all the objects and their pairwise visual relationships in a given image. Although SGG has achieved remarkable progress over the last few years, almost all existing SGG models follow the same training paradigm: they treat both object and predicate classification in SGG as a single-label classification problem, and the ground-truths are one-hot target labels. However, this prevalent training paradigm has overlooked two characteristics of current SGG datasets: 1) For positive samples, some specific subject-object instances may have multiple reasonable predicates. 2) For negative samples, there are numerous missing annotations. Regardless of the two characteristics, SGG models are easy to be confused and make wrong predictions. To this end, we propose a novel model-agnostic Label Semantic Knowledge Distillation (LS-KD) for unbiased SGG. Specifically, LS-KD dynamically generates a “soft” label for each subject-object instance by fusing a predicted Label Semantic Distribution (LSD) with its original one-hot target label. LSD reflects the correlations between this instance and multiple predicate categories. Meanwhile, we propose two different strategies to predict LSD: iterative self-KD and synchronous self-KD. Extensive ablations and results on three SGG tasks have attested to the superiority and generality of our proposed LS-KD, which can consistently achieve decent trade-off performance between different predicate categories. Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Wenxiao Wang 0001, Jian Shao 0001, Anan Liu, Yi Yang 0001, Long Chen 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | GSSF: Generalized Structural Sparse Function for Deep Cross-Modal Metric LearningabstractCross-modal metric learning is a prominent research topic that bridges the semantic heterogeneity between vision and language. Existing methods frequently utilize simple cosine or complex distance metrics to transform the pairwise features into a similarity score, which suffers from an inadequate or inefficient capability for distance measurements. Consequently, we propose a Generalized Structural Sparse Function to dynamically capture thorough and powerful relationships across modalities for pair-wise similarity learning while remaining concise but efficient. Specifically, the distance metric delicately encapsulates two formats of diagonal and block-diagonal terms, automatically distinguishing and highlighting the cross-channel relevancy and dependency inside a structured and organized topology. Hence, it thereby empowers itself to adapt to the optimal matching patterns between the paired features and reaches a sweet spot between model complexity and capability. Extensive experiments on cross-modal and two extra uni-modal retrieval tasks (image-text retrieval, person re-identification, fine-grained image retrieval) have validated its superiority and flexibility over various popular retrieval frameworks. More importantly, we further discover that it can be seamlessly incorporated into multiple application scenarios, and demonstrates promising prospects from Attention Mechanism to Knowledge Distillation in a plug-and-play manner. Haiwen Diao, Ying Zhang 0021, Shang Gao 0012, Jiawen Zhu 0003, Long Chen 0016, Huchuan Lu |
IEEE Trans. Image Process. | 5 |
| 2024 | In Defense of Clip-Based Video Relation DetectionabstractVideo Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be broadly categorized into bottom-up and top-down paradigms, depending on their approach to classifying relations. Bottom-up methods follow a clip-based approach where they classify relations of short clip tubelet pairs and then merge them into long video relations. On the other hand, top-down methods directly classify long video tubelet pairs. While recent video-based methods utilizing video tubelets have shown promising results, we argue that the effective modeling of spatial and temporal context plays a more significant role than the choice between clip tubelets and video tubelets. This motivates us to revisit the clip-based paradigm and explore the key success factors in VidVRD. In this paper, we propose a Hierarchical Context Model (HCM) that enriches the object-based spatial context and relation-based temporal context based on clips. We demonstrate that using clip tubelets can achieve superior performance compared to most video-based methods. Additionally, using clip tubelets offers more flexibility in model designs and helps alleviate the limitations associated with video tubelets, such as the challenging long-term object tracking problem and the loss of temporal information in long-term tubelet feature compression. Extensive experiments conducted on two challenging VidVRD benchmarks validate that our HCM achieves a new state-of-the-art performance, highlighting the effectiveness of incorporating advanced spatial and temporal context modeling within the clip-based paradigm. Long Chen 0016, Wei Ji 0008, Xiaoyu Yue, Roger Zimmermann |
IEEE Trans. Image Process. | 2 |
| 2024 | Improving Reference-Based Distinctive Image Captioning with Contrastive RewardsabstractDistinctive Image Captioning (DIC)—generating distinctive captions that describe the unique details of a target image—has received considerable attention over the last few years. A recent DIC method proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-Based DIC (Ref-DIC). It aims to force the generated captions to distinguish between the target image and the reference image. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects “ towel ” and “ toilet ” while all reference images are without them, then a simple caption “ A bathroom with a towel and a toilet ” is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at the object-/attribute-level (vs. scene-level). Second, to generate distinctive captions, we develop a Transformer-based Ref-DIC baseline TransDIC . It not only extracts visual features from the target image but also encodes the differences between objects in the target and reference images. Taking one step further, we propose a stronger TransDIC \({++}\) , which consists of an extra contrastive learning module to make full use of the reference images. This new module is model-agnostic, which can be easily incorporated into various Ref-DIC architectures. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC \({++}\) can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics. Yangjun Mao, Jun Xiao 0001, Meng Cao 0002, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Iterative Proposal Refinement for Weakly-Supervised Video GroundingabstractWeakly-Supervised Video Grounding (WSVG) aims to localize events of interest in untrimmed videos with only video-level annotations. To date, most of the state-of-the-art WSVG methods follow a two-stage pipeline, i.e., firstly generating potential temporal proposals and then grounding with these proposal candidates. Despite the recent progress, existing proposal generation methods suffer from two draw-backs: 1) lack of explicit correspondence modeling; and 2) partial coverage of complex events. To this end, we propose a novel IteRative prOposal refiNement network (dubbed as IRON) to gradually distill the prior knowledge into each proposal and encourage proposals with more complete coverage. Specifically, we set up two lightweight distillation branches to uncover the cross-modal correspondence on both the semantic and conceptual levels. Then, an iterative Label Propagation (LP) strategy is devised to prevent the network from focusing excessively on the most discriminative events instead of the whole sentence content. Precisely, during each iteration, the proposal with the minimal distillation loss and its adjacent ones are regarded as the positive samples, which refines proposal confidence scores in a cascaded manner. Extensive experiments and ablation studies on two challenging WSVG datasets have attested to the effectiveness of our IRON. The code will be available at https://github.com/mengcaopku/IRON. Meng Cao 0002, Fangyun Wei, Can Xu 0002, Xiubo Geng, Long Chen 0016, Can Zhang 0001, Yuexian Zou, Tao Shen 0001, Daxin Jiang |
CVPR | 5 |
| 2023 | Compositional Feature Augmentation for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to detect all the visual relation tripletsin a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recent years. However, due to the ubiquitous long-tailed predicate distributions, today’s SGG models are still easily biased to the head predicates. Currently, the most prevalent debiasing solutions for SGG are re-balancing methods, e.g., changing the distributions of original training samples. In this paper, we argue that all existing re-balancing strategies fail to increase the diversity of the relation triplet features of each predicate, which is critical for robust SGG. To this end, we propose a novel Compositional Feature Augmentation (CFA) strategy, which is the first unbiased SGG work to mitigate the bias issue from the perspective of increasing the diversity of triplet features. Specifically, we first decompose each relation triplet feature into two components: intrinsic feature and extrinsic feature, which correspond to the intrinsic characteristics and extrinsic contexts of a relation triplet, respectively. Then, we design two different feature augmentation modules to enrich the feature diversity of original relation triplets by replacing or mixing up either their intrinsic or extrinsic features from other samples. Due to its model-agnostic nature, CFA can be seamlessly incorporated into various SGG frameworks. Extensive ablations have shown that CFA achieves a new state-of-the-art performance on the trade-off between different metrics. Lin Li 0065, Guikun Chen, Jun Xiao 0001, Yi Yang 0001, Chunping Wang 0001, Long Chen 0016 |
ICCV | 6 |
| 2023 | Video Scene Graph Generation from Single-Frame Weak Supervision
Jun Xiao 0001, Long Chen 0016 |
ICLR | 3 |
| 2023 | Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Kaifeng Gao, Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Qianru Sun |
ICLR | 2 |
| 2023 | TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Xudong Lin 0003, Guangxing Han, Shih-Fu Chang |
ICLR | 4 |
| 2023 | Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes
Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001 |
ICLR | 3 |
| 2023 | Discrepancy-Guided Reconstruction Learning for Image Forgery DetectionabstractIn this paper, we propose a novel image forgery detection paradigm for boosting the model learning capacity on both forgery-sensitive and genuine compact visual patterns. Compared to the existing methods that only focus on the discrepant-specific patterns (\eg, noises, textures, and frequencies), our method has a greater generalization. Specifically, we first propose a Discrepancy-Guided Encoder (DisGE) to extract forgery-sensitive visual patterns. DisGE consists of two branches, where the mainstream backbone branch is used to extract general semantic features, and the accessorial discrepant external attention branch is used to extract explicit forgery cues. Besides, a Double-Head Reconstruction (DouHR) module is proposed to enhance genuine compact visual patterns in different granular spaces. Under DouHR, we further introduce a Discrepancy-Aggregation Detector (DisAD) to aggregate these genuine compact visual patterns, such that the forgery detection capability on unknown patterns can be improved. Extensive experimental results on four challenging datasets validate the effectiveness of our proposed method against state-of-the-art competitors. Zenan Shi, Haipeng Chen 0002, Long Chen 0016 |
IJCAI | 3 |
| 2023 | Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement LearningabstractExploration strategy plays an important role in reinforcement learning, especially in sparse-reward tasks. In cooperative multi-agent reinforcement learning~(MARL), designing a suitable exploration strategy is much more challenging due to the large state space and the complex interaction among agents. Currently, mainstream exploration methods in MARL either contribute to exploring the unfamiliar states which are large and sparse, or measuring the interaction among agents with high computational costs. We found an interesting phenomenon that different kinds of exploration plays a different role in different MARL scenarios, and choosing a suitable one is often more effective than designing an exquisite algorithm. In this paper, we propose a exploration method that incorporate the \underline{C}uri\underline{O}sity-based and \underline{IN}fluence-based exploration~(COIN) which is simple but effective in various situations. First, COIN measures the influence of each agent on the other agents based on mutual information theory and designs it as intrinsic rewards which are applied to each individual value function. Moreover, COIN computes the curiosity-based intrinsic rewards via prediction errors which are added to the extrinsic reward. For integrating the two kinds of intrinsic rewards, COIN utilizes a novel framework in which they complement each other and lead to a sufficient and effective exploration on cooperative MARL tasks. We perform extensive experiments on different challenging benchmarks, and results across different scenarios show the superiority of our method. Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Fei Wu 0001, Jun Xiao 0001, Long Chen 0016 |
NeurIPS | 7 |
| 2023 | Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsabstractPretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs within an image. However, naively utilizing CLIP with prevalent class-based prompts for zero-shot VRD has several weaknesses, e.g., it struggles to distinguish between different fine-grained relation types and it neglects essential spatial information of two objects. To this end, we propose a novel method for zero-shot VRD: RECODE, which solves RElation detection via COmposite DEscription prompts. Specifically, RECODE first decomposes each predicate category into subject, object, and spatial components. Then, it leverages large language models (LLMs) to generate description-based prompts (or visual cues) for each component. Different visual cues enhance the discriminability of similar relation categories from different perspectives, which significantly boosts performance in VRD. To dynamically fuse different cues, we further introduce a chain-of-thought method that prompts LLMs to generate reasonable weights for different visual cues. Extensive experiments on four VRD benchmarks have demonstrated the effectiveness and interpretability of RECODE. Lin Li 0065, Jun Xiao 0001, Guikun Chen, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
NeurIPS | 6 |
| 2023 | Federated unsupervised representation learningabstractTo leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins. Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Counterfactual Samples Synthesizing and Training for Robust Visual Question AnsweringabstractToday's VQA models still tend to capture superficial linguistic correlations in the training set and fail to generalize to the test set with different QA distributions. To reduce these language biases, recent VQA works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on diagnostic benchmarks for out-of-distribution testing. However, due to the complex model design, ensemble-based methods are unable to equip themselves with two indispensable characteristics of an ideal VQA model: 1) Visual-explainable: The model should rely on the right visual regions when making decisions. 2) Question-sensitive: The model should be sensitive to the linguistic variations in questions. To this end, we propose a novel model-agnostic Counterfactual Samples Synthesizing and Training (CSST) strategy. After training with CSST, VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. Specifically, CSST is composed of two parts: Counterfactual Samples Synthesizing (CSS) and Counterfactual Samples Training (CST). CSS generates counterfactual samples by carefully masking critical objects in images or words in questions and assigning pseudo ground-truth answers. CST not only trains the VQA models with both complementary samples to predict respective ground-truth answers, but also urges the VQA models to further distinguish the original samples and superficially similar counterfactual ones. To facilitate the CST training, we propose two variants of supervised contrastive loss for VQA, and design an effective positive and negative sample selection mechanism based on CSS. Extensive experiments have shown the effectiveness of CSST. Particularly, by building on top of model LMH+SAR (Clark et al. 2019), (Si et al. 2021), we achieve record-breaking performance on all out-of-distribution benchmarks (e.g., VQA-CP v2, VQA-CP v1, and GQA-OOD). Long Chen 0016, Yuhang Zheng 0003, Yulei Niu, Hanwang Zhang, Jun Xiao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and ApproachabstractTemporal Sentence Grounding in Videos (TSGV) , which aims to ground a natural language sentence that indicates complex human activities in an untrimmed video, has drawn widespread attention over the past few years. However, recent studies have found that current benchmark datasets may have obvious moment annotation biases, enabling several simple baselines even without training to achieve state-of-the-art (SOTA) performance. In this paper, we take a closer look at existing evaluation protocols for TSGV, and find that both the prevailing dataset splits and evaluation metrics are the devils that lead to untrustworthy benchmarking. Therefore, we propose to re-organize the two widely-used datasets, making the ground-truth moment distributions different in the training and test splits, i.e., out-of-distribution (OOD) test. Meanwhile, we introduce a new evaluation metric “dR@ n ,IoU= m ” that discounts the basic recall scores especially with small IoU thresholds, so as to alleviate the inflating evaluation caused by biased datasets with a large proportion of long ground-truth moments. New benchmarking results indicate that our proposed evaluation protocols can better monitor the research progress in TSGV. Furthermore, we propose a novel causality-based Multi-branch Deconfounding Debiasing (MDD) framework for unbiased moment prediction. Specifically, we design a multi-branch deconfounder to eliminate the effects caused by multiple confounders with causal intervention. In order to help the model better align the semantics between sentence queries and video moments, we enhance the representations during feature encoding. Specifically, for textual information, the query is parsed into several verb-centered phrases to obtain a more fine-grained textual feature. For visual information, the positional information has been decomposed from the moment features to enhance the representations of moments with diverse locations. Extensive experiments demonstrate that our proposed approach can achieve competitive results among existing SOTA approaches and outperform the base model with great gains. Xiaohan Lan, Yitian Yuan, Xin Wang 0019, Long Chen 0016, Zhi Wang 0001, Lin Ma 0002, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language MatchingabstractThe prevailing framework for matching multimodal inputs is based on a two-stage process: (1) detecting proposals with an object detector and (2) matching text queries with proposals. Existing two-stage solutions mostly focus on the matching step. In this article, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., query-agnostic), hoping that the proposals contain all instances mentioned in the text query (i.e., query-aware). Due to this mismatch, chances are that proposals relevant to the text query are suppressed during the filtering process, which in turn bounds the matching performance. To this end, we propose VL-NMS, which is the first method to yield query-aware proposals at the first stage. VL-NMS regards all mentioned instances as critical objects and introduces a lightweight module to predict a score for aligning each proposal with a critical object. These scores can guide the NMS operation to filter out proposals irrelevant to the text query, increasing the recall of critical objects, and resulting in a significantly improved matching performance. Since VL-NMS is agnostic to the matching step, it can be easily integrated into any state-of-the-art two-stage matching method. We validate the effectiveness of VL-NMS on three multimodal matching tasks, namely referring expression grounding, phrase grounding, and image-text matching. Extensive ablation studies on several baselines and benchmarks consistently demonstrate the superiority of VL-NMS. Chenchi Zhang, Jun Xiao 0001, Hanwang Zhang, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Rethinking the Two-Stage Framework for Grounded Situation RecognitionabstractGrounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g.,buying) and detecting all corresponding semantic roles (e.g.,agent and goods), is an essential step towards “human-like” event understanding. Since each verb is associated with a specific set of semantic roles, all existing GSR methods resort to a two-stage framework: predicting the verb in the first stage and detecting the semantic roles in the second stage. However, there are obvious drawbacks in both stages: 1) The widely-used cross-entropy (XE) loss for object recognition is insufficient in verb classification due to the large intra-class variation and high inter-class similarity among daily activities. 2) All semantic roles are detected in an autoregressive manner, which fails to model the complex semantic relations between different roles. To this end, we propose a novel SituFormerfor GSR which consists of a Coarse-to-Fine Verb Model (CFVM) and a Transformer-based Noun Model (TNM). CFVM is a two-step verb prediction model: a coarse-grained model trained with XE loss first proposes a set of verb candidates, and then a fine-grained model trained with triplet loss re-ranks these candidates with enhanced verb features (not only separable but also discriminative). TNM is a transformer-based semantic role detection model, which detects all roles parallelly. Owing to the global relation modeling ability and flexibility of the transformer decoder, TNM can fully explore the statistical dependency of the roles. Extensive validations on the challenging SWiG benchmark show that SituFormer achieves a new state-of-the-art performance with significant gains under various metrics. Code is available at https://github.com/kellyiss/SituFormer. Long Chen 0016, Wei Ji 0008, Xiaoyu Yue, Tat-Seng Chua |
AAAI | 2 |
| 2022 | Rethinking the Evaluation of Unbiased Scene Graph Generation
Long Chen 0016, Jian Shao 0001, Shaoning Xiao, Songyang Zhang 0004, Jun Xiao 0001 |
BMVC | 2 |
| 2022 | Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsabstractToday's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The ground-truth predicate labels for proposals are partially correct. 2) They break the high-order relations among different predicate instances of a same subject-object pair. 3) VidSGG performance is upper-bounded by the quality of the proposals. To this end, we propose a new classification-then-grounding framework for VidSGG, which can avoid all the three overlooked drawbacks. Meanwhile, under this framework, we reformulate the video scene graphs as temporal bipartite graphs, where the entities and predicates are two types of nodes with time slots, and the edges denote different semantic roles between these nodes. This formulation takes full advantage of our new framework. Accordingly, we further propose a novel BIpartite Graph based SGG model: BIG. It consists of a classification stage and a grounding stage, where the former aims to classify the categories of all the nodes and the edges, and the latter tries to localize the temporal location of each relation instance. Extensive ablations on two VidSGG datasets have attested to the effectiveness of our framework and BIG. Code is available at https://github.com/Dawn-LX/VidSGG-BIG. Kaifeng Gao, Long Chen 0016, Yulei Niu, Jian Shao 0001, Jun Xiao 0001 |
CVPR | 2 |
| 2022 | Few-Shot Object Detection with Fully Cross-TransformerabstractFew-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and calculate the similarity between image regions and few-shot examples for detection. However, in previous works, the interaction between the two branches is only restricted in the detection head, while leaving the remaining hundreds of layers for separate feature extraction. Inspired by the recent work on vision transformers and vision-language transformers, we propose a novel Fully Cross-Transformer based model (FCT) for FSOD by incorporating cross-transformer into both the feature backbone and detection head. The asymmetric-batched cross-attention is proposed to aggregate the key information from the two branches with different batch sizes. Our model can improve the few-shot similarity learning between the two branches by introducing the multi-level interactions. Comprehensive experiments on both PASCAL VOC and MSCOCO FSOD benchmarks demonstrate the effectiveness of our model. Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Shih-Fu Chang |
CVPR | 4 |
| 2022 | The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationabstractUnbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that both assumptions are inapplicable to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions, and these noisy samples actually harm the training of unbiased SGG models. To this end, we propose a novel model-agnostic NoIsy label CorrEction strategy for SGG: NICE. NICE can not only detect noisy samples but also reassign more high-quality predicate labels to them. After the NICE training, we can obtain a cleaner version of SGG dataset for model training. Specifically, NICE consists of three components: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, we formulate this task as an out-of-distribution detection problem, and assign pseudo labels to all detected noisy negative samples. Then, in Pos-NSD, we use a clustering-based algorithm to divide all positive samples into multiple sets, and treat the samples in the noisiest set as noisy positive samples. Lastly, in NSC, we use a simple but effective weighted KNN to reassign new predicate labels to noisy positive samples. Extensive results on different backbones and tasks have attested to the effectiveness and generalization abilities of each component of NICE. Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jun Xiao 0001 |
CVPR | 2 |
| 2022 | Rethinking Data Augmentation for Robust Visual Question Answering
Long Chen 0016, Yuhang Zheng 0003, Jun Xiao 0001 |
ECCV (36) | 1 |
| 2022 | Explicit Image Caption Editing
Zhen Wang 0004, Long Chen 0016, Guangxing Han, Yulei Niu, Jian Shao 0001, Jun Xiao 0001 |
ECCV (36) | 2 |
| 2022 | Weakly-Supervised Temporal Article GroundingabstractLong Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Long Chen 0016, Yulei Niu, Brian Chen 0001, Xudong Lin 0003, Guangxing Han, Christopher Thomas 0004, Hammad A. Ayyubi, Heng Ji 0001, Shih-Fu Chang |
EMNLP | 1 |
| 2022 | Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample PerspectivesabstractReasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at different levels of abstraction.Existing efforts mainly focus on designing sophisticated architectures while utilizing frame-or object-level visual representations.In this paper, we reconsider the multi-modal alignment in VideoQA from feature and sample perspectives to achieve better performance.From the view of feature, we break down the video into trajectories and first leverage trajectory feature in VideoQA to enhance the alignment between two modalities.Moreover, we adopt a heterogeneous graph architecture and design a hierarchical framework to align both trajectory-level and frame-level visual feature with language feature.In addition, we found that VideoQA models are largely dependent on language priors and always neglect visuallanguage interactions.Thus, two effective yet portable training augmentation strategies are designed to strengthen the cross-modal correspondence ability of our model from the view of sample.Extensive results show that our method outperforms all state-of-the-art models on the challenging NExT-QA benchmark.* Long Chen is the corresponding author.… objects trajectories Q: Why did the boy in orange hold a ball on his head?A: want to throw the ball. Shaoning Xiao, Long Chen 0016, Kaifeng Gao, Yi Yang 0001, Jun Xiao 0001 |
EMNLP | 2 |
| 2022 | CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang 0001, Long Chen 0016, Binbin Lin 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ICLR | 3 |
| 2022 | Deconfounded Value Decomposition for Multi-Agent Reinforcement LearningabstractValue decomposition (VD) methods have been widely used in cooperative multi-agent reinforcement learning (MARL), where credit assignment plays an important role in guiding the agents’ decentralized execution. In this paper, we investigate VD from a novel perspective of causal inference. We first show that the environment in existing VD methods is an unobserved confounder as the common cause factor of the global state and the joint value function, which leads to the confounding bias on learning credit assignment. We then present our approach, deconfounded value decomposition (DVD), which cuts off the backdoor confounding path from the global state to the joint value function. The cut is implemented by introducing the trajectory graph, which depends only on the local trajectories, as a proxy confounder. DVD is general enough to be applied to various VD methods, and extensive experiments show that DVD can consistently achieve significant performance gains over different state-of-the-art VD methods on StarCraft II and MACO benchmarks. Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Furui Liu, Long Chen 0016, Changjie Fan, Fei Wu 0001, Jun Xiao 0001 |
ICML | 5 |
| 2022 | Correspondence Matters for Video Referring Expression ComprehensionabstractWe investigate the problem of video Referring Expression Comprehension (REC), which aims to localize the referent objects described in the sentence to visual regions in the video frames. Despite the recent progress, existing methods suffer from two problems: 1) inconsistent localization results across video frames; 2) confusion between the referent and contextual objects. To this end, we propose a novel Dual Correspondence Network (dubbed as DCNet) which explicitly enhances the dense associations in both the inter-frame and cross-modal manners. Firstly, we aim to build the inter-frame correlations for all existing instances within the frames. Specifically, we compute the inter-frame patch-wise cosine similarity to estimate the dense alignment and then perform the inter-frame contrastive learning to map them close in feature space. Secondly, we propose to build the fine-grained patch-word alignment to associate each patch with certain words. Due to the lack of this kind of detailed annotations, we also predict the patch-word correspondence through the cosine similarity. Extensive experiments demonstrate that our DCNet achieves state-of-the-art performance on both video and image REC benchmarks. Furthermore, we conduct comprehensive ablation studies and thorough analyses to explore the optimal model designs. Notably, our inter-frame and cross-modal contrastive losses are plug-and-play functions and are applicable to any video REC architectures. For example, by building on top of Co-grounding, we boost the performance by 1.48% absolute improvement on [email protected] for VID-Sentence dataset. Meng Cao 0002, Ji Jiang, Long Chen 0016, Yuexian Zou |
ACM Multimedia | 3 |
| 2022 | Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph GenerationabstractRecently, increasing efforts have been focused on Weakly Supervised Scene Graph Generation (WSSGG). The mainstream solution for WSSGG typically follows the same pipeline: they first align text entities in the weak image-level supervisions (e.g., unlocalized relation triplets or captions) with image regions, and then train SGG models in a fully-supervised manner with aligned instance-level "pseudo" labels. However, we argue that most existing WSSGG works only focus on object-consistency, which means the grounded regions should have the same object category label as text entities. While they neglect another basic requirement for an ideal alignment: interaction-consistency, which means the grounded region pairs should have the same interactions (i.e., visual relations) as text entity pairs. Hence, in this paper, we propose to enhance a simple grounding module with both object-aware and interaction-aware knowledge to acquire more reliable pseudo labels. To better leverage these two types of knowledge, we regard them as two teachers and fuse their generated targets to guide the training process of our grounding module. Specifically, we design two different strategies to adaptively assign weights to different teachers by assessing their reliability on each training sample. Extensive experiments have demonstrated that our method consistently improves WSSGG performance on various kinds of weak supervision. Long Chen 0016, Yi Yang 0001, Jun Xiao 0001 |
ACM Multimedia | 2 |
| 2022 | Rethinking the Reference-based Distinctive Image CaptioningabstractDistinctive Image Captioning (DIC) --- generating distinctive captions that describe the unique details of a target image --- has received considerable attention over the last few years. A recent DIC work proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-based DIC (Ref-DIC). It aims to make the generated captions can tell apart the target and reference images. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects "towel'' and "toilet'' while all reference images are without them, then a simple caption "A bathroom with a towel and a toilet'' is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at object-/attribute- level (vs. scene-level). Secondly, to generate distinctive captions, we develop a strong Transformer-based Ref-DIC baseline, dubbed as TransDIC. It not only extracts visual features from the target image, but also encodes the differences between objects in the target and reference images. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics. Yangjun Mao, Long Chen 0016, Zhihong Jiang, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 2 |
| 2022 | Respecting Transfer Gap in Knowledge DistillationabstractKnowledge distillation (KD) is essentially a process of transferring a teacher model's behavior, e.g., network response, to a student model. The network response serves as additional supervision to formulate the machine domain, which uses the data collected from the human domain as a transfer set. Traditional KD methods hold an underlying assumption that the data collected in both human domain and machine domain are both independent and identically distributed (IID). We point out that this naive assumption is unrealistic and there is indeed a transfer gap between the two domains. Although the gap offers the student model external knowledge from the machine domain, the imbalanced teacher knowledge would make us incorrectly estimate how much to transfer from teacher to student per sample on the non-IID transfer set. To tackle this challenge, we propose Inverse Probability Weighting Distillation (IPWD) that estimates the propensity of a training sample belonging to the machine domain, and assigns its inverse amount to compensate for under-represented samples. Experiments on CIFAR-100 and ImageNet demonstrate the effectiveness of \ours~for both two-stage distillation and one-stage self-distillation. Yulei Niu, Long Chen 0016, Hanwang Zhang |
NeurIPS | 2 |
| 2022 | Deep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Feifei Shao, Long Chen 0016, Jian Shao 0001, Wei Ji 0008, Shaoning Xiao, Lu Ye, Yueting Zhuang, Jun Xiao 0001 |
Neurocomputing | 2 |
| 2022 | Deep Motion Prior for Weakly-Supervised Temporal Action LocalizationabstractWeakly-Supervised Temporal Action Localization (WSTAL) aims to localize actions in untrimmed videos with only video-level labels. Currently, most state-of-the-art WSTAL methods follow a Multi-Instance Learning (MIL) pipeline: producing snippet-level predictions first and then aggregating to the video-level prediction. However, we argue that existing methods have overlooked two important drawbacks: 1) inadequate use of motion information and 2) the incompatibility of prevailing cross-entropy training loss. In this paper, we analyze that the motion cues behind the optical flow features are complementary informative. Inspired by this, we propose to build a context-dependent motion prior, termed as motionness. Specifically, a motion graph is introduced to model motionness based on the local motion carrier (e.g., optical flow). In addition, to highlight more informative video snippets, a motion-guided loss is proposed to modulate the network training conditioned on motionness scores. Extensive ablation studies confirm that motionness efficaciously models action-of-interest, and the motion-guided loss leads to more accurate results. Besides, our motion-guided loss is a plug-and-play loss function and is applicable with existing WSTAL methods. Without loss of generality, based on the standard MIL pipeline, our method achieves new state-of-the-art performance on three challenging benchmarks, including THUMOS'14, ActivityNet v1.2 and v1.3. Meng Cao 0002, Can Zhang 0001, Long Chen 0016, Zheng Shou 0001, Yuexian Zou |
IEEE Trans. Image Process. | 3 |
| 2021 | Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression GroundingabstractThe prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to one of the proposals. Existing two-stage solutions mostly focus on the grounding step, which aims to align the expressions with the proposals. In this paper, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., expression-agnostic), hoping that the proposals contain all right instances in the expression (i.e., expression-aware). Due to this mismatch, current two-stage methods suffer from a severe performance drop between detected and ground-truth proposals. To this end, we propose Ref-NMS, which is the first method to yield expression-aware proposals at the first stage. Ref-NMS regards all nouns in the expression as critical objects, and introduces a lightweight module to predict a score for aligning each box with a critical object. These scores can guide the NMS operation to filter out the boxes irrelevant to the expression, increasing the recall of critical objects, resulting in a significantly improved grounding performance. Since Ref- NMS is agnostic to the grounding step, it can be easily integrated into any state-of-the-art two-stage method. Extensive ablation studies on several backbones, benchmarks, and tasks consistently demonstrate the superiority of Ref-NMS. Codes are available at: https://github.com/ChopinSharp/ref-nms. Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shih-Fu Chang |
AAAI | 1 |
| 2021 | Boundary Proposal Network for Two-stage Natural Language Video LocalizationabstractWe aim to address the problem of Natural Language Video Localization (NLVL) — localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories: 1) anchor-based approach: it first pre-defines a series of video segment candidates (e.g., by sliding window), and then does classification for each candidate; 2) anchor-free approach: it directly predicts the probabilities for each video frame as a boundary or intermediate frame inside the positive segment. However, both kinds of one-stage approaches have inherent drawbacks: the anchor-based approach is susceptible to the heuristic rules, further limiting the capability of handling videos with variant length. While the anchor-free approach fails to exploit the segment-level interaction thus achieving inferior results. In this paper, we propose a novel Boundary Proposal Network (BPNet), a universal two-stage framework that gets rid of the issues mentioned above. Specifically, in the first stage, BPNet utilizes an anchor-free model to generate a group of high-quality candidate video segments with their boundaries. In the second stage, a visual-language fusion layer is proposed to jointly model the multi-modal interaction between the candidate and the language query, followed by a matching score rating layer that outputs the alignment score for each candidate. We evaluate our BPNet on three challenging NLVL benchmarks (i.e., Charades-STA, TACoS and ActivityNet-Captions). Extensive experiments and ablative studies on these datasets demonstrate that the BPNet outperforms the state-of-the-art methods. Shaoning Xiao, Long Chen 0016, Songyang Zhang 0004, Wei Ji 0008, Jian Shao 0001, Lu Ye, Jun Xiao 0001 |
AAAI | 2 |
| 2021 | Human-Like Controllable Image Captioning With Verb-Specific Semantic RolesabstractControllable Image Captioning (CIC) — generating image descriptions following designated control signals — has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerning objective properties, such as contents of interest or descriptive patterns. However, we argue that almost all existing objective control signals have overlooked two indispensable characteristics of an ideal control signal: 1) Event-compatible: all visual contents referred to in a single sentence should be compatible with the described activity. 2) Sample-suitable: the control signals should be suitable for a specific image sample. To this end, we propose a new control signal for CIC: Verb-specific Semantic Roles (VSR). VSR consists of a verb and some semantic roles, which represents a targeted activity and the roles of entities involved in this activity. Given a designated VSR, we first train a grounded semantic role labeling (GSRL) model to identify and ground all entities for each role. Then, we propose a semantic structure planner (SSP) to learn human-like descriptive semantic structures. Lastly, we use a role-shift captioning model to generate the captions. Extensive experiments and ablations demonstrate that our framework can achieve better controllability than several strong base-lines on two challenging CIC benchmarks. Besides, we can generate multi-level diverse captions easily. The code is available at: https://github.com/mad-red/VSR-guided-CIC. Long Chen 0016, Zhihong Jiang, Jun Xiao 0001, Wei Liu 0005 |
CVPR | 1 |
| 2021 | On Pursuit of Designing Multi-modal Transformer for Video GroundingabstractVideo grounding aims to localize the temporal segment corresponding to a sentence query from an untrimmed video.Almost all existing video grounding methods fall into two frameworks: 1) Top-down model: It predefines a set of segment candidates and then conducts segment classification and regression.2) Bottomup model: It directly predicts frame-wise probabilities of the referential segment boundaries.However, all these methods are not end-to-end, i.e., they always rely on some time-consuming post-processing steps to refine predictions.To this end, we reformulate video grounding as a set prediction task and propose a novel end-toend multi-modal Transformer model, dubbed as GTR.Specifically, GTR has two encoders for video and language encoding, and a crossmodal decoder for grounding prediction.To facilitate the end-to-end training, we use a Cubic Embedding layer to transform the raw videos into a set of visual tokens.To better fuse these two modalities in the decoder, we design a new Multi-head Cross-Modal Attention.The whole GTR is optimized via a Many-to-One matching loss.Furthermore, we conduct comprehensive studies to investigate different model design choices.Extensive results on three benchmarks have validated the superiority of GTR.All three typical GTR variants achieve recordbreaking performance on all datasets and metrics, with several times faster inference speed.Our project is available at GTR. Meng Cao 0002, Long Chen 0016, Zheng Shou 0001, Can Zhang 0001, Yuexian Zou |
EMNLP (1) | 2 |
| 2021 | Natural Language Video Localization with Learnable Moment ProposalsabstractGiven an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query.To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment candidates and then find out the best-matching one.2) proposal-free models directly predict two temporal boundaries of the referential moment from frames.Currently, almost all the propose-and-rank methods have inferior performance than proposal-free counterparts.In this paper, we argue that propose-and-rank approach is underestimated due to the predefined manners: 1) Hand-designed rules are hard to guarantee the complete coverage of targeted segments.2) Densely sampled candidate moments cause redundant computation and degrade the performance of ranking process.To this end, we propose a novel model termed LP-Net (Learnable Proposal Network for NLVL) with a fixed set of learnable moment proposals.The position and length of these proposals are dynamically adjusted during training process.Moreover, a boundary-aware loss has been proposed to leverage frame-level information and further improve the performance.Extensive ablations on two challenging NLVL benchmarks have demonstrated the effectiveness of LPNet over existing state-of-the-art methods 1 . Shaoning Xiao, Long Chen 0016, Jian Shao 0001, Yueting Zhuang, Jun Xiao 0001 |
EMNLP (1) | 2 |
| 2021 | Accelerate CNNs from Three Dimensions: A Comprehensive Pruning FrameworkabstractMost neural network pruning methods, such as filter-level and layer-level prunings, prune the network model along one dimension (depth, width, or resolution) solely to meet a computational budget. However, such a pruning policy often leads to excessive reduction of that dimension, thus inducing a huge accuracy loss. To alleviate this issue, we argue that pruning should be conducted along three dimensions comprehensively. For this purpose, our pruning framework formulates pruning as an optimization problem. Specifically, it first casts the relationships between a certain model’s accuracy and depth/width/resolution into a polynomial regression and then maximizes the polynomial to acquire the optimal values for the three dimensions. Finally, the model is pruned along the three optimal dimensions accordingly. In this framework, since collecting too much data for training the regression is very time-costly, we propose two approaches to lower the cost: 1) specializing the polynomial to ensure an accurate regression even with less training data; 2) employing iterative pruning and fine-tuning to collect the data faster. Extensive experiments show that our proposed algorithm surpasses state-of-the-art pruning algorithms and even neural architecture search-based algorithms. Wenxiao Wang 0001, Minghao Chen 0001, Shuai Zhao 0006, Long Chen 0016, Jinming Hu, Haifeng Liu 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ICML | 4 |
| 2021 | Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningabstractCentralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process is credit assignment, which aims to deduce the contributions of each agent according to the global rewards. Existing credit assignment methods focus on either decomposing the joint value function into individual value functions or measuring the impact of local observations and actions on the global value function. These approaches lack a thorough consideration of the complicated interactions among multiple agents, leading to an unsuitable assignment of credit and subsequently mediocre results on MARL. We propose Shapley Counterfactual Credit Assignment, a novel method for explicit credit assignment which accounts for the coalition of agents. Specifically, Shapley Value and its desired properties are leveraged in deep MARL to credit any combinations of agents, which grants us the capability to estimate the individual credit for each agent. Despite this capability, the main technical difficulty lies in the computational complexity of Shapley Value who grows factorially as the number of agents. We instead utilize an approximation method via Monte Carlo sampling, which reduces the sample complexity while maintaining its effectiveness. We evaluate our method on StarCraft II benchmarks across different scenarios. Our method outperforms existing cooperative MARL algorithms significantly and achieves the state-of-the-art, with especially large margins on tasks with more severe difficulties. Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Furui Liu, Long Chen 0016, Fei Wu 0001, Jun Xiao 0001 |
KDD | 5 |
| 2021 | Video Relation Detection via Tracklet based Visual TransformerabstractVideo Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA[7] and deepSORT [27] to generate tracklet proposals. Then we perform VidVRD in a tracklet-based manner without any pre-cutting operations. Specifically, we design a tracklet-based visual Transformer. It contains a temporal-aware decoder which performs feature interactions between the tracklets and learnable predicate query embeddings, and finally predicts the relations. Experimental results strongly demonstrate the superiority of our method, which outperforms other methods by a large margin on the Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021. Codes are released at https://github.com/Dawn-LX/VidVRD-tracklets. Kaifeng Gao, Long Chen 0016, Jun Xiao 0001 |
ACM Multimedia | 2 |
| 2021 | Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based ExplanationabstractInterpreting model knowledge is an essential topic to improve human understanding of deep black-box models. Traditional methods contribute to providing intuitive instance-wise explanations which allocating importance scores for low-level features (e.g, pixels for images). To adapt to the human way of thinking, one strand of recent researches has shifted its spotlight to mining important concepts. However, these concept-based interpretation methods focus on computing the contribution of each discovered concept on the class level and can not precisely give instance-wise explanations. Besides, they consider each concept as an independent unit, and ignore the interactions among concepts. To this end, in this paper, we propose a novel COncept-based NEighbor Shapley approach (dubbed as CONE-SHAP) to evaluate the importance of each concept by considering its physical and semantic neighbors, and interpret model knowledge with both instance-wise and class-wise explanations. Thanks to this design, the interactions among concepts in the same image are fully considered. Meanwhile, the computational complexity of Shapley Value is reduced from exponential to polynomial. Moreover, for a more comprehensive evaluation, we further propose three criteria to quantify the rationality of the allocated contributions for the concepts, including coherency, complexity, and faithfulness. Extensive experiments and ablations have demonstrated that our CONE-SHAP algorithm outperforms existing concept-based methods and simultaneously provides precise explanations for each instance and class. Jiahui Li 0003, Kun Kuang 0001, Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 4 |
| 2020 | Rethinking the Bottom-Up Framework for Query-Based Video LocalizationabstractIn this paper, we focus on the task query-based video localization, i.e., localizing a query in a long and untrimmed video. The prevailing solutions for this problem can be grouped into two categories: i) Top-down approach: It pre-cuts the video into a set of moment candidates, then it does classification and regression for each candidate; ii) Bottom-up approach: It injects the whole query content into each video frame, then it predicts the probabilities of each frame as a ground truth segment boundary (i.e., start or end). Both two frameworks have respective shortcomings: the top-down models suffer from heavy computations and they are sensitive to the heuristic rules, while the performance of bottom-up models is behind the performance of top-down counterpart thus far. However, we argue that the performance of bottom-up framework is severely underestimated by current unreasonable designs, including both the backbone and head network. To this end, we design a novel bottom-up model: Graph-FPN with Dense Predictions (GDP). For the backbone, GDP firstly generates a frame feature pyramid to capture multi-level semantics, then it utilizes graph convolution to encode the plentiful scene relationships, which incidentally mitigates the semantic gaps in the multi-scale feature pyramid. For the head network, GDP regards all frames falling in the ground truth segment as the foreground, and each foreground frame regresses the unique distances from its location to bi-directional boundaries. Extensive experiments on two challenging query-based video localization tasks (natural language video localization and video relocalization), involving four challenging benchmarks (TACoS, Charades-STA, ActivityNet Captions, and Activity-VRL), have shown that GDP surpasses the state-of-the-art top-down models. Long Chen 0016, Chujie Lu, Siliang Tang, Jun Xiao 0001, Chilie Tan |
AAAI | 1 |
| 2020 | Counterfactual Samples Synthesizing for Robust Visual Question AnsweringabstractDespite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the train set and fail to generalize to the test set with different QA distributions. To reduce the language biases, several recent works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on VQA-CP. However, since the complexity of design, current methods are unable to equip the ensemble-based models with two indispensable characteristics of an ideal VQA model: 1) visual-explainable: the model should rely on the right visual regions when making decisions. 2) question-sensitive: the model should be sensitive to the linguistic variations in question. To this end, we propose a model-agnostic Counterfactual Samples Synthesizing (CSS) training scheme. The CSS generates numerous counterfactual training samples by masking critical objects in images or words in questions, and assigning different ground-truth answers. After training with the complementary samples (ie, the original and generated samples), the VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. In return, the performance of these models is further boosted. Extensive ablations have shown the effectiveness of CSS. Particularly, by building on top of the model LMH, we achieve a record-breaking performance of 58.95% on VQA-CP v2, with 6.5% gains. Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shiliang Pu, Yueting Zhuang |
CVPR | 1 |
| 2020 | Hierarchical Fashion Graph Network for Personalized Outfit RecommendationabstractFashion outfit recommendation has attracted increasing attentions from online shopping services and fashion communities.Distinct from other scenarios (e.g., social networking or content sharing) which recommend a single item (e.g., a friend or picture) to a user, outfit recommendation predicts user preference on a set of well-matched fashion items. Hence, performing high-quality personalized outfit recommendation should satisfy two requirements -- 1) the nice compatibility of fashion items and 2) the consistence with user preference. However, present works focus mainly on one of the requirements and only consider either user-outfit or outfit-item relationships, thereby easily leading to suboptimal representations and limiting the performance. Xiang Wang 0010, Xiangnan He 0001, Long Chen 0016, Jun Xiao 0001, Tat-Seng Chua |
SIGIR | 4 |
| 2020 | Hierarchical Temporal Fusion of Multi-grained Attention Features for Video Question Answering
Shaoning Xiao, Yunan Ye, Long Chen 0016, Shiliang Pu, Zhou Zhao 0001, Jian Shao 0001, Jun Xiao 0001 |
Neural Process. Lett. | 4 |
| 2019 | DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video LocalizationabstractChujie Lu, Long Chen, Chilie Tan, Xiaolin Li, Jun Xiao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chujie Lu, Long Chen 0016, Chilie Tan, Jun Xiao 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Counterfactual Critic Multi-Agent Training for Scene Graph GenerationabstractScene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects. For example, ``person'' on ``bike'' can help to determine the relationship ``ride'', which in turn contributes to the confidence of the two objects. However, we argue that the visual context is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes should not be penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach. CMAT is a multi-agent policy gradient method that frames objects into cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the predictions of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art performance by significant gains under various settings and metrics. Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Xiangnan He 0001, Shiliang Pu, Shih-Fu Chang |
ICCV | 1 |
| 2019 | Learning Using Privileged Information for Food RecognitionabstractFood recognition for user-uploaded images is crucial in visual diet tracking, an emerging application linking multimedia and healthcare domains. However, it is challenging due to the various visual appearances of food images. This is caused by different conditions when taking the photos, such as angles, distances, light conditions, food containers, and background scenes. To alleviate such a semantic gap, this paper presents a cross-modal alignment and transfer network (ATNet), which is motivated by the paradigm of learning using privileged information (LUPI). It additionally utilizes the ingredients in food images as an "intelligent teacher" in the training stage to facilitate cross-modal information passing. Specifically, ATNet first uses a pair of synchronized autoencoders to build the base image and ingredient channels for information flow. Subsequently, the information passing is enabled through a two-stage cross-modal interaction. The first stage of interaction adopts a two-step method, called partial heterogeneous transfer, to 1) alleviate the intrinsic heterogeneity between images and ingredients and 2) align them in a shared space to make their carried information about food classes interact. In the second stage, ATNet learns to map the visual embeddings of images to the ingredient channel for food recognition from the view of "teacher''. This leads a refined recognition by a multi-view fusion. Experiments on two real-world datasets show that ATNet can be incorporated with any state-of-the-art CNN models to consistently improve their performance. Lei Meng 0001, Long Chen 0016, Xun Yang 0001, Dacheng Tao, Hanwang Zhang, Chunyan Miao, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2018 | Zero-Shot Visual Recognition Using Semantics-Preserving Adversarial Embedding NetworksabstractWe propose a novel framework called Semantics-Preserving Adversarial Embedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test images and their classes are both unseen during training. SP-AEN aims to tackle the inherent problem - semantic loss - in the prevailing family of embedding-based ZSL, where some semantics would be discarded during training if they are non-discriminative for training classes, but could become critical for recognizing test classes. Specifically, SP-AEN prevents the semantic loss by introducing an independent visual-to-semantic space embedder which disentangles the semantic space into two subspaces for the two arguably conflicting objectives: classification and reconstruction. Through adversarial learning of the two subspaces, SP-AEN can transfer the semantics from the reconstructive subspace to the discriminative one, accomplishing the improved zero-shot recognition of unseen classes. Comparing with prior works, SP-AEN can not only improve classification but also generate photo-realistic images, demonstrating the effectiveness of semantic preservation. On four popular benchmarks: CUB, AWA, SUN and aPY, SP-AEN considerably outperforms other state-of-the-art methods by an absolute performance difference of 12.2%, 9.3%, 4.0% and 3.6% in terms of harmonic mean values [62]. Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Wei Liu 0005, Shih-Fu Chang |
CVPR | 1 |
| 2017 | SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image CaptioningabstractVisual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN encoding an input image. However, we argue that such spatial attention does not necessarily conform to the attention mechanism - a dynamic feature extractor that combines contextual fixations over time, as CNN features are naturally spatial, channel-wise and multi-layer. In this paper, we introduce a novel convolutional neural network dubbed SCA-CNN that incorporates Spatial and Channel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN dynamically modulates the sentence generation context in multi-layer feature maps, encoding where (i.e., attentive spatial locations at multiple layers) and what (i.e., attentive channels) the visual attention is. We evaluate the proposed SCA-CNN architecture on three benchmark image captioning datasets: Flickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN significantly outperforms state-of-the-art visual attention-based image captioning methods. Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Liqiang Nie, Jian Shao 0001, Wei Liu 0005, Tat-Seng Chua |
CVPR | 1 |
| 2017 | Video Question Answering via Attribute-Augmented Attention Network LearningabstractVideo Question Answering is a challenging problem in visual information retrieval, which provides the answer to the referenced video content according to the question. However, the existing visual question answering approaches mainly tackle the problem of static image question, which may be ineffectively for video question answering due to the insufficiency of modeling the temporal dynamics of video contents. In this paper, we study the problem of video question answering by modeling its temporal dynamics with frame-level attention mechanism. We propose the attribute-augmented attention network learning framework that enables the joint frame-level attribute detection and unified video representation learning for video question answering. We then incorporate the multi-step reasoning process for our proposed attention network to further improve the performance. We construct a large-scale video question answering dataset. We conduct the experiments on both multiple-choice and open-ended video question answering tasks to show the effectiveness of the proposed method. Yunan Ye, Zhou Zhao 0001, Long Chen 0016, Jun Xiao 0001, Yueting Zhuang |
SIGIR | 4 |