EDBT 2026 Demo / reviewers in the wild / expert
Peng Jin 0001
dblp:83/6151-1
· DBLP profile ↗
25ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0001-9287-6410ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 10 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 18 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Next Patch Prediction for AutoRegressive Visual GenerationabstractAutoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autoregressive image generation and extend it to a novel Next Patch Prediction (NPP) paradigm. Our key idea is to group and aggregate image tokens into patch tokens with higher information density. By using patch tokens as a more compact input sequence, the autoregressive model is trained to predict the next patch, significantly reducing computational costs. To further exploit the natural hierarchical structure of image data, we propose a multi-scale coarse-to-fine patch grouping strategy. With this strategy, the training process begins with a large patch size and ends with vanilla NTP where the patch size is 1x1, thus maintaining the original inference process without modifications. Extensive experiments across a diverse range of model sizes demonstrate that NPP could reduce the training cost to around 0.6 times while improving image generation quality by up to 1.0 FID score on the ImageNet 256x256 generation benchmark. Notably, our method retains the original autoregressive model architecture without introducing additional trainable parameters or specifically designing a custom image tokenizer, offering a flexible and plug-and-play solution for enhancing autoregressive visual generation. Yatian Pang, Peng Jin 0001, Bin Lin 0014, Chaoran Feng 0001, Zhenyu Tang 0004, Liuhan Chen, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
AAAI | 2 |
| 2026 | MoE-LLaVA: Mixture of Experts for Large Vision-Language ModelsabstractRecently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA. Bin Lin 0014, Zhenyu Tang 0004, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin 0001, Munan Ning, Jiebo Luo 0001, Li Yuan 0007 |
IEEE Trans. Multim. | 7 |
| 2026 | Disentangled Concept Matching for Text-video Retrieval through Perception ImitationabstractText-video retrieval plays a pivotal role in cross-modal tasks, aiming to match textual descriptions with corresponding video content accurately. Existing methods often employ fine-grained feature matching to improve retrieval accuracy, but such approaches consume extensive computational resources. Conversely, coarse-grained feature matching between entire sentences and videos offers computational efficiency but may overlook the heterogeneous semantic concepts embedded within the data. To overcome these challenges, we develop the Disentangled Concept Matching (DCM) framework, designed as an imitation of human semantic perception processes. The framework utilizes Disentangled Representation Learning (DRL) to divide coarse-grained features into distinct semantic concepts represented as latent factors, effectively generating finer-grained features while reducing computational demands. To improve the accuracy of retrieval, we first propose the Composed Spatial-temporal Module (CSTM) to optimize the quality of multimodal feature extraction. Utilizing a branch-structured temporal modeling approach, CSTM effectively enhances the DCM model’s comprehension of video content and temporal information, leading to the extraction of refined video features. Second, building on the optimized features, we propose the Adaptive Pooling Module (APM) to measure the confidence level of each latent factor matching during the process of decoupling concepts. APM enhances the fidelity of text and video concepts, thereby further ensuring the accuracy of matching after decoupling. With CSTM and APM, DCM accurately matches latent factors in lower dimensions, achieving significant improvements in computing efficiency and retrieval performance. Our experimental evaluations across standard datasets, namely MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that the DCM framework achieves state-of-the-art performance, with Recall@1 scores of 48.7%, 25.6%, 48.4%, 45.0%, and 48.6%, respectively. Compared to our previous model, the DCM framework shows improvements of 2.54%, 0.08%, 2.11%, 6.89%, and 6.35%, respectively. Peng Jin 0001, Chunyu Zou, Ziyao Zhang 0003, Jie Chen 0001, Wen Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationabstractTemporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage the image-text pretraining model for recognizing object instances by separately aligning each frame with class texts. As a result, the separation breaks the instance movement context of videos and requires a lot of inference overhead. To tackle these issues, we propose BridgeText Alignment (BTA) to link frame-level instance representations as a Brownian Bridge. On one hand, we can calculate the global descriptor of a Brownian bridge for capturing instance dynamics, which enables extra considering temporal information rather than only static information of each frame for aligning with texts. On the other hand, according to the goal-conditioned property of the Brownian bridge, we can estimate the middle frame features via the start and the end frame features so the global feature calculation of a Brownian bridge only needs to infer a few frames, which largely reduces inference overhead. We term our overall pipeline as BriVIS. Following the training settings of previous works, BriVIS surpasses the SOTA (OV2Seg) by a clear margin. For example, on the challenging large-vocabulary datasets (BURST, LVVIS), BriVIS achieves 5.7 and 20.9 mAP, which exhibits +2.2∼+6.7 mAP improvement compared to OV2Seg. Furthermore, after training via BTA, using only the head and the tail frames for alignment improves the speed by 32% (2.77 → 1.88 s/iter) while just decreasing the performance by 0.2 mAP (21.1 → 20.9 mAP). Zesen Cheng, Kehan Li 0002, Hao Li 0073, Peng Jin 0001, Xiawu Zheng, Jie Chen 0001 |
AAAI | 4 |
| 2025 | MUSE: Mamba Is Efficient Multi-scale Learner for Text-video RetrievalabstractText-Video Retrieval (TVR) aims to align and associate relevant video content with corresponding natural language queries. Most existing TVR methods are based on large-scale pre-trained vision-language models (e.g., CLIP). However, due to CLIP's inherent plain structure, few TVR methods explore the multi-scale representations which offer richer contextual information for a more thorough understanding. To this end, we propose MUSE, a multi-scale mamba with linear computational complexity for efficient cross-resolution modeling. Specifically, the multi-scale representations are generated by applying a feature pyramid on the last single-scale feature map. Then, we employ the Mamba structure as an efficient multi-scale learner to jointly learn scale-wise representations. Furthermore, we conduct comprehensive studies to investigate different model structures and designs. Extensive results on three popular benchmarks have validated the superiority of MUSE. Meng Cao 0002, Jinfa Huang, Ruyang Liu, Peng Jin 0001, Ge Li 0002, Xiaodan Liang |
AAAI | 5 |
| 2025 | LlaVA-CoT: Let Vision Language Models Reason Step-By-Step
Guowei Xu 0001, Peng Jin 0001, Ziang Wu, Yibing Song, Lichao Sun 0001, Li Yuan 0007 |
ICCV | 2 |
| 2025 | MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation ExpertsabstractIn this work, we aim to simultaneously enhance the effectiveness and efficiency of Mixture-of-Experts (MoE) methods. To achieve this, we propose MoE++, a general and heterogeneous MoE framework that integrates both Feed-Forward Network (FFN) and zero-computation experts. Specifically, we introduce three types of zero-computation experts: the zero expert, copy expert, and constant expert, which correspond to discard, skip, and replace operations, respectively. This design offers three key advantages: (i) **Low Computing Overhead**: Unlike the uniform mixing mechanism for all tokens within vanilla MoE, MoE++ allows each token to engage with a dynamic number of FFNs, be adjusted by constant vectors, or even skip the MoE layer entirely. (ii) **High Performance**: By enabling simple tokens to utilize fewer FFN experts, MoE++ allows more experts to focus on challenging tokens, thereby unlocking greater performance potential than vanilla MoE. (iii) **Deployment Friendly**: Given that zero-computation experts have negligible parameters, we can deploy all zero-computation experts on each GPU, eliminating the significant communication overhead and expert load imbalance associated with FFN experts distributed across different GPUs. Moreover, we leverage gating residuals, enabling each token to consider the pathway taken in the previous layer when selecting the appropriate experts. Extensive experimental results demonstrate that MoE++ achieves better performance while delivering 1.1$\sim$2.1$\times$ expert forward throughput compared to a vanilla MoE model of the same size, which lays a solid foundation for developing advanced and efficient MoE-related models. Peng Jin 0001, Li Yuan 0007, Shuicheng Yan |
ICLR | 1 |
| 2025 | MoH: Multi-Head Attention as Mixture-of-Head AttentionabstractIn this work, we upgrade the multi-head attention mechanism, the core of the Transformer model, to reduce computational costs while maintaining or surpassing the previous accuracy level. We show that multi-head attention can be expressed in the summation form. Drawing on the insight that not all attention heads hold equal significance, we propose Mixture-of-Head attention (MoH), a new architecture that treats attention heads as experts in the Mixture-of-Experts (MoE) mechanism. MoH has two significant advantages: First, MoH enables each token to select the appropriate attention heads, enhancing inference efficiency without compromising accuracy or increasing the number of parameters. Second, MoH replaces the standard summation in multi-head attention with a weighted summation, introducing flexibility to the attention mechanism and unlocking extra performance potential. Extensive experiments on ViT, DiT, and LLMs demonstrate that MoH outperforms multi-head attention by using only 50%$\sim$90% of the attention heads. Moreover, we demonstrate that pre-trained multi-head attention models, such as LLaMA3-8B, can be further continue-tuned into our MoH models. Notably, MoH-LLaMA3-8B achieves an average accuracy of 64.0% across 14 benchmarks, outperforming LLaMA3-8B by 2.4% by utilizing only 75% of the attention heads. We believe the proposed MoH is a promising alternative to multi-head attention and provides a strong foundation for developing advanced and efficient attention-based models. Peng Jin 0001, Li Yuan 0007, Shuicheng Yan |
ICML | 1 |
| 2025 | Orthogonal Subspace Decomposition for Generalizable AI-Generated Image DetectionabstractDetecting AI-generated images (AIGIs), such as natural images or face images, has become increasingly important yet challenging. In this paper, we start from a new perspective to excavate the reason behind the failure generalization in AIGI detection, named the asymmetry phenomenon, where a naively trained detector tends to favor overfitting to the limited and monotonous fake patterns, causing the feature space to become highly constrained and low-ranked, which is proved seriously limiting the expressivity and generalization. One potential remedy is incorporating the pre-trained knowledge within the vision foundation models (higher-ranked) to expand the feature space, alleviating the model's overfitting to fake. To this end, we employ Singular Value Decomposition (SVD) to decompose the original feature space into two orthogonal subspaces. By freezing the principal components and adapting only the remained components, we preserve the pre-trained knowledge while learning fake patterns. Compared to existing full-parameters and LoRA-based tuning methods, we explicitly ensure orthogonality, enabling the higher rank of the whole feature space, effectively minimizing overfitting and enhancing generalization. We finally identify a crucial insight: our method implicitly learns a vital prior that fakes are actually derived from the real, indicating a hierarchical relationship rather than independence. Modeling this prior, we believe, is essential for achieving superior generalization. Our codes are publicly available at https://github.com/YZY-stack/Effort-AIGI-Detection. Zhiyuan Yan 0002, Jiangming Wang, Peng Jin 0001, Ke-Yue Zhang, Chengchun Liu, Shen Chen 0004, Taiping Yao, Shouhong Ding, Baoyuan Wu, Li Yuan 0007 |
ICML | 3 |
| 2025 | Hierarchical Banzhaf Interaction for General Video-Language Representation LearningabstractMultimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning representations using global semantic interactions between pre-defined video-text pairs. However, to enhance and refine such coarse-grained global interactions, more detailed interactions are necessary for fine-grained multimodal learning. In this study, we introduce a creative approach that models video-text as game players using multivariate cooperative game theory to handle uncertainty during fine-grained semantic interactions with diverse granularity, flexible combination, and vague intensity. Specifically, we design the Hierarchical Banzhaf Interaction to simulate the finegrained correspondence between video clips and textual words from hierarchical perspectives. Furthermore, to mitigate the bias in calculations within Banzhaf Interaction, we propose reconstructing the representation through a fusion of single-modal and crossmodal components. This reconstructed representation ensures fine granularity comparable to that of the single-modal representation, while also preserving the adaptive encoding characteristics of cross-modal representation. Additionally, we extend our original structure into a flexible encoder-decoder framework, enabling the model to adapt to various downstream tasks. Extensive experiments on commonly used text-video retrieval, video-question answering, and video captioning benchmarks, with superior performance, validate the effectiveness and generalization of our method. The code is available at https://github.com/jpthu17/HBI. Peng Jin 0001, Hao Li 0073, Li Yuan 0007, Shuicheng Yan, Jie Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Parallel Vertex Diffusion for Unified Visual GroundingabstractUnified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling object box and contour prediction and provide a text-powered interface to vast related multi-modal tasks, e.g., visual question answering and captioning. However, these methods typically generate vertexes sequentially through autoregression, which is prone to be trapped in error accumulation and heavy computation, especially for high-dimension sequence generation in complex scenarios. In this paper, we develop Parallel Vertex Diffusion (PVD) based on the parallelizability of diffusion models to accurately and efficiently generate vertexes in a parallel and scalable manner. Since the coordinates fluctuate greatly, it typically encounters slow convergence when training diffusion models without geometry constraints. Therefore, we consummate our PVD by two critical components, i.e., center anchor mechanism and angle summation loss, which serve to normalize coordinates and adopt a differentiable geometry descriptor from the point-in-polygon problem of computational geometry to constrain the overall difference of prediction and label vertexes. These innovative designs empower our PVD to demonstrate its superiority with state-of-the-art performance across various grounding tasks. Zesen Cheng, Kehan Li 0002, Peng Jin 0001, Siheng Li, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
AAAI | 3 |
| 2024 | Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video UnderstandingabstractLarge language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multi-modal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language model capable of comprehending and engaging in conver-sations involving images and videos through a unified visual representation. Specifically, we employ a set of dynamic visual tokens to uniformly represent images and videos. This representation framework empowers the model to ef-ficiently utilize a limited number of visual tokens to simul-taneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Moreover, we leverage a multi-scale representation, enabling the model to perceive both high-level seman-tic concepts and low-level visual details. Notably, Chat-UniVi is trained on a mixed dataset containing both images and videos, allowing direct application to tasks involving both mediums without requiring any modifications. Exten-sive experimental results demonstrate that Chat- UniVi con-sistently outperforms even existing methods exclusively de-signed for either images or videos. Code is available at https://github.com/PKu-Yuan Group/Chat-UniVi. Peng Jin 0001, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, Li Yuan 0007 |
CVPR | 1 |
| 2024 | Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Runyi Yu 0002, Chang Liu 0047, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (25) | 1 |
| 2024 | FreestyleRet: Retrieving Images from Style-Diversified Queries
Hao Li 0073, Yanhao Jia, Peng Jin 0001, Zesen Cheng, Kehan Li 0002, Jialu Sui, Chang Liu 0047, Li Yuan 0007 |
ECCV (23) | 3 |
| 2024 | Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting
Junwu Zhang, Zhenyu Tang 0004, Yatian Pang, Xinhua Cheng, Peng Jin 0001, Yida Wei, Munan Ning, Li Yuan 0007 |
ECCV (25) | 5 |
| 2024 | Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionabstractLarge Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding.Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models.However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers.In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other.As a result, Video-LLaVA outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively.Additionally, our Video-LLaVA also achieves superior performances on a broad range of 9 image benchmarks.Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos.We aim for this work to provide modest insights into the multi-modal inputs for the LLM. Bin Lin 0014, Jiaxi Cui, Munan Ning, Peng Jin 0001, Li Yuan 0007 |
EMNLP | 6 |
| 2023 | Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningabstractContrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challenging shell-breaking interactions for fine-grained cross-modal learning. In this paper, we creatively model video-text as game players with multivariate cooperative game theory to wisely handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. Concretely, we propose Hierarchical Banzhaf Interaction (HBI) to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast. To efficiently realize the cooperative game of multiple video frames and multiple text words, the proposed method clusters the original video frames (text words) and computes the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we achieve cooperative games at different semantic levels. Extensive experiments on commonly used text-video retrieval and video-question answering bench-marks with superior performances justify the efficacy of our HBI. More encouragingly, it can also serve as a visualization tool to promote the understanding of cross-modal interaction, which have a far-reaching impact on the community. Project page is available at https://jpthu17.github.io/HBI/. Peng Jin 0001, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu 0030, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
CVPR | 1 |
| 2023 | DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelabstractExisting text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-distribution data. To address this limitation, we creatively tackle this task from a generative viewpoint and model the correlation between the text and the video as their joint probability p(candidates,query). This is accomplished through a diffusion-based text-video retrieval framework (Diffusion-Ret), which models the retrieval task as a process of gradually generating joint distribution from noise. During training, DiffusionRet is optimized from both the generation and discrimination perspectives, with the generator being optimized by generation loss and the feature extractor trained with contrastive loss. In this way, DiffusionRet cleverly leverages the strengths of both generative and discriminative methods. Extensive experiments on five commonly used text-video retrieval benchmarks, including MSRVTT, LSMDC, MSVD, ActivityNet Captions, and DiDeMo, with superior performances, justify the efficacy of our method. More encouragingly, without any modification, DiffusionRet even performs well in out-domain retrieval settings. We believe this work brings fundamental insights into the related fields. Code is available at https://github.com/jpthu17/DiffusionRet. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Xiangyang Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
ICCV | 1 |
| 2023 | Multi-granularity Interaction Simulation for Unsupervised Interactive SegmentationabstractInteractive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-oriented interactions with manually labeled object masks. In this work, we reveal that informative interactions can be made by simulation with semantic-consistent yet diverse region exploration in an unsupervised paradigm. Concretely, we introduce a Multi-granularity Interaction Simulation (MIS) approach to open up a promising direction for unsupervised interactive segmentation. Drawing on the high-quality dense features produced by recent self-supervised models, we propose to gradually merge patches or regions with similar features to form more extensive regions and thus, every merged region serves as a semantic-meaningful multi-granularity proposal. By randomly sampling these proposals and simulating possible interactions based on them, we provide meaningful interaction at multiple granularities to teach the model to understand interactions. Our MIS significantly outperforms non-deep learning unsupervised methods and is even comparable with some previous deep-supervised methods without any annotation. Kehan Li 0002, Yian Zhao, Zhennan Wang 0001, Zesen Cheng, Peng Jin 0001, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
ICCV | 5 |
| 2023 | WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image SegmentationabstractThe top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturbed by Inferior Positive (IP) errors due to the lack of prior object information. Nevertheless, we discover that two types of methods are highly complementary for restraining respective weaknesses but the direct average combination leads to harmful interference. In this context, we build Win-win Cooperation (WiCo) to exploit complementary nature of two types of methods on both interaction and integration aspects for achieving a win-win improvement. For the interaction aspect, Complementary Feature Interaction (CFI) introduces prior object information to bottom-up branch and provides fine-grained information to top-down branch for complementary feature enhancement. For the integration aspect, Gaussian Scoring Integration (GSI) models the gaussian performance distributions of two branches and weighted integrates results by sampling confident scores from the distributions. With our WiCo, several prominent bottom-up and top-down combinations achieve remarkable improvements on three common datasets with reasonable extra costs, which justifies effectiveness and generality of our method. Zesen Cheng, Peng Jin 0001, Hao Li 0073, Kehan Li 0002, Siheng Li, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 2 |
| 2023 | Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set AlignmentabstractText-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings. For disentangled conceptualization, we divide the coarse feature into multiple latent factors related to semantic concepts. For set-to-set alignment, where a set of visual concepts correspond to a set of textual concepts, we propose an adaptive pooling method to aggregate semantic concepts to address the partial matching. In particular, since we encode concepts independently in only a few dimensions, DiCoSA is superior at efficiency and granularity, ensuring fine-grained interactions using a similar computational complexity as coarse-grained alignment. Extensive experiments on five datasets, including MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that our method outperforms the existing state-of-the-art methods. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Jinfa Huang, Zhennan Wang 0001, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 1 |
| 2023 | TG-VQA: Ternary Game of Video Question AnsweringabstractVideo question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-grained visual-linguistic alignments. In this work, we innovatively resort to game theory, which can simulate complicated relationships among multiple players with specific interaction strategies, e.g., video, question, and answer as ternary players, to achieve fine-grained alignment for VideoQA task. Specifically, we carefully design a VideoQA-specific interaction strategy to tailor the characteristics of VideoQA, which can mathematically generate the fine-grained visual-linguistic alignment label without label-intensive efforts. Our TG-VQA outperforms existing state-of-the-art by a large margin (more than 5%) on long-term and short-term VideoQA datasets, verifying its effectiveness and generalization ability. Thanks to the guidance of game-theoretic interaction, our model impressively convergences well on limited data (10^4 videos), surpassing most of those pre-trained on large-scale data (10^7 videos). Hao Li 0073, Peng Jin 0001, Zesen Cheng, Songyang Zhang 0001, Kai Chen 0026, Zhennan Wang 0001, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 2 |
| 2023 | Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsabstractMost text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-trained weights are available at https://github.com/jpthu17/GraphMotion. Peng Jin 0001, Yang Wu 0001, Yanbo Fan, Zhongqian Sun, Wei Yang 0019, Li Yuan 0007 |
NeurIPS | 1 |
| 2023 | Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question AnsweringabstractText-based Visual Question Answering (TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial reasoning between texts and objects is crucial in TextVQA. However, existing approaches are constrained within 2D spatial information learned from the input images and rely on transformer-based architectures to reason implicitly during the fusion process. Under this setting, these 2D spatial reasoning approaches cannot distinguish the fine-grained spatial relations between visual objects and scene texts on the same image plane, thereby impairing the interpretability and performance of TextVQA models. In this paper, we introduce 3D geometric information into the spatial reasoning process to capture the contextual knowledge of key objects step-by-step. Specifically, (i) we propose a relation prediction module for accurately locating the region of interest of critical objects; (ii) we design a depth-aware attention calibration module for calibrating the OCR tokens' attention according to critical objects. Extensive experiments show that our method achieves state-of-the-art performance on TextVQA and ST-VQA datasets. More encouragingly, our model surpasses others by clear margins of 5.7% and 12.1% on questions that involve spatial reasoning in TextVQA and ST-VQA valid split. Besides, we also verify the generalizability of our model on the text-based image captioning task. Hao Li 0073, Jinfa Huang, Peng Jin 0001, Guoli Song, Qi Wu 0001, Jie Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsabstractMost video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to the semantic similarities of text-video pairs. However, such learned shared latent spaces are not often optimal, and the modality gap between visual and textual representation can not be fully eliminated. In this paper, we propose Expectation-Maximization Contrastive Learning (EMCL) to learn compact video-and-language representations. Specifically, we use the Expectation-Maximization algorithm to find a compact set of bases for the latent space, where the features could be concisely represented as the linear combinations of these bases. Such feature decomposition of video-and-language representations reduces the rank of the latent space, resulting in increased representing power for the semantics. Extensive experiments on three benchmark text-video retrieval datasets prove that our EMCL can learn more discriminative video-and-language representations than previous methods, and significantly outperform previous state-of-the-art methods across all metrics. More encouragingly, the proposed method can be applied to boost the performance of existing approaches either as a jointly training layer or an out-of-the-box inference module with no extra training, making it easy to be incorporated into any existing methods. Peng Jin 0001, Jinfa Huang, Xian Wu 0001, Shen Ge, Guoli Song, David A. Clifton, Jie Chen 0001 |
NeurIPS | 1 |