EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Liu 0001
dblp:12/8228-1
· DBLP profile ↗
60ranked-venue papers
15as first author
45since 2021 · last 2026
0000-0001-9940-6366ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 11 first-author · 32 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 23 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Navigating Truth in Multimodal Fact-checking via Retrieval- and Reasoning-Enhanced Large Language ModelsabstractRecent studies show that claims incorporating both text and images spread more effectively than those with text alone, presenting significant challenges for multimodal fact-checking. The rapid development of Multi-modal Large Language Models (MLLMs) has greatly advanced research in this field, enabling stronger performance. However, existing MLLM-based fact-checking methods fail to fully exploit visual evidence, and their reliance on rigid fine-tuning templates limits context-aware explanations and leads to weak deep reasoning. To address these limitations, we propose FACTCOMPASS, a novel framework that combines reasoning-aware fine-tuning with large-scale rule-based reinforcement learning and incorporates a semantic- and knowledge-enhanced retrieval module to strengthen deep reasoning and improve evidence utilization. This framework enhances evidence retrieval by obtaining semantically relevant evidence images, enriching the contextual understanding of claim-related images, and refining textual evidence at the knowledge level. To further enhance reasoning, we introduce a self-refining reinforcement fine-tuning strategy: (1) distilling GPT-4o's reasoning from partially fact-checking data for cold-start Chain-of-Thought learning; (2) activating reasoning across broader datasets using prior knowledge and rejection sampling; (3) applying Group Relative Policy Optimization to explore diverse reasoning paths and optimize factual consistency. Extensive experiments have demonstrated the effectiveness of the proposed framework. Fanrui Zhang, Qiang Zhang 0051, Chuanhao Li 0001, Jiaxin Ai, Yukang Feng, Zizhen Li, Kaipeng Zhang, Jiawei Liu 0001, Zhengjun Zha |
WWW | 9 |
| 2026 | Mamba-Driven Comprehensive Context Learning for Zero-Shot HOI Detection
Jiawei Liu 0001, Yongchao Xu, Sen Tao, Yuexuan Qi, Zhengjun Zha |
Int. J. Comput. Vis. | 1 |
| 2026 | Boosting Active Prompt Learning via Discriminative Self-Training Dual-Curriculum Learning
Sen Tao, Jiawei Liu 0001, Yongchao Xu, Bingyu Hu, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2026 | Frequency-Guided Multi-Perspective Prompt Learning for Visible-Infrared Person Re-IdentificationabstractPre-trained vision-language models enhance single-modal person re-identification by leveraging high-level semantic knowledge through prompt learning. However, in visible-in-frared person re-identification, existing prompt learning frameworks primarily emphasize the spatial encoding of modality-shared pedestrian features, inevitably entangling identity-specific characteristics with modality-specific details. Consequently, they struggle to fully disentangle and exploit critical modality-specific cues from the entangled information and fail to construct robust modality-invariant discriminative representations by effectively integrating identity-specific information with features from both modalities, thereby hindering effective cross-modality alignment. To address these limitations, we propose a novel Frequency-Guided Multi-Perspective Prompt Learning (FGMP) framework for visible-infrared person re-identification, which decouples identity-specific and modality-specific information via frequency-guided feature disentanglement, establishes discriminative pedestrian semantics and global modality prototypes through multi-perspective prompts, and performs cross-modality feature fusion, collectively reducing modality discrepancies and facilitating robust pedestrian representation. Specifically, FGMP introduces the Frequency Domain Gated Decoupler (FDGD), which leverages the inherent discreteness of frequency representations to decompose pedestrian features into modality-shared and modality-specific components, overcoming the limitations of prior spatial-centric approaches that entangle identity-specific cues with modality-specific information. Building on this, a heterogeneous prompt learning strategy is employed to generate ID-specific and modality-specific prompts, enriching instance-discriminative information and global modality prototypes by activating CLIP’s inherent multi-modal knowledge and forming complementary high-level semantic interactions. Finally, a multi-modal feature fusion module is designed to seamlessly integrate representations from different modalities and identity-specific characteristics, yielding unified and discriminative pedestrian representations. Extensive experiments conducted on three widely used datasets demonstrate the effectiveness of the proposed FGMP. Jiawei Liu 0001, Yufei Zheng, Guozhi Zhao, Zhengjun Zha |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI DetectionabstractHuman-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative learning, while they fail to identify rare and complex interaction categories effectively as well. In this work, we propose a novel Efficient Mamba-based Disentangled Progressive Learning (HOIMamba) for HOI Detection to absorb the advantages of the existing three approaches and adaptively aggregate multi-level interaction semantics guided by cross-task bidirectional information contexts. Specifically, HOIMamba builds an efficient and effective decoder through cascaded Low-Rank Adaptations (LoRAs), with high efficiency, thorough decoupling of tasks, and good multi-task collaborative learning. Furthermore, to alleviate the recognition problem of interactions in difficult HOI samples, a novel Mamba-based comprehensive progressive learning strategy with Cross-enhance Mamba (CEM) blocks and Detection Context Propagation (DCP) blocks is designed to gradually excavate interaction-related discriminative cues from four levels. CEM blocks automatically aggregate context to generate diverse task-shared semantics and simultaneously realize the cross-task interaction between human and object branches, guiding the interaction branch to extract more expressive HOI representation. DCP blocks further transfer the comprehensive interaction context to human and object branches to achieve rich and effective information exchange, facilitating the model to discover more HOI instances. Extensive experimental results on two standard benchmarks demonstrate the effectiveness of our HOIMamba. Yongchao Xu, Jiawei Liu 0001, Sen Tao, Qiang Zhang 0051, Zhengjun Zha |
AAAI | 2 |
| 2025 | Multi-granularity and Multi-modal Prompt Learning for Person Re-Identification
Jiawei Liu 0001, Guozhi Zhao, Fanrui Zhang, Zhengjun Zha |
CVM (3) | 2 |
| 2025 | Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image CaptioningabstractGenerating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this paper, we introduce a detailed caption benchmark, termed as CompreCap, to evaluate the visual context from a directed scene graph view. Concretely, we first manually segment the image into semantically meaningful regions (i.e., semantic segmentation mask) according to common-object vocabulary, while also distinguishing attributes of objects within all those regions. Then directional relation labels of these objects are annotated to compose a directed scene graph that can well encode rich compositional information of the image. Based on our directed scene graph, we develop a pipeline to assess the generated detailed captions from LVLMs on multiple levels, including the object-level coverage, the accuracy of attribute descriptions, the score of key relationships, etc. Experimental results on the CompreCap dataset confirm that our evaluation method aligns closely with human evaluation scores across LVLMs. We have released the code and the dataset here to support the community. Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu 0001, Wei Zhai, Yang Cao 0010, Yujun Shen, Zhengjun Zha |
CVPR | 6 |
| 2025 | Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video GenerationabstractSora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask2DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask2DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject/. Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang, Jiawei Liu 0001, SiYu Zhou 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
CVPR | 5 |
| 2025 | AR-Diffusion: Asynchronous Video Generation with Auto-Regressive DiffusionabstractThe task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inconsistencies between training and inference, leading to issues such as error accumulation, while synchronous diffusion models are limited by their reliance on rigid sequence length. To address these issues, we introduce Auto-Regressive Diffusion (AR-Diffusion), a novel model that combines the strengths of auto-regressive and diffusion models for flexible, asynchronous video generation. Specifically, our approach leverages diffusion to gradually corrupt video frames in both training and inference, reducing the discrepancy between these phases. Inspired by auto-regressive generation, we incorporate a non-decreasing constraint on the corruption timesteps of individual frames, ensuring that earlier frames remain clearer than subsequent ones. This setup, together with temporal causal attention, enables flexible generation of videos with varying lengths while preserving temporal coherence. In addition, we design two specialized timestep schedulers: the FoPP scheduler for balanced timestep sampling during training, and the AD scheduler for flexible timestep differences during inference, supporting both synchronous and asynchronous generation. Extensive experiments demonstrate the superiority of our proposed method, which achieves competitive and state-of-the-art results across four challenging benchmarks.1 2 Mingzhen Sun, Weining Wang 0001, Jiawei Liu 0001, Wanquan Feng, Shanshan Lao, SiYu Zhou 0002, Jing Liu 0001 |
CVPR | 4 |
| 2025 | Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time AdaptationabstractTest-time adaptation using vision- language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single- task test- time adaptation scenarios and have not effectively explored the issue of multi- task adaptation. To address this practical problem, we propose a novel Hierarchical Knowledge Prompt Tuning (HKPT) method, which achieves joint adaptation to multiple target domains by mining more comprehensive source domain discriminative knowledge and hierarchically modeling task- specific and task- shared knowledge. Specifically, HKPT constructs a CLIP prompt distillation framework that utilizes the broader source domain knowledge of large teacher CLIP to guide prompt tuning for lightweight student CLIP from multiple views during testing. Meanwhile, HKPT establishes task- specific dual dynamic knowledge graph to capture fine- grained contextual knowledge from continuous test data. To fully exploit the complementarity among multiple target tasks, HKPT employs an adaptive task grouping strategy for achieving intertask knowledge sharing. Furthermore, HKPT can seamlessly transfer to basic single- task test- time adaptation scenarios while maintaining robust performance. Extensive experimental results in both multi- task and single- task testtime adaptation settings demonstrate that our HKPT significantly outperforms state- of- the- art methods. Qiang Zhang 0051, Mengsheng Zhao, Jiawei Liu 0001, Fanrui Zhang, Yongchao Xu, Zhengjun Zha |
CVPR | 3 |
| 2025 | I2VControl: Disentangled and Unified Video Motion Synthesis ControlabstractMotion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We rethink camera control, object dragging, and motion brush, reformulating all tasks into a consistent representation based on point trajectories, each managed by a dedicated formulation. Accordingly, we propose a spatial partitioning strategy, where each unit is assigned to a concomitant control category, enabling diverse control types to be dynamically orchestrated within a single synthesis pipeline without conflicts. Furthermore, we design an adapter structure that functions as a plug-in for pre-trained models and is agnostic to specific model architectures. We conduct extensive experiments, achieving excellent performance on various control tasks, and our method further facilitates user-driven creative combinations, enhancing innovation and creativity. Project page: https://wanquanf.github.io/I2VControl . Wanquan Feng, Tianhao Qi, Jiawei Liu 0001, Mingzhen Sun, Pengqi Tu, Tianxiang Ma, Songtao Zhao, SiYu Zhou 0002 |
ICCV | 3 |
| 2025 | I2VControl-Camera: Precise Video Camera Control with Adjustable Motion StrengthabstractVideo generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However, existing camera control methods still suffer from several limitations, including control precision and the neglect of the control for subject motion dynamics. In this work, we propose I2VControl-Camera, a novel camera control method that significantly enhances controllability while providing adjustability over the strength of subject motion. To improve control precision, we employ point trajectory in the camera coordinate system instead of only extrinsic matrix information as our control signal. To accurately control and adjust the strength of subject motion, we explicitly model the higher-order components of the video trajectory expansion, not merely the linear terms, and design an operator that effectively represents the motion strength. We use an adapter architecture that is independent of the base model structure. Experiments on static and dynamic scenes show that our framework outperformances previous methods both quantitatively and qualitatively. Project page: https://wanquanf.github.io/I2VControlCamera. Wanquan Feng, Jiawei Liu 0001, Pengqi Tu, Tianhao Qi, Mingzhen Sun, Tianxiang Ma, Songtao Zhao, SiYu Zhou 0002 |
ICLR | 2 |
| 2025 | Learnable Frequency Decomposition for Image Forgery Detection and LocalizationabstractConcern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic changes caused by image tampering. Observations indicate that manipulation traces are especially prominent in phase components and span both low and high-frequency bands. Based on these findings, we propose a forensic frequency decomposition network (F2D-Net), which incorporates deep Fourier transforms and leverages both phase information and high and low-frequency components to enhance IFDL. Specifically, F2D-Net consists of the Spectral Decomposition Subnetwork (SDSN) and the Frequency Separation Subnetwork (FSSN). The former decomposes the image into amplitude and phase, focusing on learning the semantic content in the phase spectrum to identify forged objects, thus improving forgery detection accuracy. The latter further adaptively decomposes the output of the SDSN to obtain corresponding high and low frequencies, and applies a divide-and-conquer strategy to refine each frequency band, mitigating the optimization difficulties caused by coupled forgery traces across different frequencies, thereby better capturing the pixels belonging to the forged object to improve localization accuracy. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image forgery detection and localization techniques both qualitatively and quantitatively. Dong Li 0055, Jiaying Zhu, Yidi Liu, Xin Lu 0008, Xueyang Fu, Jiawei Liu 0001, Aiping Liu, Zhengjun Zha |
IJCAI | 6 |
| 2025 | Dual Uncertainty-Guided Feature Alignment Learning for Text-Based Person RetrievalabstractText-based person retrieval (TBPR) aims to retrieve pedestrian images based on textual descriptions, facing challenges due to the inherent heterogeneity and uncertainty between visual and textual modalities. Most existing methods focus on addressing heterogeneity while neglecting the issue of uncertainty. To tackle the uncertainty arising from the diverse textual expressions, including both structural and semantic content variations, we propose a novel Dual Uncertainty-Guided Feature Alignment Learning (DUAL) approach, utilizing instance-level and identity-level uncertainty estimations to mitigate these impacts. Specifically, for the uncertainty caused by textual structure variations, DUAL first introduces an uncertainty Gaussian modeling module that represents image and text features as Gaussian distributions in a learnable manner, and estimates instance-level uncertainty coefficients to quantify structural differences within the text. Subsequently, DUAL leverages ShareGPT4V to standardize the text structure, dynamically aligning the original text features with structure-invariant generated text features through adaptive knowledge distillation guided by the instance-level uncertainty coefficients, effectively reducing structural diversity's impact while minimizing noise. Moreover, for the uncertainty caused by the diversity of textual semantic content, DUAL designs an alignment loss that utilizes identity-level uncertainty coefficients, estimated via a Gaussian Mixture Model based on the distances between image and text features of the same identity, effectively mitigating the impact of semantic content diversity. Experimental results demonstrate that DUAL outperforms existing methods on TBPR benchmarks, highlighting its superiority in multimodal person retrieval. Yufei Zheng, Jiawei Liu 0001, Bingyu Hu, Zikun Wei, Zhengjun Zha |
ACM Multimedia | 2 |
| 2025 | Fact-R1: Towards Explainable Video Misinformation Detection with Deep ReasoningabstractThe rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100,000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification. Fanrui Zhang, Qiang Zhang 0051, Jun Chen 0005, Sinbadliu, Junxiong Lin, Jiahong Yan, Jiawei Liu 0001, Zhengjun Zha |
NeurIPS | 8 |
| 2025 | Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceabstractThe dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coordinated editing, leading to less reliable outcomes. To tackle these challenges, we propose CARML, a novel retrieval-augmented editing framework that integrates conflict-aware dynamic retrieval with multi-level implicit and explicit guidance for reliable lifelong multimodal editing. Specifically, CARML introduces intra-modal uncertainty and inter-modal conflict quantification to dynamically integrate multi-channel retrieval results, so as to pinpoint the most relevant knowledge to the incoming edit samples. Afterwards, an edit scope classifier discerns whether the edit sample semantically aligns with the edit scope of the retrieved knowledge. If deemed in-scope, CARML refines the retrieved knowledge into information-rich continuous prompt prefixes, serving as the implicit knowledge guide. These prefixes not only include static knowledge prompt that capture key textual semantics but also incorporate token-level, context-aware dynamic prompt to explore fine-grained cross-modal associations between the edit sample and retrieved knowledge. To further enhance reliability, CARML incorporates a "hard correction" mechanism, leveraging explicit label knowledge to adjust the model’s output logits. Extensive experiments across multiple MLLMs and datasets indicate the superior performance of CARML in lifelong multimodal editing scenarios. Qiang Zhang 0051, Fanrui Zhang, Jiawei Liu 0001, Junjun He, Zhengjun Zha |
NeurIPS | 3 |
| 2025 | Advancing Visible-Infrared Person Re-Identification: Synergizing Visual-Textual Reasoning and Cross-Modal Feature AlignmentabstractVisible-infrared person re-identification (VI-ReID) is a critical cross-modality fine-grained classification task with significant implications for public safety and security applications. Existing VI-ReID methods primarily focus on extracting modality-invariant features for person retrieval. However, due to the inherent lack of texture information in infrared images, these modality-invariant features tend to emphasize global contexts. Consequently, individuals with similar silhouettes are often misidentified, posing potential risks to security systems and forensic investigations. To address this problem, this paper innovatively introduces natural language descriptions to learn the global-local contexts for VI-ReID. Specifically, we design a framework that jointly optimizes visible-infrared alignment plus (VIAP) and visual-textual reasoning (VTR), and introduces local-global joint measure (LJM) to enhance the metric, while proposing a human-LLM collaborative approach to incorporate textual descriptions into existing cross-modal person re-identification datasets. VIAP achieves cross-modal alignment between RGB and IR. It can explicitly utilize designed frequency-aware modality alignment and relationship-reinforced fusion to explore the potential of local cues in global features and modality-invariant information. VTR proposes pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. LJM proposes introducing local feature distances into the measure stage metric to enhance the relevance of matching using fine-grained information. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches. The dataset is publicly available athttps://github.com/qyx596/vireid-caption. Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Toward Effective and Transferable Detection for Multi-Modal Fake News in the Social Media StreamabstractThe rapid proliferation of multimedia fake news on social media has raised significant concerns in recent years. Existing studies on fake news detection predominantly adopt an instance-based paradigm, where the detector evaluates a single post to determine its veracity. Despite notable advancements achieved in this domain, we argue that the instance-based approach is misaligned with real-world deployment scenarios. In practice, detectors typically operate on servers that process incoming posts in temporal order, striving to assess their authenticity promptly. Instance-based detectors lack awareness of temporal information and contextual relationships between surrounding posts, therefore fail to capture long-range dependencies from the timeline. To bridge this gap, we introduce a more practical stream-based multi-modal fake news detection paradigm, which assumes that social media posts arrive continuously over time and allows the utilization of previously seen posts to aid in the classification of incoming ones. To enable effective and transferable fake news detection under this novel paradigm, we propose maintaining historical knowledge as a collection of incremental high-level forgery patterns. Based on this principle, we design a novel framework called Incremental Forgery Pattern Learning and Clues Refinement (IPLCR). IPLCR incrementally learns high-level forgery patterns as the stream evolves, leveraging this knowledge to improve the detection of newly arrived posts. At the core of IPLCR is the Incremental Forgery Pattern Bank (IPB), which dynamically summarizes historical posts into a set of latent forgery patterns. IPB is designed to continuously incorporate timely knowledge and actively discard obsolete information, even during inference. When a new post arrives, IPLCR retrieves the most relevant forgery pattern knowledge from IPB and refines the clues for fake news detection. The refined clues are subsequently incorporated into IPB to enrich its knowledge base. Extensive experiments validate IPLCR's effectiveness as a robust stream-based detector. Moreover, IPLCR addresses several critical issues relevant to industrial applications, including seamless context transfer and efficient model upgrading, making it a practical solution for realworld deployment. Jiawei Liu 0001, Zhengjun Zha |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Noise-Resistance Learning via Multi-Granularity Consistency for Unsupervised Domain Adaptive Person Re-IdentificationabstractUnsupervised domain adaptive person re-identification aims at adapting the re-identification model trained on a labeled source domain to an unlabeled target domain. The mainstream pipeline alternates between clustering-based pseudo-label prediction and representation learning, but the imperfect interaction between these steps generates noisy pseudo labels that diminish the model’s effectiveness. Previous methods reduce noisy pseudo labels impact by assessing consistency only at a single granularity level, overlooking multi-level confidence analysis for better feature representation. To address the issue, we propose a novel multi-granularity consistency network (MGCN) to perform the noise-resistance learning across different granularity consistency perspectives, including prototype-wise consistency, triplet-wise consistency and list-wise consistency, to suppress the contribution of noisy samples simultaneously. Specifically, the prototype-wise consistency leverages the prototypical output affinity between teacher and student networks to evaluate the reliability of pseudo label of a target sample, thus reducing their negative impact on identity classification loss. Triplet-wise consistency focuses on the triplet distance discrepancies between the teacher and student networks to retain the reliable and informative samples that satisfy triplet distance constraint in the triplet loss, thereby facilitating more effective model training and improved performance in the target domain. Furthermore, the list-wise consistency uses accurate list-wise similarity rankings from the teacher’s memory bank to select more dependable neighboring samples in the student’s memory bank, pulling these closer in feature space to alleviate the detrimental effects of noisy labels in contrastive loss. Based on the multi-granularity consistency, MGCN evaluates the credibility of pseudo labels and adjusts their impact across three re-ID losses for effective domain adaptation. Experimental results demonstrate that our proposed method achieves significant improvements over the existing methods on multiple benchmarks. Yangchun Zhu, Yufei Zheng, Jiawei Liu 0001, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | DEADiff: An Efficient Stylization Diffusion Model with Disentangled RepresentationsabstractThe diffusion-based text-to-image model harbors im-mense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transfer-ring styles. In this paper, we introduce DEADiff to address this issue using the following two strategies: 1) a mecha-nism to decouple the style and semantics of reference images. The decoupled feature representations are first extracted by Q-Formers which are instructed by different text descriptions. Then they are injected into mutually exclusive subsets of cross-attention layers for better disentanglement. 2) A non-reconstructive learning method. The Q-Formers are trained using paired images rather than the identical target, in which the reference image and the ground-truth image are with the same style or semantics. We show that DEADiff attains the best visual stylization results and optimal balance between the text controllability inherent in the text-to-image model and style similarity to the reference image, as demonstrated both quantitatively and qualitatively. Our project page is https://tianhao-qi.github.io/DEADiff‘/. Tianhao Qi, Shancheng Fang, Yanze Wu, Hongtao Xie 0001, Jiawei Liu 0001, Lang Chen, Yongdong Zhang 0001 |
CVPR | 5 |
| 2024 | Noise-Assisted Prompt Learning for Image Forgery Detection and Localization
Dong Li 0055, Jiaying Zhu, Xueyang Fu, Xun Guo 0001, Yidi Liu, Jiawei Liu 0001, Zhengjun Zha |
ECCV (11) | 7 |
| 2024 | Joint Visual-Textual Reasoning and Visible-Infrared Modality Alignment for Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality fine-grained classification task. Existing approaches for VI-ReID mainly explore modality-invariant features for person retrieval. However, modality-invariant features pay more attention to global contexts, due to the lack of texture information in infrared images. This leads to a person with similar silhouette often being misidentified. Targeting this problem, this paper innovatively introduces natural language specification to learn global-local contexts for VI-ReID. Specifically, our framework jointly optimizes visible-infrared alignment (VIA) and visual-textual reasoning (VTR). VIA achieves cross-modal between RGB and IR. It can explicitly utilize designed modality-guided alignment and relationship-reinforced fusion to explore the potential of local cues in global features. VTR proposes the pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches. Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002 |
ICME | 4 |
| 2024 | Natural Language-centered Inference Network for Multi-modal Fake News Detection
Qiang Zhang 0051, Jiawei Liu 0001, Fanrui Zhang, Zhengjun Zha |
IJCAI | 2 |
| 2024 | Cross-Modal Semantic Alignment Learning for Text-Based Person Search
Wenjun Gan, Jiawei Liu 0001, Yangchun Zhu, Guozhi Zhao, Zhengjun Zha |
MMM (1) | 2 |
| 2024 | ESCNet: Entity-enhanced and Stance Checking Network for Multi-modal Fact-CheckingabstractRecently, misinformation incorporating both texts and images has been disseminated more effectively than those containing text alone on social media, raising significant concerns for multi-modal fact-checking. Existing research makes contributions to multi-modal feature extraction and interaction, but fails to fully enhance the valuable semantic representations or excavate the intricate entity information. Besides, existing multi-modal fact-checking datasets are primarily focused on English and merely concentrate on a single type of misinformation, thereby neglecting a comprehensive summary and coverage of various types of misinformation. Taking these factors into account, we construct the first large-scale Chinese Multi-modal Fact-Checking (CMFC) dataset which encompasses 46,000 claims. The CMFC covers all types of misinformation for fact-checking and is divided into two sub-datasets, Collected Chinese Multi-modal Fact-Checking (CCMF) and Synthetic Chinese Multi-modal Fact-Checking (SCMF). To establish baseline performance, we propose a novel Entity-enhanced and Stance Checking Network (ESCNet), which includes Multi-modal Feature Extraction Module, Stance Transformer, and Entity-enhanced Encoder. The ESCNet jointly models stance semantic reasoning features and knowledge-enhanced entity pair features, in order to simultaneously learn effective semantic-level and knowledge-level claim representations. Our work offers the first step and establishes a benchmark for evidence-based, multi-type, multi-modal fact-checking. Fanrui Zhang, Jiawei Liu 0001, Qiang Zhang 0051, Yongchao Xu, Zhengjun Zha |
WWW | 2 |
| 2024 | Exert Diversity and Mitigate Bias: Domain Generalizable Person Re-identification with a Comprehensive Benchmark
Bingyu Hu, Jiawei Liu 0001, Yufei Zheng, Kecheng Zheng, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2024 | Adaptive Texture and Spectrum Clue Mining for Generalizable Face Forgery DetectionabstractAlthough existing face forgery detection methods achieve satisfactory performance under closed within-dataset scenario where training and testing sets are created by the same manipulation technique, they are vulnerable to samples created by unseen manipulation techniques under cross-dataset scenario. To solve this problem, in this work, we propose a novel adaptive texture and spectral clue mining (ATSC) approach for generalizable face forgery detection. It adaptively adjusts the parameters depended on input images to mine specific intrinsic forgery clues on both spatial and frequency domains. Specifically, ATSC customizes a Texture Clue Mining Module and a Spectrum Clue Selecting Module. The former module exploits instance-aware dynamic convolution on spatial domain to dynamically assemble multiple parallel convolutional kernels based on the learned image-dependent attention maps for effectively capturing subtle texture artifacts on spatial domain. A customized attention loss is also applied to the attention maps as supervision to precisely localize forgery artifacts whilst retain useful background information from suspicious and non-suspicious regions, which drives the module to explore all potential crucial clues and learning robust texture-related forgery feature. Moreover, the latter module applies adaptive frequency filtering mechanism on DCT-based frequency signals, which selects frequency information of interest to capture refined spectrum clues on frequency domain in an input-adaptive manner. Equipped with the above two modules, ATSC can learn more generalizable forgery features for face forgery detection. Extensive experimental results demonstrate the superior generalization ability of the proposed ATSC over various state-of-the-art methods on the challenging benchmarks. Jiawei Liu 0001, Yang Wang 0015, Zhengjun Zha |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Unleashing Knowledge Potential of Source Hypothesis for Source-Free Domain AdaptationabstractSource-Free Domain Adaptation (SFDA) task aims to transfer knowledge from a labeled source domain to a label-scarce target domain, in which the source data can not be accessed but only a pre-trained source model and unlabeled target data are available during adaptation. Previous methods for source model adaptation rely on hypothesis transfer learning that trains the feature extractor to learn target features aligned to the distribution of source features while freezing the source classifier. However, reusing only the source classifier without exploring the comprehensive knowledge of the source model can lead to biased feature alignment. To this end, we propose a novel method called Transformer-bAsed thorouGh Source HypOthesis Transfer (TagSHOT) framework to effectively unleash the thorough knowledge potential of pre-trained source hypothesis. Specifically, our approach delves into the correlation coefficient among CLS/patch tokens across different Transformer layers, uncovering the concealed insights within the pre-trained source model and constructing a comprehensive source hypothesis. By tailoring the target feature alignment to this thorough source hypothesis, our model facilitates the adaptation of a broader range of classification-related knowledge to the target domain. Furthermore, we introduce a Salient Token Extension (STE) module, designed to capture the target-specific discriminative information by propagating the salient information among tokens. This mechanism enriches our model's ability to understand and incorporate target-specific nuances. Extensive experiments have been conducted to validate the effectiveness of our method, which outperforms state-of-the-art approaches by a large margin. Bingyu Hu, Jiawei Liu 0001, Kecheng Zheng, Zhengjun Zha |
IEEE Trans. Multim. | 2 |
| 2024 | Sounding Video Generator: A Unified Framework for Text-Guided Sounding Video GenerationabstractAs a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded. In this work, we concentrate on a rarely investigated problem of text-guided sounding video generation and propose the Sounding Video Generator (SVG), a unified framework for generating realistic videos along with audio signals. Specifically, we present the SVG-VQGAN to transform visual frames and audio mel-spectrograms into discrete tokens. SVG-VQGAN applies a novel hybrid contrastive learning method to model inter-modal and intra-modal consistency and improve the quantized representations. A cross-modal attention module is employed to extract associated features of visual frames and audio signals for contrastive learning. Then, a Transformer-based decoder is used to model associations between texts, visual frames, and audio signals at token level for auto-regressive sounding video generation. AudioSet-Cap, a human annotated text-video-audio paired dataset, is produced for training SVG. Experimental results demonstrate the superiority of our method when compared with existing text-to-video generation methods as well as audio generation methods on Kinetics and VAS datasets. Jiawei Liu 0001, Weining Wang 0001, Jing Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Edge-aware Regional Message Passing Controller for Image Forgery LocalizationabstractDigital image authenticity has promoted research on image forgery localization. Although deep learning-based methods achieve remarkable progress, most of them usually suffer from severe feature coupling between the forged and authentic regions. In this work, we propose a two-step Edge-aware Regional Message Passing Controlling strategy to address the above issue. Specifically, the first step is to account for fully exploiting the edge information. It consists of two core designs: context-enhanced graph construction and threshold-adaptive differentiable binarization edge algorithm. The former assembles the global semantic information to distinguish the features between the forged and authentic regions, while the latter stands on the output of the former to provide the learnable edges. In the second step, guided by the learnable edges, a region message passing controller is devised to weaken the message passing between the forged and authentic regions. In this way, our ERMPC is capable of explicitly modeling the inconsistency between the forged and authentic regions and enabling it to perform well on refined forged images. Extensive experiments on several challenging benchmarks show that our method is superior to state-of-the-art image forgery localization methods qualitatively and quantitatively. Dong Li 0055, Jiaying Zhu, Menglu Wang 0003, Jiawei Liu 0001, Xueyang Fu, Zhengjun Zha |
CVPR | 4 |
| 2023 | WL-MSR: Watch and Listen for Multimodal Subtitle RecognitionabstractVideo subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition (WL-MSR) framework to obtain comprehensive video subtitles, by fusing the information provided by Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models. Specifically, we build a Transformer model with mask and crop strategies and multi-level identity embeddings to aggregate both the textual results and features of the two modalities. To pre-filter out the noise items in OCR results before fusion, we adopt an OCR filter based on ASR results and confidence scores of OCR. By combining these techniques, our solution wins the 2nd place in Multimodal Subtitle Recognition Challenge on ICPR2022. Jiawei Liu 0001, Weining Wang 0001, Xingjian He, Jing Liu 0001 |
ICASSP | 1 |
| 2023 | Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language ModelsabstractPrompt tuning and adapter tuning have shown great potential in transferring pre-trained vision-language models (VLMs) to various downstream tasks. In this work, we design a new type of tuning method, termed as regularized mask tuning, which masks the network parameters through a learnable selection. Inspired by neural pathways, we argue that the knowledge required by a downstream task already exists in the pre-trained weights but just gets concealed in the upstream pre-training stage. To bring the useful knowledge back into light, we first identify a set of parameters that are important to a given downstream task, then attach a binary mask to each parameter, and finally optimize these masks on the downstream data with the parameters frozen. When updating the mask, we introduce a novel gradient dropout strategy to regularize the parameter selection, in order to prevent the model from forgetting old knowledge and overfitting the downstream data. Experimental results on 11 datasets demonstrate the consistent superiority of our method over previous alternatives. It is noteworthy that we manage to deliver 18.73% performance improvement compared to the zero-shot CLIP via masking an average of only 2.56% parameters. Furthermore, our method is synergistic with most existing parameter-efficient tuning methods and can boost the performance on top of them. Project page can be found here. Kecheng Zheng, Ruili Feng, Kai Zhu 0004, Jiawei Liu 0001, Deli Zhao, Zhengjun Zha, Wei Chen 0001, Yujun Shen |
ICCV | 5 |
| 2023 | ED-T2V: An Efficient Training Framework for Diffusion-based Text-to-Video GenerationabstractDiffusion models have achieved remarkable performance on image generation. However, It is difficult to reproduce this success on video generation because of expensive training cost. In fact, pretrained image generation models have already acquired visual generation capabilities and could be utilized for video generation. Thus, we propose an Efficient training framework for Diffusion-based Text-to-Video generation (ED-T2V), which is built on a pretrained text-to-image generation model. To model the temporal dynamic information, we propose temporal transformer blocks with novel identity attention and temporal cross-attention. ED-T2V has the following advantages: 1) most of the parameters of pretrained model are frozen to inherit the generation capabilities and reduce the training cost; 2) the identity attention requires the currently generated frame to attend to all positions of its previous frame, thus providing an efficient way to keep main content consistent across frames and enable movement generation; 3) the temporal cross-attention is proposed to construct associations between textual descriptions and multiple video tokens in the time dimension, which could better model video movement than traditional cross-attention methods. With the aforementioned benefits, ED-T2V not only significantly reduces the training cost of video diffusion models, but also has excellent generation fidelity and controllability. Jiawei Liu 0001, Weining Wang 0001, Jing Liu 0001 |
IJCNN | 1 |
| 2023 | ECENet: Explainable and Context-Enhanced Network for Muti-modal Fact verificationabstractRecently, falsified claims incorporating both text and images have been disseminated more effectively than those containing text alone, raising significant concerns for multi-modal fact verification. Existing research makes contributions to multi-modal feature extraction and interaction, but fails to fully utilize and enhance the valuable and intricate semantic relationships between distinct features. Moreover, most detectors merely provide a single outcome judgment and lack an inference process or explanation. Taking these factors into account, we propose a novel Explainable and Context-Enhanced Network (ECENet) for multi-modal fact verification, making the first attempt to integrate multi-clue feature extraction, multi-level feature reasoning, and justification (explanation) generation within a unified framework. Specifically, we propose an Improved Coarse- and Fine-grained Attention Network, equipped with two types of level-grained attention mechanisms, to facilitate a comprehensive understanding of contextual information. Furthermore, we propose a novel justification generation module via deep reinforcement learning that does not require additional labels. In this module, a sentence extractor agent measures the importance between the query claim and all document sentences at each time step, selecting a suitable amount of high-scoring sentences to be rewritten as the explanation of the model. Extensive experiments demonstrate the effectiveness of the proposed method. Fanrui Zhang, Jiawei Liu 0001, Qiang Zhang 0051, Esther Sun, Zhengjun Zha |
ACM Multimedia | 2 |
| 2023 | Hierarchical Semantic Enhancement Network for Multimodal Fake News DetectionabstractThe explosion of multimodal fake news content on social media has sparked widespread concern. Existing multimodal fake news detection methods have made significant contributions to the development of this field, but fail to adequately exploit the potential semantic information of images and ignore the noise embedded in news entities, which severely limits the performance of the models. In this paper, we propose a novel Hierarchical Semantic Enhancement Network (HSEN) for multimodal fake news detection by learning text-related image semantic and precise news high-order knowledge semantic information. Specifically, to complement the image semantic information, HSEN utilizes textual entities as the prompt subject vocabulary and applies reinforcement learning to discover the optimal prompt format for generating image captions specific to the corresponding textual entities, which contain multi-level cross-modal correlation information. Moreover, HSEN extracts visual and textual entities from image and text, and identifies additional visual entities from image captions to extend image semantic knowledge. Based on that, HSEN exploits an adaptive hard attention mechanism to automatically select strongly related news entities and remove irrelevant noise entities to obtain precise high-order knowledge semantic information, while generating attention mask for guiding cross-modal knowledge interaction. Extensive experiments show that our method outperforms state-of-the-art methods. Qiang Zhang 0051, Jiawei Liu 0001, Fanrui Zhang, Zhengjun Zha |
ACM Multimedia | 2 |
| 2022 | Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-identificationabstractRGB-infrared person re-identification is an emerging cross-modality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a novel modality-adaptive mixup and invariant decomposition (MID) approach for RGB-infrared person re-identification towards learning modality-invariant and discriminative representations. MID designs a modality-adaptive mixup scheme to generate suitable mixed modality images between RGB and infrared images for mitigating the inherent modality discrepancy at the pixel-level. It formulates modality mixup procedure as Markov decision process, where an actor-critic agent learns dynamical and local linear interpolation policy between different regions of cross-modality images under a deep reinforcement learning framework. Such policy guarantees modality-invariance in a more continuous latent space and avoids manifold intrusion by the corrupted mixed modality samples. Moreover, to further counter modality discrepancy and enforce invariant visual semantics at the feature-level, MID employs modality-adaptive convolution decomposition to disassemble a regular convolution layer into modality-specific basis layers and a modality-shared coefficient layer. Extensive experimental results on two challenging benchmarks demonstrate superior performance of MID over state-of-the-art methods. Zhipeng Huang 0014, Jiawei Liu 0001, Liang Li 0003, Kecheng Zheng, Zhengjun Zha |
AAAI | 2 |
| 2022 | Debiased Batch Normalization via Gaussian Process for Generalizable Person Re-identificationabstractGeneralizable person re-identification aims to learn a model with only several labeled source domains that can perform well on unseen domains. Without access to the unseen domain, the feature statistics of the batch normalization (BN) layer learned from a limited number of source domains is doubtlessly biased for unseen domain. This would mislead the feature representation learning for unseen domain and deteriorate the generalizaiton ability of the model. In this paper, we propose a novel Debiased Batch Normalization via Gaussian Process approach (GDNorm) for generalizable person re-identification, which models the feature statistic estimation from BN layers as a dynamically self-refining Gaussian process to alleviate the bias to unseen domain for improving the generalization. Specifically, we establish a lightweight model with multiple set of domain-specific BN layers to capture the discriminability of individual source domain, and learn the corresponding parameters of the domain-specific BN layers. These parameters of different source domains are employed to deduce a Gaussian process. We randomly sample several paths from this Gaussian process served as the BN estimations of potential new domains outside of existing source domains, which can further optimize these learned parameters from source domains, and estimate more accurate Gaussian process by them in return, tending to real data distribution. Even without a large number of source domains, GDNorm can still provide debiased BN estimation by using the mean path of the Gaussian process, while maintaining low computational cost during testing. Extensive experiments demonstrate that our GDNorm effectively improves the generalization ability of the model on unseen domain. Jiawei Liu 0001, Zhipeng Huang 0014, Liang Li 0003, Kecheng Zheng, Zhengjun Zha |
AAAI | 1 |
| 2022 | Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-IdentificationabstractImage-to-video person re-identification aims to retrieve the same pedestrian as the image-based query from a video-based gallery set. Existing methods treat it as a cross-modality retrieval task and learn the common latent embeddings from image and video modalities, which are both less effective and efficient due to large modality gap and redundant feature learning by utilizing all video frames. In this work, we first regard this task as point-to-set matching problem identical to human decision process, and propose a novel Temporal Complementarity-Guided Reinforcement Learning (TCRL) approach for image-to-video person re-identification. TCRL employs deep reinforcement learning to make sequential judgments on dynamically selecting suitable amount of frames from gallery videos, and accumulate adequate temporal complementary information among these frames by the guidance of the query image, towards balancing efficiency and accuracy. Specifically, TCRL formulates point-to-set matching procedure as Markov decision process, where a sequential judgement agent measures the uncertainty between the query image and all historical frames at each time step, and verifies that sufficient complementary clues are accumulated for judgment (same or different) or one more frames are requested to assist judgment. Moreover, TCRL maintains a sequential feature extraction module with complementary residual detectors to dynamically suppress redundant salient regions and thoroughly mine diverse complementary clues among these selected frames for enhancing frame-level representation. Extensive experiments demonstrate the superiority of our method. Jiawei Liu 0001, Kecheng Zheng, Qibin Sun, Zhengjun Zha |
CVPR | 2 |
| 2022 | JPEG Compression-aware Image Forgery LocalizationabstractImage forgery localization, which aims to find suspicious regions tampered with splicing, copy-move or removal manipulations, has attracted increasing attention. Existing image forgery localization methods have made great progress on public datasets. However, these methods suffer a severe performance drop when the forged images are JPEG compressed, which is widely applied in social media transmission. To tackle this issue, we propose a wavelet-based compression representation learning scheme for the specific JPEG-resistant image forgery localization. Specifically, to improve the performance against JPEG compression, we first learn the abstract representations to distinguish various compression levels through wavelet integrated contrastive learning strategy. Then, based on the learned representations, we introduce a JPEG compression-aware image forgery localization network to flexibly handle forged images compressed with various JPEG quality factors. Moreover, a boundary correction branch is designed to alleviate the edge artifacts caused by JPEG compression. Extensive experiments demonstrate the superiority of our method to existing state-of-the-art approaches, not only on standard datasets, but also on the JPEG forged images with multiple compression quality factors. Menglu Wang 0003, Xueyang Fu, Jiawei Liu 0001, Zhengjun Zha |
ACM Multimedia | 3 |
| 2022 | Evaluating Effects of Background Stories on Graph PerceptionabstractA graph is an abstract model that represents relations among entities, for example, the interactions between characters in a novel. A background story endows entities and relations with real-world meanings and describes the semantics and context of the abstract model, for example, the actual story that the novel presents. Considering practical experience and prior research, human viewers who are familiar with the background story of a graph and those who do not know the background story may perceive the same graph differently. However, no previous research has adequately addressed this problem. This research article thus presents an evaluation that investigated the effects of background stories on graph perception. Three hypotheses that focused on the role of visual focus areas, graph structure identification, and mental model formation on graph perception were formulated and guided three controlled experiments that evaluated the hypotheses using real-world graphs with background stories. An analysis of the resulting experimental data, which compared the performance of participants who read and did not read the background stories, obtained a set of instructive findings. First, having knowledge about a graph's background story influences participants' focus areas during interactive graph explorations. Second, such knowledge significantly affects one's ability to identify community structures but not high degree and bridge structures. Third, this knowledge influences graph recognition under blurred visual conditions. These findings can bring new considerations to the design of storytelling visualizations and interactive graph explorations. Ying Zhao 0001, Jingcheng Shi, Jiawei Liu 0001, Jian Zhao 0010, Wenzhi Zhang, Kangyi Chen, Xin Zhao 0025, Chunyao Zhu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in VideosabstractVideo-based person re-identification aims to match pedestrians from video sequences across non-overlapping camera views. The key factor for video person re-identification is to effectively exploit both spatial and temporal clues from video sequences. In this work, we propose a novel Spatial-Temporal Correlation and Topology Learning framework (CTL) to pursue discriminative and robust representation by modeling cross-scale spatial-temporal correlation. Specifically, CTL utilizes a CNN backbone and a key-points estimator to extract semantic local features from human body at multiple granularities as graph nodes. It explores a context-reinforced topology to construct multi-scale graphs by considering both global contextual information and physical connections of human body. Moreover, a 3D graph convolution and a cross-scale graph convolution are designed, which facilitate direct cross-spacetime and cross-scale information propagation for capturing hierarchical spatial-temporal dependencies and structural information. By jointly performing the two convolutions, CTL effectively mines comprehensive clues that are complementary with appearance information to enhance representational capacity. Extensive experiments on two video benchmarks have demonstrated the effectiveness of the proposed method and the state-of-the-art performance. Jiawei Liu 0001, Zhengjun Zha, Kecheng Zheng, Qibin Sun |
CVPR | 1 |
| 2021 | Adversarial Disentanglement and Correlation Network for Rgb-Infrared Person Re-IdentificationabstractRGB-infrared person re-identification is a challenging task for intelligent video surveillance. Compared to traditional person re-identification, it concerns the additional modality discrepancy between RGB and infrared images originated from the different imaging processes of spectrum cameras, as well as the pedestrian’s appearance discrepancy. In this work, we propose a novel Adversarial Disentanglement and Correlation Network (ADCNet) towards learning modality-invariant and discriminative representations of pedestrians for RGB-infrared person re-identification. ADCNet consists of a feature disentanglement network and a feature alignment network. The feature disentanglement network is designed with an auto-encoder and an adversarial learning module to unify the representations for images across modalities, and the feature alignment network is developed with multiple second-order correlation blocks to employ the second-order non-local position-wise operations for refining the representations and shrinking the intra-modality variations. Extensive experimental results on two challenging benchmarks have demonstrated the effectiveness of the proposed method. Bingyu Hu, Jiawei Liu 0001, Zhengjun Zha |
ICME | 2 |
| 2021 | MM21 Pre-training for Video Understanding Challenge: Video Captioning with Pretraining TechniquesabstractThe quality of video representation directly decides the performance of video related tasks, for both understanding and generation. In this paper, we propose single-modality pretrained feature fusion technique which is composed of reasonable multi-view feature extraction method and designed multi-modality feature fusion strategy. We conduct comprehensive ablation studies on MSR-VTT dataset to demonstrate the effectiveness of proposed method and it surpasses the state-of-the-art methods on both MSR-VTT and VATEX datasets. We further propose the multi-modality pretrained model finetuning technique and dataset augmentation scheme to improve the model's generalization capability. Based on these two proposed pretraining techniques and dataset augmentation scheme, we win the first place in the video captioning track of the MM21 pretraining for video understanding challenge. Dongze Hao, Jiawei Liu 0001, Zijia Zhao, Longteng Guo, Jing Liu 0001 |
ACM Multimedia | 5 |
| 2021 | Cluster and Scatter: A Multi-grained Active Semi-supervised Learning Framework for Scalable Person Re-identificationabstractActive learning has recently attracted increasing attention in the task of person re-identification, due to its unique scalability that not only maximally reduces the annotation cost but also retains the satisfying performance. Although some preliminary active learning methods have been explored in scalable person re-identification task, they have the following two problems: 1) the inefficiency in the selection process of image pairs due to the huge search space, and 2) the ineffectiveness caused by ignoring the impact of unlabeled data in model training. Considering that, we propose a Multi-grained Active Semi-Supervised learning framework, named MASS, to address the scalable person re-identification problem existing in the practical scenarios. Specifically, we firstly design a cluster-scatter procedure to alleviate the inefficiency problem, which consists of two components: cluster step and scatter step. The cluster step shrinks the search space into individual small clusters by a coarse-grained clustering method, and the subsequent scatter step further mines the hard distinguished image pairs from unlabelled set to purify the learned clusters by a novel centrality-based adaptive purification strategy. Afterward, we introduce a customized purification loss for the purified clustering, which utilizes the complementary information in both labeled and unlabeled data to optimize the model for solving the ineffectiveness problem. The cluster-scatter procedure and the model optimization are performed in an iterative fashion to achieve the promising performance while greatly reducing the annotation cost. Extensive experimental results have demonstrated that MASS can even achieve a competitive performance with fully supervised methods in the case of extremely less annotation requirements. Bingyu Hu, Zhengjun Zha, Jiawei Liu 0001, Xierong Zhu, Hongtao Xie 0001 |
ACM Multimedia | 3 |
| 2021 | Pose-Guided Feature Learning with Knowledge Distillation for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) aims to match person images with occlusion. It is fundamentally challenging because of the serious occlusion which aggravates the misalignment problem between images. At the cost of incorporating a pose estimator, many works introduce pose information to alleviate the misalignment in both training and testing. To achieve high accuracy while preserving low inference complexity, we propose a network named Pose-Guided Feature Learning with Knowledge Distillation (PGFL-KD), where the pose information is exploited to regularize the learning of semantics aligned features but is discarded in testing. PGFL-KD consists of a main branch (MB), and two pose-guided branches, e.g., a foreground-enhanced branch (FEB), and a body part semantics aligned branch (SAB). The FEB intends to emphasise the features of visible body parts while excluding the interference of obstructions and background (e.g., foreground feature alignment). The SAB encourages different channel groups to focus on different body parts to have body part semantics aligned representation. To get rid of the dependency on pose information when testing, we regularize the MB to learn the merits of the FEB and SAB through knowledge distillation and interaction-based training. Extensive experiments on occluded, partial, and holistic ReID tasks show the effectiveness of our proposed network. Kecheng Zheng, Cuiling Lan, Wenjun Zeng 0001, Jiawei Liu 0001, Zhizheng Zhang 0004, Zhengjun Zha |
ACM Multimedia | 4 |
| 2020 | Multi-Scale Spatial-Temporal Integration Convolutional Tube for Human Action RecognitionabstractApplying multi-scale representations leads to consistent performance improvements on a wide range of image recognition tasks. However, with the addition of the temporal dimension in video domain, directly obtaining layer-wise multi-scale spatial-temporal features will add a lot extra computational cost. In this work, we propose a novel and efficient Multi-Scale Spatial-Temporal Integration Convolutional Tube (MSTI) aiming at achieving accurate recognition of actions with lower computational cost. It firstly extracts multi-scale spatial and temporal features through the multi-scale convolution block. Considering the interaction of different-scales representations and the interaction of spatial appearance and temporal motion, we employ the cross-scale attention weighted blocks to perform feature recalibration by integrating multi-scale spatial and temporal features. An end-to-end deep network, MSTI-Net, is also presented based on the proposed MSTI tube for human action recognition. Extensive experimental results show that our MSTI-Net significantly boosts the performance of existing convolution networks and achieves state-of-the-art accuracy on three challenging benchmarks, i.e., UCF-101, HMDB-51 and Kinetics-400, with much fewer parameters and FLOPs. Haoze Wu 0003, Jiawei Liu 0001, Xierong Zhu, Meng Wang 0001, Zhengjun Zha |
IJCAI | 2 |
| 2020 | Co-Saliency Spatio-Temporal Interaction Network for Person Re-Identification in VideosabstractPerson re-identification aims at identifying a certain pedestrian across non-overlapping camera networks. Video-based person re-identification approaches have gained significant attention recently, expanding image-based approaches by learning features from multiple frames. In this work, we propose a novel Co-Saliency Spatio-Temporal Interaction Network (CSTNet) for person re-identification in videos. It captures the common salient foreground regions among video frames and explores the spatial-temporal long-range context interdependency from such regions, towards learning discriminative pedestrian representation. Specifically, multiple co-saliency learning modules within CSTNet are designed to utilize the correlated information across video frames to extract the salient features from the task-relevant regions and suppress background interference. Moreover, multiple spatial-temporal interaction modules within CSTNet are proposed, which exploit the spatial and temporal long-range context interdependencies on such features and spatial-temporal information correlation, to enhance feature representation. Extensive experiments on two benchmarks have demonstrated the effectiveness of the proposed method. Jiawei Liu 0001, Zhengjun Zha, Xierong Zhu |
IJCAI | 1 |
| 2020 | Dual Context-Aware Refinement Network for Person SearchabstractPerson search has recently gained increasing attention as the novel task of localizing and identifying a target pedestrian from a gallery of non-cropped scene images. Its performance depends on accurate person detection and re-identification simultaneously by learning effective representations. In this work, we propose a novel dual context-aware refinement network (DCRNet) for person search, which jointly explores two kinds of contexts including intra-instance context and inter-instance context to learn discriminative representation. Specifically, an intra-instance context module is designed to refine the representation for the bounding box of a pedestrian by leveraging its surrounding regions covering the same pedestrian and its accessories, which contain abundant complementary visual appearance of pedestrians. Moreover, an inter-instance context module is proposed to expand the instance-level feature for the bounding box of a pedestrian, by utilizing the rich scene contexts of neighboring co-travelers across images. These two modules are built on top of a joint detection and feature learning framework, i.e., Faster R-CNN. Extensive experimental results on two challenging datasets have demonstrated the effectiveness of DCRNet with significant performance improvements over state-of-the-art methods. Jiawei Liu 0001, Zhengjun Zha, Richang Hong, Meng Wang 0001, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2020 | Hierarchical Gumbel Attention Network for Text-based Person SearchabstractText-based person search aims to retrieve the pedestrian images that best match a given textual description from gallery images. Previous methods utilize the soft-attention mechanism to infer the semantic alignments between the regions of image and the corresponding words in sentence. However, these methods may fuse the irrelevant multi-modality features together which cause matching redundancy problem. In this work, we propose a novel hierarchical Gumbel attention network for text-based person search via Gumbel top-k re-parameterization algorithm. Specifically, it adaptively selects the strong semantically relevant image regions and words/phrases from images and texts for precise alignment and similarity calculation. This hard selection strategy is able to fuse the strong-relevant multi-modality features for alleviating the problem of matching redundancy. Meanwhile, a Gumbel top-k re-parameterization algorithm is designed as a low-variance, unbiased gradient estimator to handle the discreteness problem of hard attention mechanism by an end-to-end manner. Moreover, a hierarchical adaptive matching strategy is employed by the model from three different granularities, i.e., word-level, phrase-level, and sentence-level, towards fine-grained matching. Extensive experimental results demonstrate the state-of-the-art performance. Compared the existed best method, we achieve the 8.24% Rank-1 and 7.6% mAP relative improvements in the text-to-image retrieval task, and 5.58% Rank-1 and 6.3% mAP relative improvements in the image-to-text retrieval task on CUHK-PEDES dataset, respectively. Kecheng Zheng, Wu Liu 0005, Jiawei Liu 0001, Zhengjun Zha, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2020 | ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosabstractThe attention mechanism has been widely applied to enhance pedestrian representation for person re-identification in videos. However, most existing methods learn the spatial and temporal attention separately, and thus ignore the correlation between them. In this work, we propose a novel Adaptive Spatio-Temporal Attention Network (ASTA-Net) to adaptively aggregate the spatial and temporal attention features into discriminative pedestrian representation for person re-identification in videos. Specifically, multiple Adaptive Spatio-Temporal Fusion modules within ASTA-Net are designed for exploring precise spatio-temporal attention on multi-level feature maps. They first obtain the preliminary spatial and temporal attention features via the spatial semantic relations for each frame and temporal dependencies among inconsecutive frames, then adaptively aggregate the preliminary attention features on the basis of their correlation. Moreover, an Adjacent-Frame Motion module is designed to explicitly extract motion patterns according to the feature-level variation among adjacent frames. Extensive experiments on the three widely-used datasets, i.e., MARS, iLIDS-VID and PRID2011, have demonstrated the effectiveness of the proposed approach. Xierong Zhu, Jiawei Liu 0001, Haoze Wu 0003, Meng Wang 0001, Zhengjun Zha |
ACM Multimedia | 2 |
| 2020 | A Structured Graph Attention Network for Vehicle Re-IdentificationabstractVehicle re-identification aims to identify the same vehicle across different surveillance cameras and plays an important role in public security. Existing approaches mainly focus on exploring informative regions or learning an appropriate distance metric. However, they not only neglect the inherent structured relationship between discriminative regions within an image, but also ignore the extrinsic structured relationship among images. The inherent and extrinsic structured relationships are crucial to learning effective vehicle representation. In this paper, we propose a Structured Graph ATtention network (SGAT) to fully exploit these relationships and allow the message propagation to update the features of graph nodes. SGAT creates two graphs for one probe image. One is an inherent structured graph based on the geometric relationship between the landmarks that can use features of their neighbors to enhance themselves. The other is an extrinsic structured graph guided by the attribute similarity to update image representations. Experimental results on two public vehicle re-identification datasets including VeRi-776 and VehicleID have shown that our proposed method achieves significant improvements over the state-of-the-art methods. Yangchun Zhu, Zhengjun Zha, Tianzhu Zhang 0001, Jiawei Liu 0001, Jiebo Luo 0001 |
ACM Multimedia | 4 |
| 2020 | Adversarial Attribute-Text Embedding for Person Search With Natural Language QueryabstractThe newly emerging task of person search with natural language query aims at retrieving the target pedestrian by a text description of the pedestrian. It is more applicable compared to person search with image/video query, i.e., person re-identification. In this paper, we propose a novel Adversarial Attribute-Text Embedding (AATE) network for person search with text query. In particular, a cross-modal adversarial learning module is proposed to learn discriminative and modality-invariant visual-textual features. It consists of a cross-modal learner and a modality discriminator, playing a min-max game in an adversarial learning way. The former is to improve intra-modality discrimination and inter-modality invariance towards confusing the modality discriminator. The latter is to distinguish the features from different modalities and boost the learning of modality-invariant features. Moreover, a visual attribute graph convolutional network is proposed to learn visual attributes of pedestrians, which possess better descriptiveness, interpretability and robustness compared to pedestrian appearance features. A hierarchical text embedding network, consisting of multi-stacked bidirectional LSTMs and a textual attention block, is developed to extract effective textual features from text descriptions of pedestrians. Extensive experimental results on two challenging benchmarks, have demonstrated the effectiveness of the proposed approach. Zhengjun Zha, Jiawei Liu 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Adaptive Transfer Network for Cross-Domain Person Re-IdentificationabstractRecent deep learning based person re-identification approaches have steadily improved the performance for benchmarks, however they often fail to generalize well from one domain to another. In this work, we propose a novel adaptive transfer network (ATNet) for effective cross-domain person re-identification. ATNet looks into the essential causes of domain gap and addresses it following the principle of "divide-and-conquer". It decomposes the complicated cross-domain transfer into a set of factor-wise sub-transfers, each of which concentrates on style transfer with respect to a certain imaging factor, e.g., illumination, resolution and camera view etc. An adaptive ensemble strategy is proposed to fuse factor-wise transfers by perceiving the affect magnitudes of various factors on images. Such "decomposition-and-ensemble" strategy gives ATNet the capability of precise style transfer at factor level and eventually effective transfer across domains. In particular, ATNet consists of a transfer network composed by multiple factor-wise CycleGANs and an ensemble CycleGAN as well as a selection network that infers the affects of different factors on transferring each image. Extensive experimental results on three widely-used datasets, i.e., Market-1501, DukeMTMC-reID and PRID2011 have demonstrated the effectiveness of the proposed ATNet with significant performance improvements over state-of-the-art methods. Jiawei Liu 0001, Zhengjun Zha, Richang Hong, Meng Wang 0001 |
CVPR | 1 |
| 2019 | Mutually Reinforced Spatio-Temporal Convolutional Tube for Human Action RecognitionabstractRecent works use 3D convolutional neural networks to explore spatio-temporal information for human action recognition. However, they either ignore the correlation between spatial and temporal features or suffer from high computational cost by spatio-temporal features extraction. In this work, we propose a novel and efficient Mutually Reinforced Spatio-Temporal Convolutional Tube (MRST) for human action recognition. It decomposes 3D inputs into spatial and temporal representations, mutually enhances both of them by exploiting the interaction of spatial and temporal information and selectively emphasizes informative spatial appearance and temporal motion, meanwhile reducing the complexity of structure. Moreover, we design three types of MRSTs according to the different order of spatial and temporal information enhancement, each of which contains a spatio-temporal decomposition unit, a mutually reinforced unit and a spatio-temporal fusion unit. An end-to-end deep network, MRST-Net, is also proposed based on the MRSTs to better explore spatio-temporal information in human actions. Extensive experiments show MRST-Net yields the best performance, compared to state-of-the-art approaches. Haoze Wu 0003, Jiawei Liu 0001, Zhengjun Zha, Zhenzhong Chen 0001, Xiaoyan Sun 0001 |
IJCAI | 2 |
| 2019 | Deep Adversarial Graph Attention Convolution Network for Text-Based Person SearchabstractThe newly emerging text-based person search task aims at retrieving the target pedestrian by a query in natural language with fine-grained description of a pedestrian. It is more applicable in reality without the requirement of image/video query of a pedestrian, as compared to image/video based person search, i.e., person re-identification. In this work, we propose a novel deep adversarial graph attention convolution network (A-GANet) for text-based person search. The A-GANet exploits both textual and visual scene graphs, consisting of object properties and relationships, from the text queries and gallery images of pedestrians, towards learning informative textual and visual representations. It learns an effective joint textual-visual latent feature space in adversarial learning manner, bridging modality gap and facilitating pedestrian matching. Specifically, the A-GANet consists of an image graph attention network, a text graph attention network and an adversarial learning module. The image and text graph attention networks are designed with a novel graph attention convolution layer, which effectively exploits graph structure in the learning of textual and visual features, leading to precise and discriminative representations. An adversarial learning module is developed with a feature transformer and a modality discriminator, to learn a joint textual-visual feature space for cross-modality matching. Extensive experimental results on two challenging benchmarks, i.e., CUHK-PEDES and Flickr30k datasets, have demonstrated the effectiveness of the proposed method. Jiawei Liu 0001, Zhengjun Zha, Richang Hong, Meng Wang 0001, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2019 | Adaptive Alignment Network for Person Re-identification
Xierong Zhu, Jiawei Liu 0001, Hongtao Xie 0001, Zhengjun Zha |
MMM (2) | 2 |
| 2019 | Dense 3D-Convolutional Neural Network for Person Re-Identification in VideosabstractPerson re-identification aims at identifying a certain pedestrian across non-overlapping multi-camera networks in different time and places. Existing person re-identification approaches mainly focus on matching pedestrians on images; however, little attention has been paid to re-identify pedestrians in videos. Compared to images, video clips contain motion patterns of pedestrians, which is crucial to person re-identification. Moreover, consecutive video frames present pedestrian appearance with different body poses and from different viewpoints, providing valuable information toward addressing the challenge of pose variation, occlusion, and viewpoint change, and so on. In this article, we propose a Dense 3D-Convolutional Network (D3DNet) to jointly learn spatio-temporal and appearance representation for person re-identification in videos. The D3DNet consists of multiple three-dimensional (3D) dense blocks and transition layers. The 3D dense blocks enlarge the receptive fields of visual neurons in both spatial and temporal dimensions, leading to discriminative appearance representation as well as short-term and long-term motion patterns of pedestrians without the requirement of an additional motion estimation module. Moreover, we formulate a loss function consisting of an identification loss and a center loss to minimize intra-class variance and maximize inter-class variance simultaneously, toward addressing the challenge of large intra-class variance and small inter-class variance. Extensive experiments on two real-world video datasets of person identification, i.e., MARS and iLIDS-VID, have shown the effectiveness of the proposed approach. Jiawei Liu 0001, Zhengjun Zha, Xuejin Chen, Zilei Wang, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Spatiotemporal-Textual Co-Attention Network for Video Question AnsweringabstractVisual Question Answering (VQA) is to provide a natural language answer for a pair of an image or video and a natural language question. Despite recent progress on VQA, existing works primarily focus on image question answering and are suboptimal for video question answering. This article presents a novel Spatiotemporal-Textual Co-Attention Network (STCA-Net) for video question answering. The STCA-Net jointly learns spatially and temporally visual attention on videos as well as textual attention on questions. It concentrates on the essential cues in both visual and textual spaces for answering question, leading to effective question-video representation. In particular, a question-guided attention network is designed to learn question-aware video representation with a spatial-temporal attention module. It concentrates the network on regions of interest within the frames of interest across the entire video. A video-guided attention network is proposed to learn video-aware question representation with a textual attention module, leading to fine-grained understanding of question. The learned video and question representations are used by an answer predictor to generate answers. Extensive experiments on two challenging datasets of video question answering, i.e., MSVD-QA and MSRVTT-QA, have shown the effectiveness of the proposed approach. Zhengjun Zha, Jiawei Liu 0001, Tianhao Yang, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | CA3Net: Contextual-Attentional Attribute-Appearance Network for Person Re-IdentificationabstractPerson re-identification aims to identify the same pedestrian across non-overlapping camera views. Deep learning techniques have been applied for person re-identification recently, towards learning representation of pedestrian appearance. This paper presents a novel Contextual-Attentional Attribute-Appearance Network ($\rm CA^3Net$) for person re-identification. The $\rm CA^3Net$ simultaneously exploits the complementarity between semantic attributes and visual appearance, the semantic context among attributes, visual attention on attributes as well as spatial dependencies among body parts, leading to discriminative and robust pedestrian representation. Specifically, an attribute network within $\rm CA^3Net$ is designed with an Attention-LSTM module. It concentrates the network on latent image regions related to each attribute as well as exploits the semantic context among attributes by a LSTM module. An appearance network is developed to learn appearance features from the full body, horizontal and vertical body parts of pedestrians with spatial dependencies among body parts. The $\rm CA^3Net$ jointly learns the attribute and appearance features in a multi-task learning manner, generating comprehensive representation of pedestrians. Extensive experiments on two challenging benchmarks, i.e., Market-1501 and DukeMTMC-reID datasets, have demonstrated the effectiveness of the proposed approach. Jiawei Liu 0001, Zhengjun Zha, Hongtao Xie 0001, Zhiwei Xiong, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2016 | Multi-Scale Triplet CNN for Person Re-IdentificationabstractPerson re-identification aims at identifying a certain person across non-overlapping multi-camera networks. It is a fundamental and challenging task in automated video surveillance. Most existing researches mainly rely on hand-crafted features, resulting in unsatisfactory performance. In this paper, we propose a multi-scale triplet convolutional neural network which captures visual appearance of a person at various scales. We propose to optimize the network parameters by a comparative similarity loss on massive sample triplets, addressing the problem of small training set in person re-identification. In particular, we design a unified multi-scale network architecture consisting of both deep and shallow neural networks, towards learning robust and effective features for person re-identification under complex conditions. Extensive evaluation on the real-world Market-1501 dataset have demonstrated the effectiveness of the proposed approach. Jiawei Liu 0001, Zhengjun Zha, Q. I. Tian, Dong Liu 0002, Ting Yao 0003, Qiang Ling 0001, Tao Mei 0001 |
ACM Multimedia | 1 |