VLDB 2026 Research / reviewers in the wild / expert
Qi Chu 0001
dblp:52/9077-1
· DBLP profile ↗
83ranked-venue papers
2as first author
74since 2021 · last 2026
0000-0003-3028-0755ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 2 first-author · 59 since 2021Artificial intelligence and machine learning · 37 · 2 first-author · 31 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flora: Effortless Context Construction to Arbitrary Length and ScaleabstractEffectively handling long contexts is challenging for Large Language Models (LLMs) due to the rarity of long texts, high computational demands, and substantial forgetting of short-context abilities. Recent approaches have attempted to construct long contexts for instruction tuning, but these methods often require LLMs or human interventions, which are both costly and limited in length and diversity. Also, the drop in short-context performances of present long-context LLMs remains significant. In this paper, we introduce Flora, an effortless (human/LLM-free) long-context construction strategy. Flora can markedly enhance the long-context performance of LLMs by arbitrarily assembling short instructions based on categories and instructing LLMs to generate responses based on long-context meta-instructions. This enables Flora to produce contexts of arbitrary length and scale with rich diversity, while only slightly compromising short-context performance. Experiments on Llama3-8B-Instruct and QwQ-32B show that LLMs enhanced by Flora excel in three long-context benchmarks while maintaining strong performances in short-context tasks. Zhentao Tan, Xiaofan Bo, Qi Chu 0001, Jieping Ye |
AAAI | 6 |
| 2026 | MagicPaint: Operate Anything for Image Inpainting with Diffusion ModelabstractRecent diffusion-based models have significantly improved inpainting quality. However, existing methods struggle with multi-task inpainting due to conflicting optimization objectives, and current datasets are typically limited to task-specific scenarios, hindering joint training. To address these challenges, we propose MagicPaint, a unified diffusion-based inpainting model that supports object addition, removal, and unconditional inpainting across both text and image modalities. MagicPaint semantically decouples operation types and target content by learnable tokens in MMToken Module, effectively reconciling conflicting optimization objectives and enabling robust multi-task, multi-modal inpainting. Besides, a novel inpainting paradigm named MagicMask, encodes operating intent directly into the mask and applies a mask loss for spatially precise supervision. In addition, existing inpainting datasets are insufficient for multi-task and multi-modal scenarios, limiting the capability of inpainting models. Thus, we further introduce a new dataset comprising 2.1M image tuples. It is dedicatedly designed to support diverse inpainting scenarios and significantly improves upon existing datasets, particularly in object removal. Through efforts from both model and data perspectives, MagicPaint enables users to operate anything—add, remove or inpaint content which is specified through either text or image modalities in a seamless and unified manner. Extensive experiments demonstrate that MagicPaint achieves state-of-the-art performance across three key tasks (i.e., text-guided addition, image-guided addition, and object removal) and produces outputs with superior visual consistency and contextual fidelity compared to existing methods. Qinhong Yang, Dongdong Chen 0001, Qi Chu 0001, Qiankun Liu 0001, Zhentao Tan, Xulin Li, Huamin Feng, Nenghai Yu |
AAAI | 3 |
| 2026 | When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use BehaviorsabstractChenghao Yang, Yuning Zhang, Zhoufutu Wen, Tao Gong, Jiaheng Liu, Qi Chu, Nenghai Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhoufutu Wen, Qi Chu 0001, Nenghai Yu |
ACL (1) | 6 |
| 2026 | PVDI: Preserving Vital and Disrupting Irrelevant Latent Attentions for Robust Backdoor Defense
Junchi Chen, Qi Chu 0001, Nenghai Yu, Dongmei Zhang 0001, Bin B. Zhu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | ABDP: Adversarial Backdoor Detection and PurificationabstractIn domains like driving and healthcare, deep learning models often rely on large, diverse datasets that can inadvertently harbor backdoor attacks. In this paper, we propose ABDP, a novel post-processing defense method, to effectively remove backdoor contamination from datasets and generate clean models without relying on any pre-existing clean data. ABDP capitalizes on the intrinsic link between untargeted adversarial attacks and backdoor attacks to detect the presence of backdoor attacks within trained models and ascertain their target labels. Subsequently, it trains a clean model capable of recognizing all labels except the target label, thus treating poisoned data as in-distribution and clean data of the target label as out-of-distribution. This distinction enables the identification of backdoor poisoned data. Finally, ABDP applies unlearning techniques to effectively eradicate the backdoor from the model. Extensive experimental evaluations across diverse datasets and against multiple backdoor attack scenarios validate the robustness and state-of-the-art performance of our approach. By employing ABDP for data and model cleansing, the attack success rate of resulting models is reduced to 1% or less, while retaining approximately 70% or more clean data at a true positive rate of 0.01 false positive rate. Notably, ABDP exhibits no adverse impact when applied to purely clean datasets, owing to its ability to detect backdoor presence in models before cleansing. Thus, our proposed method achieves state-of-the-art performance in cleansing both backdoor-poisoned data and backdoor models. The code for ABDP will be made available upon publication of the paper. Bin B. Zhu, Qi Chu 0001, Nenghai Yu, Dongmei Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningabstractIn 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. For example, MAMP shows that instead of following the prevalent masked joint reconstruction, explicit masked motion reconstruction is key to the success of learning effective feature representation for 3D action recognition. However, we find that if we make a simple and effective change to the reconstructed target of masked joint reconstruction, masked joint reconstruction can achieve the same results as masked motion reconstruction. The devil is in the special characteristic of 3D skeleton data and the normalization process of training targets. We need to dig for all effective information of targets during normalization. Besides, considering that mask data reconstruction focuses more on learning local relations in input data for fulfilling the reconstruction task, instead of modeling the relation among samples, we further employ contrastive learning to learn more discriminative 3D action representations. We show that contrastive learning can consistently boost the performance of model pre-trained by masked joint prediction under various settings, especially in the semi-supervised setting that has a very limited number of labeled samples. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed pre-training strategy achieves state-of-the-art results without bells and whistles. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 2 |
| 2025 | Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingabstractOpen-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive image-text pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing semantic prototypes in the construction stage and image-to-image matching (i.e., prototype matching) during testing. However, these methods often struggle to effectively capture the visual characteristics of categories and fail to utilize local features during prototype matching. To deal with these problems, we propose a novel training-free framework for OVSS that constructs diverse prototypes and performs fine-grained sub-region matching. Specifically, our method leverages Large Language Models (LLMs) to guide support image generation by descriptions of different attributes of categories and employs coarse-fine clustering to obtain diverse and robust part-level prototypes in the construction stage. During testing, we propose a sub-region matching method, which assigns part-level prototypes to sub-regions utilizing optimal transport, to fully utilize local image features among part-level prototypes. Extensive experiments demonstrate the effectiveness of our method and show that our method achieves state-of-the-art performance, outperforming previous methods across five datasets. Xuanpu Zhao, Dianmo Sheng, Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 6 |
| 2025 | UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype DiscoveryabstractRecent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which often exhibit quality issues resulting from various sources and inaccurate labeling. How to extract more robust representations from these examples has always been one of the goals of in-context visual learning. In response, we propose UNICL-SAM, to better model the example distribution and extract robust representations to help in-context segmentation. We incorporate an uncertainty probabilistic module to quantify each example’s reliability during both the training and testing phases. Utilizing this uncertainty estimation, we introduce an uncertainty-guided graph augmentation and feature refinement strategy, aimed at mitigating the impact of high-uncertainty regions to enhance the learning of robust representations. Subsequently, we construct prototypes for each example by aggregating part information, thereby creating reliable in-context instruction that effectively represents fine-grained local semantics. This approach serves as a valuable complement to traditional global pooling features. Experimental results demonstrate the effectiveness of the proposed framework, underscoring its potential for real-world applications. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Wenbin Tu, Shengwei Xu, Nenghai Yu |
CVPR | 5 |
| 2025 | Training an Anti-KD Model that Cannot Teach Students via Similarity DisruptionabstractKnowledge Distillation (KD) aims to enhance the performance of student models by transferring knowledge from teacher models. While reaping the benefits of KD, the intellectual property risks associated with it cannot be ignored. Even if models are released without training data or provided as a service, potential adversaries can still clone the target model using KD. To mitigate the risks, some researchers propose training the anti-KD model that cannot teach student models. However, we find existing methods cannot defend against representation-based KD. To address the knowledge leakage from representations, we introduce Similarity Disruption (SD). SD increases the distance between the representation similarity matrices of our anti-KD model and the normal model, thereby reducing the effective information in the representation space. Extensive experiments demonstrate the proposed method can effectively defend against representation-based KD. Qi Chu 0001, Bin Liu 0016, Quanchen Zou, Deyue Zhang, Nenghai Yu |
ICASSP | 2 |
| 2025 | CMGait: Enhancing Cross-Modality Gait Recognition between LiDAR and RGB through Contrastive Identity-consistent Feature AggregationabstractCombination usage of LiDAR and RGB cameras for gait recognition can achieve cross space recognition and privacy protection. In addition, the widespread application of LiDAR cameras with 3D geometry information and the large amount of RGB gaits has led to the demand for cross-modality gait recognition on LiDAR and RGB modalities. To address the challenge of cross-modality recognition, we proposed a novel cross-modality gait recognition paradigm called CMGait. The key innovations include a novel projection method for transforming LiDAR point clouds into depth maps, feature alignment modules, Transformer-based identity encoders, and an embedding distance fusion method with similarity matrices based contrastive learning. Experimental results showcase state-of-the-art performance with Rank-1 accuracy of 62.8% and 66.1% for different directions in cross-modality gait recognition. Ablation experiments validate the effectiveness of the proposed methods, highlighting advancements in feature alignment and modality fusion techniques. Yubo Wang 0011, Bin Liu 0016, Jixiang Niu, Qi Chu 0001, Nenghai Yu |
ICASSP | 5 |
| 2025 | FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and Segmentation
Qi Chu 0001, Bin Liu 0016, Wei Zhou 0021, Nenghai Yu |
ICCV | 2 |
| 2025 | Exploiting Feature Gating and Injection For Multi-modal Manipulation Detection and Grounding
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Nenghai Yu |
ICIG (3) | 7 |
| 2025 | Exploring Generalized Features For LLM-Generated Text Detection
Jiazhen Wang, Bin Liu 0016, Changtao Miao, Qi Chu 0001, Quanchen Zou, Deyue Zhang, Nenghai Yu |
ICIG (3) | 6 |
| 2025 | Multimodal Consistency-Driven Deepfake Detection
Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICIG (2) | 3 |
| 2025 | Remote Sensing Target Detector with Multi Scale Attention MechanismabstractMost of the existing rotation detection models focus on solving problems such as feature misalignment and boundary discontinuity, but ignore the use of contextual information in remote sensing images. However, context information plays a vital role in the accurate detection of small instance targets. Especially when the receptive field of the model is limited, it is very easy to cause deviations in the detection results. Therefore, we propose a Remote Sensing Target Detector with Multi-scale Attention Mechanism Network(MAMNet) that gradually enhances the features in the region of interest by utilizing the spatial, local, and global information of the image input features, thereby supplementing the attention mechanism at multiple scales, better identifying targets, and improving the detection performance of the model. Our model was experimented on DOTAv1.0, DOTAv1.5, HRSC2016 and has achieved the state of the art on DOTAv1.0. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICIP | 3 |
| 2025 | Adversarial Examples Detection Based on Adversarial Attack SensitivityabstractDeep neural networks have found widespread application in critical fields but remain vulnerable to adversarial attacks. Existing detection methods aim to achieve defense without modifying the model, but they generally struggle with generalization to unseen attacks. To address this limitation, we investigate the underlying principles of max-loss and min-distance adversarial attacks and uncover a strong positive correlation between perturbation magnitude, prediction confidence, and the distance to the decision boundary. Building on this insight, we introduce Adversarial Detection via Adversarial Sensitivity (ADAS), a novel approach that detects adversarial attacks by analyzing the sensitivity of a model's predictions to perturbation magnitude. ADAS estimates the distance to the decision boundary through sensitivity analysis by simulating adversarial attacks on input samples, identifying anomalies indicative of adversarial manipulation. Extensive experiments demonstrate the robustness and generalizability of ADAS across diverse and previously unseen adversarial attack scenarios, establishing its efficacy as a versatile and reliable detection framework. Cong Ming 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICME | 4 |
| 2025 | Towards Anytime Retrieval: A Benchmark for Anytime Person Re-IdentificationabstractIn real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and only provide training and evaluation for specific scenarios. Therefore, we investigate a new task called Anytime Person Re-identification (AT-ReID), which aims to achieve effective retrieval in multiple scenarios based on variations in time. To address the AT-ReID problem, we collect the first large-scale dataset, AT-USTC, which contains 135k images of individuals wearing multiple clothes captured by RGB and IR cameras. Our data collection spans over an entire year and 270 volunteers were photographed on average 29.1 times across different dates or scenes, 4-15 times more than current datasets, providing conditions for follow-up investigations in AT-ReID. Further, to tackle the new challenge of multi-scenario retrieval, we propose a unified model named Uni-AT, which comprises a multi-scenario ReID (MS-ReID) framework for scenario-specific features learning, a Mixture-of-Attribute-Experts (MoAE) module to alleviate inter-scenario interference, and a Hierarchical Dynamic Weighting (HDW) strategy to ensure balanced training across all scenarios. Extensive experiments show that our model leads to satisfactory results and exhibits excellent generalization to all scenarios. Xulin Li, Yan Lu 0001, Bin Liu 0016, Qinhong Yang, Qi Chu 0001, Mang Ye, Nenghai Yu |
IJCAI | 7 |
| 2025 | Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and LocalizationabstractWith the advancement of face manipulation technology, forgery images in multi-face scenarios are gradually becoming a more complex and realistic challenge. Despite this, detection and localization methods for such multi-face manipulations remain underdeveloped. Traditional manipulation localization methods either indirectly derive detection results from localization masks, resulting in limited detection performance, or employ a naive two-branch structure to simultaneously obtain detection and localization results, which cannot effectively benefit the localization capability due to limited interaction between the two tasks. This paper proposes a new framework, namely MoNFAP, specifically tailored for multi-face manipulation detection and localization. The MoNFAP primarily introduces two novel modules: the Forgery-aware Unified Predictor (FUP) Module and the Mixture-of-Noises Module (MNM). The proposed FUP integrates detection and localization tasks using a token learning strategy and multiple forgery-aware transformers, which facilitates the use of classification information to enhance localization capability. Furthermore, to mitigate the interference from general semantic object information, we propose the MNM that leverages multiple noise extractors based on the mixture of experts concept. This allows the MNM to learn semantic-agnostic forgery features from general RGB features, further boosting the performance of our proposed framework. Finally, we establish a comprehensive benchmark for multi-face detection and localization, and the proposed MoNFAP achieves significant performance. The code is available: https://github.com/miaoct/MoNFAP. Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Honggang Hu, Nenghai Yu |
ACM Multimedia | 2 |
| 2025 | MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World ScenariosabstractRapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (MFFI ) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates 50 different forgery methods and contains 1024K image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at https://github.com/inclusionConf/MFFI. Changtao Miao, Weiwei Feng, Qi Chu 0001, Jianshu Li, Yunfeng Diao, Wei Zhou 0021, Joey Tianyi Zhou, Xiaoshuai Hao |
ACM Multimedia | 6 |
| 2025 | Towards Good Generalizations for Diffusion Generated Image Detection Using Multiple Reconstruction Contrastive LearningabstractA striking proficiency of diffusion models in producing and manipulating images with an unprecedented level of realism has unquestionably elicited concerns. Many methods have been proposed to detect generated images. In particular, recent studies reveal that autoencoder reconstruction error can serve as an effective indicator for distinguishing authentic and synthetic images, since most generative models adopt analogous encoder-decoder operation. However, the reliance on a single autoencoder reconstruction error provides only limited information, which is insufficient for comprehensively capturing discriminative features, resulting in restricted generalization performance. In this paper, we propose Multiple Reconstruction Contrastive Learning (MRCL), which leverages multiple reconstruction residuals to enhance the generalizability of generated image detection. Specifically, MRCL applies Dinov2-ViT with LoRA fine-tuning to extract fine-grained feature representations of origin images and their multiple VAE reconstructions. In addition, a Residual Dense Fusion module is designed to effectively combine multiple VAE reconstruction residuals. Further, a contrastive learning strategy is adopted to guide the distance of origin images and VAE reconstruction representations. Extensive experimental results demonstrate the superior generalization performance of the proposed MRCL. Wanyi Zhuang, Qi Chu 0001, Changtao Miao, Nenghai Yu |
ACM Multimedia | 2 |
| 2025 | GUPNet++: Geometry Uncertainty Propagation Network for Monocular 3D Object DetectionabstractGeometry plays a significant role in monocular 3D object detection. It can be used to estimate object depth by using the perspective projection between object's physical size and 2D projection in the image plane, which can introduce mathematical priors into deep models. However, this projection process also introduces error amplification, where the error of the estimated height is amplified and reflected into the projected depth. It leads to unreliable depth inferences and also impairs training stability. To tackle this problem, we propose a novel Geometry Uncertainty Propagation Network (GUPNet++) by modeling geometry projection in a probabilistic manner. This ensures depth predictions are well-bounded and associated with a reasonable uncertainty. The significance of introducing such geometric uncertainty is two-fold: (1). It models the uncertainty propagation relationship of the geometry projection during training, improving the stability and efficiency of the end-to-end model learning. (2). It can be derived to a highly reliable confidence to indicate the quality of the 3D detection result, enabling more reliable detection inference. Experiments show that the proposed approach not only obtains (state-of-the-art) SOTA performance in image-based monocular 3D detection but also demonstrates superiority in efficacy with a simplified framework. The code and model will be released at https://github.com/SuperMHP/GUPNet_Plus. Yan Lu 0001, Xinzhu Ma, Lei Yang 0045, Tianzhu Zhang 0001, Qi Chu 0001, Tong He 0001, Yonghui Li 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Context-Aware Weakly Supervised Image Manipulation Localization With SAM RefinementabstractMalicious image manipulation poses societal risks, increasing the importance of effective image manipulation detection methods. Recent approaches in image manipulation detection have largely been driven by fully supervised approaches, which require labor-intensive pixel-level annotations. Thus, it is essential to explore weakly supervised image manipulation localization methods that only require image-level binary labels for training. However, existing weakly supervised image manipulation methods overlook the importance of edge information for accurate localization, leading to suboptimal localization performance. To address this, we propose a Context-Aware Boundary Localization (CABL) module to aggregate boundary features and learn context-inconsistency for localizing manipulated areas. Furthermore, by leveraging Class Activation Mapping (CAM) and Segment Anything Model (SAM), we introduce the CAM-Guided SAM Refinement (CGSR) module to generate more accurate manipulation localization maps. By integrating two modules, we present a novel weakly supervised framework based on a dual-branch Transformer-CNN architecture. Our method achieves outstanding localization performance across multiple datasets. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IEEE Signal Process. Lett. | 3 |
| 2025 | Bootstrapping Audio-Visual Video Segmentation by Strengthening Audio CuesabstractHow to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual video segmentation (AVS) task has been proposed, aiming to segment the sounding objects in video frames under the guidance of audio cues. However, most existing AVS methods are hindered by a modality imbalance where the visual features tend to dominate those of the audio modality, due to a unidirectional and insufficient integration of audio cues. This imbalance skews the feature representation towards the visual aspect, impeding the learning of joint audio-visual representations and potentially causing segmentation inaccuracies. To address this issue, we propose AVSAC. Our approach features a Bidirectional Audio-Visual Decoder (BAVD) with integrated bidirectional bridges, enhancing audio cues and fostering continuous interplay between audio and visual modalities. This bidirectional interaction narrows the modality imbalance, facilitating more effective learning of integrated audio-visual representations. Additionally, we present a strategy for audio-visual frame-wise synchrony as fine-grained guidance of BAVD. This strategy enhances the share of auditory components in visual features, contributing to a more balanced audio-visual representation learning. Extensive experiments show that our method has state-of-the-art performance on several AVS public benchmarks. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Le Lu 0001, Jieping Ye |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Zig-RiR: Zigzag RWKV-in-RWKV for Efficient Medical Image SegmentationabstractMedical image segmentation has made significant strides with the development of basic models. Specifically, models that combine CNNs with transformers can successfully extract both local and global features. However, these models inherit the transformer's quadratic computational complexity, limiting their efficiency. Inspired by the recent Receptance Weighted Key Value (RWKV) model, which achieves linear complexity for long-distance modeling, we explore its potential for medical image segmentation. While directly applying vision-RWKV yields suboptimal results due to insufficient local feature exploration and disrupted spatial continuity, we propose a novel nested structure, Zigzag RWKV-in-RWKV (Zig-RiR), to address these issues. It consists of Outer and Inner RWKV blocks to adeptly capture both global and local features without disrupting spatial continuity. We treat local patches as "visual sentences" and use the Outer Zig-RWKV to explore global information. Then, we decompose each sentence into sub-patches ("visual words") and use the Inner Zig-RWKV to further explore local information among words, at negligible computational cost. We also introduce a Zigzag-WKV attention mechanism to ensure spatial continuity during token scanning. By aggregating visual word and sentence features, our Zig-RiR can effectively explore both global and local information while preserving spatial continuity. Experiments on four medical image segmentation datasets of both 2D and 3D modalities demonstrate the superior accuracy and efficiency of our method, outperforming the state-of-the-art method 14.4 times in speed and reducing GPU memory usage by 89.5% when testing on ${1024} \times {1024}$ high-resolution medical images. Our code is available at https://github.com/txchen-USTC/Zig-RiR. Zhentao Tan, Qi Chu 0001, Nenghai Yu, Le Lu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Multi-spectral Class Center Network for Face Manipulation LocalizationabstractAs Deepfake content proliferates online, advancing face manipulation forensics has become crucial. To combat this emerging threat, previous methods mainly focus on studying how to distinguish authentic and manipulated face images. Although impressive, image-level classification lacks explainability and is limited to specific application scenarios, spurring recent research on pixel-level prediction for face manipulation forensics. However, existing forgery localization methods suffer from exploring frequency-based forgery traces in the localization network. In this paper, we observe that multi-frequency spectrum information is effective for identifying tampered regions. To this end, a novel Multi-spectral Class Center Network (MSCCNet) is proposed for face manipulation localization. Specifically, we design a Multi-spectral Class Center (MSCC) module to learn more generalizable and multi-frequency features. Based on the features of different frequency bands, the MSCC module collects multi-spectral class centers and computes pixel-to-class relations. Applying multi-spectral class-level representations suppresses the semantic information of the visual concepts which is insensitive to manipulated regions of forgery images. Furthermore, we propose a Multi-level Features Aggregation (MFA) module to employ more low-level forgery artifacts and structural textures. Meanwhile, we conduct a comprehensive localization benchmark based on pixel-level FF++ and Dolos datasets. Experimental results quantitatively and qualitatively demonstrate the effectiveness and superiority of the proposed MSCCNet. We expect this work to inspire more studies on pixel-level face manipulation localization. The codes are available. Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Bin Liu 0016, Honggang Hu, Nenghai Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is critical to national security and has been extensively applied in military areas. ISTD aims to segment small target pixels from background. Most ISTD networks focus on designing feature extraction blocks or feature fusion modules, but rarely describe the ISTD process from the feature map evolution perspective. In the ISTD process, the network attention gradually shifts towards target areas. We abstract this process as the directional movement of feature map pixels to target areas through convolution, pooling and interactions with surrounding pixels, which can be analogous to the movement of thermal particles constrained by surrounding variables and particles. In light of this analogy, we propose Thermal Conduction-Inspired Transformer (TCI-Former) based on the theoretical principles of thermal conduction. According to thermal conduction differential equation in heat dynamics, we derive the pixel movement differential equation (PMDE) in the image domain and further develop two modules: Thermal Conduction-Inspired Attention (TCIA) and Thermal Conduction Boundary Module (TCBM). TCIA incorporates finite difference method with PMDE to reach a numerical approximation so that target body features can be extracted. To further remove errors in boundary areas, TCBM is designed and supervised by boundary masks to refine target body features with fine boundary details. Experiments on IRSTD-1k and NUAA-SIRST demonstrate the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 3 |
| 2024 | MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsabstractGenerating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only a single modality of the control signal, which limits their application in the real digital human industry. This paper presents a Motion General-Purpose generaTor (MotionGPT) that can use multimodal control signals, e.g., text and single-frame poses, for generating consecutive human motions by treating multimodal signals as special input tokens in large language models (LLMs). Specifically, we first quantize multimodal control signals into discrete codes and then formulate them in a unified prompt instruction to ask the LLMs to generate the motion answer. Our MotionGPT demonstrates a unified human motion generation model with multimodal control signals by tuning a mere 0.4% of LLM parameters. To the best of our knowledge, MotionGPT is the first method to generate human motion by multimodal control signals, which we hope can shed light on this new direction. Visit our webpage at https://qiqiapink.github.io/MotionGPT/. Bin Liu 0016, Shixiang Tang, Yan Lu 0001, Lu Chen 0001, Lei Bai 0001, Qi Chu 0001, Nenghai Yu, Wanli Ouyang |
AAAI | 8 |
| 2024 | Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationabstractText-to-Image person re-identification (TI-ReID) aims to retrieve the images of target identity according to the given textual description. The existing methods in TI-ReID focus on aligning the visual and textual modalities through contrastive feature alignment or reconstructive masked language modeling (MLM). However, these methods parameterize the image/text instances as deterministic embeddings and do not explicitly consider the inherent uncertainty in pedestrian images and their textual descriptions, leading to limited image-text relationship expression and semantic alignment. To address the above problem, in this paper, we propose a novel method that unifies multi-modal uncertainty modeling and semantic alignment for TI-ReID. Specifically, we model the image and textual feature vectors of pedestrian as Gaussian distributions, where the multi-granularity uncertainty of the distribution is estimated by incorporating batch-level and identity-level feature variances for each modality. The multi-modal uncertainty modeling acts as a feature augmentation and provides richer image-text semantic relationship. Then we present a bi-directional cross-modal circle loss to more effectively align the probabilistic features between image and text in a self-paced manner. To further promote more comprehensive image-text semantic alignment, we design a task that complements the masked language modeling, focusing on the cross-modality semantic recovery of global masked token after cross-modal interaction. Extensive experiments conducted on three TI-ReID datasets highlight the effectiveness and superiority of our method over state-of-the-arts. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
AAAI | 4 |
| 2024 | Towards More Unified In-Context Visual UnderstandingabstractThe rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently, ICL has been employed in visual understanding tasks, such as semantic segmentation and image captioning, yielding promising results. However, existing visual ICL framework can not enable producing content across multiple modalities, whicd limits their potential usage scenarios. To address this issue, we present a new ICLframeworkfor visual understanding with multi-modal output enabled. First, we quantize and embed both text and visual prompt into a unified representational space, structured as interleaved in-context sequences. Then a decoder-only sparse transformer architecture is employed to perform generative modeling on them, facilitating in-context learning. Thanks to this design, the model is capable of handling in-context vision understanding tasks with multimodal output in a unified pipeline. Experimental re-sults demonstrate that our model achieves competitive performance compared with specialized models and previous ICL baselines. Overall, our research takes a further step toward unified multimodal in-context learning. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Jianmin Bao, Bin Liu 0016, Shengwei Xu, Nenghai Yu |
CVPR | 5 |
| 2024 | Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and GroundingabstractAI-synthesized text and images have gained significant attention, particularly due to the widespread dissemination of multi-modal manipulations on the internet, which has resulted in numerous negative impacts on society. Existing methods for multi-modal manipulation detection and grounding primarily focus on fusing vision-language features to make predictions, while overlooking the importance of modality-specific features, leading to sub-optimal results. In this paper, we construct a simple and novel transformer-based framework for multi-modal manipulation detection and grounding tasks. Our framework simultaneously explores modality-specific features while preserving the capability for multi-modal alignment. To achieve this, we introduce visual/language pre-trained encoders and dual-branch cross-attention (DCA) to extract and fuse modality-unique features. Furthermore, we design decoupled fine-grained classifiers (DFC) to enhance modality-specific feature mining and mitigate modality competition. Moreover, we propose an implicit manipulation query (IMQ) that adaptively aggregates global contextual cues within each modality using learnable queries, thereby improving the discovery of forged details. Extensive experiments on the DGM4dataset demonstrate the superior performance of our proposed model compared to state-of-the-art approaches. Jiazhen Wang, Bin Liu 0016, Changtao Miao, Wanyi Zhuang, Qi Chu 0001, Nenghai Yu |
ICASSP | 6 |
| 2024 | Delving Deeper Into Vulnerable Samples in Adversarial TrainingabstractRecently, vulnerable samples have been shown to be crucial for improving adversarial training performance. Our analysis on existing vulnerable samples mining methods indicate that existing methods have two problems: 1) valuable connections among different pairs of natural samples and their adversarial counterparts are ignored; 2) parts of vulnerable samples are unconsidered. To better leverage vulnerable samples, we propose INter PAir ConstrainT (INPACT) and Vulnerable Aware adveRsarial Training (VART) to address these drawbacks respectively. INPACT assesses adversarial risk with more comprehensive regularization on sample relationships, which takes both inter and intra connections of natural/adversarial sample pairs into consideration. Meanwhile VART makes full use of all vulnerable samples, including notable proportion neglected by existing instance re-weighting strategies. Extensive experiments on different datasets and backbones demonstrate the effectiveness of the proposed method. Qi Chu 0001, Shubin Xu, Nenghai Yu |
ICASSP | 3 |
| 2024 | Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationabstractLarge-scale Vision Transformers have achieved promising performance on downstream tasks through feature pre-training. However, the performance of vanilla lightweight Vision Transformers (ViTs) is still far from satisfactory compared to that of recent lightweight CNNs or hybrid networks. In this paper, we aim to unlock the potential of vanilla lightweight ViTs by exploring the adaptation of the widely-used re-parameterization technology to ViTs for improving learning ability during training without increasing the inference cost. The main challenge comes from the fact that CNNs perfectly complement with re-parameterization over convolution and batch normalization, while vanilla Transformer architectures are mainly comprised of linear and layer normalization layers. We propose to incorporate the nonlinear ensemble into linear layers by expanding the depth of the linear layers with batch normalization and fusing multiple linear features with hierarchical representation ability through a pyramid structure. We also discover and solve a new transformer-specific distribution rectification problem caused by multi-branch re-parameterization. Finally, we propose our Two-Dimensional Re-parameterized Linear module (TDRL) for ViTs. Under the popular self-supervised pre-training and supervised fine-tuning strategy, our TDRL can be used in these two stages to enhance both generic and task-specific representation. Experiments demonstrate that our proposed method not only boosts the performance of vanilla Vit-Tiny on various vision tasks to new state-of-the-art (SOTA) but also shows promising generality ability on other networks. Code will be available. Zhentao Tan, Qi Chu 0001, Le Lu 0001, Nenghai Yu, Jieping Ye |
ICLR | 4 |
| 2024 | MFMS: Learning Modality-Fused and Modality-Specific Features for Deepfake Detection and Localization TasksabstractThis paper presents a summary of the proposed solution to the AV-Deepfake1M competition. Deepfake technology is developing fast, and realistic generation techniques of audio and videos have aroused public concerns. With this background, the AV-Deepfake1M competition aims to address the problem of audio-video Deepfake and provides a large-scale dataset named AV-Deepfake1M to boost the research in this area. In this paper, we present our solutions which have achieved top performance in this competition. We also provide more detailed experiments to prove the effectiveness of the modules used in our methods. Changtao Miao, Jianshu Li, Wenzhong Deng, Weibin Yao, Zhe Li 0081, Bingyu Hu, Weiwei Feng, Qi Chu 0001 |
ACM Multimedia | 11 |
| 2024 | Detect Text Forgery with Non-forged Image Features: A Framework for Detection and Grounding of Image-Text Manipulation
Changtao Miao, Qi Chu 0001, Dianmo Sheng, Jiazhen Wang, Bin Liu 0016, Nenghai Yu |
PRCV (11) | 3 |
| 2024 | Transformer Based Pluralistic Image Completion With Reduced Information LossabstractTransformer based methods have achieved great success in image inpainting recently. However, we find that these solutions regard each pixel as a token, thus suffering from an information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration. 2) They quantize 2563RGB values to a small number (such as 512) of quantized color values. The indices of quantized pixels are used as tokens for the inputs and prediction targets of the transformer. To mitigate these issues, we propose a new transformer based framework called “PUT”. Specifically, to avoid input downsampling while maintaining computation efficiency, we design a patch-based auto-encoder P-VQVAE. The encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by input quantization, an Un-quantized Transformer is applied. It directly takes features from the P-VQVAE encoder as input without any quantization and only regards the quantized tokens as prediction targets.Furthermore, to make the inpainting process more controllable, we introduce semantic and structural conditions as extra guidance. Extensive experiments show that our method greatly outperforms existing transformer based methods on image fidelity and achieves much higher diversity and better fidelity than state-of-the-art pluralistic inpainting methods on complex large-scale datasets (e.g., ImageNet). Codes are available athttps://github.com/liuqk3/PUT. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Ying Fu 0001, Qi Chu 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Feature Preservation and Shape Cues Assist Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims to segment small target pixels from infrared images and has extensive applications in many fields. Despite multiple progress, challenges remain as present methods still easily suffer from missed detection. Also, present methods are not sensitive enough to irregular target shapes. We argue that the main reason is that some informative small target features get lost during the aggressive downsampling in the encoder without effective recovery. In this article, we propose a new network with a dual-branch encoder-decoder structure for ISTD to address the two challenges. Specifically, to better preserve small target body features for more accurate target locations, we propose to maintain a relatively high resolution of feature maps in one encoder branch. For the other encoder branch, we gradually enlarge feature channels while shrinking resolutions and devise Perona-Malik diffusion (PMD) blocks to preserve shape cues inspired by the shape-preserving effect of PMD in denoising. The encoded high-resolution target body features and high-channel shape cues actually complement each other, so we design channel-resolution interact modules (CRIMs) to combine them. In the decoder, we propose orthogonal central difference fusion (OCDF) that relies on mining contrast differences to further refine shape-aware ISTD quality. Experiments on NUAA-SIRST and IRSTD-1k prove the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small-Target DetectionabstractRecently, infrared small-target detection (ISTD) has made significant progress, thanks to the development of basic models. Specifically, the models combining CNNs with Transformers can successfully extract both local and global features. However, the disadvantage of the Transformer is also inherited, that is, the quadratic computational complexity to sequence length. Inspired by the recent basic model with linear complexity for long-distance modeling, Mamba, we explore the potential of this state-space model (SSM) for ISTD tasks in terms of effectiveness and efficiency in the article. However, directly applying Mamba achieves suboptimal performances due to the insufficient harnessing of local features, which are imperative for detecting small targets. Instead, we tailor a nested structure, Mamba-in-Mamba (MiM-ISTD), for efficient ISTD. It consists of Outer and Inner Mamba blocks to adeptly capture both global and local features. Specifically, we treat the local patches as “visual sentences” and use the Outer Mamba to explore the global information. We then decompose each visual sentence into subpatches as “visual words” and use the Inner Mamba to further explore the local information among words in the visual sentence with negligible computational costs. By aggregating the visual word and visual sentence features, our MiM-ISTD can effectively explore both global and local information. Experiments on NUAA-SIRST and IRSTD-1k show the superior accuracy and efficiency of our method. Specifically, MiM-ISTD is$8\times $faster than the SOTA method and reduces GPU memory usage by 62.2% when testing on$2048 \times 2048$images, overcoming the computation and memory constraints on high-resolution infrared images. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Jieping Ye |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Exploring the Application of Large-Scale Pre-Trained Models on Adverse Weather RemovalabstractImage restoration under adverse weather conditions (e.g., rain, snow, and haze) is a fundamental computer vision problem that has important implications for various downstream applications. Distinct from early methods that are specially designed for specific types of weather, recent works tend to simultaneously remove various adverse weather effects based on either spatial feature representation learning or semantic information embedding. Inspired by various successful applications incorporating large-scale pre-trained models (e.g., CLIP), in this paper, we explore their potential benefits for leveraging large-scale pre-trained models in this task based on both spatial feature representation learning and semantic information embedding aspects: 1) spatial feature representation learning, we design a Spatially Adaptive Residual (SAR) encoder to adaptively extract degraded areas. To facilitate training of this model, we propose a Soft Residual Distillation (CLIP-SRD) strategy to transfer spatial knowledge from CLIP between clean and adverse weather images; 2) semantic information embedding, we propose a CLIP Weather Prior (CWP) embedding module to enable the network to adaptively respond to different weather conditions. This module integrates the sample-specific weather priors extracted by the CLIP image encoder with the distribution-specific information (as learned by a set of parameters) and embeds these elements using a cross-attention mechanism. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art performance under various and severe adverse weather conditions. The code will be made available. Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Le Lu 0001, Jieping Ye, Nenghai Yu |
IEEE Trans. Image Process. | 4 |
| 2024 | Joint Identity-Aware Mixstyle and Graph-Enhanced Prototype for Clothes-Changing Person Re-IdentificationabstractIn recent years, considerable progress has been witnessed in the person re-identification (Re-ID). However, in a more realistic long-term scenario, the appearance shift arising from the clothes-changing inevitably deteriorates the conventional methods that heavily depend on the clothing color. Although the current clothes-changing person Re-ID methods introduce external human knowledge (i.e, contour, mask) and sophisticated feature decoupling strategy to alleviate the clothing shift, they still face the risk of overfitting to clothing due to the limited clothing diversity of training set. To more efficiently and effectively promote the clothes-irrelevant feature learning, we present a novel joint Identity-aware Mixstyle and Graph-enhanced Prototype method for clothes-changing person Re-ID. Specifically, by treating the cloth-changing as fine-grained domain/style shift, the identity-aware mixstyle (IMS) is proposed from the perspective of domain generalization, which mixes the instance-level feature statistics of samples within each identity to synthesize novel and diverse clothing styles, while retaining the correspondence between synthesized samples and latent label space. By incorporating the IMS module, the more diverse styles can be exploited to train a clothing-shift robust model. To further reduce the feature discrepancy caused by clothing variations, the graph-enhanced prototype constraint (GEP) module is proposed to explore the graph similarity structure of style-augmented samples across memory bank to build informative and robust prototypes, which serve as powerful exemplars for better clothing-irrelevant metric learning. The two modules are integrated into a joint learning framework and benefit each other. The extensive experiments conducted on clothes-changing person Re-ID datasets validate the superiority and effectiveness of our method. In addition, our method also shows good universality and corruption robustness on other Re-ID tasks. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2023 | BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is indispensable in remote sensing and military surveillance. Existing ISTD methods can discover regularly-shaped and clear objects well, but tend to overlook the tough-to-detect ones, such as targets with irregular shapes or blurry boundaries, causing inaccurate segmentation and missed detection. Considering that boundary areas assemble rich uncertainty information, we propose the Boundary-Aware Uncertainty Enhanced Network (BAUENet), where Uncertainty Enhanced Context Refinement (UECR) and Adaptive Feature Fusion Modules (AFFM) are devised to address this problem. Specifically, UECR extracts spatial contexts and refines them with uncertain area maps derived from backbone intermediate outputs, so as to distinguish boundary areas from other regions. AFFM adaptively aggregates cross-level features via balancing low-level details and high-level semantics for finer boundary preservation in both channel and spatial dimensions during up-sampling feature fusion. Experiments on several public datasets demonstrate the effectiveness of the proposed method, especially for irregular shape and blurry boundary cases. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICASSP | 2 |
| 2023 | Dual-Feature Enhancement for Weakly Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WTAL) aims at localizing actions in untrimmed videos with only video-level labels. Most existing methods embrace a "localization by classification" paradigm and adopt a model that pre-trained with recognition task for feature extraction. The gap between recognition and localization tasks leads to inferior performance. Some recent works attempt to utilize feature enhancement to obtain better feature for localization and boost the performance to some extent. However, they are limited to intra-video information exploiting, while ignoring meaningful inter-video information in the dataset. In this paper, we propose a novel Dual-Feature Enhancement (DFE) method for WTAL, which can utilize both intra-and inter-video information. For intra-video, a local feature enhancement module is designed to promote the feature interaction along the temporal dimension within each video. For inter-video information, a global memory module is firstly designed to learn the representations for different categories across different videos. Then, a global feature enhancement module is used to enhance the video features with the help of those global representations in the memory. Besides, to reduce the extra computational cost caused by global enhancement module in the inference stage, a distillation loss is applied to enforce the local branch to learn the information from global branch, so the global enhancement module could be removed during inference. The proposed method achieves state-of-the-art performance on popular benchmarks. Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICASSP | 3 |
| 2023 | Dual-Uncertainty Guided Curriculum Learning and Part-Aware Feature Refinement for Domain Adaptive Person Re-IdentificationabstractUnsupervised Domain Adaptative person re-identification (UDA ReID) aims to transfer the knowledge of pre-trained model from labeled source domain to unlabeled target domain. Although the current clustering-based methods have achieved promising success, they neglect the tolerance of the model to cope with different-level noise, which may cause the model to memorize some incorrect patterns caused by label noise and overfit on them rapidly in the early stages. In this paper, we introduce a novel Dual Uncertainty guided Curriculum Learning (DUCL) method to tackle the above problems. Specifically, the reliability-based curriculum allocation is proposed to enforce the sample adaptation in an easy-to-hard manner, which is further assisted by a novel dual-uncertainty re-weighting strategy to alleviate the influence of label noise. In addition, we design Part-aware Feature Refinement (PAFR) to enhance the discrimination of model and thereby acquiring more reliable pseudo-labels. Specifically, the part-aware attention maps are exploited in the PAFR to integrate fine-grained semantics into holistic representation. Extensive experiments have validated the superiority of the proposed method. Zhangping Liu, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICASSP | 4 |
| 2023 | Evopose: A Recursive Transformer for 3D Human Pose Estimation with Kinematic Structure PriorsabstractTransformer is popular in recent 3D human pose estimation, which utilizes long-term modeling to lift 2D keypoints into the 3D space. However, current transformer-based methods do not fully exploit the prior knowledge of the human skeleton provided by the kinematic structure. In this paper, we propose a novel transformer-based model EvoPose to introduce the human body prior knowledge for 3D human pose estimation effectively. Specifically, a Structural Priors Representation (SPR) module represents human priors as structural features carrying rich body patterns, e.g. joint relationships. The structural features are interacted with 2D pose sequences and help the model to achieve more informative spatiotemporal features. Moreover, a Recursive Refinement (RR) module is applied to refine the 3D pose outputs by utilizing estimated results and further injects human priors simultaneously. Extensive experiments demonstrate the effectiveness of EvoPose which achieves a new state of the art on two most popular benchmarks, Human3.6M and MPI-INF-3DHP. Yan Lu 0001, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICASSP | 5 |
| 2023 | Enhancing Adversarial Transferability from the Perspective of Input Loss Landscape
Yinhu Xu, Qi Chu 0001, Zixiang Luo, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 2 |
| 2023 | Revisiting TENT for Test-Time Adaption Semantic Segmentation and Classification Head Adjustment
Xuanpu Zhao, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu |
ICIG (3) | 2 |
| 2023 | ABMNet: Coupling Transformer with CNN Based on Adams-Bashforth-Moulton Method for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims at segmenting the small targets from infrared images, which has wide applications in military surveillance. Present methods are mainly based on CNN and focus on modelling locality while ignoring global dependencies, which are indispensable because the local areas similar to small targets always spread over most of the background, causing heavy target ambiguity. Recently, RKformer [1] has combined local features with global dependencies and further introduced Runge-Kutta method, a one-step Ordinary Differential Equation (ODE) solver, to ISTD and performed well. However, the method simply fuses features from original transformer and residual blocks by naive concatenation, causing insufficient feature interaction. Also, it inevitably brings effective information loss, which greatly impairs ambiguous target features. To address above problems and target ambiguity, we introduce Adams-Bashforth-Moulton method and propose ABMNet, which has (1) multi-step memory and self-rectification mechanisms, guaranteeing more sufficient information usage and more accurate detection, (2) and achieves more sufficient interaction of both local and global information. Experiments on MDFA and IRSTD-1k demonstrate the superiority of our method. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICME | 2 |
| 2023 | X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionabstractCopy-Paste is a simple and effective data augmentation strategy for instance segmentation. By randomly pasting object instances onto new background images, it creates new training data for free and significantly boosts the segmentation performance, especially for rare object categories. Although diverse, high-quality object instances used in Copy-Paste result in more performance gain, previous works utilize object instances either from human-annotated instance segmentation datasets or rendered from 3D object models, and both approaches are too expensive to scale up to obtain good diversity. In this paper, we revisit Copy-Paste at scale with the power of newly emerged zero-shot recognition models (e.g., CLIP) and text2image models (e.g., StableDiffusion). We demonstrate for the first time that using a text2image model to generate images or zero-shot recognition model to filter noisily crawled images for different object categories is a feasible way to make Copy-Paste truly scalable. To make such success happen, we design a data acquisition and processing framework, dubbed ``X-Paste", upon which a systematic study is conducted. On the LVIS dataset, X-Paste provides impressive improvements over the strong baseline CenterNet2 with Swin-L as the backbone. Specifically, it archives +2.6 box AP and +2.1 mask AP gains on all classes and even more significant gains with +6.8 box AP +6.5 mask AP on long-tail classes. Dianmo Sheng, Jianmin Bao, Dongdong Chen 0001, Dong Chen 0003, Fang Wen 0001, Lu Yuan 0001, Ce Liu 0001, Wenbo Zhou 0004, Qi Chu 0001, Weiming Zhang 0001, Nenghai Yu |
ICML | 10 |
| 2023 | Fluid Dynamics-Inspired Network for Infrared Small Target DetectionabstractMost infrared small target detection (ISTD) networks focus on building effective neural blocks or feature fusion modules but none describes the ISTD process from the image evolution perspective. The directional evolution of image pixels influenced by convolution, pooling and surrounding pixels is analogous to the movement of fluid elements constrained by surrounding variables ang particles. Inspired by this, we explore a novel research routine by abstracting the movement of pixels in the ISTD process as the flow of fluid in fluid dynamics (FD). Specifically, a new Fluid Dynamics-Inspired Network (FDI-Net) is devised for ISTD. Based on Taylor Central Difference (TCD) method, the TCD feature extraction block is designed, where convolution and Transformer structures are combined for local and global information. The pixel motion equation during the ISTD process is derived from the Navier–Stokes (N-S) equation, constructing a N-S Refinement Module that refines extracted features with edge details. Thus, the TCD feature extraction block determines the primary movement direction of pixels during detection, while the N-S Refinement Module corrects some skewed directions of the pixel stream to supplement the edge details. Experiments on IRSTD-1k and SIRST demonstrate that our method achieves SOTA performance in terms of evaluation metrics. Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IJCAI | 2 |
| 2023 | Semantic Probability Distribution Modeling for Diverse Semantic Image SynthesisabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level or even instance-level multimodal results, still remains a challenge. In this article, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at both semantics and instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. To further extend the user interaction function of the proposed method, we also introduce sketches into the network. In addition, specially designed generator modules, Progressive Growing Module and Multi-Scale Refinement Module, can be used as a general module to improve the performance of complex scene generation. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Codes are available at https://github.com/tzt101/INADE.git. Zhentao Tan, Qi Chu 0001, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qiankun Liu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | F2Trans: High-Frequency Fine-Grained Transformer for Face Forgery DetectionabstractIn recent years, face forgery detectors have aroused great interest and achieved impressive performance, but they are still struggling with generalization and robustness. In this work, we explore taking full advantage of the fine-grained forgery traces in both spatial and frequency domains to alleviate this issue. Specifically, we propose a novel High-Frequency Fine-Grained Transformer (F2Trans) network which contains two important components, namely Central Difference Attention (CDA) and High-frequency Wavelet Sampler (HWS). The premier CDA module is capable of capturing invariant fine-grained manipulation patterns by aggregating both pixel-level intensity and gradient information of the query to generate key and value pairs. Subsequently, the proposed HWS discards the low-frequency components of wavelet transformation and hierarchically explores high-frequency forgery cues of feature maps, which prevents model confusion caused by low-frequency components and pays attention to local frequency information. In addition, HWS can be employed as a special pooling layer for the F2Trans architecture to produce hierarchical feature representations in the spatial-frequency domain. Extensive experiments on multiple popular benchmarks demonstrate the generalization and robustness of the specially designed F2Trans framework is well-tailored for face forgery detection when confronting the cross-dataset, cross-manipulation, and unseen perturbations. Changtao Miao, Zichang Tan, Qi Chu 0001, Huan Liu 0030, Honggang Hu, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | AutoMA: Towards Automatic Model Augmentation for Transferable Adversarial AttacksabstractRecent adversarial attack works attempt to improve the transferability by applying various differentiable transformations on input images. Considering the differentiable transformations and the original model together as a new model, these methods can be regarded as model augmentation that effectively derives an ensemble of models from the single original model. Despite their impressive performance, the model augmentation policies used in these methods are manually designed by experimental attempts, leaving the design of model augmentation policy an open question. In this paper, we propose an Automatic Model Augmentation (AutoMA) approach to find a strong model augmentation policy for transferable adversarial attacks. Specifically, we design a discrete search space that contains various diffierentiable transformations with different parameters and adopt reinforcement learning to search for the strong augmentation policy. The sampled augmentation policies together with the rewards they obtain during the searching process reveal several valuable observations for designing more powerful attacks using model augmentation policy:1) Augmentation transformations on color space are less effective; 2) The transformation type diversity matters; and 3) Using small distortion for geometric transformations while larger distortion for intensity transformations.Extensive experiments show that the augmentation policy found by AutoMA achieves superior performance than existing manually designed policies in a wide range of cases. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Multim. | 2 |
| 2022 | Affinity-Aware Relation Network for Oriented Object Detection in Aerial Images
Tingting Fang, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ACCV (5) | 4 |
| 2022 | Reduce Information Loss in Transformers for Pluralistic Image InpaintingabstractTransformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer from information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration, incurring information loss and extra misalignment for the boundaries of masked regions. 2) They quantize 2563RGB pixels to a small number (such as 512) of quantized pixels. The indices of quantized pixels are used as tokens for the inputs and prediction targets of transformer. Although an extra CNN network is used to upsample and refine the low-resolution results, it is difficult to retrieve the lost information back. To keep input information as much as possible, we propose a new transformer based framework “PUT”. Specifically, to avoid input downsampling while maintaining the computation efficiency, we design a patch-based auto-encoder P-VQVAE, where the encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by quantization, an Un-Quantized Transformer (UQ-Transformer) is applied, which directly takes the features from P-VQVAE encoder as input without quantization and regards the quantized tokens only as prediction targets. Extensive experiments show that PUT greatly outperforms state-of-the-art methods on image fidelity, especially for large masked regions and complex large-scale datasets. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan 0001, Nenghai Yu |
CVPR | 4 |
| 2022 | Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification
Xulin Li, Yan Lu 0001, Bin Liu 0016, Guojun Yin, Qi Chu 0001, Jinyang Huang, Feng Zhu 0006, Rui Zhao 0001, Nenghai Yu |
ECCV (26) | 6 |
| 2022 | UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection
Wanyi Zhuang, Qi Chu 0001, Zhentao Tan, Qiankun Liu 0001, Changtao Miao, Zixiang Luo, Nenghai Yu |
ECCV (5) | 2 |
| 2022 | Towards Intrinsic Common Discriminative Features Learning for Face Forgery Detection Using Adversarial LearningabstractExisting face forgery detection methods usually treat face forgery detection as a binary classification problem and adopt deep convolution neural networks to learn discriminative features. The ideal discriminative features should be only related to the real/fake labels of facial images. However, we observe that the features learned by vanilla classification networks are correlated to unnecessary properties, such as forgery methods and facial identities. Such phenomenon would limit forgery detection performance especially for the generalization ability. Motivated by this, we propose a novel method which utilizes adversarial learning to eliminate the negative effect of different forgery methods and facial identities, which helps classification network to learn intrinsic common discriminative features for face forgery detection. To leverage data lacking ground truth label of facial identities, we design a special identity discriminator based on similarity information derived from off-the-shelf face recognition model. Extensive experiments demonstrate the effectiveness of the proposed method under both intra-dataset and cross-dataset evaluation settings. Wanyi Zhuang, Qi Chu 0001, Changtao Miao, Bin Liu 0016, Nenghai Yu |
ICME | 2 |
| 2022 | Cloth-Aware Center Cluster Loss for Cloth-Changing Person Re-identification
Xulin Li, Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
PRCV (1) | 4 |
| 2022 | Multi-view Geometry Distillation for Cloth-Changing Person ReID
Hanlei Yu, Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu |
PRCV (1) | 4 |
| 2022 | Online multi-object tracking with unsupervised re-identification learning and occlusion estimation
Qiankun Liu 0001, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Bin Liu 0016, Lei Zhang 0001, Nenghai Yu |
Neurocomputing | 3 |
| 2022 | Efficient Semantic Image Synthesis via Class-Adaptive NormalizationabstractSpatially-adaptive normalization (SPADE) is remarkably successful recently in conditional semantic image synthesis in T. Park et al. 2019 which modulates the normalized activation with spatially-varying transformations learned from semantic layouts, to prevent the semantic information from being washed away. Despite its impressive performance, a more thorough understanding of the advantages inside the box is still highly demanded to help reduce the significant computation and parameter overhead introduced by this novel structure. In this paper, from a return-on-investment point of view, we conduct an in-depth analysis of the effectiveness of this spatially-adaptive normalization and observe that its modulation parameters benefit more from semantic-awareness rather than spatial-adaptiveness, especially for high-resolution input masks. Inspired by this observation, we propose class-adaptive normalization (CLADE), a lightweight but equally-effective variant that is only adaptive to semantic class. In order to further improve spatial-adaptiveness, we introduce intra-class positional map encoding calculated from semantic layouts to modulate the normalization parameters of CLADE and propose a truly spatially-adaptive variant of CLADE, namely CLADE-ICPE. Through extensive experiments on multiple challenging datasets, we demonstrate that the proposed CLADE can be generalized to different SPADE-based methods while achieving comparable generation quality compared to SPADE, but it is much more efficient with fewer extra parameters and lower computational cost. The code and pretrained models are available at https://github.com/tzt101/CLADE.git. Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Menglei Chai, Jing Liao 0001, Mingming He, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hierarchical Frequency-Assisted Interactive Networks for Face Manipulation DetectionabstractRecently, face manipulation techniques have caused increasing trust concerns in our society. Although current face manipulation detection methods achieve impressive performance regarding intra-dataset evaluation, they are struggling to improve the generalization and robustness ability. To address this issue, we propose a novel Hierarchical Frequency-assisted Interactive Networks (HFI-Net) to explore comprehensive frequency-related forgery cues for face manipulation detection. At first, we formulate HFI-Net as a dual-branch network to take full advantage of both CNN and transformer for capturing local details and global context information, respectively. Considering the forged faces are easy to show flaws in the frequency domain, a novel Frequency-based Feature Refinement (FFR) module is proposed to learn frequency-based attention from RGB features. FFR module emphasizes forgery cues and suppresses the pristine semantics information by keeping middle-high frequency features while discarding the low-frequency ones. Based on FFR, we further develop a co-sharing Global-Local Interaction (GLI) module to conduct frequency-assisted interactions while capturing complementarity among dual branches. Lastly, we further implement the GLI module in each stage of the network to effectively explore multi-level frequency artifacts. Extensive experiments are conducted on several popular benchmarks including FaceForensics++, Celeb-DF, DeepFake-TIMIT, DFDC, UADFV, and DeeperForensics-1.0, which shows that our model outperforms the state-of-the-art, especially in unseen datasets, manipulations, and perturbations evaluation. Changtao Miao, Zichang Tan, Qi Chu 0001, Nenghai Yu, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Temporal ROI Align for Video Object RecognitionabstractVideo object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame. However, ROI Align, as one of the most core procedures of video detectors, still remains extracting features from a single-frame feature map for proposals, making the extracted ROI features lack temporal information from videos. In this work, considering the features of the same object instance are highly similar among frames in a video, a novel Temporal ROI Align operator is proposed to extract features from other frames feature maps for current frame proposals by utilizing feature similarity. The proposed Temporal ROI Align operator can extract temporal information from the entire video for proposals. We integrate it into single-frame video detectors and other state-of-the-art video detectors, and conduct quantitative experiments to demonstrate that the proposed Temporal ROI Align operator can consistently and significantly boost the performance. Besides, the proposed Temporal ROI Align can also be applied into video instance segmentation. Kai Chen 0026, Xinjiang Wang, Qi Chu 0001, Feng Zhu 0006, Dahua Lin, Nenghai Yu, Huamin Feng |
AAAI | 4 |
| 2021 | Joint Color-irrelevant Consistency Learning and Identity-aware Modality Adaptation for Visible-infrared Cross Modality Person Re-identificationabstractVisible-infrared cross modality person re-identification (VI-ReID) is a core but challenging technology in the 24-hours intelligent surveillance system. How to eliminate the large modality gap lies in the heart of VI-ReID. Conventional methods mainly focus on directly aligning the heterogeneous modalities into the same space. However, due to the unbalanced color information between the visible and infrared images, the features of visible images tend to overfit the clothing color information, which would be harmful to the modality alignment. Besides, these methods mainly align the heterogeneous feature distributions in dataset-level while ignoring the valuable identity information, which may cause the feature misalignment of some identities and weaken the discrimination of features. To tackle above problems, we propose a novel approach for VI-ReID. It learns the color-irrelevant features through the color-irrelevant consistency learning (CICL) and aligns the identity-level feature distributions by the identity-aware modality adaptation (IAMA). The CICL and IAMA are integrated into a joint learning framework and can promote each other. Extensive experiments on two popular datasets SYSU-MM01 and RegDB demonstrate the superiority and effectiveness of our approach against the state-of-the-art methods. Bin Liu 0016, Qi Chu 0001, Yan Lu 0001, Nenghai Yu |
AAAI | 3 |
| 2021 | Diverse Semantic Image Synthesis via Probability Distribution ModelingabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level multimodal results, still remains a challenge. In this paper, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at semantic or even instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Code will be available at https://github.com/tzt101/INADE.git Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
CVPR | 5 |
| 2021 | ISNet: Integrate Image-Level and Semantic-Level Context for Semantic SegmentationabstractCo-occurrent visual pattern makes aggregating contextual information a common paradigm to enhance the pixel representation for semantic image segmentation. The existing approaches focus on modeling the context from the perspective of the whole image, i.e., aggregating the image-level contextual information. Despite impressive, these methods weaken the significance of the pixel representations of the same category, i.e., the semantic-level contextual information. To address this, this paper proposes to augment the pixel representations by aggregating the image-level and semantic-level contextual information, respectively. First, an image-level context module is designed to capture the contextual information for each pixel in the whole image. Second, we aggregate the representations of the same category for each pixel where the category regions are learned under the supervision of the ground-truth segmentation. Third, we compute the similarities between each pixel representation and the image-level contextual information, the semantic-level contextual information, respectively. At last, a pixel representation is augmented by weighted aggregating both the image-level contextual information and the semantic-level contextual information with the similarities as the weights. Integrating the image-level and semantic-level context allows this paper to report state-of-the-art accuracy on four benchmarks, i.e., ADE20K, LIP, COCOStuff and Cityscapes1. Zhenchao Jin, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
ICCV | 3 |
| 2021 | Mining Contextual Information Beyond Image for Semantic SegmentationabstractThis paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations of the pixels of the corresponding class beyond the input image. To address this, this paper proposes to mine the contextual information beyond individual images to further augment the pixel representations. We first set up a feature memory module, which is updated dynamically during training, to store the dataset-level representations of various categories. Then, we learn class probability distribution of each pixel representation under the supervision of the ground-truth segmentation. At last, the representation of each pixel is augmented by aggregating the dataset-level representations based on the corresponding class probability distribution. Furthermore, by utilizing the stored dataset-level representations, we also propose a representation consistent learning strategy to make the classification head better address intra-class compactness and inter-class dispersion. The proposed method could be effortlessly incorporated into existing segmentation frameworks (e.g., FCN, PSPNet, OCRNet and DeepLabV3) and brings consistent performance improvements. Mining contextual information beyond image allows us to report state-of-the-art performance on various benchmarks: ADE20K, LIP, Cityscapes and COCO-Stuff1. Zhenchao Jin, Dongdong Yu, Qi Chu 0001, Changhu Wang, Jie Shao 0006 |
ICCV | 4 |
| 2021 | Improve Unsupervised Pretraining for Few-label TransferabstractUnsupervised pretraining has achieved great success and many recent works have shown unsupervised pretraining can achieve comparable or even slightly better transfer performance than supervised pretraining on downstream target datasets. But in this paper, we find this conclusion may not hold when the target dataset has very few labeled samples for finetuning, i.e., few-label transfer. We analyze the possible reason from the clustering perspective: 1) The clustering quality of target samples is of great importance to few-label transfer; 2) Though contrastive learning is essential to learn how to cluster, its clustering quality is still inferior to supervised pretraining due to lack of label supervision. Based on the analysis, we interestingly discover that only involving some unlabeled target domain into the unsupervised pretraining can improve the clustering quality, subsequently reducing the transfer performance gap with supervised pretraining. This finding also motivates us to propose a new progressive few-label transfer algorithm for real applications, which aims to maximize the transfer performance under a limited annotation budget. To support our analysis and proposed method, we conduct extensive experiments on nine different target datasets. Experimental results show our proposed method can significantly boost the few-label transfer performance of unsupervised pretraining. Suichan Li, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Lei Zhang 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 6 |
| 2021 | Geometry Uncertainty Projection Network for Monocular 3D Object DetectionabstractGeometry Projection is a powerful depth estimation method in monocular 3D object detection. It estimates depth dependent on heights, which introduces mathematical priors into the deep model. But projection process also introduces the error amplification problem, in which the error of the estimated height will be amplified and reflected greatly at the output depth. This property leads to uncontrollable depth inferences and also damages the training efficiency. In this paper, we propose a Geometry Uncertainty Projection Network (GUP Net) to tackle the error amplification problem at both inference and training stages. Specifically, a GUP module is proposed to obtains the geometry-guided uncertainty of the inferred depth, which not only provides high reliable confidence for each depth but also benefits depth learning. Furthermore, at the training stage, we propose a Hierarchical Task Learning strategy to reduce the instability caused by error amplification. This learning algorithm monitors the learning situation of each task by a proposed indicator and adaptively assigns the proper loss weights for different tasks according to their pre-tasks situation. Based on that, each task starts learning only when its pre-tasks are learned well, which can significantly improve the stability and efficiency of the training process. Extensive experiments demonstrate the effectiveness of the proposed method. The overall model can infer more reliable object depth than existing methods and outperforms the state-of-the-art image-based monocular 3D detectors by 3.74% and 4.7% AP40of the car and pedestrian categories on the KITTI benchmark. The code and model will be released at https://github.com/SuperMHP/GUPNet. Yan Lu 0001, Xinzhu Ma, Lei Yang 0045, Tianzhu Zhang 0001, Qi Chu 0001, Wanli Ouyang |
ICCV | 6 |
| 2021 | Towards More Powerful Multi-column Convolutional Network for Crowd Counting
Jiabin Zhang, Qi Chu 0001, Weihai Li, Bin Liu 0016, Weiming Zhang 0001, Nenghai Yu |
ICIG (1) | 2 |
| 2021 | Deepfake Video Detection Using 3D-Attentional Inception Convolutional Neural NetworkabstractThe current spike of deepfake techniques has received considerable attention due to security concerns. To mitigate the potential risks brought by deepfake techniques, many detection methods have been proposed. However, most existing works merely leverage spatial information from separate frames and ignore valuable inter-frame temporal information. In this paper, we propose a deepfake detection scheme that uses 3D-attentional inception network. The proposed model encompasses both spatial and temporal information simultaneously with the 3D kernels. Furthermore, the channel and spatial-temporal attention modules are applied to improve detection capabilities. Comprehensive experiments demonstrate that our scheme outperforms state-of-the-art methods. Changlei Lu, Bin Liu 0016, Wenbo Zhou 0004, Qi Chu 0001, Nenghai Yu |
ICIP | 4 |
| 2021 | Content-Independent Online Handwriting Verification Based on Multi-Modal FusionabstractUser identity authentication is essencial for ensuring information security. With the widespread use of electronic devices, online handwriting verification becomes more important in identity authentication based on biometrics and widely used in financial, commercial, and forensic fields. In this paper, we propose a multi-path feature fusion network for multi-modal fusion of static and dynamic handwriting obtained by electronic devices to intensify the handwriting verification. Since traditional handwritten signature verification, of which the handwritten content just the writer’s name, is vulnerable to skilled forgery attacks, we propose a content-independent handwriting verification scheme to solve this problem. We also build a handwriting dataset with approximately 5400 samples of 30 individuals’ handwriting, which contributes to extracting content-independent handwriting style features. We test our method on widely used BiosecurID dataset and our dataset. The experimental results demonstrate the feasibility of the proposed method. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Zhenchao Jin, Nenghai Yu |
ICME | 5 |
| 2021 | Efficient Open-Set Adversarial Attacks on Deep Face RecognitionabstractDifferent from close-set classification task, deep face recognition models are often used in open-set scenarios, where the models need to handle arbitrary faces. Open-set adversarial attacks can identify the vulnerability of deep face recognition models. Compared to time-consuming iterative gradient-based methods, generator-based methods can produce adversarial examples with only one forward pass, which greatly improves attack efficiency. However, existing generator-based attack methods need to train an individual model for each target identity and can only generate a fixed perturbation pattern regardless of different attack intensity constraints, which is impractical and sub-optimal for open-set adversarial attacks. In this paper, we propose an efficient generator-based Single Model ARbitrary Target (SMART) approach for open-set adversarial attacks against deep face recognition models. Given an arbitrary source-target face image pair, SMART first generates an additive perturbation and then adds it to the source image to obtain the final adversarial face image. After the training with various source-target pairs randomly sampled on large scale face images, SMART could effectively learn inherent perturbation patterns for arbitrary source-target face images pairs. Besides, we also propose a novel Constraint-aware Adversarial Decoder (CAD) module, which makes SMART the first generator-based method that could produce adaptive adversarial patterns according to different constraints on attack intensity. Extensive experimental results in various settings demonstrate the effectiveness of the proposed method. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
ICME | 2 |
| 2021 | Towards Generalizable and Robust Face Manipulation Detection via Bag-of-featureabstractOver the past several years, to solve the problem of malicious abuse of facial manipulation technology, face manipulation detection technology has obtained considerable attention and achieved remarkable progress. However, most existing methods have very impoverished generalization ability and robustness. In this paper, we propose a novel method for face manipulation detection, which can improve the generalization ability and ro-bustness by bag-of-feature. Specifically, we extend Transformers using bag-of-feature approach to encode inter-patch relation-ships, allowing it to learn forgery features without any additional mask supervision. Extensive experiments demonstrate that our method can outperform competing for state-of-the-art methods on FaceForensics++, Celeb-DF and DeeperForensics-l.0 datasets. Changtao Miao, Qi Chu 0001, Weihai Li, Wanyi Zhuang, Nenghai Yu |
VCIP | 2 |
| 2021 | Real Time Video Object Segmentation in Compressed DomainabstractMany of the recent methods for semi-supervised video object segmentation are still far from being applicable for real time applications due to their slow inference speed. Therefore, we explore a propagation based segmentation method in compressed domain to accelerate inference speed in this paper. In particular, we only extract the features of I-frames by traditional deep convolutional neural network and produce the features of P-frames through information flow propagation. In the process of feature propagation, we propose two effective components to enhance the representation ability of simply warped features in terms of appearance and location. Specifically, we propose a residual supplement module to supplement appearance information which is lost in direct warping and a spatial attention module that can mine extra spatial saliency to provide the location information of the specified object. Besides, we propose a metric based decoder module which consists of a feature match module and a multi-level refinement module to transform information from semantic representation to shape segmentation mask. Extensive experiments on several video datasets demonstrate that the proposed method can achieve comparable accuracy while much faster inference speed when compared to the state-of-the-art algorithms. Zhentao Tan, Bin Liu 0016, Qi Chu 0001, Hangshi Zhong, Weihai Li, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | DASOT: A Unified Framework Integrating Data Association and Single Object Tracking for Online Multi-Object TrackingabstractIn this paper, we propose an online multi-object tracking (MOT) approach that integrates data association and single object tracking (SOT) with a unified convolutional network (ConvNet), named DASOTNet. The intuition behind integrating data association and SOT is that they can complement each other. Following Siamese network architecture, DASOTNet consists of the shared feature ConvNet, the data association branch and the SOT branch. Data association is treated as a special re-identification task and solved by learning discriminative features for different targets in the data association branch. To handle the problem that the computational cost of SOT grows intolerably as the number of tracked objects increases, we propose an efficient two-stage tracking method in the SOT branch, which utilizes the merits of correlation features and can simultaneously track all the existing targets within one forward propagation. With feature sharing and the interaction between them, data association branch and the SOT branch learn to better complement each other. Using a multi-task objective, the whole network can be trained end-to-end. Compared with state-of-the-art online MOT methods, our method is much faster while maintaining a comparable performance. Qi Chu 0001, Wanli Ouyang, Bin Liu 0016, Feng Zhu 0006, Nenghai Yu |
AAAI | 1 |
| 2020 | Density-Aware Graph for Deep Semi-Supervised Visual RecognitionabstractSemi-supervised learning (SSL) has been extensively studied to improve the generalization ability of deep neural networks for visual recognition. To involve the unlabelled data, most existing SSL methods are based on common density-based cluster assumption: samples lying in the same high-density region are likely to belong to the same class, including the methods performing consistency regularization or generating pseudo-labels for the unlabelled images. Despite their impressive performance, we argue three limitations exist: 1) Though the density information is demonstrated to be an important clue, they all use it in an implicit way and have not exploited it in depth. 2) For feature learning, they often learn the feature embedding based on the single data sample and ignore the neighborhood information. 3) For label-propagation based pseudo-label generation, it is often done offline and difficult to be end-to-end trained with feature learning. Motivated by these limitations, this paper proposes to solve the SSL problem by building a novel density-aware graph, based on which the neighborhood information can be easily leveraged and the feature learning and label propagation can also be trained in an end-to-end way. Specifically, we first propose a new Density-aware Neighborhood Aggregation(DNA) module to learn more discriminative features by incorporating the neighborhood information in a density-aware manner. Then a novel Density-ascending Path based Label Propagation(DPLP) module is proposed to generate the pseudo-labels for unlabeled samples more efficiently according to the feature distribution characterized by density. Finally, the DNA module and DPLP module evolve and improve each other end-to-end. Extensive experiments demonstrate the effectiveness of the newly proposed density-aware graph based SSL framework and our approach can outperform current state-of-the-art methods by a large margin. Suichan Li, Bin Liu 0016, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Nenghai Yu |
CVPR | 4 |
| 2020 | Cross-Modality Person Re-Identification With Shared-Specific Feature TransferabstractCross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning modality-shared representation by embedding different modalities into a same feature space, lowering the upper bound of feature distinctiveness. In this paper, we tackle the above limitation by proposing a novel cross-modality shared-specific feature transfer algorithm (termed cm-SSFT) to explore the potential of both the modality-shared information and the modality-specific characteristics to boost the reidentification performance. We model the affinities of different modality samples according to the shared features and then transfer both shared and specific features among and across modalities. We also propose a complementary feature learning strategy including modality adaption, project adversarial learning and reconstruction enhancement to learn discriminative and complementary shared and specific features of each modality, respectively. The entire cmSSFTalgorithm can be trained in an end-to-end manner. We conducted comprehensive experiments to validate the superiority ofthe overall algorithm and the effectiveness ofeach component. The proposed algorithm significantly outperforms state-of-the-arts by 22.5% and 19.3% mAP on the two mainstream benchmark datasets SYSU-MM01 and RegDB, respectively. Yan Lu 0001, Bin Liu 0016, Tianzhu Zhang 0001, Baopu Li, Qi Chu 0001, Nenghai Yu |
CVPR | 6 |
| 2020 | GSM: Graph Similarity Model for Multi-Object TrackingabstractThe popular tracking-by-detection paradigm for multi-object tracking (MOT) focuses on solving data association problem, of which a robust similarity model lies in the heart. Most previous works make effort to improve feature representation for individual object while leaving the relations among objects less explored, which may be problematic in some complex scenarios. In this paper, we focus on leveraging the relations among objects to improve robustness of the similarity model. To this end, we propose a novel graph representation that takes both the feature of individual object and the relations among objects into consideration. Besides, a graph matching module is specially designed for the proposed graph representation to alleviate the impact of unreliable relations. With the help of the graph representation and the graph matching module, the proposed graph similarity model, named GSM, is more robust to the occlusion and the targets sharing similar appearance. We conduct extensive experiments on challenging MOT benchmarks and the experimental results demonstrate the effectiveness of the proposed method. Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IJCAI | 2 |
| 2020 | SAFNet: A Semi-Anchor-Free Network With Enhanced Feature Pyramid for Object DetectionabstractIn recent years, the field of object detection has made significant progress. The success of most of the state-of-the-art object detectors is derived from the use of feature pyramid and the carefully designed anchor boxes. However, the current methods of constructing feature pyramid usually blindly integrate multi-scale representations on each feature hierarchy. Furthermore, these detectors also suffer from some drawbacks brought by the hand-designed anchors. To mitigate the adverse effects caused thereby, we introduce a one-stage object detector, named as the semi-anchor-free network with enhanced feature pyramid (SAFNet). Specifically, to better construct feature pyramid, we propose a novel enhanced feature pyramid generation paradigm, which mainly consists of two modules, i.e., adaptive feature fusion module (AFFM) and self-enhanced module (SEM). The paradigm adaptively integrates multi-scale representations in a non-linear method meanwhile suppress the redundant semantic information for each pyramid level, such that a clean and enhanced feature pyramid could be obtained. In addition, an adaptive anchor generator (AAG) is designed to yield fewer but more suitable anchor boxes for each input image. Benefiting from the enhanced feature pyramid, AAG is capable of generating more accurate anchor boxes by introducing few priors. Thus, AAG has the ability to alleviate the drawbacks caused by the preset anchor hyper-parameters and helps to decrease the computation cost. Extensive experiments demonstrate the effectiveness of our approach. Profited from the proposed modules, SAFNet significantly boosts the detection performance, i.e., achieving 2 points and 2.1 points higher Average Precision (AP) than RetinaNet (our baseline) on PASCAL VOC and MS COCO respectively. Codes will be publicly available soon. Zhenchao Jin, Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
IEEE Trans. Image Process. | 3 |
| 2020 | MichiGAN: multi-input-conditioned hair image generation for portrait editingabstractDespite the recent success of face image generation with GANs, conditional hair editing remains challenging due to the under-explored complexity of its geometry and appearance. In this paper, we present MichiGAN (Multi-Input-Conditioned Hair Image GAN), a novel conditional image generation method for interactive portrait hair manipulation. To provide user control over every major hair visual factor, we explicitly disentangle hair into four orthogonal attributes, including shape, structure, appearance, and background. For each of them, we design a corresponding condition module to represent, process, and convert user inputs, and modulate the image generation pipeline in ways that respect the natures of different visual attributes. All these condition modules are integrated with the backbone generator to form the final end-to-end network, which allows fully-conditioned hair generation from multiple user inputs. Upon it, we also build an interactive portrait hair editing system that enables straightforward manipulation of hair by projecting intuitive and high-level user inputs such as painted masks, guiding strokes, or reference photos to well-defined condition representations. Through extensive experiments and evaluations, we demonstrate the superiority of our method regarding both result quality and user controllability. Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Lu Yuan 0001, Sergey Tulyakov, Nenghai Yu |
ACM Trans. Graph. | 5 |
| 2019 | Using multi-label classification to improve object detection
Bin Liu 0016, Qi Chu 0001, Nenghai Yu |
Neurocomputing | 3 |
| 2017 | Online Multi-object Tracking Using CNN-Based Single Object Tracker with Spatial-Temporal Attention MechanismabstractIn this paper, we propose a CNN-based framework for online MOT. This framework utilizes the merits of single object trackers in adapting appearance models and searching for target in the next frame. Simply applying single object tracker for MOT will encounter the problem in computational efficiency and drifted results caused by occlusion. Our framework achieves computational efficiency by sharing features and using ROI-Pooling to obtain individual features for each target. Some online learned target-specific CNN layers are used for adapting the appearance model for each target. In the framework, we introduce spatial-temporal attention mechanism (STAM) to handle the drift caused by occlusion and interaction among targets. The visibility map of the target is learned and used for inferring the spatial attention map. The spatial attention map is then applied to weight the features. Besides, the occlusion status can be estimated from the visibility map, which controls the online updating process via weighted loss on training samples with different occlusion statuses in different frames. It can be considered as temporal attention mechanism. The proposed algorithm achieves 34.3% and 46.0% in MOTA on challenging MOT15 and MOT16 benchmark dataset respectively. Qi Chu 0001, Wanli Ouyang, Hongsheng Li 0001, Xiaogang Wang 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 1 |
| 2016 | Consistent matching based on boosted salience channels for group re-identificationabstractAssociating groups of people across non-overlapping camera views is an important but unsolved problem. Compared with the similar person re-identification task, group re-identification introduces some new challenges, such as significant deformation in uncontrolled directions, great intra-group occlusions and so on. In this paper, we propose a novel patch matching based framework for group re-identification. Discriminative salience channels are learned to filter out highly unreliable and non-informative patch matches between two group images, while retain true matches undergoing appearance variations. The resulting candidate correspondences are further explored by the proposed consistent matching process, which prefers coherent matches in true group image pairs. The effectiveness of our approach is validated on two group re-identification datasets: ZeCSS and i-LIDS MCTS. It outperforms state-of-the-art methods on both datasets. Feng Zhu 0006, Qi Chu 0001, Nenghai Yu |
ICIP | 2 |