VLDB 2026 Research / reviewers in the wild / expert
Wei Zhai
dblp:189/3967
· DBLP profile ↗
61ranked-venue papers
13as first author
51since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 7 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E-MaT: Event-oriented Mamba for Egocentric Point TrackingabstractEgocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods struggle in highly dynamic settings—a common challenge in first-person perspectives, where the head-mounted camera undergoes frequent and abrupt rotations, resulting in high angular velocities, motion blur, and large inter-frame displacements. In contrast, event cameras capture motion at microsecond temporal resolution, naturally avoiding blur and delivering low-latency, high-fidelity cues crucial for egocentric point tracking. Moreover, rapid egocentric motion disrupts local smoothness, breaking the assumption that spatially adjacent regions share similar motion. Event dynamics expose global motion trends, guiding coherent modeling and consistent feature flow. Therefore, this paper proposes a mamba-based tracking framework that constructs feature modeling paths aligned with the dominant motion trend extracted from events, and modulates feature propagation along these paths based on local motion intensity, enhancing stability by suppressing unreliable signals and emphasizing consistent cues. Additionally, a motion-adaptive suppression module enhances temporal robustness by adaptively suppressing correlation features based on motion intensity variations, mitigating the effects of intensity fluctuations and partial observability. To facilitate research in this domain, a multimodal dataset named DVS-EgoPoints with both events and videos for egocentric point tracking is collected. Experiments on the DVS-EgoPoints dataset and a simulation benchmark demonstrate superior performance over state-of-the-art methods, especially under challenging motion and occlusion conditions. Wei Zhai, Yang Cao 0010, Bin Li 0025, Zhengjun Zha |
AAAI | 2 |
| 2026 | AnchorSeg: Language Grounded Query Banks for Reasoning SegmentationabstractRui Qian, Chuanhang Deng, Qiang Huang, Jian Xiong, Mingxuan Li, Yingbo Zhou, Wei Zhai, Jintao Chen, Dejing Dou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Qian 0002, Chuanhang Deng, Wei Zhai, Dejing Dou |
ACL (1) | 7 |
| 2026 | Visual-Geometric Collaborative Guidance for Affordance Learning
Hongchen Luo, Wei Zhai, Jiao Wang 0002, Yang Cao 0010, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2026 | SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for UAVabstractUncrewed aerial vehicles (UAV) are increasingly deployed to assist humans in diverse tasks, where understanding human intentions is critical to effective collaboration. Referring expression comprehension (REC) links language to visual targets, allowing UAV to recognize human-intended targets of interest, thereby supporting subsequent actions. However, existing REC research is almost exclusively confined to ground-based scenarios, leaving aerial scenarios largely unexplored. In this paper, we formally define UAV-based REC as a new research problem and highlight its unique challenges, including abundant background interference, small target size, and complex referring relations. To enable systematic study, we introduce SkyFind, a large-scale dataset with one million high-quality target-expression pairs, providing a solid foundation. In addition, we propose AerialREC, a baseline framework that reduces background interference in UAV imagery by searching for a potential target region before localization. We establish benchmark results on SkyFind using ten representative REC methods and validate the effectiveness of the AerialREC framework. Guanbo Wu, Xueyang Fu, Kean Liu, Xin Lu 0008, Chengjie Ge, Wei Zhai, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) enables the inspection of unseen objects by bridging textual prompts and visual features, showing great potential in flexible manufacturing. While existing ZSAD methods rely on predefined prompts and struggle with unseen defects, Multimodal Large Language Models (MLLMs) offer promising solutions through their generative and interpretative capabilities. However, adapting MLLMs to Industrial Anomaly Detection (IAD) remains challenging due to fine-grained anomaly patterns and subtle visual distinctions. We propose VMAD (Visual-enhanced MLLM Anomaly Detection), a framework that enriches MLLM with visual IAD knowledge through two key components: a Defect-Sensitive Structure Learning scheme that transfers patch-similarities for improved discrimination, and a Locality-enhanced Token Compression that leverages multi-level local features for fine-grained detection. We also introduce RIAD, a comprehensive IAD dataset with detailed anomaly annotations. Extensive experiments on MVTec-AD, Visa, WFDD, and RIAD demonstrate VMAD’s superior performance. The dataset and code will be publicly available at https://github.com/denghuilin-cyber/VMAD. Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Corrections to "VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection"abstractIn the above article [1], an earlier draft of Fig. 6 was inadvertently included. The correct Fig. 6 is presented on the next page.Fig. 6.Zero-shot anomaly segmentation on MVTec-AD, WFDD, and ViSA datasets. Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Toward Better De-Raining Generalization via Rainy Characteristics Memorization and ReplayabstractCurrent image de-raining methods primarily learn from a limited dataset, leading to inadequate performance in varied real-world rainy conditions. To tackle this, we introduce a new framework that enables networks to progressively expand their de-raining knowledge base by tapping into a growing pool of datasets, significantly boosting their adaptability. Drawing inspiration from the human brain's ability to continually absorb and generalize from ongoing experiences, our approach borrows the mechanism of the complementary learning system. Specifically, we first deploy generative adversarial networks (GANs) to capture and retain the unique features of new data, mirroring the hippocampus's role in learning and memory. Then, the de-raining network is trained with both existing and GAN-synthesized data, mimicking the process of hippocampal replay and interleaved learning. Furthermore, we employ knowledge distillation with the replayed data to replicate the synergy between the neocortex's activity patterns triggered by hippocampal replays and the preexisting neocortical knowledge. This comprehensive framework empowers the de-raining network to accumulate knowledge from various datasets, continually enhancing its performance on previously unseen rainy scenes. Our testing on three benchmark de-raining networks confirms the framework's effectiveness. It not only facilitates continual knowledge accumulation across six datasets but also surpasses state-of-the-art methods in generalizing to new real-world scenarios. Our code is available at https://github.com/wangkunyu241/CLGID. Xueyang Fu, Chengzhi Cao, Chengjie Ge, Wei Zhai, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support HotlinesabstractMental health is a significant global public health issue, and psychological support hotlines play a crucial role in providing mental health assistance and identifying suicide risks at an early stage. However, the emotional expressions conveyed during these calls remain underexplored in current research. This study introduces a novel method that combines pitch acoustic features with deep learning-based features to analyze and understand emotions expressed during hotline interactions. Using data from China's largest psychological support hotline, which includes 105 subjects, our method achieved an F1-score of 79.13% for negative binary emotion classification. Additionally, the proposed approach was validated on an open dataset for multi-class emotion classification, where it demonstrated better performance compared to the state-of-the-art methods. To explore its clinical relevance, we applied the model to analysis the frequency of negative emotions and the rate of emotional change in the conversation, comparing 46 subjects with suicidal behavior to those without. While the suicidal group exhibited more frequent emotional changes than the non-suicidal group, the difference was not statistically significant. Importantly, our findings suggest that emotional fluctuation intensity and frequency could serve as novel features for psychological assessment scales and suicide risk prediction. The proposed method provides valuable insights into emotional dynamics and has the potential to advance early intervention and improve suicide prevention strategies through integration with clinical tools and assessments. The source code is publicly available at: https://github.com/Sco-field/Speechemotionrecognition/tree/main. Han Wang 0059, Jianqiang Li 0002, Qing Zhao 0005, Zhonglong Chen, Changwei Song, Yuning Huang, Wei Zhai, Yongsheng Tong, Guanghui Fu |
COMPSAC | 8 |
| 2025 | Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image CaptioningabstractGenerating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this paper, we introduce a detailed caption benchmark, termed as CompreCap, to evaluate the visual context from a directed scene graph view. Concretely, we first manually segment the image into semantically meaningful regions (i.e., semantic segmentation mask) according to common-object vocabulary, while also distinguishing attributes of objects within all those regions. Then directional relation labels of these objects are annotated to compose a directed scene graph that can well encode rich compositional information of the image. Based on our directed scene graph, we develop a pipeline to assess the generated detailed captions from LVLMs on multiple levels, including the object-level coverage, the accuracy of attribute descriptions, the score of key relationships, etc. Experimental results on the CompreCap dataset confirm that our evaluation method aligns closely with human evaluation scores across LVLMs. We have released the code and the dataset here to support the community. Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu 0001, Wei Zhai, Yang Cao 0010, Yujun Shen, Zhengjun Zha |
CVPR | 7 |
| 2025 | GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance GroundingabstractOpen-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages that depict interactions with 3D geometries to introduce external interaction priors. However, they are still vulnerable to a limited semantic space by failing to leverage implied invariant geometries and potential interaction intentions. Normally, humans address complex tasks through multi-step reasoning and respond to diverse situations by leveraging associative and analogical thinking. In light of this, we propose GREAT (GeometRy-intEntion collAboraTive inference) for Open-Vocabulary 3D Object Affordance Grounding, a novel framework that mines the object invariant geometry attributes and performs analogically reason in potential interaction scenarios to form affordance knowledge, fully combining the knowledge with both geometries and visual contents to ground 3D object affordance. Besides, we introduce the Point Image Affordance Dataset v2 (PIADv2), the largest 3D object affordance dataset at present to support the task. Extensive experiments demonstrate the effectiveness and superiority of GREAT. The code and dataset are available at https://yawen-shao.github.io/GREAT/. Yawen Shao, Wei Zhai, Yuhang Yang 0002, Hongchen Luo, Yang Cao 0010, Zhengjun Zha |
CVPR | 2 |
| 2025 | Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningabstractContinual test-time adaptive object detection (CTTA-OD) aims to online adapt a source pre-trained detector to everchanging environments during inference under continuous domain shifts. Most existing CTTA-OD methods prioritize effectiveness while overlooking computational efficiency, which is crucial for resource-constrained scenarios. In this paper, we propose an efficient CTTA-OD method via pruning. Our motivation stems from the observation that not all learned source features are beneficial; certain domain-sensitive feature channels can adversely affect target domain performance. Inspired by this, we introduce a sensitivity-guided channel pruning strategy that quantifies each channel based on its sensitivity to domain discrepancies at both image and instance levels. We apply weighted sparsity regularization to selectively suppress and prune these sensitive channels, focusing adaptation efforts on invariant ones. Additionally, we introduce a stochastic channel reactivation mechanism to restore pruned channels, enabling recovery of potentially useful features and mitigating the risks of early pruning. Extensive experiments on three benchmarks show that our method achieves superior adaptation performance while reducing computational overhead by 12% in FLOPs compared to the recent SOTA method. Xueyang Fu, Xin Lu 0008, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha |
CVPR | 6 |
| 2025 | Improved Video VAE for Latent Video Diffusion ModelabstractVariational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI’s Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial compression, this paper presents two astonishing findings: (1) The initialization from a well-trained image VAE with the same latent dimensions is not an optimal scheme. (2) The adoption of causal reasoning leads to unequal information interactions and unbalanced performance between frames. To alleviate these problems, we propose a keyframe-based temporal compression (KTC) architecture and a group causal convolution (GCConv) module to further improve video VAE (IV-VAE). Specifically, the KTC architecture divides the latent space into two branches, in which one half completely inherits the compression prior of keyframes from a lower-dimension image VAE while the other half involves temporal compression through 3D group causal convolution, reducing temporal-spatial conflicts and accelerating the convergence speed of video VAE. The GC-Conv in the above 3D half uses standard convolution within each frame group to ensure inter-frame equivalence, and employs causal logical padding between groups to maintain flexibility in processing variable frame video. Extensive experiments on five benchmarks demonstrate the SOTA video reconstruction and generation abilities of our IV-VAE. Pingyu Wu, Kai Zhu 0004, Yu Liu 0063, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
CVPR | 5 |
| 2025 | MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic ModelingabstractRecent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, we have identified that recent methods suffer from loss of image information during understanding task, due to either image discretization or diffusion de-noising steps. To address this issue, we propose a novel Multi-Modal Auto-Regressive (MMAR) probabilistic modeling framework. Unlike discretization line of method, MMAR takes in continuous-valued image tokens to avoid information loss in an efficient way. Differing from diffusion-based approaches, we disentangle the diffusion process from auto-regressive backbone model by employing a lightweight diffusion head on top each auto-regressed image patch embedding. In this way, when the model transits from image generation to understanding through text generation, the backbone model’s hidden representation of the image is not limited to the last denoising step. To successfully train our method, we also propose a theoretically proven technique that addresses the numerical stability issue and a training strategy that balances the generation and understanding task goals. Extensive evaluations on 18 image understanding benchmarks show that MMAR significantly outperforms most of the existing joint multi-modal models, surpassing the method that employs pre-trained CLIP vision encoder. Meanwhile, MMAR is able to generate high quality images. We also show that our method is scalable with larger data and model size. Jian Yang 0003, Dacheng Yin, Yizhou Zhou, Fengyun Rao, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
CVPR | 5 |
| 2025 | MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social MediaabstractWei Zhai, Nan Bai, Qing Zhao, Jianqiang Li, Fan Wang, Hongzhi Qi, Meng Jiang, Xiaoqin Wang, Bing Xiang Yang, Guanghui Fu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wei Zhai, Nan Bai, Qing Zhao 0005, Jianqiang Li 0002, Hongzhi Qi, Bing Xiang Yang, Guanghui Fu |
EMNLP | 1 |
| 2025 | MATE: Motion-Augmented Temporal Consistency for Event-Based Point Tracking
Wei Zhai, Yang Cao 0010, Bin Li 0025, Zhengjun Zha |
ICCV | 2 |
| 2025 | Emotive: Event-Guided Trajectory Modeling for 3D Motion Estimation
Zengyu Wan, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
ICCV | 2 |
| 2025 | SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assetsabstract3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall into several paradigms: optimization-based and feed-forward (both single-view regression and multi-view generation with reconstruction). However, they are limited by slow speed, low quality, cascade reasoning, and ambiguity in mapping low-dimensional planes to high-dimensional space due to occlusion and invisibility, respectively. Furthermore, existing 3D human assets remain small-scale, insufficient for large-scale training. To address these challenges, we propose a latent space generation paradigm for 3D human digitization, which involves compressing multi-view images into Gaussians via a UV-structured VAE, along with DiT-based conditional generation, we transform the ill-posed low-to-high-dimensional mapping problem into a learnable distribution shift, which also supports end-to-end inference. In addition, we employ the multi-view optimization approach combined with synthetic data to construct the HGS-1M dataset, which contains $1$ million 3D Gaussian assets to support the large-scale training. Experimental results demonstrate that our paradigm, powered by large-scale training, produces high-quality 3D human Gaussians with intricate textures, facial details, and loose clothing deformation. Yuhang Yang 0002, Fengqi Liu, Yixing Lu, Pingyu Wu, Wei Zhai, Ran Yi 0002, Yang Cao 0010, Lizhuang Ma, Zhengjun Zha, Junting Dong |
ICCV | 6 |
| 2025 | HERO: Human Reaction Generation from VideosabstractHuman reaction generation represents a significant research domain for interactive AI, as humans constantly interact with their surroundings. Previous works focus mainly on synthesizing the reactive motion given a human motion sequence. This paradigm limits interaction categories to human-human interactions and ignores emotions that may influence reaction generation. In this work, we propose to generate 3D human reactions from RGB videos, which involves a wider range of interaction categories and naturally provides information about expressions that may reflect the subject's emotions. To cope with this task, we present HERO, a simple yet powerful framework for Human rEaction geneRation from videOs. HERO considers both global and frame-level local representations of the video to extract the interaction intention, and then uses the extracted interaction intention to guide the synthesis of the reaction. Besides, local visual representations are continuously injected into the model to maximize the exploitation of the dynamic properties inherent in videos. Furthermore, the ViMo dataset containing paired Video-Motion data is collected to support the task. In addition to human-human interactions, these video-motion pairs also cover animal-human interactions and scene-human interactions. Extensive experiments demonstrate the superiority of our methodology. The code and dataset will be publicly available at https://jackyu6.github.io/HERO. Chengjun Yu, Wei Zhai, Yuhang Yang 0002, Yang Cao 0010, Zhengjun Zha |
ICCV | 2 |
| 2025 | ViewPoint: Panoramic Video Generation with Pretrained Diffusion ModelsabstractPanoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspective data, which constitutes the majority of the training data for modern diffusion models. In this paper, we propose a novel framework utilizing pretrained perspective video models for generating panoramic videos. Specifically, we design a novel panorama representation named ViewPoint map, which possesses global spatial continuity and fine-grained visual details simultaneously. With our proposed Pano-Perspective attention mechanism, the model benefits from pretrained perspective priors and captures the panoramic spatial correlations of the ViewPoint map effectively. Extensive experiments demonstrate that our method can synthesize highly dynamic and spatially consistent panoramic videos, achieving state-of-the-art performance and surpassing previous methods. Zixun Fang, Kai Zhu 0004, Yu Liu 0063, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 5 |
| 2025 | EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian SplattingabstractScene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vision, record pixel-wise intensity changes asynchronously with high temporal resolution and low latency, providing valuable scene and motion information in blind inter-frame intervals. In this paper, we introduce the event cameras to aid scene construction from a casually captured video for the first time, and propose Event-Aided Free-Trajectory 3DGS, called EF-3DGS, which seamlessly integrates the advantages of event cameras into 3DGS through three key components. First, we leverage the Event Generation Model (EGM) to fuse events and frames, enabling continuous supervision between discrete frames. Second, we extract motion information through Contrast Maximization (CMax) of warped events, which calibrates camera poses and provides gradient-domain constraints for 3DGS. Third, to address the absence of color information in events, we combine photometric bundle adjustment (PBA) with a Fixed-GS training strategy that separates structure and color optimization, effectively ensuring color consistency across different views. We evaluate our method on the public Tanks and Temples benchmark and a newly collected real-world dataset, RealEv-DAVIS. Our method achieves up to 3dB higher PSNR and 40% lower Absolute Trajectory Error (ATE) compared to state-of-the-art methods under challenging high-speed scenarios. Bohao Liao, Wei Zhai, Zengyu Wan, Zhixin Cheng, Wenfei Yang, Yang Cao 0010, Tianzhu Zhang 0001, Zhengjun Zha |
NeurIPS | 2 |
| 2025 | PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) aims to online adapt a pre-trained model to changing environments during inference. Most existing methods focus on exploiting target data, while overlooking another crucial source of information, the pre-trained weights, which encode underutilized domain-invariant priors. This paper takes the geometric attributes of pre-trained weights as a starting point, systematically analyzing three key components: magnitude, absolute angle, and pairwise angular structure. We find that the pairwise angular structure remains stable across diverse corrupted domains and encodes domain-invariant semantic information, suggesting it should be preserved during adaptation. Based on this insight, we propose PAID (Pairwise Angular Invariant Decomposition), a prior-driven CTTA method that decomposes weight into magnitude and direction, and introduces a learnable orthogonal matrix via Householder reflections to globally rotate direction while preserving the pairwise angular structure. During adaptation, only the magnitudes and the orthogonal matrices are updated. PAID achieves consistent improvements over recent SOTA methods on four widely used CTTA benchmarks, demonstrating that preserving pairwise angular structure offers a simple yet effective principle for CTTA. Our code is available at https://github.com/wangkunyu241/PAID. Xueyang Fu, Yuanfei Bao, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha |
NeurIPS | 6 |
| 2024 | Hypercorrelation Evolution for Video Class-Incremental LearningabstractVideo class-incremental learning aims to recognize new actions while restricting the catastrophic forgetting of old ones, whose representative samples can only be saved in limited memory. Semantically variable subactions are susceptible to class confusion due to data imbalance. While existing methods address the problem by estimating and distilling the spatio-temporal knowledge, we further explores that the refinement of hierarchical correlations is crucial for the alignment of spatio-temporal features. To enhance the adaptability on evolved actions, we proposes a hierarchical aggregation strategy, in which hierarchical matching matrices are combined and jointly optimized to selectively store and retrieve relevant features from previous tasks. Meanwhile, a correlation refinement mechanism is presented to reinforce the bias on informative exemplars according to online hypercorrelation distribution. Experimental results demonstrate the effectiveness of the proposed method on three standard video class-incremental learning benchmarks, outperforming state-of-the-art methods. Code is available at: https://github.com/Lsen991031/HCE Sen Liang, Kai Zhu 0004, Wei Zhai, Yang Cao 0010 |
AAAI | 3 |
| 2024 | LEMON: Learning 3D Human-Object Interaction Relation from 2D ImagesabstractLearning 3D human-object interaction relation is piv-otal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements, e.g., human contact, object affordance, and human-object spatial relation, primarily from the perspective of either the human or the object. Which underexploit certain correlations between the interaction counterparts (human and object), and struggle to address the uncertainty in interactions. Actually, objects' functionalities potentially affect humans' interaction intentions, which reveals what the interaction is. Mean-while, the interacting humans and objects exhibit matching geometric structures, which presents how to interact. In light of this, we propose harnessing these inherent correlations between interaction counterparts to mitigate the uncertainty and jointly anticipate the above interaction el-ements in 3D space. To achieve this, we present LEMON (LEarning 3D huMan-Object iNteraction relation), a unified model that mines interaction intentions of the counter-parts and employs curvatures to guide the extraction of ge-ometric correlations, combining them to anticipate the interaction elements. Besides, the 3D Interaction Relation dataset (3DIR) is collected to serve as the test bed for training and evaluation. Extensive experiments demonstrate the superiority of LEMON over methods estimating each element in isolation. The code and dataset are available at https://yyvhang.github.io/LEMON. Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Zhengjun Zha |
CVPR | 2 |
| 2024 | Bidirectional Progressive Transformer for Interaction Intention Anticipation
Zichen Zhang 0022, Hongchen Luo, Wei Zhai, Yang Cao 0010, Yu Kang 0001 |
ECCV (59) | 3 |
| 2024 | Building Earthquake Damage Recognition Performance of Frequency Domain Texture FeaturesabstractBuilding collapse arising from destructive earthquakes is often the primary cause of casualties and economic loss. Building damage assessment is one of the top priorities in earthquake emergency work. Quad-polarimetric synthetic aperture radar (PolSAR) data not only has the advantages of radar imaging being neither exposed to sunlight nor blocked by clouds, but also contains the most abundant information of the four polarimetric channels. In many cases, the texture feature even outperforms other kinds of features. The texture features of buildings include not only spatial texture but also frequency texture. We proposed a parameter called the sector texture feature of the Fourier amplitude spectrum (STFFAS) to describe frequency-domain texture features based on the Fourier amplitude spectrum for building damage recognition. Our experimental results show that the recognition performance of the frequency texture feature performs satisfactorily for building damage information extraction. Wei Zhai, Yaxin Bi, Jianqing Du, Gangyu Yang |
IGARSS | 1 |
| 2024 | UniDense: Unleashing Diffusion Models with Meta-Routers for Universal Few-Shot Dense PredictionabstractUniversal few-shot dense prediction requires a versatile model capable of learning any dense prediction task from limited labeled images, which necessitates the model to possess efficient adaptation abilities. Prevailing few-shot learning methods rely on efficient fine-tuning of model weights for few-shot adaptation, which carries the risk of disrupting the pre-trained knowledge and lacks the capability to extract task-specific knowledge contained in the pre-trained model. To overcome these limitations, our paper approaches universal few-shot dense prediction from a novel perspective. Unlike conventional fine-tuning techniques that use all model parameters and modify a specific set of weights for few-shot adaptation, our method focuses on selecting task-relevant computation pathways of the pre-trained model while keeping the model weights frozen. Building upon this idea, we introduce a novel framework UniDense for universal few-shot dense prediction. First, we construct a versatile MoE (Mixture of Experts) architecture for dense prediction based on the Stable Diffusion model. We then utilize episodes-based meta-learning to train a set of routers for this MoE model, called Meta-Routers, which act as hyper-networks responsible for selecting computation blocks relevant to each task. We demonstrate that fine-tuning these meta-routers enables efficient few-shot adaptation of the entire model. Moreover, for each few-shot task, we leverage support samples to extract a task embedding, which serves as a conditioning factor for meta-routers. This strategy allows meta-routers to dynamically adapt themselves for different few-shot task, leading to improved adaptation performance. Experiments on a challenging variant of Taskonomy dataset with 10 dense prediction tasks demonstrate the superiority of our approach. Lintao Dong, Wei Zhai, Zhengjun Zha |
ACM Multimedia | 2 |
| 2024 | EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsabstractUnderstanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction specifically manifests in 3D space is also crucial, which links the perception and operation. Existing methods primarily leverage observations of HOI to capture interaction regions from an exocentric view. However, incomplete observations of interacting parties in the egocentric view introduce ambiguity between visual observations and interaction contents, impairing their efficacy. From the egocentric view, humans integrate the visual cortex, cerebellum, and brain to internalize their intentions and interaction concepts of objects, allowing for the pre-formulation of interactions and making behaviors even when interaction regions are out of sight. In light of this, we propose harmonizing the visual appearance, head motion, and 3D object to excavate the object interaction concept and subject intention, jointly inferring 3D human contact and object affordance from egocentric videos. To achieve this, we present EgoChoir, which links object structures with interaction contexts inherent in appearance and head motion to reveal object affordance, further utilizing it to model human contact. Additionally, a gradient modulation is employed to adopt appropriate clues for capturing interaction regions across various egocentric scenarios. Moreover, 3D contact and affordance are annotated for egocentric videos collected from Ego-Exo4D and GIMO to support the task. Extensive experiments on them demonstrate the effectiveness and superiority of EgoChoir. Yuhang Yang 0002, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 2 |
| 2024 | SOS-1K: A Fine-Grained Suicide Risk Classification Dataset for Chinese Social Media AnalysisabstractIn the social media, users frequently express personal emotions, a subset of which may indicate potential suicidal tendencies. The implicit and varied forms of expression in internet language complicate accurate and rapid identification of suicidal intent on social media, thus creating challenges for timely intervention efforts. The development of deep learning models for suicide risk detection is a promising solution, but there is a notable lack of relevant datasets, especially in the Chinese context. To address this gap, this study presents a Chinese social media dataset designed for fine-grained suicide risk classification, focusing on indicators such as expressions of suicide intent, methods of suicide, and urgency of timing. Seven pre-trained models were evaluated in two tasks: high and low suicide risk, and fine-grained suicide risk classification on a level of 0 to 10. In our experiments, deep learning models show good performance in distinguishing between high and low suicide risk, with the best model achieving an F1 score of 88.39%. However, the results for fine-grained suicide risk classification were still unsatisfactory, with the best weighted F1 score of 50.89%. To address the issues of data imbalance and limited dataset size, we investigated both traditional and advanced, large language model based data augmentation techniques, demonstrating that data augmentation can enhance this model performance by up to 4.65% points in F1-score. Notably, the Chinese MentalBERT model, which was pre-trained on psychological domain data, shows superior performance in both tasks. This study provides valuable insights for automatic identification of suicidal individuals, facilitating timely psychological intervention on social media platforms. The source code and data are publicly available at: https://github.com/HongzhiQ/FineGrainedSuicideDetection. Hongzhi Qi, Hanfei Liu, Jianqiang Li 0002, Qing Zhao 0005, Wei Zhai, Tian Yu He, Bing Xiang Yang, Guanghui Fu |
SMC | 5 |
| 2024 | Grounded Affordance from Exocentric View
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2024 | Background Activation Suppression for Weakly Supervised Object Localization and Semantic Segmentation
Wei Zhai, Pingyu Wu, Kai Zhu 0004, Yang Cao 0010, Feng Wu 0001, Zhengjun Zha |
Int. J. Comput. Vis. | 1 |
| 2024 | On Exploring Multiplicity of Primitives and Attributes for Texture Recognition in the WildabstractTexture recognition is a challenging visual task since its multiple primitives or attributes can be perceived from the texture image under different spatial contexts. Existing approaches predominantly built upon CNN incorporate rich local descriptors with orderless aggregation to capture invariance to the spatial layout. However, these methods ignore the inherent structure relation organized by primitives and the semantic concept described by attributes, which are critical cues for texture representation. In this paper, we propose a novel Multiple Primitives and Attributes Perception network (MPAP) that extracts features by modeling the relation of bottom-up structure and top-down attribute in a multi-branch unified framework. A bottom-up process is first proposed to capture the inherent relation of various primitive structures by leveraging structure dependency and spatial order information. Then, a top-down process is introduced to model the latent relation of multiple attributes by transferring attribute-related features between adjacent branches. Moreover, an augmentation module is devised to bridge the gap between high-level attributes and low-level structure features. MPAP can learn representation through jointing bottom-up and top-down processes in a mutually reinforced manner. Experimental results on six challenging texture datasets demonstrate the superiority of MPAP over state-of-the-art methods in terms of accuracy, robustness, and efficiency. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Haiyong Xie 0001, Dacheng Tao, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Event-Based Optical Flow via Transforming Into Motion-Dependent ViewabstractEvent cameras respond to temporal dynamics, helping to resolve ambiguities in spatio-temporal changes for optical flow estimation. However, the unique spatio-temporal event distribution challenges the feature extraction, and the direct construction of motion representation through the orthogonal view is less than ideal due to the entanglement of appearance and motion. This paper proposes to transform the orthogonal view into a motion-dependent one for enhancing event-based motion representation and presents a Motion View-based Network (MV-Net) for practical optical flow estimation. Specifically, this motion-dependent view transformation is achieved through the Event View Transformation Module, which captures the relationship between the steepest temporal changes and motion direction, incorporating these temporal cues into the view transformation process for feature gathering. This module includes two phases: extracting the temporal evolution clues by central difference operation in the extraction phase and capturing the motion pattern by evolution-guided deformable convolution in the perception phase. Besides, the MV-Net constructs an eccentric downsampling process to avoid response weakening from the sparsity of events in the downsampling stage. The whole network is trained end-to-end in a self-supervised manner, and the evaluations conducted on four challenging datasets reveal the superior performance of the proposed model compared to state-of-the-art (SOTA) methods. Zengyu Wan, Ganchao Tan, Yang Wang 0015, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
IEEE Trans. Image Process. | 4 |
| 2024 | Learning Visual Affordance Grounding From Demonstration VideosabstractVisual affordance grounding aims to segment all possible interaction regions between people and objects from an image/video, which benefits many applications, such as robot grasping and action recognition. Prevailing methods predominantly depend on the appearance feature of the objects to segment each region of the image, which encounters the following two problems: 1) there are multiple possible regions in an object that people interact with and 2) there are multiple possible human interactions in the same object region. To address these problems, we propose a hand-aided affordance grounding network (HAG-Net) that leverages the aided clues provided by the position and action of the hand in demonstration videos to eliminate the multiple possibilities and better locate the interaction regions in the object. Specifically, HAG-Net adopts a dual-branch structure to process the demonstration video and object image data. For the video branch, we introduce hand-aided attention to enhance the region around the hand in each video frame and then use the long short-term memory (LSTM) network to aggregate the action features. For the object branch, we introduce a semantic enhancement module (SEM) to make the network focus on different parts of the object according to the action classes and utilize a distillation loss to align the output features of the object branch with that of the video branch and transfer the knowledge in the video branch to the object branch. Quantitative and qualitative evaluations on two challenging datasets show that our method has achieved state-of-the-art results for affordance grounding. The source code is available at: https://github.com/lhc1224/HAG-Net. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Exploring Tuning Characteristics of Ventral Stream's Neurons for Few-Shot Image ClassificationabstractHuman has the remarkable ability of learning novel objects by browsing extremely few examples, which may be attributed to the generic and robust feature extracted in the ventral stream of our brain for representing visual objects. In this sense, the tuning characteristics of ventral stream's neurons can be useful prior knowledge to improve few-shot classification. Specifically, we computationally model two groups of neurons found in ventral stream which are respectively sensitive to shape cues and color cues. Then we propose the hierarchical feature regularization method with these neuron models to regularize the backbone of a few-shot model, thus making it produce more generic and robust features for few-shot classification. In addition, to simulate the tuning characteristic that neuron firing at a higher rate in response to foreground stimulus elements compared to background elements, which we call belongingness, we design a foreground segmentation algorithm based on the observation that the foreground object usually does not appear at the edge of the picture, then multiply the foreground mask with the backbone of few-shot model. Our method is model-agnostic and can be applied to few-shot models with different backbones, training paradigms and classifiers. Lintao Dong, Wei Zhai, Zhengjun Zha |
AAAI | 2 |
| 2023 | Uncertainty-Aware Optimal Transport for Semantically Coherent Out-of-Distribution DetectionabstractSemantically coherent out-of-distribution (SCOOD) detection aims to discern outliers from the intended data distribution with access to unlabeled extra set. The coexistence of in-distribution and out-of-distribution samples will exacerbate the model overfitting when no distinction is made. To address this problem, we propose a novel uncertainty-aware optimal transport scheme. Our scheme consists of an energy-based transport (ET) mechanism that estimates the fluctuating cost of uncertainty to promote the assignment of semantic-agnostic representation, and an inter-cluster extension strategy that enhances the discrimination of semantic property among different clusters by widening the corresponding margin distance. Furthermore, a T-energy score is presented to mitigate the magnitude gap between the parallel transport and classifier branches. Extensive experiments on two standard SCOOD benchmarks demonstrate the above-par OOD detection performance, outperforming the state-of-the-art methods by a margin of 27.69% and 34.4% on FPR@95, respectively. Code is available at https://github.com/LuFan3/IET-OOD. Kai Zhu 0004, Wei Zhai, Kecheng Zheng, Yang Cao 0010 |
CVPR | 3 |
| 2023 | Leverage Interactive Affinity for Affordance LearningabstractPerceiving potential “action possibilities” (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assignment paradigm and presume that there is a unique relationship between functional region and affordance label, yielding poor performance when adapting to unseen environments with large appearance variations. In this paper, we propose to leverage interactive affinity for affordance learning, i.e. extracting interactive affinity from human-object interaction and transferring it to non-interactive objects. Interactive affinity, which represents the contacts between different parts of the human body and local regions of the target object, can provide inherent cues of interconnectivity between humans and objects, thereby reducing the ambiguity of the perceived action possibilities. Specifically, we propose a pose-aided interactive affinity learning framework that exploits human pose to guide the network to learn the interactive affinity from human-object interactions. Particularly, a keypoint heuristic perception (KHP) scheme is devised to exploit the keypoint association of human pose to alleviate the uncertainties due to interaction diversities and contact occlusions. Besides, a contact-driven affordance learning (CAL) dataset is constructed by collecting and labeling over 5, 000 images. Experimental results demonstrate that our method outperforms the representative models regarding objective metrics and visual quality. Code and dataset: github.com/lhc1224/PIAL-Net. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
CVPR | 2 |
| 2023 | Spatial-Aware Token for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging task aiming to localize objects with only image-level supervision. Recent works apply visual transformer to WSOL and achieve significant success by exploiting the long-range feature dependency in self-attention mechanism. However, existing transformer-based methods synthesize the classification feature maps as the localization map, which leads to optimization conflicts between classification and localization tasks. To address this problem, we propose to learn a task-specific spatial-aware token (SAT) to condition localization in a weakly supervised manner. Specifically, a spatial token is first introduced in the input space to aggregate representations for localization task. Then a spatial aware attention module is constructed, which allows spatial token to generate foreground probabilities of different patches by querying and to extract localization knowledge from the classification task. Besides, for the problem of sparse and unbalanced pixel-level supervision obtained from the image-level label, two spatial constraints, including batch area loss and normalization loss, are designed to compensate and enhance this supervision. Experiments show that the proposed SAT achieves state-of-the-art performance on both CUB-200 and ImageNet, with 98.45% and 73.13% GT-known Loc, respectively. Even under the extreme setting of using only 1 image per class from ImageNet for training, SAT already exceeds the SOTA method by 2.1% GT-known Loc. Code and models are available at https://github.com/wpy1999/SAT. Pingyu Wu, Wei Zhai, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
ICCV | 2 |
| 2023 | Grounding 3D Object Affordance from 2D Interactions in ImagesabstractGrounding 3D object affordance seeks to locate objects’ "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotations to declare interactive regions of interest on the object and establishing a mapping between the regions and affordances. However, the essence of learning object affordance is to understand how to use it, and the manner that detaches interactions is limited in generalization. Normally, humans possess the ability to perceive object affordances in the physical world through demonstration images or videos. Motivated by this, we introduce a novel task setting: grounding 3D object affordance from 2D interactions in images, which faces the challenge of anticipating affordance through interactions of different sources. To address this problem, we devise a novel Interaction-driven 3D Affordance Grounding Network (IAG), which aligns the region feature of objects from different sources and models the interactive contexts for 3D object affordance grounding. Besides, we collect a Point-Image Affordance Dataset (PIAD) to support the proposed task. Comprehensive experiments on PIAD demonstrate the reliability of the proposed task and the superiority of our method. The project is available at https://github.com/yyvhang/IAGNet. Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
ICCV | 2 |
| 2023 | Building Earthquake Damage Recognition Performance of Texture Features from SAR Image in Frequency Domain and Spatial DomainabstractThe post-earthquake polarimetric SAR (PolSAR) data is an efficient measure for identifying building earthquake damage as the texture feature extracted from SAR data is a very effective indicator in identifying different damage statuses of buildings in earthquake areas. In many cases, the texture feature even outperforms other kinds of features. The texture features of buildings include not only spatial texture but also frequency texture. However, most studies presently only focus on the texture features in the spatial domain and ignore texture features in the frequency domain. To investigate the effectiveness of the texture in the frequency domain, in this study we proposed a new texture feature derived from the frequency domain, we conducted a comparative analysis of building earthquake damage identification performance based on the two texture features drawn from the frequency and spatial domains. Our experimental results show that the recognition performance of the frequency texture feature performs better than the spatial one. Wei Zhai, Yaxin Bi, Guiyu Zhu, Jianqing Du |
IGARSS | 1 |
| 2023 | Location-Free Camouflage Generation NetworkabstractCamouflage is a common visual phenomenon, which refers to hiding the foreground objects into the background images, making them briefly invisible to the human eye. Previous work has typically been implemented by an iterative optimization process. However, these methods struggle in 1) efficiently generating camouflage images using foreground and background with flexible structure; 2) camouflaging foreground objects to regions with multiple appearances (e.g.the junction of the vegetation and the mountains), which limit their practical application. To address these problems, this paper proposes a novelLocation-freeCamouflageGenerationNetwork(LCG-Net) that fuse high-level features of foreground and background image, and generate result by one inference. Specifically, aPosition-aligned Structure Fusion(PSF) module is devised to guide structure feature fusion based on the point-to-point structure similarity of foreground and background, and introduce local appearance features point-by-point. To retain the necessary identifiable features, a new immerse loss is adopted under our pipeline, while a background patch appearance loss is utilized to ensure that the hidden objects look continuous and natural at regions with multiple appearances. Experiments show that our method has results as satisfactory as state-of-the-art in the single-appearance regions and are less likely to be completely invisible, but far exceed the quality of the state-of-the-art in the multi-appearance regions. Moreover, our method is hundreds of times faster than previous methods. Benefitting from the unique advantages of our method, we provide some downstream applications for camouflage generation, which show its potential. The related code and dataset will be released athttps://github.com/Tale17/LCG-Net. Wei Zhai, Yang Cao 0010, Zhengjun Zha |
IEEE Trans. Multim. | 2 |
| 2023 | Deep Texton-Coherence Network for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging visual task since the appearance and morphology of foreground objects and background regions are highly similar in nature. Recent CNN-based studies gradually integrated the high-level semantic information and the low-level local features of images through hierarchical and progressive structures to achieve camouflaged object detection. However, these methods ignore thespatial statistical propertiesof the local context, which is a critical cue for distinguishing and describing camouflaged objects. To address this problem, we propose a novel Deep Texton-Coherence Network (DTC-Net) that leverages the spatial organization of textons in the foreground and background regions as discriminative cues for camouflaged object detection. Specifically, a Local Bilinear module (LB) is devised to obtain the robust representation of texton to trivial details and illumination changes, by replacing the classic first-order linearization operations with bilinear second-order statistical operations in the convolution process. Next, these texton representations are associated with a Spatial Coherence Organization module (SCO) to capture irregular spatial coherence via a deformable convolutional strategy, and then the descriptions of the textons extracted by the LB module are used as weights to suppress features that are spatially adjacent but have different representations. Finally, the texton-coherence representation is integrated with the original features at different levels to achieve camouflaged object detection. Evaluation on the three most challenging camouflaged object detection datasets demonstrats the superiority of the proposed model when compared to the state-of-the-art methods. Furthermore, our ablation studies and performance analyses demonstrate the effectiveness of the texton-coherence module. Wei Zhai, Yang Cao 0010, Haiyong Xie 0001, Zhengjun Zha |
IEEE Trans. Multim. | 1 |
| 2022 | Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental LearningabstractNon-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a novel self-sustaining representation expansion scheme. Our scheme consists of a structure reorganization strategy that fuses main-branch expansion and side-branch updating to maintain the old features, and a main-branch distillation scheme to transfer the invariant knowledge. Furthermore, a prototype selection mechanism is proposed to enhance the discrimination between the old and new classes by selectively incorporating new samples into the distillation process. Extensive experiments on three benchmarks demonstrate significant incremental performance, outperforming the state-of-the-art methods by a margin of 3%, 3% and 6%, respectively. Kai Zhu 0004, Wei Zhai, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
CVPR | 2 |
| 2022 | Learning Affordance Grounding from Exocentric ImagesabstractAffordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability that transform the various exocentric interactions to invariant egocentric affordance so as to counter the impact of interactive diversity. To empower an agent with such ability, this paper proposes a task of affordance grounding from exocentric view, i.e., given exocentric human-object interaction and egocentric object images, learning the affordance knowledge of the object and transferring it to the egocentric image using only the affordance label as supervision. To this end, we devise a cross-view knowledge transfer framework that extracts affordance-specific features from exocentric interactions and enhances the perception of affordance regions by preserving affordance correlation. Specifically, an Affordance Invariance Mining module is devised to extract specific clues by minimizing the intra-class differences originated from interaction habits in exocentric images. Besides, an Affordance Co-relation Preserving strategy is presented to perceive and localize affordance by aligning the co-relation matrix of predicted results between the two views. Particularly, an affordance grounding dataset named AGD20K is constructed by collecting and labeling over 20K images from 36 affordance categories. Experimental results demonstrate that our method outperforms the representative models in terms of objective metrics and visual quality. Code: github.com/lhc1224/Cross-View-AG. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
CVPR | 2 |
| 2022 | Background Activation Suppression for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects using only image-level labels. Recently a new paradigm has emerged by generating a foreground prediction map (FPM) to achieve localization task. Existing FPM-based methods use cross-entropy (CE) to evaluate the foreground prediction map and to guide the learning of generator. We argue for using activation value to achieve more efficient learning. It is based on the experimental observation that, for a trained network, CE converges to zero when the foreground mask covers only part of the object region. While activation value increases until the mask expands to the object boundary, which indicates that more object areas can be learned by using activation value. In this paper, we propose a Background Activation Suppression (BAS) method. Specifically, an Activation Map Constraint module (AMC) is designed to facilitate the learning of generator by suppressing the background activation value. Meanwhile, by using the foreground region guidance and the area constraint, BAS can learn the whole region of the object. In the inference phase, we consider the prediction maps of different categories together to obtain the final localization results. Extensive experiments show that BAS achieves significant and consistent improvement over the baseline methods on the CUB-200-2011 and ILSVRC datasets. Code and models are available at github.com/wpy1999IBAS. Pingyu Wu, Wei Zhai, Yang Cao 0010 |
CVPR | 2 |
| 2022 | Exploring Figure-Ground Assignment Mechanism in Perceptual OrganizationabstractPerceptual organization is a challenging visual task that aims to perceive and group the individual visual element so that it is easy to understand the meaning of the scene as a whole. Most recent methods building upon advanced Convolutional Neural Network (CNN) come from learning discriminative representation and modeling context hierarchically. However, when the visual appearance difference between foreground and background is obscure, the performance of existing methods degrades significantly due to the visual ambiguity in the discrimination process. In this paper, we argue that the figure-ground assignment mechanism, which conforms to human vision cognitive theory, can be explored to empower CNN to achieve a robust perceptual organization despite visual ambiguity. Specifically, we present a novel Figure-Ground-Aided (FGA) module to learn the configural statistics of the visual scene and leverage it for the reduction of visual ambiguity. Particularly, we demonstrate the benefit of using stronger supervisory signals by teaching (FGA) module to perceive configural cues, \ie, convexity and lower region, that human deem important for the perceptual organization. Furthermore, an Interactive Enhancement Module (IEM) is devised to leverage such configural priors to assist representation learning, thereby achieving robust perception organization with complex visual ambiguities. In addition, a well-founded visual segregation test is designed to validate the capability of the proposed FGA mechanism explicitly. Comprehensive evaluation results demonstrate our proposed FGA mechanism can effectively enhance the capability of perception organization on various baseline models. Nevertheless, the model augmented via our proposed FGA mechanism also outperforms state-of-the-art approaches on four challenging real-world applications. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Zhengjun Zha |
NeurIPS | 1 |
| 2022 | One-Shot Object Affordance Detection in the Wild
Wei Zhai, Hongchen Luo, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2022 | Robust Object Detection via Adversarial Novel Style ExplorationabstractDeep object detection models trained on clean images may not generalize well on degraded images due to the well-known domain shift issue. This hinders their application in real-life scenarios such as video surveillance and autonomous driving. Though domain adaptation methods can adapt the detection model from a labeled source domain to an unlabeled target domain, they struggle in dealing with open and compound degradation types. In this paper, we attempt to address this problem in the context of object detection by proposing a robust object Detector via Adversarial Novel Style Exploration (DANSE). Technically, DANSE first disentangles images into domain-irrelevant content representation and domain-specific style representation under an adversarial learning framework. Then, it explores the style space to discover diverse novel degradation styles that are complementary to those of the target domain images by leveraging a novelty regularizer and a diversity regularizer. The clean source domain images are transferred into these discovered styles by using a content-preserving regularizer to ensure realism. These transferred source domain images are combined with the target domain images and used to train a robust degradation-agnostic object detection model via adversarial domain adaptation. Experiments on both synthetic and real benchmark scenarios confirm the superiority of DANSE over state-of-the-art methods. Wen Wang 0009, Jing Zhang 0037, Wei Zhai, Yang Cao 0010, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2021 | Self-Promoted Prototype Refinement for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning is to recognize the new classes given few samples and not forget the old classes. It is a challenging task since representation optimization and prototype reorganization can only be achieved under little supervision. To address this problem, we propose a novel incremental prototype learning scheme. Our scheme consists of a random episode selection strategy that adapts the feature representation to various generated incremental episodes to enhance the corresponding extensibility, and a self-promoted prototype refinement mechanism which strengthens the expression ability of the new classes by explicitly considering the dependencies among different classes. Particularly, a dynamic relation projection module is proposed to calculate the relation matrix in a shared embedding space and leverage it as the factor for bootstrapping the update of prototypes. Extensive experiments on three benchmark datasets demonstrate the above-par incremental performance, outperforming state-of-the-art methods by a margin of 13%, 17% and 11%, respectively. Kai Zhu 0004, Yang Cao 0010, Wei Zhai, Zhengjun Zha |
CVPR | 3 |
| 2021 | One-Shot Affordance DetectionabstractAffordance detection refers to identifying the potential action possibilities of objects in an image, which is an important ability for robot perception and manipulation. To empower robots with this ability in unseen scenarios, we consider the challenging one-shot affordance detection problem in this paper, i.e., given a support image that depicts the action purpose, all objects in a scene with the common affordance should be detected. To this end, we devise a One-Shot Affordance Detection (OS-AD) network that firstly estimates the purpose and then transfers it to help detect the common affordance from all candidate images. Through collaboration learning, OS-AD can capture the common characteristics between objects having the same underlying affordance and learn a good adaptation capability for perceiving unseen affordances. Besides, we build a Purpose-driven Affordance Dataset (PAD) by collecting and labeling 4k images from 31 affordance and 72 object categories. Experimental results demonstrate the superiority of our model over previous representative ones in terms of both objective metrics and visual quality. The benchmark suite is at ProjectPage. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
IJCAI | 2 |
| 2021 | A tri-attention enhanced graph convolutional network for skeleton-based action recognitionabstractAbstract Skeleton‐based action recognition has recently attracted a lot of research interests due to its advantage in computational efficiency. Some recent work building upon Graph Convolutional Networks (GCNs) has shown promising performance in this task by modelling intrinsic spatial correlations between skeleton joints. However, these methods only consider local properties of action sequences in the spatial‐temporal domain, and consequently, are limited in distinguishing complex actions with similar local movements. To address this problem, a novel tri‐attention module (TAM) is proposed to guide GCNs to perceive significant variations across local movements. Specifically, the devised TAM is implemented in three steps: i) A dimension permuting unit is proposed to characterise skeleton action sequences in three different domains: body poses, joint trajectories, and evolving projections. ii) A global statistical modelling unit is introduced to aggregate the first‐order and second‐order properties of global contexts to perceive the significant movement variations of each domain. iii) A fusion unit is presented to integrate the features of these three domains together and leverage as orientation for graph convolution at each layer. Through these three steps, significant‐variation frames, joints, and channels can be enhanced. We conduct extensive experiments on two large‐scale benchmark datasets, NTU RGB‐D and Kinetics‐Skeleton. Experimental results demonstrate that the proposed TAM can be easily plugged into existing GCNs and achieve comparable performance with the state‐of‐the‐art methods. Wei Zhai, Yang Cao 0010 |
IET Comput. Vis. | 2 |
| 2021 | One-Shot Texture Retrieval Using Global Grouping MetricabstractTexture retrieval is widely used in the fields of fashion and e-commerce. This paper presents the problem of one-shot texture retrieval: given an example of a new reference texture, we aim to detect and segment all pixels of the same texture category within an arbitrary image. To address this problem, an OS-TR network is proposed to encode both reference and query images into a texture representation space, and a better comparison is made based on the global grouping information. Because the learned texture representation should be invariant to the spatial layout while preserving the rough semantic concepts, we introduce an adaptive directionality-aware module to finely discriminate the orderless texture details. To make full use of the global context information given only a few examples, we incorporate a grouping-attention mechanism into the relation network, resulting in the per-channel modulation of the local relation features. Extensive experiments on two benchmark datasets (i.e., the DTD and ADE20K dataset) and real scenarios demonstrate that our proposed method can achieve above-par segmentation performance and robust generalization across domains. Kai Zhu 0004, Yang Cao 0010, Wei Zhai, Zhengjun Zha |
IEEE Trans. Multim. | 3 |
| 2020 | Deep Structure-Revealed Network for Texture RecognitionabstractTexture recognition is a challenging visual task since various primitives along with their arrangements can be recognized from a same texture image when perceiving with different contexts. Some recent work building on CNNs exploits orderless aggregating to provide invariance to spatial arrangements. However, these methods ignore the inherent structural property of textures, which is a critical cue for distinguishing and describing texture images in the wild. To address this problem, we propose a novel Deep Structure-Revealed Network (DSR-Net) that leverages spatial dependency among the captured primitives as structural representation for texture recognition. Specifically, a primitive capturing module (PCM) is devised to generate multiple primitives from eight directional spatial contexts, in which deep features are firstly extracted under the constrains of direction map and then encoded based on the similarities of neighborhood. Next, these primitives are associated with a dependence learning module (DLM) to generate structural representation, in which a two-way collaborative relationship strategy is introduced to perceive the spatial dependencies among multiple primitives. At last, the structure-revealed texture representations are integrated with spatial ordered information to achieve real-world texture recognition. Evaluation on the five most challenging texture recognition datasets has demonstrated the superiority of the proposed model against state-of-the-art methods. The structure-revealed performances of DSR-Net are further verified on some extensive experiments, including fine-grained classification and semantic segmentation. Wei Zhai, Yang Cao 0010, Zhengjun Zha, Haiyong Xie 0001, Feng Wu 0001 |
CVPR | 1 |
| 2020 | Deep Inhomogeneous Regularization For Transfer LearningabstractFine-tuning is an effective transfer learning method to achieve ideal performance on target task with limited training data. Some recent works regularize parameters of deep neural networks for better knowledge transfer. However, these methods enforce homogeneous penalties for all parameters, resulting in catastrophic forgetting or negative transfer. To address this problem, we propose a novel Inhomogeneous Regularization (IR) method that imposes a strong regularization on parameters of transferable convolutional filters to tackle catastrophic forgetting and alleviate the regularization on parameters of less transferable filters to tackle negative transfer. Moreover, we use the decaying averaged deviation of parameters from the start point (pre-trained parameters) to accurately measure the transferability of each filter. Evaluation on the three challenging benchmarks datasets has demonstrated the superiority of the proposed model against state-of-the-art methods. Wen Wang 0009, Wei Zhai, Yang Cao 0010 |
ICIP | 2 |
| 2020 | Self-Supervised Tuning for Few-Shot SegmentationabstractFew-shot segmentation aims at assigning a category label to each image pixel with few annotated samples. It is a challenging task since the dense prediction can only be achieved under the guidance of latent features defined by sparse annotations. Existing meta-learning based method tends to fail in generating category-specifically discriminative descriptor when the visual features extracted from support images are marginalized in embedding space. To address this issue, this paper presents an adaptive tuning framework, in which the distribution of latent features across different episodes is dynamically adjusted based on a self-segmentation scheme, augmenting category-specific descriptors for label prediction. Specifically, a novel self-supervised inner-loop is firstly devised as the base learner to extract the underlying semantic features from the support image. Then, gradient maps are calculated by back-propagating self-supervised loss through the obtained features, and leveraged as guidance for augmenting the corresponding elements in the embedding space. Finally, with the ability to continuously learn from different episodes, an optimization-based meta-learner is adopted as outer loop of our proposed framework to gradually refine the segmentation results. Extensive experiments on benchmark PASCAL-5i and COCO-20i datasets demonstrate the superiority of our proposed method over state-of-the-art. Kai Zhu 0004, Wei Zhai, Yang Cao 0010 |
IJCAI | 2 |
| 2019 | Deep Multiple-Attribute-Perceived Network for Real-World Texture RecognitionabstractTexture recognition is a challenging visual task as multiple perceptual attributes may be perceived from the same texture image when combined with different spatial context. Some recent works building upon Convolutional Neural Network (CNN) incorporate feature encoding with orderless aggregating to provide invariance to spatial layouts. However, these existing methods ignore visual texture attributes, which are important cues for describing the real-world texture images, resulting in incomplete description and inaccurate recognition. To address this problem, we propose a novel deep Multiple-Attribute-Perceived Network (MAP-Net) by progressively learning visual texture attributes in a mutually reinforced manner. Specifically, a multi-branch network architecture is devised, in which cascaded global contexts are learned by introducing similarity constraint at each branch, and leveraged as guidance of spatial feature encoding at next branch through an attribute transfer scheme. To enhance the modeling capability of spatial transformation, a deformable pooling strategy is introduced to augment the spatial sampling with adaptive offsets to the global context, leading to perceive new visual attributes. An attribute fusion module is then introduced to jointly utilize the perceived visual attributes and the abstracted semantic concepts at each branch. Experimental results on the five most challenging texture recognition datasets have demonstrated the superiority of the proposed model against the state-of-the-arts. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Zhengjun Zha |
ICCV | 1 |
| 2019 | One-Shot Texture Retrieval with Global Context MetricabstractIn this paper, we tackle one-shot texture retrieval: given an example of a new reference texture, detect and segment all the pixels of the same texture category within an arbitrary image. To address this problem, we present an OS-TR network to encoding both reference patch and query image, leading to achieve texture segmentation towards the reference category. Unlike the existing texture encoding methods that integrate CNN with orderless pooling, we propose a directionality-aware network to capture the texture variations at each direction, resulting in spatially invariant representation. To segment new categories given only few examples, we incorporate a self-gating mechanism into relation network to exploit global context information for adjusting per-channel modulation weights of local relation features. Extensive experiments on benchmark texture datasets and real scenarios demonstrate the above-par segmentation performance and robust generalization across domains of our proposed method. Kai Zhu 0004, Wei Zhai, Zhengjun Zha, Yang Cao 0010 |
IJCAI | 2 |
| 2019 | PixTextGAN: structure aware text image synthesis for license plate recognitionabstractRapid progress on text image recognition has been achieved with the development of deep‐learning techniques. However, it is still a great challenge to achieve a comprehensive license plate recognition in the real scenes, since there are no publicly available large diverse datasets for the training of deep learning models. This paper aims at synthesising of license plate images with generative adversarial networks (GAN), refraining from collecting a vast amount of labelled data. The authors thus propose a novel PixTextGAN that leverages a controllable architecture that generates specific character structures for different text regions to generate synthetic license plate images with reasonable text details. Specifically, a comprehensive structure‐aware loss function is presented to preserve the key characteristic of each character region and thus to achieve appearance adaption for better recognition. Qualitative and quantitative experiments demonstrate the superiority of authors’ proposed method in text image synthetisation over state‐of‐the‐art GANs. Further experimental results of license plate recognition on ReId and CCPD dataset demonstrate that using the synthesised images by PixTextGAN can greatly improve the recognition accuracy. Shilian Wu, Wei Zhai, Yang Cao 0010 |
IET Image Process. | 2 |
| 2018 | A Generative Adversarial Network Based Framework for Unsupervised Visual Surface InspectionabstractVisual surface inspection is a challenging task due to the highly inconsistent appearance of the target surfaces and the abnormal regions. Most of the state-of-the-art methods are highly dependent on the labelled training samples, which are difficult to collect in practical industrial applications. To address this problem, we propose a generative adversarial network based framework for unsupervised surface inspection. The generative adversarial network is trained to generate the fake images analogous to the normal surface images. It implies that a well-trained GAN indeed learns a good representation of the normal surface images in a latent feature space. And consequently, the discriminator of GAN can serve as a naturally one-class classifier. We use the first three conventional layer of the discriminator as the feature extractor, whose response is sensitive to the abnormal regions. Particularly, a multi-scale fusion strategy is adopted to fuse the responses of the three convolution layers and thus improve the segmentation performance of abnormal detection. Various experimental results demonstrate the effectiveness of our proposed method. Wei Zhai, Yang Cao 0010, Zengfu Wang |
ICASSP | 1 |
| 2018 | Co-occurrent Structural Edge Detection for Color-Guided Depth Map Super-Resolution
Wei Zhai, Yang Cao 0010, Zhengjun Zha |
MMM (1) | 2 |
| 2017 | Building damage information investigation from a single post-earthquake PolSAR image based on the fusion of multiple texture featuresabstractOnly using the post-earthquake PolSAR imagery to interpret collapsed buildings information is a rapid and effective disaster investigation means, which is also easy and fast for implementation of the earthquake damage assessment. This work is focused on rapid building earthquake damage information detection in urban areas using a single post-earthquake PolSAR image. In this paper, the Precision Weighted Multi-feature Fusion (PWMF) method is proposed to fuse multiple texture features for more accurate extraction of the collapsed buildings and the undamaged buildings. In addition, the algorithm of Optimization of Polarimetric Contrast Enhancement (OPCE) is employed to enhance the contrast ratio between the collapsed buildings and the oriented buildings in order to improve the extraction accuracy of the collapsed buildings and undamaged buildings. The building damage assessment is carried out at the city block scale according to the building collapse rate. Wei Zhai, Chunlin Huang, Wansheng Pei, Yan Li 0079 |
IGARSS | 1 |
| 2016 | Building damage information investigation after earthquake using single post-event PolSAR imageabstractRapidly and accurately obtaining collapsed buildings information of the earthquake-stricken areas can help to effectively guide the implementation of the emergency rescue and can reduce disaster losses and casualties. This work is focused on rapid building earthquake damage information detection in urban areas using a single post-earthquake PolSAR data. In this paper, the methods of polarization orientation angle (POA) compensation and Wishart supervised classification are employed to extract the collapsed buildings and undamaged buildings. In addition, the two parameters of the normalized difference of the dihedral component (NDDC) and the HH-HV Correlation Coefficient (ρHHHV) are proposed to improve the extraction accuracy of the collapsed buildings and undamaged buildings. The building damage assessment is carried out at the city block scale according to the building collapse rate. Wei Zhai, Huanfeng Shen, Chunlin Huang, Wansheng Pei |
IGARSS | 1 |