VLDB 2026 Research / reviewers in the wild / expert
Wei-Shi Zheng 0001
dblp:30/8399 · also Weishi Zheng 0001
· DBLP profile ↗
411ranked-venue papers
18as first author
235since 2021 · last 2026
0000-0001-8327-0003ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 279 · 7 first-author · 158 since 2021Artificial intelligence and machine learning · 264 · 17 first-author · 143 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 17 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Security and privacy · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Surround-View Fisheye Camera 3D Object DetectionabstractIn this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic pinhole-based 3D object detectors to fisheye imagery. To mitigate this, we then develop two methods that incorporate the unique geometry of fisheye images into mainstream detection frameworks: one based on the bird's-eye-view (BEV) paradigm, named FisheyeBEVDet, and the other on the query-based paradigm, named FisheyePETR. Both methods adopt spherical spatial representations to effectively capture fisheye geometry. In light of the lack of dedicated evaluation benchmarks, we release Fisheye3DOD, a new open dataset synthesized using CARLA and featuring both standard pinhole and fisheye camera arrays. Experiments on Fisheye3DOD demonstrate that our fisheye-compatible modeling improves accuracy by up to 6.2% compared to baseline methods. Changcai Li, Wenwei Lin, Zuoxun Hou, Gang Chen 0023, Wei Zhang 0009, Wei-Shi Zheng 0001 |
AAAI | 7 |
| 2026 | TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingabstractTo guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detailed, comprehensible feedback on what is done well and what can be improved. However, existing score-based action assessment methods are still far from reaching this practical scenario. To bridge this gap, we investigate a new task termed Descriptive Action Coaching (DescCoach) which requires the model to provide detailed commentary on what is done well and what can be improved beyond a simple quality score for action execution. To this end, we first build a new dataset named EE4D-DescCoach. Through an automatic annotation pipeline, our dataset goes beyond the existing action assessment datasets by providing detailed TechPoint-level commentary. Furthermore, we propose TechCoach, a new framework that explicitly incorporates TechPoint-level reasoning into the DescCoach process. The central to our method lies in the Context-aware TechPoint Reasoner, which enables TechCoach to learn TechPoint-related quality representation by querying visual context under the supervision of TechPoint-level coaching commentary. By leveraging the visual context and the TechPoint-related quality representation, a unified TechPoint-aware Action Assessor is then employed to provide the overall coaching commentary together with the quality score. Combining all of these, we establish a new benchmark for DescCoach and evaluate the effectiveness of our method through extensive experiments. Yuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin, Yu-Ming Tang, Wei-Shi Zheng 0001 |
AAAI | 6 |
| 2026 | DCAC: Dynamic Class-Aware Cache Creates Stronger Out-of-Distribution DetectorsabstractOut-of-distribution (OOD) detection remains a fundamental challenge for deep neural networks, particularly due to overconfident predictions on unseen OOD samples during testing. We reveal a key insight: OOD samples predicted as the same class, or given high probabilities for it, are visually more similar to each other than to the true in-distribution (ID) samples. Motivated by this class-specific observation, we propose DCAC (Dynamic Class-Aware Cache), a training-free, test-time calibration module that maintains separate caches for each ID class to collect high-entropy samples and calibrate the raw predictions of input samples. DCAC leverages cached visual features and predicted probabilities through a lightweight two-layer module to mitigate overconfident predictions on OOD samples. This module can be seamlessly integrated with various existing OOD detection methods across both unimodal and vision-language models while introducing minimal computational overhead. Extensive experiments on multiple OOD benchmarks demonstrate that DCAC significantly enhances existing methods, achieving substantial improvements, i.e., reducing FPR95 by 6.55% when integrated with ASH-S on ImageNet OOD benchmark. Yanqi Wu, Qichao Chen, Runhe Lai, Xinhua Lu, Jiaxin Zhuang, Zhi-Lin Zhao 0001, Wei-Shi Zheng 0001 |
AAAI | 7 |
| 2026 | Class incremental learning with task-specific batch normalization and out-of-distribution detection
Zhiping Zhou, Xuchen Xie, Yiqiao Qiu, Run Lin, Wei-Shi Zheng 0001 |
Neurocomputing | 5 |
| 2026 | Dual-modality adaptation in vision-language models for continual learning
Jiayang Zeng, Wentao Zhang 0005, Kanghao Chen, Jiantao Tan, Wei-Shi Zheng 0001 |
Neural Networks | 6 |
| 2026 | Momentor++: Advancing Video Large Language Models With Fine-Grained Long Video ReasoningabstractLarge Language Models (LLMs) exhibit remarkable proficiency in understanding and managing text-based tasks.Many works try to transfer these capabilities to the video domain, which are referred to as Video-LLMs. However, current Video-LLMs can only grasp the coarse-grained semantics and are unable to efficiently handle tasks involving the comprehension or localization of specific video segments. To address these challenges, we propose Momentor, a Video-LLM designed to perform fine-grained temporal understanding tasks. To facilitate the training of Momentor, we develop an automatic data generation engine to build Moment-10M, a large-scale video instruction dataset with segment-level instruction data. Building upon the foundation of the previously published Momentor and the Moment-10M dataset, we further extend this work by introducing a Spatio-Temporal Token Consolidation (STTC) method, which can merge redundant visual tokens spatio-temporally in a parameter-free manner, thereby significantly promoting computational efficiency while preserving fine-grained visual details. We integrate STTC with Momentor to develop Momentor++ and validate its performance on various benchmarks. Momentor demonstrates robust capabilities in fine-grained temporal understanding and localization. Further, Momentor++ excels in efficiently processing and analyzing extended videos with complex events, showcasing marked advancements in handling extensive temporal contexts. Juncheng Li 0006, Minghe Gao, Xiangnan He 0001, Siliang Tang, Wei-Shi Zheng 0001, Jun Xiao 0001, Meng Wang 0001, Tat-Seng Chua, Yueting Zhuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Human Motion Prediction via Continual Prior CompensationabstractHuman Motion Prediction (HMP) aims to predict future human poses at different moments according to observed past motion sequences. Previous approaches mainly treated the prediction of different temporal moments as a single prediction task and learned the predictions of varied moments simultaneously, which would encounter a main limitation: the learning of short-term predictions (referring to "near-future" prediction) could be hindered by the predictions of long-term (referring to "far-future" prediction) motions. In this paper, we develop a novel temporal continual learning framework called Continual Prior Compensation (CPC) to progressively train HMP models, in which we divide the prediction task of motions corresponding to varied temporal moments into several subtasks and train the model in a multi-stage manner. To mitigate the prior information forgetting in the progressive training, we further introduce a learnable random variable Prior Compensation Factor (PCF) to explicitly measure the prior knowledge loss. We theoretically show that the PCF can be efficiently learned together with the model parameters by minimizing a reasonable upper bound of the objective function. The proposed CPC is further enhanced to estimate the prior information loss for each subtask and a new framework called Continual Prior Compensation++ (CPC++) with Fine-Grained Prior Compensation Factor (FGPCF) is finally developed. Our CPC and CPC++ frameworks are quite flexible and can be easily integrated with different HMP backbone models and adapted to various datasets and applications. Extensive experiments on three HMP benchmark datasets using multiple SOTA HMP backbones (PGBIG, siMLPe, MotionMixer, and LTD) demonstrate the effectiveness and flexibility of our frameworks. Jianwei Tang, Jianfang Hu, Tianming Liang, Xiaotong Lin 0002, Jiangxin Sun, Wei-Shi Zheng 0001, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Neural Prediction Errors as a Unified Cue for Abstract Visual ReasoningabstractHumans exhibit remarkable abilities in recognizing relationships and performing complex reasoning. In contrast, deep neural networks have long been critiqued for their limitations in abstract visual reasoning (AVR), a key challenge in achieving artificial general intelligence. Drawing on the well-known concept of prediction errors from neuroscience, we propose that prediction errors can serve as a unified mechanism for both supervised and self-supervised learning in AVR. In our novel supervised learning model, AVR is framed as a prediction-and-matching process, where the central component is the discrepancy (i.e., prediction error) between a predicted feature based on abstract rules and candidate features within a reasoning context. In the self-supervised model, prediction errors as a key component unify the learning and inference processes. Both supervised and self-supervised prediction-based models achieve state-of-the-art performance on a broad range of AVR datasets and task conditions. Most notably, hierarchical prediction errors in the supervised model automatically decrease during training, an emergent phenomenon closely resembling the decrease of dopamine signals observed in biological learning. These findings underscore the critical role of prediction errors in AVR and highlight the potential of leveraging neuroscience theories to advance computational models for high-level cognition in artificial intelligence. Lingxiao Yang, Xiaohua Xie, Wei-Shi Zheng 0001, Ru-Yuan Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Arbitrary-scale atmospheric downscaling with mixture of implicit neural networks trained on fixed-scale data
Teng-Yue Chen, Jie-Lan Xie, Wei Zhou 0097, Jianfang Hu, Peng-Qin Yao, Tianming Liang, Wei-Shi Zheng 0001, Pak Wai Chan |
Pattern Recognit. | 7 |
| 2026 | Dual-level data-anchor association learning via k-partite graph factorization for multi-view clustering
Jian-Sheng Wu, Yuan-Tong Cheng, Weidong Min, Wei-Shi Zheng 0001 |
Pattern Recognit. | 4 |
| 2026 | High-order Aligned Deep Complementary and View-Specific Similarity Graphs for Unsupervised Multi-View Feature Selection
Jian-Sheng Wu, Jia-Tao Yu, Jun-Yun Wu, Weidong Min, Wei-Shi Zheng 0001 |
Pattern Recognit. | 5 |
| 2026 | Slowly expanding neural network for class incremental learning
Zhengjin Xu, Xiaobin Chang, Wei-Shi Zheng 0001 |
Pattern Recognit. | 4 |
| 2026 | Terafly: A Multinode FPGA-Based Accelerator Design for Efficient Cooperative Inference in LLMsabstractIn this paper, we propose Terafly, a multi-node accelerator design tailored for efficient Large Language Model (LLM) deployment and inference. Conventional accelerator architectures struggle to effectively handle both the prefill and decode stages during inference. To address this limitation, we introduce a hybrid spatial-temporal architecture that combines the high-throughput advantages of spatial architectures with the flexibility of temporal architectures, enabling it to accommodate the diverse inference patterns of LLMs. In addition, we propose a generation framework to streamline the customization of our LLM-friendly design for various deployment scenarios. Within this framework, users can specify their requirements such as model type, target platform, and performance goals. The framework then generates multiple accelerator nodes and maps them to distinct Super Logic Regions (SLRs) within a single FPGA, enabling cooperative inference under a model parallelism scheme. Through experiments, our generated accelerator can be easily deployed on both Alveo U250 and U50lv cards, serving models ranging from OPT-350M to OPT-1.3B under various performance settings. Notably, when running OPT-1.3B using the generated dual-node accelerator on a single Alveo U50lv card, we achieve an average 1.1x speed-up and a 3.4x improvement in energy efficiency compared to the Nvidia A100 GPU. Jianing Zheng, Gang Chen 0023, Libo Huang 0002, Xin Lou 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | PLNK: Prompt Learning With Neutral Knowledge for Few-Shot Out-of-Distribution DetectionabstractRecent developments in few-shot out-of-distribution (OOD) detection have yielded remarkable performance, benefiting from large pre-trained vision-language models (VLMs). Our prior work focuses on using in-distribution (ID) knowledge as references to learn richer knowledge beyond the textual semantics of class labels, which is prone to cause the model overconfidence and results in a limited score gap between ID and OOD data. In this paper, rather than treating ID knowledge as references, we propose Prompt Learning with Neutral Knowledge (PLNK) to better differentiate ID from OOD data. Our key insight lies in leveraging diverse neutral knowledge to improve ID discrimination while alleviating the inherent model overconfidence on OOD data induced by ID knowledge, thereby capturing the notable discrepancy between ID and OOD data. By introducing neutral knowledge with a balanced degree of similarity to both ID and OOD data, we amplify the discrepancy between the learnable prompt and the references (i.e., diverse neutral knowledge) for ID data, while reducing it for OOD data. In this way, the simple yet effective PLNK framework brings a notable score gap between ID and OOD data, thereby improving OOD detection. Moreover, we incorporate the visual neutral prompt with richer semantics alongside the original text-only reference. Comprehensive experiments show that our method consistently surpasses current state-of-the-art methods. The codes will be released publicly. Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Zhiming Dai, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Double Nonconvex Tensor Robust Kernel Principal Component Analysis and Its Visual ApplicationsabstractTensor robust principal component analysis (TRPCA), as a popular linear low-rank method, has been widely applied to various visual tasks. The mathematical process of the low-rank prior is derived from the linear latent variable model. However, for nonlinear tensor data with rich information, their nonlinear structures may break through the assumption of low-rankness and lead to the large approximation error for TRPCA. Motivated by the latent low-dimensionality of nonlinear tensors, the general paradigm of the nonlinear tensor plus sparse tensor decomposition problem, called tensor robust kernel principal component analysis (TRKPCA), is first established in this paper. To efficiently tackle TRKPCA problem, two novel nonconvex regularizers the kernelized tensor Schatten- $p$ norm (KTSPN) and generalized nonconvex regularization are designed, where the former KTSPN with tighter theoretical support adequately captures nonlinear features (i.e., implicit low-rankness) and the latter ensures the sparser structural coding, guaranteeing more robust separation results. Then by integrating their strengths, we propose a double nonconvex TRKPCA (DNTRKPCA) method to achieve our expectation. Finally, we develop an efficient optimization framework via the alternating direction multiplier method (ADMM) to implement the proposed nonconvex kernel method. Experimental results on synthetic data and several real databases show the higher competitiveness of our method compared with other state-of-the-art regularization methods. The code has been released in our ResearchGate homepage: https://www.researchgate.net/publication/397181729_DNTRKPCA_code. Jianjun Wang 0003, Wei-Shi Zheng 0001, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2026 | Adaptively Fine-Tuning and Ensembling Vision-Language Model for Few-Shot Image ClassificationabstractThe main challenge in few-shot learning (FSL) is overfitting of the model to limited training data. Recently developed vision-language models like CLIP have been employed to alleviate the overfitting issue and achieved state-of-the-art FSL performance. However, such approaches heavily depend on good alignments between images and associated texts, and therefore may not work as expected when image-text alignment is challenging in downstream tasks, e.g., fine-grained and cross domain image classification. In this study, a new CLIP-based fine-tuning and inference framework is proposed to particularly help the model more accurately recognize visually similar classes and also work well in new imaging domains. During fine-tuning, a modified contrastive loss with adaptively weighted negative pairs is proposed to effectively separate visually similar classes and cluster each class more compactly. During inference, an instance level adaptive ensemble strategy is proposed, utilizing the visual prototypes to adaptively complement the prediction from the CLIP's image-text alignment. Extensive experimental evaluations demonstrate the superiority of the proposed framework, outperforming current state-of-the-art methods by a decent margin on twelve public datasets. The source code will be released publicly. Baishun Dong, Xiaoyuan Guan, Wei-Shi Zheng 0001, Tong Zhang 0017 |
IEEE Trans. Multim. | 4 |
| 2026 | Negative Semantic Guided Identity Boundary Construction for Open-World Person Re-IdentificationabstractPerson re-identification (ReID) aims to match person images of the same identity under different camera views. Conventional ReID models mainly consider a closed-world setting where person identities in query and gallery are exactly the same. However, in real-world applications, query identities and gallery identities usually do not exactly contain the same persons. Therefore, open-world ReID has been proposed to match the images of gallery identities (targets) with a large number of non-gallery identities ( non-targets). Since some non-targets are quite similar to the targets, the ReID model may make incorrect judgments when verifying these non-targets. To solve this problem, we leverage the impressive cross-modal matching capabilities of the large vision-language model (VLM) to constructNegativeSemantic guided identity boundaries for each person to develop the open-worldReIDmodel (NS-ReID). To construct the identity boundary, we propose Virtual Non-target Repulsion that utilizes negative semantics to prompt the ReID model to push virtual non-targets away from the targets. The prompts expressing negative semantics offer a different perspective to guide the training process to avoid contradictory optimization. Moreover, we propose the Dual-Boost Refinement Learning strategy to train learnable identity prompts to capture detailed identity information, which is essential for constructing the identity boundary since the variations among identities are comparatively small. These facilitate the model in constructing the wide identity boundary of each person. Extensive experiments on two benchmark ReID datasets demonstrate that our proposed NS-ReID achieves state-of-the-art performance compared with existing methods. Xiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual LearningabstractThe generation of a virtual digital avatar is a crucial research topic in the field of computer vision. Many existing works utilize Neural Radiance Fields (NeRF) to address this issue and have achieved impressive results. However, previous works assume the images of the training person are available and fixed while the appearances and poses of a subject could constantly change and increase in real-world scenarios. How to update the human avatar but also maintain the ability to render the old appearance of the person is a practical challenge. One trivial solution is to combine the existing virtual avatar models based on NeRF with continual learning methods. However, there are some critical issues in this approach: learning new appearances and poses can cause the model to forget past information, which in turn leads to a degradation in the rendering quality of past appearances, especially color bleeding issues, and incorrect human body poses. In this work, we propose a maintainable avatar (MaintaAvatar) based on neural radiance fields by continual learning, which resolves the issues by utilizing a Global-Local Joint Storage Module and a Pose Distillation Module. Overall, our model requires only limited data collection to quickly fine-tune the model while avoiding catastrophic forgetting, thus achieving a maintainable virtual avatar. The experimental results validate the effectiveness of our MaintaAvatar model. Shengbo Gu, Yu-Kun Qiu, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001 |
AAAI | 5 |
| 2025 | CLIP-RestoreX: Restore Image Structure and Perception in Exposure CorrectionabstractExposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correction method, referred to as CLIP-RestoreX, that leverages structural and perceptual priors from CLIP to tackle exposure correction. Specifically, we in CLIP-RestoreX propose to perform exposure correction by aligning CLIP-based structural and perceptual feature of the impaired image with its ground-truth image. To better restore the damaged structural information and perceptual information, we further design a frequency-domain based feature enhancement diffusion model, where we utilize the globality of Fourier transform to help reveal potential the relationship within the features. We conduct extensive experiments on several benchmark datasets. The results demonstrate that the proposed CLIP-RestoreX outperforms state-of-the-art exposure correction methods. Qing Zhang 0006, Jianfang Hu, Wei-Shi Zheng 0001 |
AAAI | 4 |
| 2025 | ParGo: Bridging Vision-Language with Partial and Global ViewsabstractThis work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous works that rely on global attention-based projectors, our ParGo bridges the representation gap between the separately pre-trained vision encoders and the LLMs by integrating global and partial views, which alleviates the overemphasis on prominent regions. To facilitate the effective training of ParGo, we collect a large-scale detail-captioned image-text dataset named ParGoCap-1M-PT, consisting of 1 million images paired with high-quality captions. Extensive experiments on several MLLM benchmarks demonstrate the effectiveness of our ParGo, highlighting its superiority in aligning vision and language modalities. Compared to conventional Q-Former projector, our ParGo achieves an improvement of 259.96 in MME benchmark. Furthermore, our experiments reveal that ParGo significantly outperforms other projectors, particularly in tasks that emphasize detail perception ability. An-Lan Wang, Bin Shan, Kun-Yu Lin, Guozhi Tang, Jingqun Tang, Wei-Shi Zheng 0001 |
AAAI | 10 |
| 2025 | Light-T2M: A Lightweight and Fast Model for Text-to-motion GenerationabstractDespite the significant role text-to-motion (T2M) generation plays across various applications, current methods involve a large number of parameters and suffer from slow inference speeds, leading to high usage costs. To address this, we aim to design a lightweight model to reduce usage costs. First, unlike existing works that focus solely on global information modeling, we recognize the importance of local information modeling in the T2M task by reconsidering the intrinsic properties of human motion, leading us to propose a lightweight Local Information Modeling Module. Second, we are the first to introduce Mamba to the T2M task, reducing the number of parameters and GPU memory demands, and we have designed a novel Pseudo-bidirectional Scan to replicate the effects of a bidirectional scan without increasing parameter count. Moreover, we propose a novel Adaptive Textual Information Injector that more effectively integrates textual information into the motion during generation. By integrating the aforementioned designs, we propose a lightweight and fast model named Light-T2M. Compared to the state-of-the-art method, MoMask, our Light-T2M model features just 10% of the parameters (4.48M vs 44.85M) and achieves a 16% faster inference time (0.152s vs 0.180s), while surpassing MoMask with an FID of 0.040 (vs. 0.045) on HumanML3D dataset and 0.161 (vs. 0.228) on KIT-ML dataset. Ling-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi Zheng 0001 |
AAAI | 4 |
| 2025 | When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired DataabstractWe present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previous work due to its inherent ambiguity and the scarcity of training data. However, we find that by properly coupling shadow removal and IID into a joint learning framework, they can reinforce each other and enable promising results on both tasks, even with unpaired training data. Our framework is comprised of an IID network for separating the shadow input image into illumination and reflectance, and an illumination recovery network for predicting shadow-free illumination with which we are able to produce the shadow removal output by recombining with the estimated reflectance. We perform extensive experiments on various benchmark datasets to demonstrate the effectiveness of our method in shadow removal, and also showcase our advantage over previous IID methods in handling images with complex shadows. Rongjia Zheng, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
AAAI | 4 |
| 2025 | Textual Prototype-Guided Continual Learning for Medical Image ClassificationabstractIntelligent diagnostic systems require continual learning (CL) to learn to diagnose new diseases. However, the systems suffer from catastrophic forgetting of old knowledge when learning knowledge of new diseases. Existing CL methods that leverage pre-trained language models (PLMs) to guide visual encoders are ineffective in the medical domain due to PLMs' limited medical knowledge. Here, we propose Textual Prototype-Guided Continual Learning (TPGCL) for effective CL. TPGCL utilizes an image captioning model to generate semantically rich disease descriptions, which are then encoded into text embeddings via a text encoder to obtain the textual prototype for each class. These prototypes guide the visual encoder during CL. In addition, a gradient-weighting mechanism merges visual adapters learned from all CL stages to preserve old knowledge and prevent model growth and adapter selection issue. Extensive experiments on three medical image datasets demonstrate TPGCL's superiority in continually learning new diseases. The source code is available at https://github.com/z1968357787/TPGCL Zhiping Zhou, Yizhe Zhang 0001, Wei-Shi Zheng 0001 |
BIBM | 4 |
| 2025 | LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsabstractRecent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset are available at https://github.com/iSEE-Laboratory/LLMDet. Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei-Shi Zheng 0001 |
CVPR | 8 |
| 2025 | Modeling Multiple Normal Action Representations for Error Detection in Procedural TasksabstractError detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common scenario where multiple, distinct actions are valid following a given sequence of executed actions. This leads to two issues: (1) the model cannot effectively detect errors using static prototypes when the inference environment or action execution distribution differs from training; and (2) the model may also use the wrong prototypes to detect errors if the ongoing action label is not the same as the predicted one. To address this problem, we propose an Adaptive Multiple Normal Action Representation (AMNAR) framework. AMNAR predicts all valid next actions and reconstructs their corresponding normal action representations, which are compared against the ongoing action to detect errors. Extensive experiments demonstrate that AMNAR achieves state-of-the-art performance, highlighting the effectiveness of AMNAR and the importance of modeling multiple valid next actions in error detection. The code is available at https://github.com/iSEE-Laboratory/AMNAR. Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia, Yu-Ming Tang, Kun-Yu Lin, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 7 |
| 2025 | Person De-reidentification: A Variation-guided Identity Shift ModelingabstractPerson re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face recognition) have developed machine unlearning techniques to address privacy concerns, such methods have not yet been explored within the Re-ID field. In this work, we pioneer the exploration of the person de-reidentification (De-ReID) problem and present its inherent challenges. In the context of ReID, De-ReID is to unlearn the knowledge about accurately matching specific persons so that these "unlearned persons" cannot be re-identified across cameras for privacy guarantee. The primary challenge is to achieve the unlearning without degrading the identity-discriminative feature embeddings to ensure the model’s utility. To address this, we formulate a De-ReID framework that utilizes a labeled dataset of un-learned persons for unlearning and an unlabeled dataset of accessible persons for knowledge preservation. Instead of unlearning based on (pseudo) identity labels, we introduce a variation-guided identity shift mechanism that unlearns the specific persons by fitting the variations in their images while preserving ReID ability on other persons by overcoming the variations in images of accessible persons. As a result, the model shifts the unlearned persons to a feature space that is vulnerable to cross-view variations. Extensive experiments on benchmarks demonstrate the superiority of our method. Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Qize Yang, Jingke Meng, Xihan Wei, Wei-Shi Zheng 0001 |
CVPR | 7 |
| 2025 | RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View ImagesabstractThis paper presents RoGSplat, a novel approach for synthesizing high-fidelity novel views of unseen human from sparse multi-view images, while requiring no cumbersome per-subject optimization. Unlike previous methods that typically struggle with sparse views with few overlappings and are less effective in reconstructing complex human geometry, the proposed method enables robust reconstruction in such challenging conditions. Our key idea is to lift SMPL vertices to dense and reliable 3D prior points representing accurate human body geometry, and then regress human Gaussian parameters based on the points. To account for possible misalignment between SMPL model and images, we propose to predict image-aligned 3D prior points by leveraging both pixel-level features and voxel-level features, from which we regress the coarse Gaussians. To enhance the ability to capture high-frequency details, we further render depth maps from the coarse 3D Gaussians to help regress fine-grained pixel-wise Gaussians. Experiments on several benchmark datasets demonstrate that our method outperforms state-of-the-art methods in novel view synthesis and cross-dataset generalization. Our code is available at https://github.com/iSEE-Laboratory/RoGSplat. Junjin Xiao, Qing Zhang 0006, Yonewei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2025 | Diffusion-based Event Generation for High-Quality Image DeblurringabstractWhile event-based deblurring have demonstrated impressive results, they are impractical for consumer photos captured by cell phones and digital cameras that are not equipped with the event sensor. To address this problem, we in this paper propose a novel deblurring framework called Event Generation Deblurring (EGDeblurring), which allows to effectively deblur an image by generating event guidance describing the motion information using a diffusion model. Specifically, we design a motion prior generation diffusion model and a feature extractor to produce prior information beneficial for deblurring, rather than generating the raw event representation. In order to achieve effective fusion of motion prior information with blurry images and produce high-quality results, we develop a regression deblurring network embedded with a dual-attention channel fusion block. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image deblurring methods. Our code is available at https://github.com/XinanXie/EGDeblurring. Xinan Xie, Qing Zhang 0006, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2025 | ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationabstractWe propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly modeling joint-level interactions is more natural and effective for generating realistic HOIs, as it directly captures the geometric and semantic relationships between joints, rather than modeling interactions in the latent pose space. To this end, ChainHOI introduces a novel joint graph to capture potential interactions with objects, and a Generative Spatiotemporal Graph Convolution Network to explicitly model interactions at the joint level. Furthermore, we propose a Kinematics-based Interaction Module that explicitly models interactions at the kinetic chain level, ensuring more realistic and biomechanically coherent motions. Evaluations on two public datasets demonstrate that ChainHOI significantly outperforms previous methods, generating more realistic, and semantically consistent HOIs. Code is available here. Ling-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu, Yu-Ming Tang, Jingke Meng, Wei-Shi Zheng 0001 |
CVPR | 7 |
| 2025 | Panorama Generation From NFoV Image Done RightabstractGenerating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitable for evaluating the distortion. In this work, we first propose a distortion-specific CLIP, named Distort-CLIP to accurately evaluate the panorama distortion and discover the "visual cheating" phenomenon in previous works (i.e., tending to improve the visual results by sacrificing distortion accuracy). This phenomenon arises because prior methods employ a single network to learn the distinct panorama distortion and content completion at once, which leads the model to prioritize optimizing the latter. To address the phenomenon, we propose PanoDecouple, a decoupled diffusion model framework, which decouples the panorama generation into distortion guidance and content completion, aiming to generate panoramas with both accurate distortion and visual appeal. Specifically, we design a DistortNet for distortion guidance by imposing panorama-specific distortion prior and a modified condition registration mechanism; and a ContentNet for content completion by imposing perspective image information. Additionally, a distortion correction loss function with Distort-CLIP is introduced to constrain the distortion explicitly. The extensive experiments validate that PanoDecouple surpasses existing methods both in distortion and visual metrics. Dian Zheng, Xiao-Ming Wu 0002, Cao Li, Chengfei Lv, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 7 |
| 2025 | Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksabstractIn this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through the theoretical framework, we point out that a class of previous methods could be mainly formulated as a loss that implicitly optimizes the forgetting term while lacking supervision for the retention term, disturbing the distribution of pre-trained model and struggling to adequately preserve knowledge of the remaining classes. To address it, we refine the retention term using "dark knowledge" and propose a mask distillation unlearning method. By applying a mask to separate forgetting logits from retention logits, our approach optimizes both the forgetting and refined retention components simultaneously, retaining knowledge of the remaining classes while ensuring thorough forgetting of the target class. Without access to the remaining data or intervention (i.e., used in some works), we achieve state-of-the-art performance across various benchmarks. What’s more, DELETE is a general solution that can be applied to various downstream tasks, including face recognition, backdoor defense, and semantic segmentation with great performance. Dian Zheng, Qijie Mo, Renjie Lu 0002, Kun-Yu Lin, Wei-Shi Zheng 0001 |
CVPR | 6 |
| 2025 | EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and CompletionabstractThis paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes only entities in the non-inpainting area as reference, avoiding the possibility to generate unpredictable sundries. Moreover, we develop two entity segmentation based metrics for quantitatively assessing the performance of object erasure, which are shown be more effective than existing metrics. Experimental results demonstrate that our approach outperforms other state-of-the-art object erasure methods. Our code and data are available at https://zyxunh.github.io/EntityErasure-ProjectPage/. Yixing Zhu, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2025 | ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng, Kun-Yu Lin, Xihan Wei, Jianfang Hu, Xiaohua Xie, Wei-Shi Zheng 0001 |
ICCV | 9 |
| 2025 | Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkabstractBimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their actions. However, we think bimanual manipulation involves not only coordinated tasks but also various uncoordinated tasks that do not require explicit cooperation during execution, such as grasping objects with the closest hand, which integrated control frameworks ignore to consider due to their enforced cooperation in the early inputs. In this paper, we propose a novel decoupled interaction framework that considers the characteristics of different tasks in bimanual manipulation. The key insight of our framework is to assign an independent model to each arm to enhance the learning of uncoordinated tasks, while introducing a selective interaction module that adaptively learns weights from its own arm to improve the learning of coordinated tasks. Extensive experiments on seven tasks in the RoboTwin dataset demonstrate that: (1) Our framework achieves outstanding performance, with a 23.5% boost over the SOTA method. (2) Our framework is flexible and can be seamlessly integrated into existing methods. (3) Our framework can be effectively extended to multi-agent manipulation tasks, achieving a 28% boost over the integrated control SOTA. (4) The performance boost stems from the decoupled design itself, surpassing the SOTA by 16.5% in success rate with only 1/6 of the model size. Jian-Jian Jiang, Xiao-Ming Wu 0002, Yi-Xiang He, Ling-An Zeng, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ICCV | 7 |
| 2025 | ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsabstractReferring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existing methods still exhibit a noticeable gap when considering all these aspects. In this work, we propose \textbf{ReferDINO}, a strong RVOS model that inherits region-level vision-language alignment from foundational visual grounding models, and is further endowed with pixel-level dense perception and cross-modal spatiotemporal reasoning. In detail, ReferDINO integrates two key components: 1) a grounding-guided deformable mask decoder that utilizes location prediction to progressively guide mask prediction through differentiable deformation mechanisms; 2) an object-consistent temporal enhancer that injects pretrained time-varying text features into inter-frame interaction to capture object-aware dynamic changes. Moreover, a confidence-aware query pruning strategy is designed to accelerate object decoding without compromising model performance. Extensive experimental results on five benchmarks demonstrate that our ReferDINO significantly outperforms previous methods (e.g., +3.9% (\mathcal{J}&\mathcal{F}) on Ref-YouTube-VOS) with real-time inference speed (51 FPS). Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Jianfang Hu |
ICCV | 5 |
| 2025 | FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionabstractPre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently. However, existing CLIP-based methods often focus on learning OOD-related knowledge to improve OOD detection, showing limited generalization or reliance on external large-scale auxiliary datasets. In this study, instead of delving into the intricate OOD-related knowledge, we propose an innovative CLIP-based framework based on Forced prompt leArning (FA), designed to make full use of the In-Distribution (ID) knowledge and ultimately boost the effectiveness of OOD detection. Our key insight is to learn a prompt (i.e., forced prompt) that contains more diversified and richer descriptions of the ID classes beyond the textual semantics of class labels. Specifically, it promotes better discernment for ID images, by forcing more notable semantic similarity between ID images and the learnable forced prompt. Moreover, we introduce a forced coefficient, encouraging the forced prompt to learn more comprehensive and nuanced descriptions of the ID classes. In this way, FA is capable of achieving notable improvements in OOD detection, even when trained without any external auxiliary datasets, while maintaining an identical number of trainable parameters as CoOp. Extensive empirical evaluations confirm our method consistently outperforms current state-of-the-art methods. Code is available at https://github.com/0xFAFA/FA. Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Wei-Shi Zheng 0001 |
ICCV | 5 |
| 2025 | monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language Navigation
Renjie Lu 0002, Hao Cheng 0012, Jingke Meng, Wei-Shi Zheng 0001 |
ICCV | 5 |
| 2025 | AffordDexGrasp: Open-Set Language-Guided Dexterous Grasp With Generalizable-Instructive Affordance
Yi-Lin Wei, Mu Lin, Jian-Jian Jiang, Xiao-Ming Wu 0002, Ling-An Zeng, Wei-Shi Zheng 0001 |
ICCV | 7 |
| 2025 | Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled Learning
Zhi-Wei Xia, Kun-Yu Lin, Yuan-Ming Li, Wei-Jin Huang, Xian-Tuo Tan, Wei-Shi Zheng 0001 |
ICCV | 6 |
| 2025 | Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow RemovalabstractWe present a diffusion-based portrait shadow removal approach that can robustly produce high-fidelity results. Unlike previous methods, we cast shadow removal as diffusion-based inpainting. To this end, we first train a shadow-independent structure extraction network on a real-world portrait dataset with various synthetic lighting conditions, which allows to generate a shadow-independent structure map including facial details while excluding the unwanted shadow boundaries. The structure map is then used as condition to train a structure-guided inpainting diffusion model for removing shadows in a generative manner. Finally, to restore the fine-scale details (e.g., eyelashes, moles and spots) that may not be captured by the structure map, we take the gradients inside the shadow regions as guidance and train a detail restoration diffusion model to refine the shadow removal result. Extensive experiments on the benchmark datasets show that our method clearly outperforms existing methods, and is effective to avoid previously common issues such as facial identity tampering, shadow residual, color distortion, structure blurring, and loss of details. Our code is available at https://github.com/wanchang-yu/Structure-Guided-Diffusion-for-Portrait-Shadow-Removal. Wanchang Yu, Qing Zhang 0006, Rongjia Zheng, Wei-Shi Zheng 0001 |
ICCV | 4 |
| 2025 | Learning Implicit Features with Flow-Infused Transformations for Realistic Virtual Try-On
Delong Zhang, Qiwei Huang, Yuanliu Liu, Wei-Shi Zheng 0001, Pengfei Xiong, Wei Zhang 0009 |
ICCV | 5 |
| 2025 | Viperson: Flexibly Generating Virtual Identity for Person Re-Identification
Xiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Zhi Ouyang, Jingke Meng, Wei-Shi Zheng 0001 |
ICCV | 6 |
| 2025 | iManip: Skill-Incremental Learning for Robotic Manipulation
Zexin Zheng, Jia-Feng Cai, Xiao-Ming Wu 0002, Yi-Lin Wei, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001 |
ICCV | 7 |
| 2025 | DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse RenderingabstractRecent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm essentially utilizes noisy images with deteriorated structure and appearance for intrinsic prediction, while it is common knowledge that structure and appearance information in an image are crucial for inverse rendering. To address this issue, we present DNF-Intrinsic, a robust yet efficient inverse rendering approach fine-tuned from a pre-trained diffusion model, where we propose to take the source image rather than Gaussian noise as input to directly predict deterministic intrinsic properties via flow matching. Moreover, we design a generative renderer to constrain that the predicted intrinsic properties are physically faithful to the source image. Experiments on both synthetic and real-world datasets show that our method clearly outperforms existing state-of-the-art methods. Rongjia Zheng, Qing Zhang 0006, Chengjiang Long, Wei-Shi Zheng 0001 |
ICCV | 4 |
| 2025 | Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI GenerationabstractWe propose a novel approach for generating text-guided human-object interactions (HOIs) that achieves explicit joint-level interaction modeling in a computationally efficient manner. Previous methods represent the entire human body as a single token, making it difficult to capture fine-grained joint-level interactions and resulting in unrealistic HOIs. However, treating each individual joint as a token would yield over twenty times more tokens, increasing computational overhead. To address these challenges, we introduce an Efficient Explicit Joint-level Interaction Model (EJIM). EJIM features a Dual-branch HOI Mamba that separately and efficiently models spatiotemporal HOI information, as well as a Dual-branch Condition Injector for integrating text semantics and object geometry into human and object motions. Furthermore, we design a Dynamic Interaction Block and a progressive masking mechanism to iteratively filter out irrelevant joints, ensuring accurate and nuanced interaction modeling. Extensive quantitative and qualitative evaluations on public datasets demonstrate that EJIM surpasses previous works by a large margin while using only 5% of the inference time. Code is available here. Guohong Huang, Ling-An Zeng, Zexin Zheng, Shengbo Gu, Wei-Shi Zheng 0001 |
ICME | 5 |
| 2025 | CosGaussian: Towards Text-to-3D Semantically Controllable 3D Object Style Transfer with Gaussian SplattingabstractWith the rapid advancement of 3D technologies, 3D editing techniques have become increasingly important, among which 3D stylization serves as a crucial tool for editing 3D surfaces. However, semantically controllable object-level 3D stylization remains an open challenge. In this article, we propose CosGaussian, a framework based on 3DGS that enables precise control of the transfer style in object-level. Firstly, by leveraging Vision-Language Models (VLMs), CosGaussian can effectively extract the stylization intentions of textual input from user and generate corresponding style reference image for subsequent stylization. Then, we introduce Adaptive Feature Transfer (AFT), a semantics-aware style transfer method that can accurately transfer the local style from style reference image to objects in 3D scenes. Experiments demonstrate that CosGaussian outperforms state-of-the-art methods in both stylization accuracy and quality, presenting a robust, user-friendly solution for semantically controllable 3D style transfer. Gaojie Wu, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2025 | Supplementary Material for "NoiseActor: A Noise-Action Collaborative Framework for Privacy-Preserving Action Recognition without Privacy Labels"abstractOur supplementary material is organized into the following sections: • Section II provides the details of the LOCATION module. • Section III provides the evaluation protocol for the SBU dataset. • Section IV provides the evaluation protocol for cross-dataset experiment on the UCF101 dataset and the VISPR dataset. • Section V provides implementation details for each datasets. • Section VI provides more anonymized frames on the SBU dataset. • Section VII provides more anonymized frames on the UCF101 dataset. • Section VIII provides details about anonymized video file. • Section IX provides ablation studies of the adapter. Xiao Li 0074, Xiao-Ming Wu 0002, Delong Zhang, Kun-Yu Lin, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
ICME | 7 |
| 2025 | Task-Oriented 6-DoF Grasp Pose Detection in CluttersabstractIn general, humans would grasp an object differently for different tasks, e.g., “grasping the handle of a knife to cut” vs. “grasping the blade to hand over”. In the field of robotic grasp pose detection research, some existing works consider this task-oriented grasping and made some progress, but they are generally constrained by low-DoF gripper type or non-cluttered setting, which is not applicable for human assistance in real life. With an aim to get more general and practical grasp models, in this paper, we investigate the problem named Task-Oriented 6-DoF Grasp Pose Detection in Clutters (TO6DGC), which extends the task-oriented problem to a more general 6-DOF Grasp Pose Detection in Cluttered (multi-object) scenario. To this end, we construct a large-scale 6-DoF task-oriented grasping dataset, 6-DoF Task Grasp (6DTG), which features 4391 cluttered scenes with over 2 million 6-DoF grasp poses. Each grasp is annotated with a specific task, involving 6 tasks and 198 objects in total. Moreover, we propose One-Stage TaskGrasp (OSTG), a strong baseline to address the TO6DGC problem. Our OSTG adopts a task-oriented point selection strategy to detect where to grasp, and a task-oriented grasp generation module to decide how to grasp given a specific task. To evaluate the effectiveness of OSTG, extensive experiments are conducted on 6DTG. The results show that our method outperforms various baselines on multiple metrics. Real robot experiments also verify that our OSTG has a better perception of the task-oriented grasp points and 6-DoF grasp poses. An-Lan Wang, Kun-Yu Lin, Yuan-Ming Li, Wei-Shi Zheng 0001 |
ICRA | 5 |
| 2025 | TacCap: A Wearable FBG-Based Tactile Sensor for Efficient Human-to-Robot Skill TransferabstractTactile sensing is essential for dexterous manipulation, yet large-scale human demonstration datasets lack tactile feedback, limiting their effectiveness in skill transfer to robots. To address this, we introduce TacCap, a wearable Fiber Bragg Grating (FBG)-based tactile sensor designed for seamless human-to-robot transfer. TacCap is lightweight, durable, and immune to electromagnetic interference, making it ideal for real-world data collection. We detail its design and fabrication, evaluate its sensitivity, repeatability, and cross-sensor consistency, and assess its effectiveness through grasp stability prediction and ablation studies. Our results demonstrate that TacCap enables transferable tactile data collection, bridging the gap between human demonstrations and robotic execution, with broad implications for fine-motor disciplines such as surgical training and musical performance. To support further research and development, we open-source our hardware design and software. Chengyi Xing, Hao Li 0076, Yi-Lin Wei, Tian-Ao Ren, Tianyu Tu, Elizabeth Schumann, Wei-Shi Zheng 0001, Mark R. Cutkosky |
IROS | 8 |
| 2025 | Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection
Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng 0001 |
MICCAI (5) | 5 |
| 2025 | Global and Local Vision-Language Alignment for Few-Shot Learning and Few-Shot OOD Detection
Xiaoyuan Guan, Wei-Shi Zheng 0001, Hao Chen 0011 |
MICCAI (5) | 3 |
| 2025 | Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationabstractThe Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches rely on deterministic and discriminative models to complete semantic maps, overlooking the inherent uncertainty in indoor layouts and limiting their ability to generalize to unseen environments. In this work, we propose GOAL, a generative flow-based framework that models the semantic distribution of indoor environments by bridging observed regions with LLM-enriched full-scene semantic maps. During training, spatial priors inferred from large language models (LLMs) are encoded as two-dimensional Gaussian fields and injected into target maps, distilling rich contextual knowledge into the flow model and enabling more generalizable completions. Extensive experiments demonstrate that GOAL achieves state-of-the-art performance on MP3D and Gibson, and shows strong generalization in transfer settings to HM3D. Badi Li, Renjie Lu 0002, Jingke Meng, Wei-Shi Zheng 0001 |
NeurIPS | 5 |
| 2025 | Multi-view Style Distillation for 3D Gaussian Splatting
Xiao Li 0074, Wei-Shi Zheng 0001 |
PRCV (10) | 3 |
| 2025 | Deep Concept Forgetting in Text-to-Image Diffusion Models
Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
PRCV (2) | 4 |
| 2025 | Transformer for Object Re-identification: A Survey
Mang Ye, Shuoyi Chen, Chenyue Li, Wei-Shi Zheng 0001, David Crandall, Bo Du 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | DiffuVolume: Diffusion Model for Volume based Stereo Matching
Dian Zheng, Xiao-Ming Wu 0002, Zuhao Liu 0002, Jingke Meng, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Protecting Feature Privacy in Person Re-IdentificationabstractPerson re-identification (ReID) is to identify the same person across non-overlapping camera views. After a decade of development, the methods based on deep networks have achieved high performance on benchmarks and become mainstream. In applications, the features of gallery images extracted by deep learning-based methods are stored to speed up the query process and protect the sensitive information contained in the images. Unfortunately, it is demonstrated that turning the images into features cannot properly protect privacy, as these features could be reversed to the corresponding images, revealing the sensitive information they contain. Therefore, for preventing privacy leakage, recent methods learn their features against some feature reversal methods, and most conventional reversal methods focus on minimizing the difference between a reconstruction and its original image. However, there could be many reasonable reconstruction results from a single feature, and the conventional reversal methods will inevitably generate reconstruction results that lie in a different distribution from one of the original images, which cannot properly assess the private information for learning to protect and thus hamper the privacy-protected feature learning. To mitigate this problem, we enforce the reconstructions to follow the same distribution as the original images by the generative adversarial network (GAN). We operate this GAN-based feature reversal module accompanied by the conventional ReID feature extraction module and form a novel GAN-based feature privacy-protected person ReID model, which is expected to protect feature privacy so as against reversal attack and maintain ReID utility. We demonstrate that optimizing ReID model to accommodate privacy protection faces a double adversarial objective and is thus challenging. As a remedy, we design a novel two-step training and lazy update strategy that alternatively optimizes the feature extraction module and stabilizes the update process of the GAN-based feature reversal module. To evaluate the efficiency of the model in balancing its ReID utility and feature privacy protection, we introduce a novel metric called utility-reversibility ratio (URR). Compared with existing privacy-protected feature extraction models, the proposed method achieves a better balance between privacy protection and person ReID performance. Extensive experiments validate that our model can effectively protect feature privacy at a tiny accuracy cost, and validate the effectiveness of our model with the emerging diffusion model. Xiao Li 0074, Yi-Xing Peng, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Human-Centric Transformer for Domain Adaptive Action RecognitionabstractWe study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-sufficient source domain to a label-free target domain. Since actions are performed by humans, it is crucial to exploit human cues in videos when recognizing actions across domains. However, existing methods are prone to losing human cues but prefer to exploit the correlation between non-human contexts and associated actions for recognition, and the contexts of interest agnostic to actions would reduce recognition performance in the target domain. To overcome this problem, we focus on uncovering human-centric action cues for domain adaptive action recognition, and our conception is to investigate two aspects of human-centric action cues, namely human cues and human-context interaction cues. Accordingly, our proposed Human-Centric Transformer (HCTransformer) develops a decoupled human-centric learning paradigm to explicitly concentrate on human-centric action cues in domain-variant video feature learning. Our HCTransformer first conducts human-aware temporal modeling by a human encoder, aiming to avoid a loss of human cues during domain-invariant video feature learning. Then, by a Transformer-like architecture, HCTransformer exploits domain-invariant and action-correlated contexts by a context encoder, and further models domain-invariant interaction between humans and action-correlated contexts. We conduct extensive experiments on three benchmarks, namely UCF-HMDB, Kinetics-NecDrone and EPIC-Kitchens-UDA, and the state-of-the-art performance demonstrates the effectiveness of our proposed HCTransformer. Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Distilling the Unknown to Unveil CertaintyabstractOut-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper presents a flexible framework for OOD knowledge distillation that extracts OOD-sensitive information from a network to develop a binary classifier capable of distinguishing between ID and OOD samples in both scenarios, with and without access to training ID data. To accomplish this, we introduce Confidence Amendment (CA), an innovative methodology that transforms an OOD sample into an ID one while progressively amending prediction confidence derived from the network to enhance OOD sensitivity. This approach enables the simultaneous synthesis of both ID and OOD samples, each accompanied by an adjusted prediction confidence, thereby facilitating the training of a binary classifier sensitive to OOD. Theoretical analysis provides bounds on the generalization error of the binary classifier, demonstrating the pivotal role of confidence amendment in enhancing OOD sensitivity. Extensive experiments spanning various datasets and network architectures confirm the efficacy of the proposed method in detecting OOD samples. Zhi-Lin Zhao 0001, Longbing Cao, Yixuan Zhang 0006, Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | A Versatile Framework for Multi-Scene Person Re-IdentificationabstractPerson Re-identification (ReID) has been extensively developed for a decade in order to learn the association of images of the same person across non-overlapping camera views. To overcome significant variations between images across camera views, mountains of variants of ReID models were developed for solving a number of challenges, such as resolution change, clothing change, occlusion, modality change, and so on. Despite the impressive performance of many ReID variants, these variants typically function distinctly and cannot be applied to other challenges. To our best knowledge, there is no versatile ReID model that can handle various ReID challenges at the same time. This work contributes to the first attempt at learning a versatile ReID model to solve such a problem. Our main idea is to form a two-stage prompt-based twin modeling framework called VersReID. Our VersReID firstly leverages the scene label to train a ReID Bank that contains abundant knowledge for handling various scenes, where several groups of scene-specific prompts are used to encode different scene-specific knowledge. In the second stage, we distill a V-Branch model with versatile prompts from the ReID Bank for adaptively solving the ReID of different scenes, eliminating the demand for scene labels during the inference stage. To facilitate training VersReID, we further introduce the multi-scene properties into self-supervised learning of ReID via a multi-scene prioris data augmentation (MPDA) strategy. Through extensive experiments, we demonstrate the success of learning an effective and versatile ReID model for handling ReID tasks under multi-scene conditions without manual assignment of scene labels in the inference stage, including general, low-resolution, clothing change, occlusion, and cross-modality scenes. Wei-Shi Zheng 0001, Junkai Yan, Yi-Xing Peng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | MultiSpectral Transformer Fusion via exploiting similarity and complementarity for robust pedestrian detection
Song Hou, Meng Yang 0001, Wei-Shi Zheng 0001, Shibo Gao |
Pattern Recognit. | 3 |
| 2025 | Distilling consistent relations for multi-source domain adaptive person re-identification
Yuqiao Xian, Yi-Xing Peng, Xing Sun 0001, Wei-Shi Zheng 0001 |
Pattern Recognit. | 4 |
| 2025 | Iterative UKF under generalized maximum correntropy criterion for intermittent observation systems with complex non-Gaussian noiseabstractThe traditional unscented Kalman filters (UKFs) under the maximum correntropy criterion provide a powerful tool for nonlinear state estimation with heavy-tailed non-Gaussian noise. Nevertheless, the above-mentioned filters may yield biased estimates because the Gaussian kernel function can only handle certain types of non-Gaussian noise. Additionally, the use of statistical linearization methods can result in approximation errors when solving linear observation equations, while the system may also experience observation data loss. Therefore, a new iterative UKF with intermittent observations under the generalized maximum correntropy criterion is proposed for systems with complex non-Gaussian noise, called GMCC-IO-IUKF. Firstly, the connection between the UKF with and without intermittent observations is established by designing a coefficient matrix including intermittent observation variables, so as to derive the UKF with intermittent observations under the maximum correntropy criterion. Secondly, for the measurement update of GMCC-IO-IUKF, a nonlinear regression augmented model that can deal with both prediction and observation errors is established using the coefficient matrix and the nonlinear function . To better adapt to different types of non-Gaussian noise, the generalized Gaussian kernel function is substituted for the traditional Gaussian kernel function. Theoretically, GMCC-IO-IUKF can achieve better estimation performance by directly employing the nonlinear function and the latest iteration value. Finally, a classical target tracking model is used to evaluate the estimation performance and feasibility of our proposed GMCC-IO-IUKF algorithm. It appears from the experiment results that our proposed GMCC-IO-IUKF can not only promote estimation precision but also handle complex non-Gaussian noise flexibly. Min Zhang 0051, Xinmin Song, Wei-Shi Zheng 0001 |
Signal Process. | 3 |
| 2025 | DAT: Dual-Branch Adapter-Tuning for Few-Shot RecognitionabstractParameter-Efficient Fine-Tuning methods based on vision-language models (such as CLIP) for few-shot learning have recently received considerable attention. However, previous works only fine-tune either the image or text branch, breaking the alignment of the original two branches, meanwhile fine-tuning both branches of the CLIP would inevitably introduce more trainable parameters and likely cause more severe over-fitting due to the limited training data. In this study, we propose a novel Dual-branch Adapter-Tuning framework (DAT), which collaboratively trains the visual adapter and textual adapter added to the two branches of the original CLIP with multiple consistency constraints. By effectively utilizing the semantically detailed class-specific prompts and outputs of the original CLIP to guide the fine-tuning of both branches, our method gains exceptional adaptation ability to the downstream few-shot learning tasks and alleviates the over-fitting issue, meanwhile maximally preserving the generalization ability of the original CLIP model. Our proposed framework has achieved superior performance on diverse datasets under various few-shot learning settings compared to the existing approaches. The source code is available athttps://github.com/SandyXi/DAT. Junxi Chen, Guangxing Wu, Hongxiang Li 0004, Jiankang Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Weakly-Supervised Temporal Action Localization by Progressive Complementary LearningabstractWeakly-Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels as supervision. A critical challenge of WSTAL is the large gap between video-level supervision and unavailable snippet-level supervision. Prevailing methods typically assign pseudo labels to snippets, but these methods suffer from significant noise caused by the pseudo snippet-level labels. In this work, we address the WSTAL from a novel category exclusion perspective, which gradually enhances the snippet-level supervision to bridge the gap. Our proposed Progressive Complementary Learning (ProCL) is inspired by the fact that, video-level labels precisely indicate the categories that all snippets surely do not belong to, which is ignored by previous works. Accordingly, we first exclude these surely non-existent categories by the deterministic complementary learning. And then, we introduce the entropy-based pseudo complementary learning that is able to exclude more categories for snippets of less ambiguity. Furthermore, for the remaining ambiguous snippets, we attempt to reduce the ambiguity by distinguishing foreground actions from the background. Extensive experimental results show that our method achieves new state-of-the-art performance on THUMOS14, ActivityNet1.3, and MultiTHUMOS benchmarks. Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin, Fa-Ting Hong, Zhongang Qi, Ying Shan, Jianfang Hu, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Progressive Human Motion Generation Based on Text and Few Motion FramesabstractAlthough existing text-to-motion (T2M) methods can produce realistic human motion from text description, it is still difficult to align the generated motion with the desired postures since using text alone is insufficient for precisely describing diverse postures. To achieve more controllable generation, an intuitive way is to allow the user to input a few motion frames describing precise desired postures. Thus, we explore a new Text-Frame-to-Motion (TF2M) generation task that aims to generate motions from text and very few given frames. Intuitively, the closer a frame is to a given frame, the lower the uncertainty of this frame is when conditioned on this given frame. Hence, we propose a novel Progressive Motion Generation (PMG) method to progressively generate a motion from the frames with low uncertainty to those with high uncertainty in multiple stages. During each stage, new frames are generated by a Text-Frame Guided Generator conditioned on frame-aware semantics of the text, given frames, and frames generated in previous stages. Additionally, to alleviate the train-test gap caused by multi-stage accumulation of incorrectly generated frames during testing, we propose a Pseudo-frame Replacement Strategy for training. Experimental results show that our PMG outperforms existing T2M generation methods by a large margin with even one given frame, validating the effectiveness of our PMG. Code is available here. Ling-An Zeng, Gaojie Wu, Ancong Wu, Jianfang Hu, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A Semantically Guided and Focused Network for Occluded Person Re-IdentificationabstractPerson re-identification (ReID) is vital for surveillance, tracking, and criminal investigations, yet occlusions often lead to partial information loss and noisy features that significantly degrade ReID performance. Recent CLIP-based occluded person ReID methods have demonstrated promising performance by leveraging cross-modal alignment, but still face two limitations: first, generic text prompts fail to capture the fine-grained semantics of specific samples; second, there is a lack of effective enhancement mechanisms for hard local features in occlusion scenarios. To overcome these limitations, we propose a Semantically Guided and Focused Network (SGFNet), which comprises three synergistic modules. First, to tackle the absence of fine-grained textual descriptions, we design a Segmentation and Text Generation (STG) module that segments pedestrian regions and generates sample-specific text features, providing detailed text descriptions and spatial information for local pedestrian regions. In addition, in order to accurately extract fine-grained features, we propose a Dual-guided Feature Refinement (DGFR) module. This module leverages a spatial attention mechanism guided by dual-semantic information to enhance discriminative fine-grained features while effectively suppressing interference from irrelevant regions. Finally, building upon the DGFR module, we further propose a Hardness-aware Semantic Focus (HASF) module. This module leverages segmentation cues to assess the difficulty of distinguishing local regions and employs a carefully designed Semantic-driven Focal Triplet loss to specifically enhance hard local feature learning, thereby improving the model’s robustness in feature extraction under occlusion scenarios. Extensive experiments demonstrate the superiority of SGFNet, achieving state-of-the-art performance on three occluded person ReID datasets while maintaining competitive results on three holistic person ReID datasets. Guorong Lin, Shunzhi Yang, Wei-Shi Zheng 0001, Zhenhua Huang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, e.g., CLIP, to inherit its generalizable recognition ability so that detectors can recognize new or novel objects. However, previous works directly align the feature space with CLIP and fail to learn the semantic knowledge effectively. In this work, we propose a hierarchical semantic distillation framework named HD-OVD to construct a comprehensive distillation process, which exploits generalizable knowledge from the CLIP model in three aspects. In the first hierarchy of HD-OVD, the detector learns fine-grainedinstance-wise semanticsfrom the CLIP image encoder by modeling relations among single objects in the visual space. Besides, we introduce text space novel-class-aware classification to help the detector assimilate the highly generalizableclass-wise semanticsfrom the CLIP text encoder, representing the second hierarchy. Lastly, abundantimage-wise semanticscontaining multi-object and their contexts are also distilled by an image-wise contrastive distillation. Benefiting from the elaborated semantic distillation in triple hierarchies, our HD-OVD inherits generalizable recognition ability from CLIP in instance, class, and image levels. Thus, we boost the novel AP on the OV-COCO dataset to 46.4% with a ResNet50 backbone, which outperforms others by a clear margin. We also conduct extensive ablation studies to analyze how each component works. Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Rethinking Temporal Context in Video-QA: A Comprehensive Study of Single-Frame Static BiasabstractVideo question answering (Video-QA) has emerged as a core task in the vision-language domain, which requires the models to understand a given video and answer textual questions related to the video. Compared to conventional image-language tasks, Video-QA is designed for improving the models' capacity of memorizing and integrating multi-frame temporal cues associated with the questions. While significant performance improvements have recently been witnessed on public benchmarks, in this work, we rethink whether these improvements truly stem from better understanding of video temporal context as expected. To this end, we accomplish a strong single-frame baseline model trained with knowledge distillation. With this model, we surprisingly find that visiting only one single frame, without incorporating multi-frame and temporal information, is sufficient to achieve state-of-the-art (SOTA) performance on multiple mainstream benchmarks. This finding reveals the prevalence of single-frame bias in current benchmarks for the first time. Around the single-frame bias, we conduct an in-depth analysis on multiple popular benchmarks, which demonstrate that: (i) merely relying on one frame is able to achieve comparable performance with SOTA temporal Video-QA models; (ii) simply ensembling the prediction scores of only 3 separate frames is able to surpass temporal SOTAs. Furthermore, we observe that most of the benchmarks are biased towards central segments, and even the latest benchmarks tailored for temporal reasoning still suffer from severe single-frame bias. In case study, we find two key properties of low-bias instances: the question emphasizes temporal dependency and contextual understanding, and the associated video content presents significant variability in scenes, actions or interactions. Through further analysis on compositional reasoning datasets, we find that constructing explicit object/event interactions upon videos to fill in well-designed temporal question templates can effectively reduce the single-frame bias during annotation. We hope our analysis helps facilitate future efforts in the field towards mitigating static bias and highlighting temporal reasoning. Tianming Liang, Jianfang Hu, Xiangyang Yu, Wei-Shi Zheng 0001, Jian-Huang Lai |
IEEE Trans. Multim. | 5 |
| 2025 | Visual Class Incremental Learning With Textual Priors Guidance Based on an Adapted Vision-Language ModelabstractAn ideal artificial intelligence (AI) system should have the capability to continually learn like humans. However, when learning new knowledge, AI systems often suffer from catastrophic forgetting of old knowledge. Although many continual learning methods have been proposed, they often ignore the issue of misclassifying similar classes and make insufficient use of textual priors of visual classes to improve continual learning performance. In this study, we propose a continual learning framework based on a pre-trained vision-language model (VLM) that does not require storing old class data. This framework utilizes parameter-efficient fine-tuning of the VLM's text encoder for constructing a shared and consistent semantic textual space throughout the continual learning process. The textual priors of visual classes are encoded by the adapted VLM's text encoder to generate discriminative semantic representations, which are then used to guide the learning of visual classes. Additionally, fake out-of-distribution (OOD) images constructed from each training image further assist in the learning of visual classes. Extensive empirical evaluations on three natural datasets and one medical dataset demonstrate the superiority of the proposed framework. Wentao Zhang 0005, Jianhui Xie, Emanuele Trucco, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Self-supervised Texture FilteringabstractDecomposing an image I into the combination of structure S and texture T components is an important problem in computational photography and image analysis. Traditional solutions are basically non-learning based, because it is difficult to construct datasets containing ground-truth decompositions or find effective structure/texture supervisions. In this article, we present a self-supervised framework for smoothing out textures while maintaining the image structures. At the core of our method is a texture-inversion observation — if structure S and texture T are well disentangled, then S-T will produce a texture-inverted image that is symmetric to the input image I=S+T and the two will be visually highly similar, while for other conditions that structure and texture are not effectively separated, the generated texture-inverted images will be less similar to the input. Based on the observation, we propose to learn texture filtering from unlabeled data by encouraging the texture inverted image generated from the filtering output to be visually more similar to the input via contrastive learning. Experiments show that our method can robustly produce high-quality texture smoothing results, and also enables various applications. Hao Jiang 0057, Rongjia Zheng, Yongwei Nie, Chunxia Xiao, Wei-Shi Zheng 0001, Qing Zhang 0006 |
ACM Trans. Graph. | 5 |
| 2025 | Towards Photorealistic Portrait Style Transfer in Unconstrained ConditionsabstractWe present a photorealistic portrait style transfer approach that allows for producing high-quality results in previously challenging unconstrained conditions, e.g., large facial perspective difference between portraits, faces with complex illumination (e.g., shadow and highlight) and occlusion, and can test without portrait parsing masks. We achieve this by developing a framework to learn robust dense correspondence across portraits for semantically aligned style transfer, where a regional style contrastive learning strategy is devised to boost the effectiveness of semantic-aware style transfer while enhancing the robustness to complex illumination. Extensive experiments demonstrate the superiority of our method. Xinbo Wang, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | FeatWalk: Enhancing Few-Shot Classification through Local View LeveragingabstractFew-shot learning is a challenging task due to the limited availability of training samples. Recent few-shot learning studies with meta-learning and simple transfer learning methods have achieved promising performance. However, the feature extractor pre-trained with the upstream dataset may neglect the extraction of certain features which could be crucial for downstream tasks. In this study, inspired by the process of human learning in few-shot tasks, where humans not only observe the whole image (`global view') but also attend to various local image regions (`local view') for comprehensive understanding of detailed features, we propose a simple yet effective few-shot learning method called FeatWalk which can utilize the complementary nature of global and local views, therefore providing an intuitive and effective solution to the problem of insufficient local information extraction from the pre-trained feature extractor. Our method can be easily and flexibly combined with various existing methods, further enhancing few-shot learning performance. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and versatility of our method.The source code is available at https://github.com/exceefind/FeatWalk. Dalong Chen, Jianjia Zhang, Wei-Shi Zheng 0001 |
AAAI | 3 |
| 2024 | TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is crucial in many real-world applications. However, intelligent models are often trained solely on in-distribution (ID) data, leading to overconfidence when misclassifying OOD data as ID classes. In this study, we propose a new learning framework which leverage simple Jigsaw-based fake OOD data and rich semantic embeddings (`anchors') from the ChatGPT description of ID knowledge to help guide the training of the image encoder. The learning framework can be flexibly combined with existing post-hoc approaches to OOD detection, and extensive empirical evaluations on multiple OOD detection benchmarks demonstrate that rich textual representation of ID knowledge and fake OOD knowledge can well help train a visual encoder for OOD detection. With the learning framework, new state-of-the-art performance was achieved on all the benchmarks. The code is available at https://github.com/Cverchen/TagFog. Jiankang Chen, Tong Zhang 0017, Wei-Shi Zheng 0001 |
AAAI | 3 |
| 2024 | Exploiting Discrepancy in Feature Statistic for Out-of-Distribution DetectionabstractRecent studies on out-of-distribution (OOD) detection focus on designing models or scoring functions that can effectively distinguish between unseen OOD data and in-distribution (ID) data. In this paper, we propose a simple yet novel ap- proach to OOD detection by leveraging the phenomenon that the average of feature vector elements from convolutional neural network (CNN) is typically larger for ID data than for OOD data. Specifically, the average of feature vector elements is used as part of the scoring function to further separate OOD data from ID data. We also provide mathematical analysis to explain this phenomenon. Experimental evaluations demonstrate that, when combined with a strong baseline, our method can achieve state-of-the-art performance on several OOD detection benchmarks. Furthermore, our method can be easily integrated into various CNN architectures and requires less computation. Source code address: https://github.com/SYSU-MIA-GROUP/statistical_discrepancy_ood. Xiaoyuan Guan, Jiankang Chen, Shenshen Bu, Wei-Shi Zheng 0001 |
AAAI | 5 |
| 2024 | Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation LearningabstractUnsupervised disentangled representation learning aims to recover semantically meaningful factors from real-world data without supervision, which is significant for model generalization and interpretability. Current methods mainly rely on assumptions of independence or informativeness of factors, regardless of interpretability. Intuitively, visually interpretable concepts better align with human-defined factors. However, exploiting visual interpretability as inductive bias is still under-explored. Inspired by the observation that most explanatory image factors can be represented by ``content + mask'', we propose a content-mask factorization network (CMFNet) to decompose an image into different groups of content codes and masks, which are further combined as content masks to represent different visual concepts. To ensure informativeness of the representations, the CMFNet is jointly learned with a generator conditioned on the content masks for reconstructing the input image. The conditional generator employs a diffusion model to leverage its robust distribution modeling capability. Our model is called the Factorized Diffusion Autoencoder (FDAE). To enhance disentanglement of visual concepts, we propose a content decorrelation loss and a mask entropy loss to decorrelate content masks in latent space and spatial space, respectively. Experiments on Shapes3d, MPI3D and Cars3d show that our method achieves advanced performance and can generate visually interpretable concept-specific masks. Source code and supplementary materials are available at https://github.com/wuancong/FDAE. Ancong Wu, Wei-Shi Zheng 0001 |
AAAI | 2 |
| 2024 | TexCIL: Text-Guided Continual Learning of Disease with Vision-Language ModelabstractCurrent intelligent diagnostic systems often catastrophically forget old knowledge when learning new diseases only from the training dataset of the new diseases. Inspired by human learning of visual classes with the effective help of language, we propose a continual learning framework based on a pre-trained visual-language model (VLM) without storing any image of previously learned diseases. In this framework, textual prior knowledge of each new disease can be obtained by utilizing the frozen VLM’s text encoder, and then used to guide the visual learning of the new disease. This framework innovatively utilizes the textual prior knowledge of all previously learned diseases as out-of-distribution (OOD) information to help differentiate currently being-learned diseases from others. Extensive empirical evaluations on both medical and natural image datasets confirm the superiority of the proposed method over existing state-of-the-art methods in continual learning of new visual classes. The source code is available at https://openi.pcl.ac.cn/OpenMedIA/TexCIL. Wentao Zhang 0005, Defeng Zhao, Wei-Shi Zheng 0001 |
BIBM | 3 |
| 2024 | Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-TrainingabstractContrastive learning has emerged as a promising paradigm for 3D open-world understanding, i.e., aligning point cloud representation to image and text embedding space individually. In this paper, we introduce Mix-Con3D, a simple yet effective method aiming to sculpt holistic 3D representation in contrastive language-image-3D pre-training. In contrast to point cloud only, we develop the 3D object-level representation from complementary perspectives, e.g., multi-view rendered images with the point cloud. Then, MixCon3D performs language-3D contrastive learning, comprehensively depicting real-world 3D objects and bolstering text alignment. Additionally, we pioneer the first thorough investigation of various training recipes for the 3D contrastive learning paradigm, building a solid baseline with improved performance. Extensive experiments conducted on three representative benchmarks reveal that our method significantly improves over the baseline, surpassing the previous state-of-the-art performance on the challenging 1,156-category Objaverse-LVIS dataset by 5.7%. The versatility of MixCon3D is showcased in applications such as text-to-3D retrieval and point cloud captioning, further evidencing its efficacy in diverse scenarios. The code is available at https://github.com/UCSC-VLAA/MixCon3D. Yipeng Gao, Zeyu Wang 0008, Wei-Shi Zheng 0001, Cihang Xie, Yuyin Zhou |
CVPR | 3 |
| 2024 | Ranking Distillation for Open-Ended Video Question Answering with Insufficient LabelsabstractThis paper focuses on open-ended video question answering, which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task, since a question may have multiple answers. However, due to annotation costs, the labels in existing benchmarks are always extremely insufficient, typically one answer per question. As a result, existing works tend to directly treat all the unlabeled answers as negative labels, leading to limited ability for generalization. In this work, we introduce a simple yet effective ranking distillation framework (RADI) to mitigate this problem without additional manual annotation. RADI employs a teacher model trained with incomplete labels to generate rankings for potential answers, which contain rich knowledge about label priority as well as label-associated visual cues, thereby enriching the insufficient labeling information. To avoid overconfidence in the imperfect teacher model, we further present two robust and parameter-free ranking distillation approaches: a pairwise approach which introduces adaptive soft margins to dynamically refine the optimization constraints on various pairwise rankings, and a listwise approach which adopts sampling-based partial listwise learning to resist the bias in teacher ranking. Extensive experiments on five popular benchmarks consistently show that both our pairwise and listwise RADIs outperform state-of-the-art methods. Further analysis demonstrates the effectiveness of our methods on the insufficient labeling problem. Tianming Liang, Chaolei Tan, Beihao Xia, Wei-Shi Zheng 0001, Jianfang Hu |
CVPR | 4 |
| 2024 | Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingabstractVideo Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal or-der from an untrimmed video. However, existing VPG approaches are heavily reliant on a considerable number of temporal labels that are laborious and time-consuming to acquire. In this work, we introduce and explore Weakly-Supervised Video Paragraph Grounding (WSVPG) to elim-inate the need of temporal annotations. Different from pre-vious weakly-supervised grounding frameworks based on multiple instance learning or reconstruction learning for two-stage candidate ranking, we propose a novel siamese learning framework that jointly learns the cross-modal feature alignment and temporal coordinate regression without timestamp labels to achieve concise one-stage localization for WSVPG. Specifically, we devise a Siamese Grounding TRansformer (SiamGTR) consisting of two weight-sharing branches for learning complementary supervision. An Aug-mentation Branch is utilized for directly regressing the tem-poral boundaries of a complete paragraph within a pseudo video, and an Inference Branch is designed to capture the order-guided feature correspondence for localizing multi-ple sentences in a normal video. We demonstrate by exten-sive experiments that our paradigm has superior practica-bility and flexibility to achieve efficient weakly-supervised or semi-supervised learning, outperforming state-of-the-art methods trained with the same or stronger supervision. Chaolei Tan, Jian-Huang Lai, Wei-Shi Zheng 0001, Jianfang Hu |
CVPR | 3 |
| 2024 | Single-View Scene Point Cloud Human Grasp GenerationabstractIn this work, we explore a novel task of generating human grasps based on single-view scene point clouds, which more accurately mirrors the typical real-world situation of observing objects from a single viewpoint. Due to the incompleteness of object point clouds and the presence of numerous scene points, the generated hand is prone to penetrating into the invisible parts of the object and the model is easily affected by scene points. Thus, we introduce S2HGrasp, a framework composed of two key modules: the Global Perception module that globally perceives partial object point clouds, and the DiffuGrasp module designed to generate high-quality human grasps based on complex inputs that include scene points. Additionally, we introduce S2HGD dataset, which comprises approximately 99,000 single-object single-view scene point clouds of 1,668 unique objects, each annotated with one human grasp. Our extensive experiments demonstrate that S2HGrasp can not only generate natural human grasps regardless of scene points, but also effectively prevent penetration between the hand and invisible parts of the object. Moreover, our model showcases strong generalization capability when applied to unseen objects. Our code and dataset are available at https://github.com/iSEE-Laboratory/S2HGrasp. Yan-Kang Wang, Chengyi Xing, Yi-Lin Wei, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2024 | NECA: Neural Customizable Human AvatarabstractHuman avatar has become a novel type of 3D asset with various applications. Ideally, a human avatar should be fully customizable to accommodate different settings and environments. In this work, we introduce NECA, an approach capable of learning versatile human representation from monocular or sparse-view videos, enabling granular customization across aspects such as pose, shadow, shape, lighting and texture. The core of our approach is to represent humans in complementary dual spaces and predict disentangled neural fields of geometry, albedo, shadow, as well as an external lighting, from which we are able to derive realistic rendering with high-frequency details via volumetric rendering. Extensive experiments demonstrate the advantage of our method over the state-of-the-art methods in photorealistic rendering, as well as various editing tasks such as novel pose synthesis and relighting. Our code is available at https://github.com/iSEE-Laboratory/NECA. Junjin Xiao, Qing Zhang 0006, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2024 | Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary AlignmentabstractWeakly-supervised action segmentation is a task of learning to partition a long video into several action segments, where training videos are only accompanied by transcripts (ordered list of actions). Most of existing methods need to infer pseudo segmentation for training by serial alignment between all frames and the transcript, which is time-consuming and hard to be parallelized while training. In this work, we aim to escape from this inefficient alignment with massive but redundant frames, and instead to directly localize a few action transitions for pseudo segmentation generation, where a transition refers to the change from an action segment to its next adjacent one in the transcript. As the true transitions are submerged in noisy boundaries due to intra-segment visual variation, we propose a novel Action-Transition-Aware Boundary Alignment (ATBA) framework to efficiently and effectively filter out noisy boundaries and detect transitions. In addition, to boost the semantic learning in the case that noise is inevitably present in the pseudo segmentation, we also introduce video-level losses to utilize the trusted video-level supervision. Extensive experiments show the effectiveness of our approach on both performance and training speed.11Code is available at https://github.com/iSEE-Laboratory/CVPR24_ATBA. Angchi Xu, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2024 | Dexterous Grasp TransformerabstractIn this work, we propose a novel discriminative frame-work for dexterous grasp generation, named Dexterous Grasp TRansformer (DGTR), capable of predicting a di-verse set of feasible grasp poses by processing the object point cloud with only one forward pass. We formulate dex-terous grasp generation as a set prediction task and design a transformer-based grasping model for it. However, we identify that this set prediction paradigm encounters sev-eral optimization challenges in the field of dexterous grasping and results in restricted performance. To address these issues, we propose progressive strategies for both the training and testing phases. First, the dynamic-static matching training (DSMT) strategy is presented to enhance the opti-mization stability during the training phase. Second, we in-troduce the adversarial-balanced test-time adaptation (AB-TTA) with a pair of adversarial losses to improve grasping quality during the testing phase. Experimental results on the DexGraspNet dataset demonstrate the capability of DGTR to predict dexterous grasp poses with both high quality and diversity. Notably, while keeping high qual-ity, the diversity of grasp poses predicted by DGTR sig-nificantly outperforms previous works in multiple metrics without any data pre-processing. Codes are available at https://github.com/iSEE-Laboratory/DGTR. Guo-Hao Xu, Yi-Lin Wei, Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2024 | Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion ModelabstractUniversal image restoration is a practical and poten-tial computer vision task for real-world applications. The main challenge of this task is handling the different degra-dation distributions at once. Existing methods mainly utilize task-specific conditions (e.g., prompt) to guide the model to learn different distributions separately, named multi-partite mapping. However, it is not suitable for universal model learning as it ignores the shared information between different tasks. In this work, we propose an advanced selective hourglass mapping strategy based on diffusion model, termed DiffUIR. Two novel considerations make our Dif-fUIR non-trivial. Firstly, we equip the model with strong condition guidance to obtain accurate generation direction of diffusion model (selective). More importantly, DiffUIR integrates a flexible shared distribution term (SDT) into the diffusion algorithm elegantly and naturally, which gradually maps different distributions into a shared one. In the reverse process, combined with SDT and strong condition guidance, DiffUIR iteratively guides the shared distribution to the task-specific distribution with high image quality (hourglass). Without bells and whistles, by only modifying the mapping strategy, we achieve state-of-the-art performance on five image restoration tasks, 22 benchmarks in the universal setting and zero-shot generalization setting. Surprisingly, by only using a lightweight model (only 0.89M), we could achieve outstanding performance. The source code and pre-trained models are available at https://github.com/iSEE-Laboratory/DiffUIR. Dian Zheng, Xiao-Ming Wu 0002, Shuzhou Yang, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 6 |
| 2024 | EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001 |
ECCV (20) | 6 |
| 2024 | PRET: Planning with Directed Fidelity Trajectory for Vision and Language Navigation
Renjie Lu 0002, Jingke Meng, Wei-Shi Zheng 0001 |
ECCV (66) | 3 |
| 2024 | Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
Qijie Mo, Yipeng Gao, Shenghao Fu, Junkai Yan, Ancong Wu, Wei-Shi Zheng 0001 |
ECCV (16) | 6 |
| 2024 | Rethinking Few-Shot Class-Incremental Learning: Learning from Yourself
Yu-Ming Tang, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
ECCV (61) | 4 |
| 2024 | An Economic Framework for 6-DoF Grasp Detection
Xiao-Ming Wu 0002, Jia-Feng Cai, Jian-Jian Jiang, Dian Zheng, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ECCV (27) | 6 |
| 2024 | DreamView: Injecting View-Specific Text Guidance Into Text-to-3D Generation
Junkai Yan, Yipeng Gao, Qize Yang, Xihan Wei, Xuansong Xie, Ancong Wu, Wei-Shi Zheng 0001 |
ECCV (25) | 7 |
| 2024 | Patch-Based Privacy Attention for Weakly-Supervised Privacy-Preserving Action RecognitionabstractPrivacy-preserving action recognition aims to prevent privacy leakage by learning to anonymize video frames for action recognition. However, apart from video-level action labels, supervised methods require costly frame-level privacy labels. To achieve privacy-preserving action recognition in the absence of privacy labels, weakly-supervised privacy-preserving action recognition proposes to merely utilize action labels for learning without using any privacy labels during training, and therefore the challenge lies in removing the privacy information without privacy annotations. Inspired by the fact that private information such as the identity of the participant is not significant to action recognition, our main idea is to utilize the attention mechanism to automatically discover the action-sensitive information in frames and remove the other information to prevent privacy leakage without using the privacy labels. Our method contains a novel patch-based privacy attention module and an action recognition module. The patch-based privacy attention module splits raw frames into patches and exploits self-attention among patches to adaptively discover action-sensitive but privacy-less information. Our patch-based privacy attention module mines the action-sensitive information from both individual frames and adjacent frames to generate anonymized frames. A distance correlation loss is introduced to enforce that the generated anonymized frames are distinguished from the original frames and contain less private information. In addition, the action recognition module learns to recognize actions based on anonymized frames. Extensive experiments demonstrate that our model can effectively alleviate privacy leakage and maintain the performance of action recognition without using privacy labels. Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Wei-Shi Zheng 0001 |
FG | 4 |
| 2024 | Out-of-Distribution Detection by Principal Component CorrespondenceabstractOut-of-distribution (OOD) detection is vital for the safe application of intelligent systems in real-world scenarios. This paper proposes an enhancement to OOD detection by leveraging the consistency in cognition between two models, both pretrained on in-distribution (ID) data. Specifically, for a given test sample, we first apply Principal Component Analysis (PCA)-based projection on the feature vectors from each model. These obtained feature vectors (with correlation between dimensions decoupled by PCA projection) are then aligned using a multiple linear mapping, which is fitted using the least squares method on the training data. We hypothesize that the regression error for OOD data will be larger than that for ID data, making it a useful metric for OOD detection. Our experimental results demonstrate the effectiveness of this method. When combined with existing robust baselines, our approach achieves state-of-the-art performance in OOD detection. Xiaoyuan Guan, Zhiyong Gan, Ling Deng, Jiankang Chen, Shenshen Bu, Chunliang Zhao, Jianfang Hu, Wei-Shi Zheng 0001 |
ICME | 10 |
| 2024 | Region Attention Fine-tuning with CLIP for Few-shot ClassificationabstractWith the advancements in visual language models such as CLIP and their strong performance in zero-shot recognition, numerous CLIP-based methods have emerged in the field of few-shot classification. However, many of them do not fully leverage the abundant feature information within the CLIP visual encoder and overlook the issue of varying region-specific importance for image classification across different datasets. To address these limitations, we present an attention pooling-based framework for few-shot fine-tuning. Our framework enables the model to learn task-specific attention weights for image regions, while also incorporating background features and a consistency constraint to enhance training. As a result, our approach outperforms the state-of-the-art approaches on 11 benchmarks, demonstrating its effectiveness. Guangxing Wu, Junxi Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001 |
ICME | 5 |
| 2024 | Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
Jia-Run Du, Kun-Yu Lin, Jingke Meng, Wei-Shi Zheng 0001 |
ICPR (16) | 4 |
| 2024 | iGrasp: An Interactive 2D-3D Framework for 6-DoF Grasp Detection
Jian-Jian Jiang, Xiao-Ming Wu 0002, Zibo Chen, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ICPR (30) | 5 |
| 2024 | Text-Guided Zero-Shot 3D Style Transfer of Neural Radiance Fields
Wei-Shi Zheng 0001 |
ICPR (8) | 2 |
| 2024 | Depth-Enhanced Alignment for Label-Free 3D Semantic Segmentation
Shangjin Xie, Zibo Chen, Zhixuan Liu, Wei-Shi Zheng 0001 |
ICPR (18) | 5 |
| 2024 | Privacy-Preserving Face Recognition with Adaptive Generative Perturbations
Delong Zhang, Yixing Peng, Ancong Wu, Wei-Shi Zheng 0001 |
ICPR (14) | 4 |
| 2024 | Fine-Grained Depth Knowledge Distillation for Cloth-Changing Person Re-identificationabstractThe mission of cloth-changing person re-identification (CC-ReID) is to discover cloth-invariant and identity-related cues, while traditional person ReID methods rely on appearance features that are biased to cloth-related cues. To tackle this cloth-biased problem, many CC-ReID methods introduced auxiliary body shape information to extract cloth-invariant features, such as 2D sketch images or the 3D Skinned Multi-Person Linear (SMPL) model. However, 2D auxiliary information lacks 3D spatial features, while the 3D SMPL model encounters challenges in capturing features at a finer granularity due to manually defined parameters. To extract fine-grained 3D shape features, we estimate depth maps that contain richer shape information and propose a Fine-grained Depth feature Mining and Distillation (FDMD) framework. We introduce a depth branch and design a fine-grained local feature interaction module to mine fine-grained 3D body shape knowledge from estimated depth maps by exploring the context of semantic-aware local body-part features. To integrate cloth-invariant depth knowledge into the appearance features, the fine-grained 3D shape features are transferred to an appearance branch by feature-space-aligned distillation. Extensive experiments demonstrate that FDMD can achieve state-of-the-art performance on three widely used CC-ReID benchmarks PRCC, Celeb-reID and LaST. Yuhan Yao 0002, Ancong Wu, Jiangqun Ni, Wei-Shi Zheng 0001 |
IJCNN | 5 |
| 2024 | DA-NAS: Learning Transferable Architecture for Unsupervised Domain Adaptation
Xiao Li 0074, Gaojie Wu, Jianjian Jiang, Wei-Shi Zheng 0001 |
KSEM (3) | 4 |
| 2024 | PTMA: Pre-trained Model Adaptation for Transfer Learning
Xiao Li 0074, Junkai Yan, Jianjian Jiang, Wei-Shi Zheng 0001 |
KSEM (1) | 4 |
| 2024 | SiFT: A Serial Framework with Textual Guidance for Federated Learning
Weizhuo Zhang, Yue Yu 0001, Wei-Shi Zheng 0001, Tong Zhang 0017 |
MICCAI (10) | 4 |
| 2024 | Deep Model Reference: Simple Yet Effective Confidence Estimation for Image Classification
Yuanhang Zheng, Yiqiao Qiu, Haoxuan Che, Hao Chen 0011, Wei-Shi Zheng 0001 |
MICCAI (10) | 5 |
| 2024 | FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution DetectorabstractOut-of-Distribution (OOD) detection is crucial when deploying machine learning models in open-world applications. The core challenge in OOD detection is mitigating the model's overconfidence on OOD data. While recent methods using auxiliary outlier datasets or synthesizing outlier features have shown promising OOD detection performance, they are limited due to costly data collection or simplified assumptions. In this paper, we propose a novel OOD detection framework FodFoM that innovatively combines multiple foundation models to generate two types of challenging fake outlier images for classifier training. The first type is based on BLIP-2's image captioning capability, CLIP's vision-language knowledge, and Stable Diffusion's image generation ability. Jointly utilizing these foundation models constructs fake outlier images which are semantically similar to but different from in-distribution (ID) images. For the second type, GroundingDINO's object detection ability is utilized to help construct pure background images by blurring foreground ID objects in ID images. The proposed framework can be flexibly combined with multiple existing OOD detection methods. Extensive empirical evaluations show that image classifiers with the help of constructed fake images can more accurately differentiate real OOD image from ID ones. New state-of-the-art OOD detection performance is achieved on multiple benchmarks. The code is available at https://github.com/Cverchen/ACMMM2024-FodFoM. Jiankang Chen, Ling Deng, Zhiyong Gan, Wei-Shi Zheng 0001 |
ACM Multimedia | 4 |
| 2024 | SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi, Wei-Yi Pei, Yexin Wang, Ying Shan, Wei-Shi Zheng 0001, Jianfang Hu |
ACM Multimedia | 9 |
| 2024 | Loc4Plan: Locating Before Planning for Outdoor Vision and Language NavigationabstractVision and Language Navigation (VLN) is a challenging task that requires agents to understand instructions and navigate to the destination in a visual environment. One of the key challenges in outdoor VLN is keeping track of which part of the instruction was completed. To alleviate this problem, previous works mainly focus on grounding the natural language to the visual input, but neglecting the crucial role of the agent's spatial position information in the grounding process. In this work, we first explore the substantial effect of spatial position locating on the grounding of outdoor VLN, drawing inspiration from human navigation. In real-world navigation scenarios, before planning a path to the destination, humans typically need to figure out their current location. This observation underscores the pivotal role of spatial localization in the navigation process. In this work, we introduce a novel framework, Locating before Planning (Loc4Plan), designed to incorporate spatial perception for action planning in outdoor VLN tasks. The main idea behind Loc4Plan is to perform the spatial localization before planning a decision action based on corresponding guidance, which comprises a block-aware spatial locating (BAL) module and a spatial-aware action planning (SAP) module. Specifically, to help the agent perceive its spatial location in the environment, we propose to learn a position predictor that measures how far the agent is from the next intersection for reflecting its position, which is achieved by the BAL module. After this locating process, we propose the PSA module to associate visual observations After the locating process, we propose the SAP module to incorporate spatial information to ground the corresponding guidance and enhance the precision of action planning. Extensive experiments on the Touchdown and map2seq datasets show that the proposed Loc4Plan outperforms the SOTA methods. Huilin Tian, Jingke Meng, Wei-Shi Zheng 0001, Yuan-Ming Li, Junkai Yan, Yunong Zhang |
ACM Multimedia | 3 |
| 2024 | PixelFade: Privacy-preserving Person Re-identification with Noise-guided Progressive Replacement
Delong Zhang, Yi-Xing Peng, Xiao-Ming Wu 0002, Ancong Wu, Wei-Shi Zheng 0001 |
ACM Multimedia | 5 |
| 2024 | Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled DataabstractThere are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorithm on training samples remains applicable to test samples. We present a novel approach called Importance Divergence (I-Div) to address the challenge of test label unavailability, enabling distribution discrepancy evaluation using only training samples. I-Div transfers the sampling patterns from the test distribution to the training distribution by estimating density and likelihood ratios. Specifically, the density ratio, informed by the selected hypothesis, is obtained by minimizing the Kullback-Leibler divergence between the actual and estimated input distributions. Simultaneously, the likelihood ratio is adjusted according to the density ratio by reducing the generalization error of the distribution discrepancy as transformed through the two ratios. Experimentally, I-Div accurately quantifies the distribution discrepancy, as evidenced by a wide range of complex data scenarios and tasks. Zhi-Lin Zhao 0001, Longbing Cao, Xuhui Fan 0001, Wei-Shi Zheng 0001 |
NeurIPS | 4 |
| 2024 | Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsabstractRecent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning them. In this work, we show that frozen foundation models can be a versatile feature enhancer, even though they are not pre-trained for object detection. Specifically, we explore directly transferring the high-level image understanding of foundation models to detectors in the following two ways. First, the class token in foundation models provides an in-depth understanding of the complex scene, which facilitates decoding object queries in the detector's decoder by providing a compact context. Additionally, the patch tokens in foundation models can enrich the features in the detector's encoder by providing semantic details. Utilizing frozen foundation models as plug-and-play modules rather than the commonly used backbone can significantly enhance the detector's performance while preventing the problems caused by the architecture discrepancy between the detector's backbone and the foundation model. With such a novel paradigm, we boost the SOTA query-based detector DINO from 49.0% AP to 51.9% AP (+2.9% AP) and further to 53.8% AP (+4.8% AP) by integrating one or two foundation models respectively, on the COCO validation set after training for 12 epochs with R50 as the detector's backbone. Code will be available. Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001 |
NeurIPS | 6 |
| 2024 | Grasp as You Say: Language-guided Dexterous Grasp GenerationabstractThis paper explores a novel task "Dexterous Grasp as You Say'' (DexGYS), enabling robots to perform dexterous grasping based on human commands expressed in natural language. However, the development of this field is hindered by the lack of datasets with natural human guidance; thus, we propose a language-guided dexterous grasp dataset, named DexGYSNet, offering high-quality dexterous grasp annotations along with flexible and fine-grained human language guidance. Our dataset construction is cost-efficient, with the carefully-design hand-object interaction retargeting strategy, and the LLM-assisted language guidance annotation system. Equipped with this dataset, we introduce the DexGYSGrasp framework for generating dexterous grasps based on human language instructions, with the capability of producing grasps that are intent-aligned, high quality and diversity. To achieve this capability, our framework decomposes the complex learning process into two manageable progressive objectives and introduce two components to realize them. The first component learns the grasp distribution focusing on intention alignment and generation diversity. And the second component refines the grasp quality while maintaining intention consistency. Extensive experiments are conducted on DexGYSNet and real world environments for validation. Yi-Lin Wei, Jian-Jian Jiang, Chengyi Xing, Xiantuo Tan, Xiao-Ming Wu 0002, Hao Li 0076, Mark R. Cutkosky, Wei-Shi Zheng 0001 |
NeurIPS | 8 |
| 2024 | Privacy-Preserving Action Recognition: A Survey
Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
PRCV (7) | 5 |
| 2024 | Dual-level feature assessment for unsupervised multi-view feature selection with latent space learning
Jian-Sheng Wu, Jun-Xiao Gong, Wei-Shi Zheng 0001 |
Inf. Sci. | 5 |
| 2024 | Cluster structure augmented deep nonnegative matrix factorization with low-rank tensor learning
Jian-Sheng Wu, Wei-Shi Zheng 0001 |
Inf. Sci. | 4 |
| 2024 | A Multi-Level Relation-Aware Transformer model for occluded person re-identification
Guorong Lin, Zhiqiang Bao, Zhenhua Huang 0001, Wei-Shi Zheng 0001, Yunwen Chen |
Neural Networks | 5 |
| 2024 | Revisiting Person Re-Identification by Camera SelectionabstractPerson re-identification (Re-ID) is a fundamental task in visual surveillance. Given a query image of the target person, conventional Re-ID focuses on the pairwise similarities between the candidate images and the query. However, conventional Re-ID does not evaluate the consistency of the retrieval results of whether the most similar images ranked in each place contain the same person, which is risky in some applications such as missing out a place where the patient passed will hinder the epidemiological investigation. In this work, we investigate a more challenging task: consistently and successfully retrieving the target person in all camera views. We define the task as continuous person Re-ID and propose a corresponding evaluation metric termed overall Rank-K accuracy. Different from the conventional Re-ID, any incorrect retrieval under an individual camera view that raises an inconsistency will fail the continuous Re-ID. Consequently, the defective cameras, in which the images are hard to be automatically associated with the images from other views, strongly degrade the performance of continuous person Re-ID. Since the camera deployment is crucial for continuous tracking across camera views, we rethink person Re-ID from the perspective of camera deployment and assess the quality of a camera network by performing continuous Re-ID. Moreover, we propose to automatically detect the defective cameras that greatly hamper the continuous Re-ID. Because brute-force search is costly when the camera network becomes complicated, we explicitly model the visual relations as well as the spatial relations among cameras and develop a relational deep Q-network to select the properly deployed cameras and the un-selected cameras are regarded as the defective cameras. Since most existing datasets do not provide topology information about the camera network, they are unsuitable for investigating the importance of spatial relations on camera selection. Thus, we collect a new dataset including 20 cameras with topology information. Compared with randomly removing cameras, the experimental results show that our method can effectively detect the defective cameras so that people could take further operations on these cameras in practice (https://www.isee-ai.cn/∼yixing/MCCPD.html). Yi-Xing Peng, Yuanxun Li, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Learning From Human Educational Wisdom: A Student-Centered Knowledge Distillation MethodabstractExisting studies on knowledge distillation typically focus on teacher-centered methods, in which the teacher network is trained according to its own standards before transferring the learned knowledge to a student one. However, due to differences in network structure between the teacher and the student, the knowledge learned by the former may not be desired by the latter. Inspired by human educational wisdom, this paper proposes a Student-Centered Distillation (SCD) method that enables the teacher network to adjust its knowledge transfer according to the student network's needs. We implemented SCD based on various human educational wisdom, e.g., the teacher network identified and learned the knowledge desired by the student network on the validation set, and then transferred it to the latter through the training set. To address the problems of current deficiency knowledge, hard sample learning and knowledge forgetting faced by a student network in the learning process, we introduce and improve Proportional-Integral-Derivative (PID) algorithms from automation fields to make them effective in identifying the current knowledge required by the student network. Furthermore, we propose a curriculum learning-based fuzzy strategy and apply it to the proposed PID control algorithm, such that the student network in SCD can actively pay attention to the learning of challenging samples after with certain knowledge. The overall performance of SCD is verified in multiple tasks by comparing it with state-of-the-art ones. Experimental results show that our student-centered distillation method outperforms existing teacher-centered ones. Shunzhi Yang, Jinfeng Yang, MengChu Zhou, Zhenhua Huang 0001, Wei-Shi Zheng 0001, Jin Ren 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | PAMI: Partition Input and Aggregate Outputs for Model Interpretation
Wentao Zhang 0005, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2024 | Continual Action Assessment via Task-Consistent Score-Discriminative Feature Distribution ModelingabstractAction Quality Assessment (AQA) is a task that tries to answer how well an action is carried out. While remarkable progress has been achieved, existing works on AQA assume that all the training data are visible for training at one time, but do not enable continual learning on assessing new technical actions. In this work, we address such a Continual Learning problem in AQA (Continual-AQA), which urges a unified model to learn AQA tasks sequentially without forgetting. Our idea for modeling Continual-AQA is to sequentially learn a task-consistent score-discriminative feature distribution, in which the latent features express a strong correlation with the score labels regardless of the task or action types. From this perspective, we aim to mitigate the forgetting in Continual-AQA from two aspects. Firstly, to fuse the features of new and previous data into a score-discriminative distribution, a novel Feature-Score Correlation-Aware Rehearsal is proposed to store and reuse data from previous tasks with limited memory size. Secondly, an Action General-Specific Graph is developed to learn and decouple the action-general and action-specific knowledge so that the task-consistent score-discriminative features can be better extracted across various tasks. Extensive experiments are conducted to evaluate the contributions of proposed components. The comparisons with the existing continual learning methods additionally verify the effectiveness and versatility of our approach. Data and code are available at https://github.com/iSEE-Laboratory/Continual-AQA. Yuan-Ming Li, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Generalized Intra-Camera Supervised Person Re-IdentificationabstractPerson re-identification (Re-ID) is to match the images of the same person from different camera views, which demands a view-invariant feature embedding. Recently, intra-camera supervised (ICS) Re-ID develops the Re-ID models without cross-view annotated data. Existing ICS methods are developed based on the assumptions, such as assuming each person in the training set appears under multiple cameras. However, there is no guarantee that the assumptions are true without cross-view annotations, and their performance degrades when the assumptions are violated. In this work, we generalize the ICS Re-ID and develop an ICS Re-ID model without the assumptions. The absence of prior assumptions and cross-view annotations poses a challenge in exploiting the discriminative information among cross-view images. To this end, we propose to mine the view-invariant relations between cross-view images for Re-ID model to exploit discriminative information and overcome the cross-view variations. Specifically, we learn composited view-aware features by compositing the identity information with different camera view information in the feature composition module. Then, we exploit the composited features to model various view-aware relations between pairwise images. By mining the common patterns among the view-aware relations, we obtain the view-invariant pairwise relation for learning. Besides, leveraging the composited view-aware features, we develop a view-aware marginal constraint for robust cross-view learning. To facilitate learning the feature composition module, we augment an auxiliary network to exploit the camera view information at the feature level. Extensive experimental results show the effectiveness of our method under different scenarios. Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Classifier-Head Informed Feature Masking and Prototype-Based Logit Smoothing for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is essential when deploying neural networks in the real world. One main challenge is that neural networks often make overconfident predictions on OOD data. In this study, we propose an effective post-hoc OOD detection method, named HIMPLoS, based on a new feature masking strategy and a novel logit smoothing strategy. Feature masking determines the important features at the penultimate layer for each in-distribution (ID) class based on the weights of the ID class in the classifier head and masks the rest features. Logit smoothing computes the cosine similarity between the feature vector of the test sample and the prototype of the predicted ID class at the penultimate layer and uses the similarity as an adaptive temperature factor on the logit to alleviate the network’s overconfidence prediction for OOD data. With these strategies, we can reduce feature activation of OOD data and enlarge the gap in OOD score between ID and OOD data. Extensive experiments on multiple standard OOD detection benchmarks demonstrate the effectiveness of our method and its compatibility with existing methods, with new state-of-the-art performance achieved from our method. The source code will be released publicly. Zhuohao Sun, Yiqiao Qiu, Zhijun Tan, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Asymmetric Mutual Learning for Unsupervised Transferable Visible-Infrared Re-IdentificationabstractVisible-infrared person re-identification (Re-ID) plays a crucial role in matching people across camera views in the darkness and normal lighting. To reduce annotation cost, it is advantageous to learn Re-ID model from unlabeled visible-infrared image pairs. However, large modality gap makes it difficult to discover the underlying cross-modality sample relations. Compared with cross-modality sample pairs in the target domain, it is easier to obtain more single-modality visible image samples from other domains. In this work, we study unsupervised transfer learning to extract modality-shared knowledge from auxiliary unlabeled visible images in a source domain and leverage this knowledge to learn cross-modality matching in the unlabeled target domain. Our framework consists of two stages: RGB-gray asymmetric mutual learning and unsupervised cross-modality self-training. In the first stage, to extract visible-infrared shared information from auxiliary unlabeled visible images, we regard RGB images and grayscale fake infrared images transformed from RGB images as two views to learn view-shared information and simultaneously preserve RGB-specific information. Based on information theoretic analysis, we learn an RGB-gray feature extractor and further introduce an auxiliary gray feature extractor to quantify RGB-gray shared knowledge. This knowledge is then transferred to the RGB-gray feature extractor without eliminating RGB-specific information. We call this process Cross-Modality Asymmetric Mutual Learning (CMAM). In the second stage, for unsupervised cross-modality self-training in the target domain, we fuse the complementary knowledge in two models by mutual learning and employ bipartite cross-modality pseudo labeling to alleviate modality gap. For a more extensive evaluation, we collected a new public multi-modality dataset, SYSU-MM02, constructed from untrimmed videos. Our method achieves the state-of-the-art performance on three benchmark datasets. Ancong Wu, Chengzhi Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Improving Knowledge Distillation via Head and Tail CategoriesabstractKnowledge distillation (KD) is a technique that transfers “dark knowledge” from a deep teacher network (teacher) to a shallow student network (student). Despite significant advances in KD, existing work has not adequately mined two crucial types of knowledge: 1) the knowledge of head categories, which represents the relationship between the target category and its similar categories. Our findings reveal that this highly similar (complex) knowledge is essential for improving student’s performance; and 2) the effectively utilized knowledge of tail categories. Existing studies often treat the non-target categories collectively without sufficiently considering the effectiveness of knowledge from tail categories. To tackle these challenges, we reformulate classical KD (ReKD) into two components: Top-KInter-class Similar Distillation (TISD) and Non-Top-KInter-class Discriminability (NTID). Firstly, TISD captures and imparts the knowledge of head categories to the student. Our experimental results have verified that TISD is particularly effective in transferring the knowledge of head categories, even in fine-grained dataset classification. Secondly, we theoretically show that the weighting coefficient of NTID increases with the probability of Top-K, leading to stronger suppression of knowledge transfer for tail categories. This observation explains why difficult samples are more informative than simple ones. To better utilize both types of knowledge, we optimize both TISD and NTID using different weighting coefficients, thereby enhancing the student’s ability to learn this valuable knowledge from both head and tail categories. Furthermore, our extensive experimental results demonstrate that ReKD achieves state-of-the-art performance on various image classification datasets, including CIFAR-100, Tiny-ImageNet, and ImageNet-1K, as well as object detection and instance segmentation using the MS-COCO dataset. Liuchi Xu, Jin Ren 0001, Zhenhua Huang 0001, Wei-Shi Zheng 0001, Yunwen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Continual Learning of Image Classes With Language Guidance From a Vision-Language ModelabstractCurrent deep learning models often catastrophically forget the knowledge of old classes when continually learning new ones. State-of-the-art approaches to continual learning of image classes often require retaining a small subset of old data to partly alleviate the catastrophic forgetting issue, and their performance would be degraded sharply when no old data can be stored due to privacy or safety concerns. In this study, inspired by human learning of visual knowledge with the effective help of language, we propose a novel continual learning framework based on a pre-trained vision-language model (VLM) without retaining any old data. Rich prior knowledge of each new image class is effectively encoded by the frozen text encoder of the VLM, which is then used to guide the learning of new image classes. The output space of the frozen text encoder is unchanged over the whole process of continual learning, through which image representations of different classes become comparable during model inference even when the image classes are learned at different times. Extensive empirical evaluations on multiple image classification datasets under various settings confirm the superior performance of our method over existing ones. The source code is available athttps://github.com/Fatflower/CIL_LG_VLM/. Wentao Zhang 0005, Yujun Huang, Weizhuo Zhang, Tong Zhang 0017, Qicheng Lao, Yue Yu 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Multimodal Action Quality AssessmentabstractAction quality assessment (AQA) is to assess how well an action is performed. Previous works perform modelling by only the use of visual information, ignoring audio information. We argue that although AQA is highly dependent on visual information, the audio is useful complementary information for improving the score regression accuracy, especially for sports with background music, such as figure skating and rhythmic gymnastics. To leverage multimodal information for AQA, i.e., RGB, optical flow and audio information, we propose a Progressive Adaptive Multimodal Fusion Network (PAMFN) that separately models modality-specific information and mixed-modality information. Our model consists of with three modality-specific branches that independently explore modality-specific information and a mixed-modality branch that progressively aggregates the modality-specific information from the modality-specific branches. To build the bridge between modality-specific branches and the mixed-modality branch, three novel modules are proposed. First, a Modality-specific Feature Decoder module is designed to selectively transfer modality-specific information to the mixed-modality branch. Second, when exploring the interaction between modality-specific information, we argue that using an invariant multimodal fusion policy may lead to suboptimal results, so as to take the potential diversity in different parts of an action into consideration. Therefore, an Adaptive Fusion Module is proposed to learn adaptive multimodal fusion policies in different parts of an action. This module consists of several FusionNets for exploring different multimodal fusion strategies and a PolicyNet for deciding which FusionNets are enabled. Third, a module called Cross-modal Feature Decoder is designed to transfer cross-modal features generated by Adaptive Fusion Module to the mixed-modality branch. Our extensive experiments validate the efficacy of the proposed method, and our method achieves state-of-the-art performance on two public datasets. Code is available at https://github.com/qinghuannn/PAMFN. Ling-An Zeng, Wei-Shi Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | SV-Learner: Support-Vector Contrastive Learning for Robust Learning With Noisy LabelsabstractNoisy-label data inevitably gives rise to confusion in various perception applications. In this work, we revisit the theory of support vector machines (SVM) which mines support vectors to build the maximum-margin hyperplane for robust classification, and propose a robust-to-noise deep learning framework, SV-Learner, including the Support Vector Contrastive Learning (SVCL) and Support Vector-based Noise Screening (SVNS). The SV-Learner mines support vectors to solve the learning problem with noisy labels (LNL) reliably. Support Vector Contrastive Learning (SVCL) adopts support vectors as positive and negative samples, driving robust contrastive learning to enlarge the feature distribution margin for learning convergent feature distributions. Support Vector-based Noise Screening (SVNS) uses support vectors with valid labels to assist in screening noisy ones from confusable samples for reliable clean-noisy sample screening. Finally, Semi-Supervised classification is performed to realize the recognition of noisy samples. Extensive experiments are evaluated on CIFAR-10, CIFAR-100, Clothing1M, and Webvision datasets, and results demonstrate the effectiveness of our proposed approach. The source code is availablehttps://github.com/yanliji/SV-Learner. Yanli Ji, Wei-Shi Zheng 0001, Wangmeng Zuo, Xiaofeng Zhu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Post-Distillation via Neural ResuscitationabstractKnowledge distillation, a widely adopted model compression technique, distils knowledge from a large teacher model to a smaller student model, with the goal of reducing the computational resources required for the student model. However, most existing distillation approaches focus on the types of knowledge and how to distil them, which neglect the student model's neuronal responses to the knowledge. In this article, we demonstrate that the kullback-leibler loss inhibits the neuronal responses in the opposite gradient direction, which injures the student model's potential during distilling. To address this problem, we introduce a principled dual-stage distillation scheme to rejuvenate all inhibited neurons at the neuronal level. In the first stage, we detect all the neurons in the student model during the standard distillation period and divide them into two parts according to their responses. In the second stage, we propose three strategies to resuscitate the neurons differently, which allows us to exploit the full potential of the student model. Through the experiments in various aspects of knowledge distillation, it is verified that the proposed approach outperforms the current state-of-the-art approaches. Our work provides a neuronal perspective for studying the response of the student model to the knowledge from the teacher model. Zhiqiang Bao, Chang-Dong Wang 0001, Wei-Shi Zheng 0001, Zhenhua Huang 0001, Yunwen Chen |
IEEE Trans. Multim. | 4 |
| 2024 | Cross-Modal Adaptive Dual Association for Text-to-Image Person RetrievalabstractText-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing work focuses on learning a latent space to narrow the modality gap and further build local correspondences between two modalities. However, these methods assume that image-to-text and text-to-image associations are modality-agnostic, resulting in suboptimal associations. In this work, we demonstrate the discrepancy between image-to-text association and text-to-image association and proposecross-modal adaptive dual association (CADA) to build fine bidirectional image-text detailed associations. Our approach features a decoder-based adaptive dual association module that enables full interaction between visual and textual modalities, enabling bidirectional and adaptive cross-modal correspondence associations. Specifically, this paper proposes a bidirectional association mechanism: Association of text Tokens to image Patches (ATP) and Association of image Regions to text Attributes (ARA). We adaptively model the ATP based on the fact that aggregating cross-modal features based on mistaken associations will lead to feature distortion. For modeling the ARA, since attributes are typically the first distinguishing cues of a person, we explore attribute-level associations by predicting the masked text phrase using the related image region. Finally, we learn the dual associations between texts and images, and the experimental results demonstrate the superiority of our dual formulation. The code used in this article will be made publicly available athttps://github.com/LinDixuan/CADA. Dixuan Lin, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Adaptive Weight Generator for Multi-Task Image Recognition by Task Grouping PromptabstractComparing to adapting the pre-trained backbone to a single image recognition task, multi-task image recognition enables the backbone to perform better when the tasks are related. An interesting research field in multi-task learning (MTL) is to learn the parameter sharing pattern among the involved tasks. Most existing works obtain the sharing pattern, ignoring the task grouping information among the involved tasks. In this work, we aim to build the task parameter sharing pattern based on automatically acquiring the task grouping information. The task grouping information together with the task specific information is then utilized to yield the task adaptive weights. Our method, called Task Grouping prompt-based Adaptive Weight generator (TGAW), consists of Prompt-based Task Representation (PTR) and Prompt-based Weight Generator (PWG). The PTR is modeled as task prompts consisting of task grouping prompt and task specific prompt. The task grouping prompt is automatically chosen from a candidate pool for each task, and tasks selecting the same grouping prompts are divided into the same group. Then, PWG generates task adaptive weights based on the task prompts. The experimental results show that TGAW achieves comparable performance with less than 30% amount of trainable parameters of the pre-trained backbone. Gaojie Wu, Ling-An Zeng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Adaptive Stage-Aware Assessment Skill Transfer for Skill DeterminationabstractSkill determination aims to evaluatehow wella participant performs a specific action. The task is rather challenging, due to the diversity of action types and the scarcity of samples. Many existing works train a skill determination model on limited samples of each action type separately. However, they neglect the skill similarities shared by different action types that can be exploited to enhance the skill determination process. How to exploit useful assessment skills from source actions to a related target action remains a challenge, and existing works have not ever found an effective way to accomplish this. In this work, we propose to achieve skill transfer for action assessment by anAdaptive Stage-aware AssessmentSkillTransfer framework (AdaST) that transfers assessment skills from source actions to different stages of a target action adaptively. A source action search scheme is proposed to select relevant source actions for each target action. Furthermore, to encourage transferring effective and non-redundant assessment skills, a consistency loss and an orthogonality loss are introduced to ensure that the transferred assessment skills do not degrade the accurate determination and it provides complementary information. Extensive experiments on three public datasets demonstrate the effectiveness of the proposed method. Jia-Hui Pan, Jibin Gao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | TwinFormer: Fine-to-Coarse Temporal Modeling for Long-Term Action RecognitionabstractThe long-term action in untrimmed video generally contains multiple sub-actions, among which various semantic patterns exist (e.g., the co-occurrence or sequentiality between sub-actions). These semantic patterns are temporally coarse, and correlated with multiple local contexts which encode the local temporal evolution of visual elements (e.g., hands, objects) in videos. The local contexts and semantic patterns form the inherent fine-to-coarse temporal structure of long-term actions, which is neglected by existing works. Accordingly, in this work we propose TwinFormer, which exploits a novel fine-to-coarse temporal modeling manner to uncover the temporal structure of long-term actions. The proposed TwinFormer consists of a pair of twin encoders with the same structural design, namely Localcontext Encoder and Semantic-pattern Encoder, and a Temporalbridged Attention to bridge the two twin encoders. The Localcontext Encoder aims to model the local contexts in the longterm action. And the Temporal-bridged Attention is designed to correlate the local contexts with semantic patterns. Furthermore, the Semantic-pattern Encoder reveals the temporal evolution of semantic patterns. Experimental results on three benchmarks demonstrate the effectiveness of the proposed model. Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Prompt-Based Learning for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning. Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 5 |
| 2024 | Online Multi-View Learning With Knowledge Registration UnitsabstractIn this work, we investigate online multi-view learning according to the multi-view complementarity and consistency principles to memorably process online multi-view data when fused across views. Online diverse features through different deep feature extractors under different views are used as input to an online learning method to privately and memorably optimize in each view for the discovery and memorization of the view-specific information. More specifically, according to the multi-view complementarity principle, a softmax-weighted reducible (SWR) loss is proposed to selectively retain credible views and neglect incredible ones for the online model's cross-view complementarity fusion. According to the multi-view consistency principle, we design a cross-view embedding consistency (CVEC) loss and a cross-view Kullback-Leibler (CVKL) divergence loss to maintain the cross-view consistency of the online model. Since the online multi-view learning setup needs to avoid repeatedly accessing online data to handle the knowledge forgetting in each view, we propose a knowledge registration unit (KRU) based on dictionary learning to incrementally register newly view-specific knowledge of online unlabeled data to the learnable and adjustable dictionary. Finally, by using the above strategies, we propose an online multi-view KRU approach and evaluate it with comprehensive experiments, thereby showing its superiority in online multi-view learning. Ancong Wu, Wei-Shi Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Class-specific Prompts in Vision Transformer for Continual Learning of New DiseasesabstractCurrent intelligent diagnosis systems are often trained to diagnose a small number of diseases and lack the ability of continually learning new disease knowledge. To have such continual learning ability, the deployed intelligent system needs to be continually updated based on training data of only new diseases, without accessing old data of previously learned diseases due to privacy concerns and challenges in data sharing across medical centers. In this study, a novel and effective prompt learning strategy is proposed to help a pretrained and fixed Vision Transformer (ViT) continually learn new diseases. In particular, a set of unique prompts for each disease are effectively learned such that discriminative disease features can be well extracted from the fixed ViT under the instructions of prompts during feature extraction, even though the ViT feature extractor is pretrained in the natural image domain. Extensive empirical evaluations on two medical image datasets and one natural image dataset demonstrate the superior performance of the proposed method. The source code is available at https://github.com/zhaoedf/CSPrompt. Defeng Zhao, Zejun Ye, Wei-Shi Zheng 0001 |
BIBM | 3 |
| 2023 | Shape-Erased Feature Learning for Visible-Infrared Person Re-IdentificationabstractDue to the modality gap between visible and infrared images with high visual ambiguity, learning diverse modality-shared semantic concepts for visible-infrared person re-identification (VI-ReID) remains a challenging problem. Body shape is one of the significant modality-shared cues for VI-ReID. To dig more diverse modality-shared cues, we expect that erasing body-shape-related semantic concepts in the learned features can force the ReID model to extract more and other modality-shared features for identification. To this end, we propose shape-erased feature learning paradigm that decorrelates modality-shared features in two orthogonal subspaces. Jointly learning shape-related feature in one subspace and shape-erased features in the orthogonal complement achieves a conditional mutual information maximization between shape-erased feature and identity discarding body shape information, thus enhancing the diversity of the learned representation explicitly. Extensive experiments on SYSU-MM01, RegDB, and HITSZ-VCM datasets demonstrate the effectiveness of our method. Ancong Wu, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2023 | AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object DetectionabstractIn this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data scarcity in the target domain leads to an extreme data imbalance between the source and target domains, which potentially causes over-adaptation in traditional feature alignment. To address the data imbalance problem, we propose an asymmetric adaptation paradigm, namely AsyFOD, which leverages the source and target instances from different perspectives. Specifically, by using target distribution estimation, the AsyFOD first identifies the target-similar source instances, which serves to augment the limited target instances. Then, we conduct asynchronous alignment between target-dissimilar source instances and augmented target instances, which is simple yet effective for alleviating the over-adaptation. Extensive experiments demonstrate that the proposed AsyFOD outperforms all state-of-the-art methods on four FSDAOD benchmarks with various environmental variances, e.g., 3.1% mAP improvement on Cityscapes-to-FoggyCityscapes and 2.9% mAP increase on Sim10k-to-Cityscapes. The code is available at https://github.com/Hlings/AsyFPD. Yipeng Gao, Kun-Yu Lin, Junkai Yan, Yaowei Wang 0001, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2023 | Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingabstractSpatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the language description, which requires effective joint modeling of spatiotemporal visuallinguistic dependencies. In this work, we propose a novel framework in which a static vision-language stream and a dynamic vision-language stream are developed to collaboratively reason the target tube. The static stream performs cross-modal understanding in a single frame and learns to attend to the target object spatially according to intraframe visual cues like object appearances. The dynamic stream models visual-linguistic dependencies across multiple consecutive frames to capture dynamic cues like motions. We further design a novel cross-stream collaborative block between the two streams, which enables the static and dynamic streams to transfer useful and complementary information from each other to achieve collaborative reasoning. Experimental results show the effectiveness of the collaboration of the two streams and our overall frame-work achieves new state-of-the-art performance on both HCSTVG and VidSTG datasets. Zihang Lin, Chaolei Tan, Jianfang Hu, Zhi Jin 0002, Tiancai Ye, Wei-Shi Zheng 0001 |
CVPR | 6 |
| 2023 | Generating Anomalies for Video Anomaly Detection with Prompt-based Feature MappingabstractAnomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in the virtual dataset but unbounded in the real world, so it reduces the generalization ability of the virtual dataset. There also exists a scene gap between virtual and real scenarios, including scene-specific anomalies (events that are abnormal in one scene but normal in another) and scene-specific attributes, such as the viewpoint of the surveillance camera. In this paper, we aim to solve the problem of the anomaly gap and scene gap by proposing a prompt-based feature mapping framework (PFMF). The PFMF contains a mapping network guided by an anomaly prompt to generate unseen anomalies with unbounded types in the real scenario, and a mapping adaptation branch to narrow the scene gap by applying domain classifier and anomaly classifier. The proposed framework outperforms the state-of-the-art on three benchmark datasets. Extensive ablation experiments also show the effectiveness of our framework design. Zuhao Liu 0002, Xiao-Ming Wu 0002, Dian Zheng, Kun-Yu Lin, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2023 | Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingabstractVideo Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex semantic relations between visual and textual modalities. Previous methods focus on modeling the contextual information between the video and text from a single-level perspective (i.e., the sentence level), ignoring rich visual-textual correspondence relations at different semantic levels, e.g., the video-word and video-paragraph correspondence. To this end, we propose a novel Hierarchical Semantic Correspondence Network (HSCNet), which explores multi-level visual-textual correspondence by learning hierarchical semantic alignment and utilizes dense supervision by grounding diverse levels of queries. Specifically, we develop a hierarchical encoder that encodes the multi-modal inputs into semantics-aligned representations at different levels. To exploit the hierarchical semantic correspondence learned in the encoder for multi-level supervision, we further design a hierarchical decoder that progressively performs finer grounding for lower-level queries conditioned on higher-level semantics. Extensive experiments demonstrate the effectiveness of HSCNet and our method significantly outstrips the state-of-the-arts on two challenging benchmarks, i.e., ActivityNet-Captions and TACoS. Chaolei Tan, Zihang Lin, Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai |
CVPR | 4 |
| 2023 | Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingabstractBinarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design different estimators instead of STE to mitigate it. However, they ignore the fact that when reducing the estimating error, the gradient stability will decrease concomitantly. These highly divergent gradients will harm the model training and increase the risk of gradient vanishing and gradient exploding. To fully take the gradient stability into consideration, we present a new perspective to the BNNs training, regarding it as the equilibrium between the estimating error and the gradient stability. In this view, we firstly design two indicators to quantitatively demonstrate the equilibrium phenomenon. In addition, in order to balance the estimating error and the gradient stability well, we revise the original straight through estimator and propose a power function based estimator, Rectified Straight Through Estimator (ReSTE for short). Comparing to other estimators, ReSTE is rational and capable of flexibly balancing the estimating error with the gradient stability. Extensive experiments on CIFAR-10 and ImageNet datasets show that ReSTE has excellent performance and surpasses the state-of-the-art methods without any auxiliary modules or losses. Xiao-Ming Wu 0002, Dian Zheng, Zuhao Liu 0002, Wei-Shi Zheng 0001 |
ICCV | 4 |
| 2023 | ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor GenerationabstractRecent sparse detectors with multiple, e.g. six, decoder layers achieve promising performance but much inference time due to complex heads. Previous works have explored using dense priors as initialization and built one-decoder-layer detectors. Although they gain remarkable acceleration, their performance still lags behind their six-decoder-layer counterparts by a large margin. In this work, we aim to bridge this performance gap while retaining fast speed. We find that the architecture discrepancy between dense and sparse detectors leads to feature conflict, hampering the performance of one-decoder-layer detectors. Thus we propose Adaptive Sparse Anchor Generator (ASAG) which predicts dynamic anchors on patches rather than grids in a sparse way so that it alleviates the feature conflict problem. For each image, ASAG dynamically selects which feature maps and which locations to predict, forming a fully adaptive way to generate image-specific anchors. Further, a simple and effective Query Weighting method eases the training instability from adaptiveness. Extensive experiments show that our method outperforms dense-initialized ones and achieves a better speed-accuracy trade-off. The code is available at https://github.com/iSEE-Laboratory/ASAG. Shenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie, Wei-Shi Zheng 0001 |
ICCV | 5 |
| 2023 | Revisit PCA-based technique for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is a desired ability to ensure the reliability and safety of intelligent systems. A scoring function is often designed to measure the degree of any new data being an OOD sample. While most designed scoring functions are based on a single source of information (e.g., the classifier’s output, logits, or feature vector), recent studies demonstrate that fusion of multiple sources may help better detect OOD data. In this study, after detailed analysis of the issue in OOD detection by the conventional principal component analysis (PCA), we propose fusing a simple regularized PCA-based reconstruction error with other source of scoring function to further improve OOD detection performance. In particular, when combined with a strong energy score-based OOD method, the regularized reconstruction error helps achieve new state-of the-art OOD detection results on multiple standard benchmarks. The code is available at https://github.com/SYSUMIA-GROUP/pca-based-out-of-distribution-detection. Xiaoyuan Guan, Zhouwu Liu, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2023 | When Prompt-based Incremental Learning Does Not Meet Strong PretrainingabstractIncremental learning aims to overcome catastrophic forgetting when learning deep networks from sequential tasks. With impressive learning efficiency and performance, prompt-based methods adopt a fixed backbone to sequential tasks by learning task-specific prompts. However, existing prompt-based methods heavily rely on strong pretraining (typically trained on ImageNet-21k), and we find that their models could be trapped if the potential gap between the pretraining task and unknown future tasks is large. In this work, we develop a learnable Adaptive Prompt Generator (APG). The key is to unify the prompt retrieval and prompt learning processes into a learnable prompt generator. Hence, the whole prompting process can be optimized to reduce the negative effects of the gap between tasks effectively. To make our APG avoid learning ineffective knowledge, we maintain a knowledge pool to regularize APG with the feature distribution of each class. Extensive experiments show that our method significantly outperforms advanced methods in exemplar-free incremental learning without (strong) pretraining. Besides, under strong pretraining, our method also has comparable performance to existing prompt-based models, showing that our method can still benefit from pretraining. Codes can be found at https://github.com/TOM-tym/APG Yu-Ming Tang, Yi-Xing Peng, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2023 | Event-Guided Procedure Planning from Instructional Videos with Text SupervisionabstractIn this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed visual states and unobserved intermediate actions, which is ignored by previous works. Specifically, this semantic gap refers to that the contents in the observed visual states are semantically different from the elements of some action text labels in a procedure. To bridge this semantic gap, we propose a novel event-guided paradigm, which first infers events from the observed states and then plans out actions based on both the states and predicted events. Our inspiration comes from that planning a procedure from an instructional video is to complete a specific event and a specific event usually involves specific actions. Based on the proposed paradigm, we contribute an Event-guided Prompting-based Procedure Planning (E3P) model, which encodes event information into the sequential modeling process to support procedure planning. To further consider the strong action associations within each event, our E3P adopts a mask-and-predict approach for relation mining, incorporating a probabilistic masking scheme for regularization. Extensive experiments on three datasets demonstrate the effectiveness of our proposed model. An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng, Wei-Shi Zheng 0001 |
ICCV | 5 |
| 2023 | Self-supervised Cross-stage Regional Contrastive Learning for Object DetectionabstractCross-stage object similarity is a vital property of generic supervised object detectors, which maintains similar feature responses to the same object across feature maps of different intermediate stages of the backbone network. Since an object can be predicted by multiple stages, this similarity is beneficial for accurate object classification and localization. Inspired by this property, we introduce Cross-stage regional Contrastive Learning (CrossCL) to learn the cross-stage object similarity during the model pre-training. Since labels are unavailable in self-supervised learning, we treat the regions sharing the same position in different stages as the same object and constrain them to have similar feature responses across stages to achieve cross-stage object similarity. The learned feature representations of CrossCL share a similar property with supervised detectors, thus showing strong transfer capability to object detection tasks. Besides, we also provide in-depth discussions, ablation studies, and visualizations to understand better how CrossCL works. Code is available at https://github.com/yanjk3/CrossCL. Junkai Yan, Lingxiao Yang, Yipeng Gao, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2023 | Grasp Region Exploration for 7-DoF Robotic Grasping in Cluttered ScenesabstractRobotic grasping is a fundamental skill for robots, but it is quite challenging in cluttered scenes. In cluttered scenes, the precise prediction of high-quality grasp configurations such as rotation and grasping width while avoiding collisions is essential. To accomplish this, the grasp detection models require the capabilities of stronger fine-grained information extracted around the grasp points. However, due to the computational resource restriction, point clouds are usually downsampled in existing networks, which inevitably make some potentially important points discarded. To overcome this problem, we propose a Grasp Region Exploration module to explore the area covered by high-quality grasps. Based on the grasp region, we enhance the point density around the grasp points to mitigate the loss of information caused by downsampling. Furthermore, we devise the Grasp Region Attention module to dynamically aggregate features of various points within the grasp region, such as the grasp point and contact points. The proposed method achieves state-of-the-art performance on the large-scale GraspNet-1Billion dataset. We also conduct real-world experiments on a Franka Emika Panda robot and show that the robot can grasp objects in cluttered scenes with a high success rate. Zibo Chen, Zhixuan Liu, Shangjin Xie, Wei-Shi Zheng 0001 |
IROS | 4 |
| 2023 | Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
Wentao Zhang 0005, Yujun Huang, Tong Zhang 0017, Qingsong Zou, Wei-Shi Zheng 0001 |
MICCAI (2) | 5 |
| 2023 | Diversifying Spatial-Temporal Perception for Video Domain GeneralizationabstractVideo domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain.
A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when recognizing target videos. To this end, we propose to perceive diverse spatial-temporal cues in videos, aiming to discover potential domain-invariant cues in addition to domain-specific cues. We contribute a novel model named Spatial-Temporal Diversification Network (STDN), which improves the diversity from both space and time dimensions of video data. First, our STDN proposes to discover various types of spatial cues within individual frames by spatial grouping. Then, our STDN proposes to explicitly model spatial-temporal dependencies between video contents at multiple space-time scales by spatial-temporal relation modeling. Extensive experiments on three benchmarks of different types demonstrate the effectiveness and versatility of our approach. Kun-Yu Lin, Jia-Run Du, Yipeng Gao, Wei-Shi Zheng 0001 |
NeurIPS | 5 |
| 2023 | Inner-Outer Aware Reconstruction Model for Monocular 3D Scene ReconstructionabstractMonocular 3D scene reconstruction aims to reconstruct the 3D structure of scenes based on posed images. Recent volumetric-based methods directly predict the truncated signed distance function (TSDF) volume and have achieved promising results. The memory cost of volumetric-based methods will grow cubically as the volume size increases, so a coarse-to-fine strategy is necessary for saving memory. Specifically, the coarse-to-fine strategy distinguishes surface voxels from non-surface voxels, and only potential surface voxels are considered in the succeeding procedure. However, the non-surface voxels have various features, and in particular, the voxels on the inner side of the surface are quite different from those on the outer side since there exists an intrinsic gap between them. Therefore, grouping inner-surface and outer-surface voxels into the same class will force the classifier to spend its capacity to bridge the gap. By contrast, it is relatively easy for the classifier to distinguish inner-surface and outer-surface voxels due to the intrinsic gap. Inspired by this, we propose the inner-outer aware reconstruction (IOAR) model. IOAR explores a new coarse-to-fine strategy to classify outer-surface, inner-surface and surface voxels. In addition, IOAR separates occupancy branches from TSDF branches to avoid mutual interference between them. Since our model can better classify the surface, outer-surface and inner-surface voxels, it can predict more precise meshes than existing methods. Experiment results on ScanNet, ICL-NUIM and TUM-RGBD datasets demonstrate the effectiveness and generalization of our model. The code is available at https://github.com/YorkQiu/InnerOuterAwareReconstruction. Yukun Qiu, Guo-Hao Xu, Wei-Shi Zheng 0001 |
NeurIPS | 3 |
| 2023 | Temporal Continual Learning with Prior Compensation for Human Motion PredictionabstractHuman Motion Prediction (HMP) aims to predict future poses at different moments according to past motion sequences. Previous approaches have treated the prediction of various moments equally, resulting in two main limitations: the learning of short-term predictions is hindered by the focus on long-term predictions, and the incorporation of prior information from past predictions into subsequent predictions is limited. In this paper, we introduce a novel multi-stage training framework called Temporal Continual Learning (TCL) to address the above challenges. To better preserve prior information, we introduce the Prior Compensation Factor (PCF). We incorporate it into the model training to compensate for the lost prior information. Furthermore, we derive a more reasonable optimization objective through theoretical derivation. It is important to note that our TCL framework can be easily integrated with different HMP backbone models and adapted to various datasets and applications. Extensive experiments on four HMP benchmark datasets demonstrate the effectiveness and flexibility of TCL. The code is available at https://github.com/hyqlat/TCL. Jianwei Tang, Jiangxin Sun, Xiaotong Lin 0002, Wei-Shi Zheng 0001, Jianfang Hu |
NeurIPS | 5 |
| 2023 | Prototypical Transformer for Weakly Supervised Action Segmentation
Xiaobin Chang, Wei Sun 0007, Wei-Shi Zheng 0001 |
PRCV (6) | 4 |
| 2023 | Unimodal-Multimodal Collaborative Enhancement for Audio-Visual Event Localization
Huilin Tian, Jingke Meng, Yuhan Yao 0002, Wei-Shi Zheng 0001 |
PRCV (6) | 4 |
| 2023 | Task-Incremental Medical Image Classification with Task-Specific Batch Normalization
Xuchen Xie, Weizhuo Zhang, Yujun Huang, Wei-Shi Zheng 0001 |
PRCV (13) | 6 |
| 2023 | Automatic Modelling for Interactive Action Assessment
Jibin Gao, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Learning to Remove Shadows from a Single Image
Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 5 |
| 2023 | A residual convolution transfer framework based on slow feature for cross-domain machinery fault diagnosis
Shubin Chen, Wei-Shi Zheng 0001, Kaiqing Luo |
Neurocomputing | 2 |
| 2023 | Class attention to regions of lesion for imbalanced medical image recognition
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
Neurocomputing | 4 |
| 2023 | Egocentric Action Recognition by Automatic Relation ModelingabstractEgocentric videos, which record the daily activities of individuals from a first-person point of view, have attracted increasing attention during recent years because of their growing use in many popular applications, including life logging, health monitoring and virtual reality. As a fundamental problem in egocentric vision, one of the tasks of egocentric action recognition aims to recognize the actions of the camera wearers from egocentric videos. In egocentric action recognition, relation modeling is important, because the interactions between the camera wearer and the recorded persons or objects form complex relations in egocentric videos. However, only a few of existing methods model the relations between the camera wearer and the interacting persons for egocentric action recognition, and moreover they require prior knowledge or auxiliary data to localize the interacting persons. In this work, we consider modeling the relations in a weakly supervised manner, i.e., without using annotations or prior knowledge about the interacting persons or objects, for egocentric action recognition. We form a weakly supervised framework by unifying automatic interactor localization and explicit relation modeling for the purpose of automatic relation modeling. First, we learn to automatically localize the interactors, i.e., the body parts of the camera wearer and the persons or objects that the camera wearer interacts with, by learning a series of keypoints directly from video data to localize the action-relevant regions with only action labels and some constraints on these keypoints. Second, more importantly, to explicitly model the relations between the interactors, we develop an ego-relational LSTM (long short-term memory) network with several candidate connections to model the complex relations in egocentric videos, such as the temporal, interactive, and contextual relations. In particular, to reduce human efforts and manual interventions needed to construct an optimal ego-relational LSTM structure, we search for the optimal connections by employing a differentiable network architecture search mechanism, which automatically constructs the ego-relational LSTM network to explicitly model different relations for egocentric action recognition. We conduct extensive experiments on egocentric video datasets to illustrate the effectiveness of our method. Haoxin Li, Wei-Shi Zheng 0001, Jianguo Zhang 0001, Haifeng Hu 0001, Jiwen Lu, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Rewarded Semi-Supervised Re-Identification on Identities Rarely Crossing Camera ViewsabstractSemi-supervised person re-identification (Re-ID) is an important approach for alleviating annotation costs when learning to match person images across camera views. Most existing works assume that training data contains abundant identities crossing camera views. However, this assumption is not true in many real-world applications, especially when images are captured in nonadjacent scenes for Re-ID in wider areas, where the identities rarely cross camera views. In this work, we operate semi-supervised Re-ID under a relaxed assumption of identities rarely crossing camera views, which is still largely ignored in existing methods. Since the identities rarely cross camera views, the underlying sample relations across camera views become much more uncertain, and deteriorate the noise accumulation problem in many advanced Re-ID methods that apply pseudo labeling for associating visually similar samples. To quantify such uncertainty, we parameterize the probabilistic relations between samples in a relation discovery objective for pseudo label training. Then, we introduce reward quantified by identification performance on a few labeled data to guide learning dynamic relations between samples for reducing uncertainty. Our strategy is called the Rewarded Relation Discovery (R$^{2}$D), of which the rewarded learning paradigm is under-explored in existing pseudo labeling methods. To further reduce the uncertainty in sample relations, we perform multiple relation discovery objectives learning to discover probabilistic relations based on different prior knowledge of intra-camera affinity and cross-camera style variation, and fuse the complementary knowledge of different probabilistic relations by similarity distillation. To better evaluate semi-supervised Re-ID on identities rarely crossing camera views, we collect a new real-world dataset called REID-CBD, and perform simulation on benchmark datasets. Experiment results show that our method outperforms a wide range of semi-supervised and unsupervised learning methods. Ancong Wu, Wenhang Ge, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | PSLT: A Light-Weight Vision Transformer With Ladder Self-Attention and Progressive ShiftabstractVision Transformer (ViT) has shown great potential for various visual tasks due to its ability to model long-range dependency. However, ViT requires a large amount of computing resource to compute the global self-attention. In this work, we propose a ladder self-attention block with multiple branches and a progressive shift mechanism to develop a light-weight transformer backbone that requires less computing resources (e.g., a relatively small number of parameters and FLOPs), termed Progressive Shift Ladder Transformer (PSLT). First, the ladder self-attention block reduces the computational cost by modelling local self-attention in each branch. In the meanwhile, the progressive shift mechanism is proposed to enlarge the receptive field in the ladder self-attention block by modelling diverse local self-attention for each branch and interacting among these branches. Second, the input feature of the ladder self-attention block is split equally along the channel dimension for each branch, which considerably reduces the computational cost in the ladder self-attention block (with nearly [Formula: see text] the amount of parameters and FLOPs), and the outputs of these branches are then collaborated by a pixel-adaptive fusion. Therefore, the ladder self-attention block with a relatively small number of parameters and FLOPs is capable of modelling long-range interactions. Based on the ladder self-attention block, PSLT performs well on several vision tasks, including image classification, objection detection and person re-identification. On the ImageNet-1 k dataset, PSLT achieves a top-1 accuracy of 79.9% with 9.2 M parameters and 1.9 G FLOPs, which is comparable to several existing models with more than 20 M parameters and 4 G FLOPs. Code is available at https://isee-ai.cn/wugaojie/PSLT.html. Gaojie Wu, Wei-Shi Zheng 0001, Yutong Lu, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning Multi-Attention Context Graph for Group-Based Re-IdentificationabstractLearning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the fact that people often walk in groups in real scenarios. In this work, we take a step further and consider employing context information for identifying groups of people, i.e., group re-id. On the one hand, group re-id is more challenging than single person re-id, since it requires both a robust modeling of local individual person appearance (with different illumination conditions, pose/viewpoint variations, and occlusions), as well as full awareness of global group structures (with group layout and group member variations). On the other hand, we believe that person re-id can be greatly enhanced by incorporating additional visual context from neighboring group members, a task which we formulate as group-aware (single) person re-id. In this paper, we propose a novel unified framework based on graph neural networks to simultaneously address the above two group-based re-id tasks, i.e., group re-id and group-aware person re-id. Specifically, we construct a context graph with group members as its nodes to exploit dependencies among different people. A multi-level attention mechanism is developed to formulate both intra-group and inter-group context, with an additional self-attention module for robust graph-level representations by attentively aggregating node-level features. The proposed model can be directly generalized to tackle group-aware person re-id using node-level representations. Meanwhile, to facilitate the deployment of deep learning models on these tasks, we build a new group re-id dataset which contains more than 3.8K images with 1.5K annotated groups, an order of magnitude larger than existing group re-id datasets. Extensive experiments on the novel dataset as well as three existing datasets clearly demonstrate the effectiveness of the proposed framework for both group-based re-id tasks. Yichao Yan, Jie Qin 0004, Bingbing Ni, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Wei-Shi Zheng 0001, Xiaokang Yang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | SATS: Self-attention transfer for continual semantic segmentation
Yiqiao Qiu, Yixing Shen, Zhuohao Sun, Yanchong Zheng, Xiaobin Chang, Wei-Shi Zheng 0001 |
Pattern Recognit. | 6 |
| 2023 | AC2AS: Activation Consistency Coupled ANN-SNN framework for fast and memory-efficient SNN training
Jianxiong Tang, Jian-Huang Lai, Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001 |
Pattern Recognit. | 5 |
| 2023 | Consistent Intra-Video Contrastive Learning With Asynchronous Long-Term Memory BankabstractUnsupervised representation learning for videos has recently achieved remarkable performance owing to the effectiveness of contrastive learning. Most works on video contrastive learning (VCL) pull all snippets from the same video into the same category, even if some of them are from different actions, leading to temporal collapse, i.e., the snippet representations of a video are invariable with the evolution of time. In this paper, we introduce a novel intra-video contrastive learning (intra-VCL) that further distinguishes intra-video actions to alleviate this issue, which includes an asynchronous long-term memory bank (that caches the representations of all snippets of each video) and mines an extra positive/negative snippet within a video based on the asynchronous long-term memory bank. In addition, since an asynchronous long-term memory bank is required for performing intra-VCL and asynchronous update of the long-term memory leads to inconsistencies when performing contrastive learning, we further propose a consistent contrastive module (CCM) to perform consistent intra-VCL. Specifically, in the CCM, we propose an intra-video self-attention refinement function to reduce the inconsistencies within the asynchronously updated representations (of all snippets of each video) in the long-term memory and an adaptive loss re-weighting to reduce unreliable self-supervision produced by inconsistent contrastive pairs. We call our method as consistent intra-VCL. Extensive experiments demonstrate the effectiveness of the proposed consistent intra-VCL, which achieves state-of-the-art performance on the standard benchmarks of self-supervised action recognition, with top-1 accuracies of 64.2% and 91.0% on HMDB-51 and UCF-101, respectively. Zelin Chen, Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Pseudo-Label Noise Prevention, Suppression and Softening for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID), including fully unsupervised ReID and unsupervised domain adaptive ReID, remains a challenge for the fields of biometrics and computer vision due to its difficulty in learning with unlabeled target domain data. Existing state-of-the-art methods, most of which generate pseudo-labels via unsupervised clustering for model optimization, are inevitably hampered by the under-explored problem of pseudo-label noise. Motivated by this, we propose a novel joint framework termed pseudo-label Noise Prevention, Suppression, and Softening (NPSS) for unsupervised person re-identification. Instead of refining generated label noise after clustering as many existing methods do, we start solving this issue from the source of pseudo-label noise by proposing a new Dynamic Camera-Adaptive Clustering (DCAC), which dynamically involves camera information to prevent noise caused by cross-camera variance, thus improving their quality during clustering. Moreover, we propose an Online Domain Union (ODU) mechanism for the classification model learning on the target domain via involving source domain data with their ground-truth labels, which effectively suppresses the indelible noisy pseudo-labels. Furthermore, we present the Self-Consistency Constraint (SCC) to soften the label noise in a single model with reduced computation and network parameter cost, which achieves intra-sample knowledge ensembling with our global-local SCC and cross-sample knowledge ensembling with our inter-instance SCC. Experiments demonstrate the effectiveness of our method as it surpasses state-of-the-art methods by a large margin on Market-1501, DukeMTMC-ReID, and MSMT17 benchmarks. The code is available at https://github.com/hjwang-824/NPSS. Haijian Wang, Meng Yang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Human Co-Parsing Guided Alignment for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) is a challenging task due to more background noises and incomplete foreground information. Although existing human parsing-based ReID methods can tackle this problem with semantic alignment at the finest pixel level, their performance is heavily affected by the human parsing model. Most supervised methods propose to train an extra human parsing model aside from the ReID model with cross-domain human parts annotation, suffering from expensive annotation cost and domain gap; Unsupervised methods integrate a feature clustering-based human parsing process into the ReID model, but lacking supervision signals brings less satisfactory segmentation results. In this paper, we argue that the pre-existing information in the ReID training dataset can be directly used as supervision signals to train the human parsing model without any extra annotation. By integrating a weakly supervised human co-parsing network into the ReID network, we propose a novel framework that exploits shared information across different images of the same pedestrian, called the Human Co-parsing Guided Alignment (HCGA) framework. Specifically, the human co-parsing network is weakly supervised by three consistency criteria, namely global semantics, local space, and background. By feeding the semantic information and deep features from the person ReID network into the guided alignment module, features of the foreground and human parts can then be obtained for effective occluded person ReID. Experiment results on two occluded and two holistic datasets demonstrate the superiority of our method. Especially on Occluded-DukeMTMC, it achieves 70.2% Rank-1 accuracy and 57.5% mAP. Shuguang Dou, Cairong Zhao, Xinyang Jiang, Shanshan Zhang 0001, Wei-Shi Zheng 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 5 |
| 2023 | Cross-Camera Trajectories Help Person Retrieval in a Camera NetworkabstractWe are concerned with retrieving a query person from multiple videos captured by a non-overlapping camera network. Existing methods often rely on purely visual matching or consider temporal constraints but ignore the spatial information of the camera network. To address this issue, we propose a pedestrian retrieval framework based on cross-camera trajectory generation that integrates both temporal and spatial information. To obtain pedestrian trajectories, we propose a novel cross-camera spatio-temporal model that integrates pedestrians' walking habits and the path layout between cameras to form a joint probability distribution. Such a cross-camera spatio-temporal model can be specified using sparsely sampled pedestrian data. Based on the spatio-temporal model, cross-camera trajectories can be extracted by the conditional random field model and further optimised by restricted non-negative matrix factorization. Finally, a trajectory re-ranking technique is proposed to improve the pedestrian retrieval results. To verify the effectiveness of our method, we construct the first cross-camera pedestrian trajectory dataset, the Person Trajectory Dataset, in real surveillance scenarios. Extensive experiments verify the effectiveness and robustness of the proposed method. Xin Zhang 0113, Xiaohua Xie, Jian-Huang Lai, Wei-Shi Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | PCCT: Progressive Class-Center Triplet Loss for Imbalanced Medical Image ClassificationabstractImbalanced training data in medical image diagnosis is a significant challenge for diagnosing rare diseases. For this purpose, we propose a novel two-stage Progressive Class-Center Triplet (PCCT) framework to overcome the class imbalance issue. In the first stage, PCCT designs a class-balanced triplet loss to coarsely separate distributions of different classes. Triplets are sampled equally for each class at each training iteration, which alleviates the imbalanced data issue and lays solid foundation for the successive stage. In the second stage, PCCT further designs a class-center involved triplet strategy to enable a more compact distribution for each class. The positive and negative samples in each triplet are replaced by their corresponding class centers, which prompts compact class representations and benefits training stability. The idea of class-center involved loss can be extended to the pair-wise ranking loss and the quadruplet loss, which demonstrates the generalization of the proposed framework. Extensive experiments support that the PCCT framework works effectively for medical image classification with imbalanced training images. On four challenging class-imbalanced datasets (two skin datasets Skin7 and Skin 198, one chest X-ray dataset ChestXray-COVID, and one eye dataset Kaggle EyePACs), the proposed approach respectively obtains the mean F1 score 86.20, 65.20, 91.32, and 87.18 over all classes and 81.40, 63.87, 82.62, and 79.09 for rare classes, achieving state-of-the-art performance and outperforming the widely used methods for the class imbalance issue. Kanghao Chen, Weixian Lei, Wei-Shi Zheng 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | DilateFormer: Multi-Scale Dilated Transformer for Visual RecognitionabstractAs ade factosolution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive field leads to quadratic computational cost. Another branch of Vision Transformers exploits local attention inspired by CNNs, which only models the interactions between patches in small neighborhoods. Although such a solution reduces the computational cost, it naturally suffers from small attended receptive fields, which may limit the performance. In this work, we explore effective Vision Transformers to pursue a preferable trade-off between the computational complexity and size of the attended receptive field. By analyzing the patch interaction of global attention in ViTs, we observe two key properties in the shallow layers, namely locality and sparsity, indicating the redundancy of global dependency modeling in shallow layers of ViTs. Accordingly, we propose Multi-Scale Dilated Attention (MSDA) to modellocalandsparsepatch interaction within the sliding window. With a pyramid architecture, we construct a Multi-Scale Dilated Transformer (DilateFormer) by stacking MSDA blocks at low-level stages and global multi-head self-attention blocks at high-level stages. Our experiment results show that our DilateFormer achieves state-of-the-art performance on various vision tasks. On ImageNet-1 K classification task, DilateFormer achieves comparable performance with 70% fewer FLOPs compared with existing state-of-the-art models. Our DilateFormer-Base achieves 85.6% top-1 accuracy on ImageNet-1 K classification task, 53.5% box mAP/46.1% mask mAP on COCO object detection/instance segmentation task and 51.1% MS mIoU on ADE20 K semantic segmentation task. Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin, Yipeng Gao, Andy Jinhua Ma, Yaowei Wang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Consistent Discrepancy Learning for Intra-Camera Supervised Person Re-IdentificationabstractSince annotating pedestrians across different views is extremely costly, intra-camera supervised person re-identification (ReID) aims to learn a ReID model from the intra-view labeled data. Under this setting, the most challenge lies in learning a view-invariant feature embedding in the absence of the cross-view annotations. Previous works focus on assigning a pseudo identity label for each image based on the feature similarity and learn view-invariant features by classification loss. However, because of the cross-view variations in lighting, background, etc., the pseudo labels are often noisy, and therefore not reliable for classification. In this paper, we explore learning a consistent discrepancy for pairwise images. Our main idea is that the discrepancy between pedestrian images should be consistent across different views regardless of view change so that it mainly depicts the identity difference. Due to the lack of cross-view annotations, we project images into different views and obtain likelihood prototypes for cross-view learning. These likelihood prototypes are used to measure the discrepancies between pairwise images under different views. And then, we propose an intra-view discrepancy preservation module to enforce the discrepancy to be view-consistent so as to encourage the model to distinguish the images based on the identities regardless of view change. Extensive experiments on multiple datasets show that our method outperforms existing related methods by clear margins and our method is comparable to supervised counterparts. Code will be made publicly available. Yi-Xing Peng, Jile Jiao, Xuetao Feng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Learning Relation Models to Detect Important People in Still ImagesabstractImportant people detection aims to identify the most important people (i.e., the people who play the main roles in scenes) in images, which is challenging since people's importance in images depends not only on their appearance but also on their interactions with others (i.e., relations among people) and their roles in the scene (i.e., relations between people and underlying events). In this work, we propose the People Relation Network (PRN) to solve this problem. PRN consists of three modules (i.e., the feature representation, relation and classification modules) to extract visual features, model relations and estimate people's importance, respectively. The relation module contains two submodules to model two types of relations, namely, the person-person relation submodule and the person-event relation submodule. The person-person relation submodule infers the relations among people from the interaction graph and the person-event relation submodule models the relations between people and events by considering the spatial correspondence between features. With the help of them, PRN can effectively distinguish important people from other individuals. Extensive experiments on the Multi-Scene Important People (MS) and NCAA Basketball Image (NCAA) datasets show that PRN achieves state-of-the-art performance and generalizes well when available data is limited. Yukun Qiu, Fa-Ting Hong, Wei-Hong Li 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept RecognitionabstractThe goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost. Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 5 |
| 2023 | Pyramid Texture FilteringabstractWe present a simple but effective technique to smooth out textures while preserving the prominent structures. Our method is built upon a key observation---the coarsest level in a Gaussian pyramid often naturally eliminates textures and summarizes the main image structures. This inspires our central idea for texture filtering, which is to progressively upsample the very low-resolution coarsest Gaussian pyramid level to a full-resolution texture smoothing result with well-preserved structures, under the guidance of each fine-scale Gaussian pyramid level and its associated Laplacian pyramid level. We show that our approach is effective to separate structure from texture of different scales, local contrasts, and forms, without degrading structures or introducing visual artifacts. We also demonstrate the applicability of our method on various applications including detail enhancement, image abstraction, HDR tone mapping, inverse halftoning, and LDR image enhancement. Code is available at https://rewindl.github.io/pyramid_texture_filtering/. Qing Zhang 0006, Hao Jiang 0057, Yongwei Nie, Wei-Shi Zheng 0001 |
ACM Trans. Graph. | 4 |
| 2022 | Lifelong Person Re-identification by Pseudo Task Knowledge PreservationabstractIn real world, training data for person re-identification (Re-ID) is collected discretely with spatial and temporal variations, which requires a model to incrementally learn new knowledge without forgetting old knowledge. This problem is called lifelong person re-identification (LReID). Variations of illumination and background for images of each task exhibit task-specific image style and lead to task-wise domain gap. In addition to missing data from the old tasks, task-wise domain gap is a key factor for catastrophic forgetting in LReID, which is ignored in existing approaches for LReID. The model tends to learn task-specific knowledge with task-wise domain gap, which results in stability and plasticity dilemma. To overcome this problem, we cast LReID as a domain adaptation problem and propose a pseudo task knowledge preservation framework to alleviate the domain gap. Our framework is based on a pseudo task transformation module which maps the features of the new task into the feature space of the old tasks to complement the limited saved exemplars of the old tasks. With extra transformed features in the task-specific feature space, we propose a task-specific domain consistency loss to implicitly alleviate the task-wise domain gap for learning task-shared knowledge instead of task-specific one. Furthermore, to guide knowledge preservation with the feature distributions of the old tasks, we propose to preserve knowledge on extra pseudo tasks which jointly distills knowledge and discriminates identity, in order to achieve a better trade-off between stability and plasticity for lifelong learning with task-wise domain gap. Extensive experiments demonstrate the superiority of our method as compared with the state-of-the-art lifelong learning and LReID methods. Wenhang Ge, Junlong Du, Ancong Wu, Yuqiao Xian, Feiyue Huang, Wei-Shi Zheng 0001 |
AAAI | 7 |
| 2022 | Expert with Outlier Exposure for Continual Learning of New DiseasesabstractCurrent intelligent diagnosis systems struggle to continually learn to diagnose more and more diseases due to catastrophic forgetting of old knowledge when learning new knowledge. Although storing small old data for subsequent continual learning can effectively help alleviate the forgetting issue, the heavy data imbalance between old classes and to-be-learned new classes in classifier training often causes biased prediction towards the new classes just learned by the updated classifier. In this study, an outlier detection technique is novelly applied to train an additional expert classifier for new classes to help alleviate the class imbalance issue and discriminate the learned new classes from old classes during inference (instead of the training phase). Specially, the stored small data of old classes are considered as outliers during training the expert classifier, such that the output probability distributions from the expert classifier are expected to be obviously different between test data of the old classes and those of the new classes. Such difference between old classes and new classes can be used to fine-tune the original output from the updated classifier which is responsible for prediction of all learned (old and new) classes. During inference, a novel ensemble strategy is proposed to combine the predictions from the updated classifier, the expert classifier, and the previously learned old classifier. The proposed learning and inference framework can be easily combined with existing continual learning strategies. Empirical evaluations on three medical image datasets and one natural image dataset show that the proposed framework can effectively improve continual learning performance. Zhengjing Xu, Kanghao Chen, Wei-Shi Zheng 0001, Zhijun Tan |
BIBM | 3 |
| 2022 | SIOD: Single Instance Annotated Per Category Per Image for Object DetectionabstractObject detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instance-level annotation, while semi-supervised object detection (SSOD) remains challenging led by the inter-image discrepancy between labeled and unlabeled data. In this study, we propose the Single Instance annotated Object Detection (SIOD), requiring only one instance annotation for each existing category in an image. Degraded from inter-task (WSOD) or inter-image (SSOD) discrepancies to the intra-image discrepancy, SIOD provides more reliable and rich prior knowledge for mining the rest of unlabeled instances and trades off the annotation cost and performance. Under the SIOD setting, we propose a simple yet effective framework, termed Dual-Mining (DMiner), which consists of a Similarity-based Pseudo Label Generating module (SPLG) and a Pixel-level Group Contrastive Learning module (PGCL). SPLG firstly mines latent instances from feature representation space to alleviate the annotation missing problem. To avoid being misled by inaccurate pseudo labels, we propose PGCL to boost the tolerance to false pseudo labels. Extensive experiments on MS COCO verify the feasibility of the SIOD setting and the superiority of the proposed method, which obtains consistent and significant improvements compared to baseline methods and achieves comparable results with fully supervised object detection (FSOD) methods with only 40% instances annotated. Code is available at https://github.com/solicucu/SIOD. Hanjun Li 0004, Xingjia Pan, Fan Tang, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2022 | Learning to Imagine: Diversify Memory for Incremental Learning using Unlabeled DataabstractDeep neural network (DNN) suffers from catastrophic forgetting when learning incrementally, which greatly limits its applications. Although maintaining a handful of samples (called “exemplars”) of each task could alleviate forgetting to some extent, existing methods are still limited by the small number of exemplars since these exemplars are too few to carry enough task-specific knowledge, and therefore the forgetting remains. To overcome this problem, we propose to “imagine” diverse counterparts of given exemplars referring to the abundant semantic-irrelevant information from unlabeled data. Specifically, we develop a learnable feature generator to diversify exemplars by adaptively generating diverse counterparts of exemplars based on semantic information from exemplars and semantically-irrelevant information from unlabeled data. We introduce semantic contrastive learning to enforce the generated samples to be semantic consistent with exemplars and perform semantic-decoupling contrastive learning to encourage diversity of generated samples. The diverse generated samples could effectively prevent DNN from forgetting when learning new tasks. Our method does not bring any extra inference cost and outperforms state-of-the-art methods on two benchmarks CIFAR-100 and ImageNet-Subset by a clear margin. Yu-Ming Tang, Yi-Xing Peng, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2022 | Likert Scoring with Grade Decoupling for Long-term Action AssessmentabstractLong-term action quality assessment is a task of evaluating how well an action is performed, namely, estimating a quality score from a long video. Intuitively, long-term actions generally involve parts exhibiting different levels of skill, and we call the levels of skill as performance grades. For example, technical highlights and faults may appear in the same long-term action. Hence, the final score should be determined by the comprehensive effect of different grades exhibited in the video. To explore this latent relationship, we design a novel Likert scoring paradigm in-spired by the Likert scale in psychometrics, in which we quantify the grades explicitly and generate the final quality score by combining the quantitative values and the corresponding responses estimated from the video, instead of performing direct regression. Moreover, we extract grade-specific features, which will be used to estimate the responses of each grade, through a Transformer decoder architecture with diverse learnable queries. The whole model is named as Grade-decoupling Likert Transformer (GDLT), and we achieve state-of-the-art results on two long-term action assessment datasets.11Project page https://isee-ai.cn/-angchi/CVPR22_GDLT.html Angchi Xu, Ling-An Zeng, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2022 | AcroFOD: An Adaptive Method for Cross-Domain Few-Shot Object Detection
Yipeng Gao, Lingxiao Yang, Yunmu Huang, Song Xie, Wei-Shi Zheng 0001 |
ECCV (33) | 6 |
| 2022 | Adversarial Partial Domain Adaptation by Cycle Inconsistency
Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001 |
ECCV (33) | 4 |
| 2022 | Improving Class Balancing at Both Feature Extractor and Classifier HeadabstractTraining data are often imbalanced across classes in practice, and such class imbalance issue often causes model predictions biased toward majority classes during inference. Different from existing solutions which employ various training strategies to alleviate the class imbalance issue, this study proposes a novel two-head model architecture to help alleviate the issue. One auxiliary classifier head helps the feature extractor of the classifier more fairly learn to extract features for each class, and the main classifier head learns in a more class-balanced manner by dividing each majority class into multiple clusters in advance and considering each cluster as a new class. Extensive empirical evaluations on four class-imbalanced image datasets showed that the proposed approach achieves state-of-the-art classification performance. Kanghao Chen, Huijuan Lu, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2022 | Learning Multi-Context Dynamic Listwise Relation for Generalizable Person Re-IdentificationabstractAlthough Person re-identification (Re-ID) has made rapid development in supervised learning and domain adaptation, it is more desirable to learn a generalizable model that can be directly applied to unseen scenes without updating. Generalizable Re-ID is challenging due to uncertain cross-camera variations in unseen target domain, such as illumination and viewpoint change, which result in visual ambiguities. Existing generalizable Re-ID methods focus on learning more generalizable features for individual instances. They ignore the context in the ranking list of the target domain. When human encounters visual ambiguities when matching pedestrians in unfamiliar scenes, comparing similar instances in the ranking list and comparing environments in different cameras can help remove ambiguities and refine matching results. This is actually exploiting contextual information in target domain, which is ignored by existing generalizable Re-ID methods. For learning contextual information to refine matching in unseen target domain, we propose a Multi-Context Dynamic Listwise Relation Network (MDLRN) to extract and aggregate instance-level and camera-level contextual features of a list of images, which can dynamically adapt the metric to unseen cross-domain scene variations. We further propose camera-specific feature perturbation (CFP) to simulate cross-camera variations in unseen target domain to improve generalization. Extensive experiments showed the superiority of our method in domain generalization. Chengzhi Lin, Ancong Wu, Wei-Shi Zheng 0001 |
ICPR | 3 |
| 2022 | Space-correlated Contrastive Representation Learning with Multiple InstancesabstractSelf-supervised contrastive learning methods have shown promising transferability in pretraining by maximizing the mutual information between two cropped regions as views from the same image. In order to effectively extract mutual information between views, the cropped regions need to be the same instance as prior hypothesis. However, the data collected in general scenes usually have multiple instances, so the two cropped regions probably contain different instances which will mislead the contrastive learning process. In this paper, we make the first attempt to exploit the spatial position relationships of the two cropped regions in self-supervised contrastive learning with images that include multiple instances. Then, we propose an effective method called Space-correlated Contrastive Learning (SpaceCL). Specifically, given two randomly cropped regions as contrastive pairs from the same image, we implement self-supervised contrastive learning by optimizing a space correspondence contrastive similarity loss. As a result, our method achieves state-of-the-art performance and remarkably outperforms other counterparts when pretrained on the COCO dataset of which images contain multiple instances. Experiments show our method outperforms ReSim with 2.6%AP on PASCAL VOC object detection, 0.8%APbband 0.6%APmkon COCO object detection and instance segmentation, 1.3%APmkon Cityscapes instance segmentation. Danming Song, Yipeng Gao, Junkai Yan, Wei Sun 0007, Wei-Shi Zheng 0001 |
ICPR | 5 |
| 2022 | TransGrasp: A Multi-Scale Hierarchical Point Transformer for 7-DoF Grasp DetectionabstractRobotic grasping pose detection that predicts the configuration of the robotic gripper for object grasping is fundamental in robot manipulation. Based on point clouds, most of the existing methods predict grasp pose with the hierarchical PointNet++ backbone, while the non-local geometric information is underexplored. In this work, we address the 7-DoF (6- DoF with the grasp width) grasp detection by introducing a one- stage Transformer-based hierarchical multi-scale model dubbed TransGrasp. Empowered by TransGrasp, the point features are enhanced via acquiring multi-scale shape awareness in the whole scene. By directly modeling the long-range relevance, our pipeline is aware of object contour to avoid collisions and able to apply analogy reasoning for long-distance geometric structures. The evaluation results on the large scale GraspNet- 1Billion dataset demonstrate the effectiveness of the proposed TransGrasp. The real robot experiments on an ABB YUMI robot with an Azure Kinect DK camera and an ABB Smart two-finger gripper show high success rates in both single object and cluttered scenes. Zhixuan Liu, Zibo Chen, Shangjin Xie, Wei-Shi Zheng 0001 |
ICRA | 4 |
| 2022 | Task-Oriented Self-supervised Learning for Anomaly Detection in Electroencephalography
Yaojia Zheng, Zhouwu Liu, Rong Mo, Wei-Shi Zheng 0001 |
MICCAI (8) | 5 |
| 2022 | PC-Dance: Posture-controllable Music-driven Dance SynthesisabstractMusic-driven dance synthesis is a task to generate high-quality dance according to the music given by the user, which has promising entertainment applications. However, most of the existing methods cannot provide an efficient and effective way for user intervention in dance generation, e.g., posture-controllable. In this work, we propose a powerful framework named PC-Dance to perform adaptive posture-controllable music-driven dance synthesis. Consisting of an music-to-dance alignment embedding network (M2D-Align) and a posture-controllable dance synthesis (PC-Syn), PC-Dance allows fine-grained control by input anchor poses efficiently without artist participation. Specifically, to relieve the cost of artist participation but ensure generating high-quality dance efficiently, a self-supervised rhythm alignment module is designed to further learn the music-to-dance alignment embedding. As for PC-Syn, we introduce an efficient scheme for adaptive motion graph construction (AMGC), which could improve the efficiency of graph-based optimization and preserve the diversity of motions. Since there is few related public dataset, we collect an MMD-ARC dataset for music-driven dance synthesis. The experimental results on MMD-ARC dataset demonstrate the effectiveness of our framework and the feasibility for dance synthesis with adaptive posture controlling. Jibin Gao, Junfu Pu, Honglun Zhang, Ying Shan, Wei-Shi Zheng 0001 |
ACM Multimedia | 5 |
| 2022 | Mixed Supervision for Instance Learning in Object Detection with Few-shot AnnotationabstractMixed supervision for object detection (MSOD) that utilizes image-level annotations and a small amount of instance-level annotations has emerged as an efficient tool by alleviating the requirement for a large amount of costly instance-level annotations and providing effective instance supervision on previous methods that only use image-level annotations. In this work, we introduce the mixed supervision instance learning (MSIL), as a novel MSOD framework to leverage a handful of instance-level annotations to provide more explicit and implicit supervision. Rather than just adding instance-level annotations directly on loss functions for detection, we aim to dig out more effective explicit and implicit relations between these two different level annotations. In particular, we firstly propose the Instance-Annotation Guided Image Classification strategy to provide explicit guidance from instance-level annotations by using positional relation to force the image classifier to focus on the proposals which contain the correct object. And then, in order to exploit more implicit interaction between the mixed annotations, an instance reproduction strategy guided by the extra instance-level annotations is developed for generating more accurate pseudo ground truth, achieving a more discriminative detector. Finally, a false target instance mining strategy is used to refine the above processing by enriching the number and diversity of training instances with the position and score information. Our experiments show that the proposed MSIL framework outperforms recent state-of-the-art mixed supervised detectors with a large margin on both the Pascal VOC2007 and the MS-COCO dataset. Chengyao Wang, Zhu Zhou, Yaowei Wang 0001, Wei-Shi Zheng 0001 |
ACM Multimedia | 6 |
| 2022 | Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalabstractCross-modal retrieval between videos and texts has gained increasing interest because of the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can have multiple different text descriptions and queries. We call it the Video-Text Correspondence Ambiguity problem. Current techniques mostly concentrate on mining local or multi-level alignment between contents of video and text (e.g., object to entity and action to verb). It is difficult for these methods to alleviate video-text correspondence ambiguity by describing a video using only one feature, which is required to be matched with multiple different text features at the same time. To address this problem, we propose a Text-Adaptive Multiple Visual Prototype Matching Model. It automatically captures multiple prototypes to describe a video by adaptive aggregation on video token features. Given a query text, the similarity is determined by the most similar prototype to find correspondence in the video, which is called text-adaptive matching. To learn diverse prototypes for representing the rich information in videos, we propose a variance loss to encourage different prototypes to attend to different contents of the video. Our method outperforms the state-of-the-art methods on four public video retrieval datasets. Chengzhi Lin, Ancong Wu, Junwei Liang 0001, Jun Zhang 0018, Wenhang Ge, Wei-Shi Zheng 0001, Chunhua Shen |
NeurIPS | 6 |
| 2022 | Discriminative Distillation to Reduce Class Confusion in Continual Learning
Changhong Zhong 0001, Zhiying Cui, Wei-Shi Zheng 0001, Hongmei Liu 0001 |
PRCV (1) | 3 |
| 2022 | Learning Multi-Scale Deep Image Prior for High-Quality Unsupervised Image DenoisingabstractAbstract Recent methods on image denoising have achieved remarkable progress, benefiting mostly from supervised learning on massive noisy/clean image pairs and unsupervised learning on external noisy images. However, due to the domain gap between the training and testing images, these methods typically have limited applicability on unseen images. Although several attempts have been made to avoid the domain gap issue by learning denoising from singe noisy image itself, they are less effective in handling real‐world noise because of assuming the noise corruptions are independent and zero mean. In this paper, we go step further beyond prior work by presenting a novel unsupervised image denoising framework trained from single noisy image without making any explicit assumptions on the noise statistics. Our approach is built upon the deep image prior (DIP), which enables diverse image restoration tasks. However, as is, the denoising performance of DIP will significantly deteriorate on nonzero‐mean noise and is sensitive to the number of iterations. To overcome this problem, we propose to utilize multi‐scale deep image prior by imposing DIP across different image scales under the constraint of a scale consistency. Experiments on synthetic and real datasets demonstrate that our method performs favorably against the state‐of‐the‐art methods for image denoising. Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Comput. Graph. Forum | 5 |
| 2022 | AR-CNN: an attention ranking network for learning urban perception
Zhetao Li, Wei-Shi Zheng 0001, Sangyoon Oh 0001, Kien Nguyen 0002 |
Sci. China Inf. Sci. | 3 |
| 2022 | Weakly supervised action anticipation without object annotations
Haoxin Li, Wei-Shi Zheng 0001 |
Frontiers Comput. Sci. | 4 |
| 2022 | Joint Bilateral-Resolution Identity Modeling for Cross-Resolution Person Re-Identification
Wei-Shi Zheng 0001, Jincheng Hong, Jiening Jiao, Ancong Wu, Xiatian Zhu, Shaogang Gong, Jiayin Qin, Jian-Huang Lai |
Int. J. Comput. Vis. | 1 |
| 2022 | Heterogeneous graph driven unsupervised domain adaptation of person re-identification
Shaochuan Lin, Jianming Lv, Zhenguo Yang, Qing Li 0001, Wei-Shi Zheng 0001 |
Neurocomputing | 5 |
| 2022 | Relaxation LIF: A gradient-based spiking neuron for direct training deep spiking neural networks
Jianxiong Tang, Jian-Huang Lai, Wei-Shi Zheng 0001, Lingxiao Yang, Xiaohua Xie |
Neurocomputing | 3 |
| 2022 | APANet: Auto-Path Aggregation for Future Instance Segmentation PredictionabstractDespite the remarkable progress achieved in conventional instance segmentation, the problem of predicting instance segmentation results for unobserved future frames remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting features of future frames. However, these methods always treat features of multiple levels (e.g., coarse-to-fine pyramid features) independently and do not exploit them collaboratively, which results in inaccurate prediction for future frames; and moreover, such a weakness can partially hinder self-adaption of a future segmentation prediction model for different input samples. To solve this problem, we propose an adaptive aggregation approach called Auto-Path Aggregation Network (APANet), where the spatio-temporal contextual information obtained in the features of each individual level is selectively aggregated using the developed "auto-path". The "auto-path" connects each pair of features extracted at different pyramid levels for task-specific hierarchical contextual information aggregation, which enables selective and adaptive aggregation of pyramid features in accordance with different videos/frames. Our APANet can be further optimized jointly with the Mask R-CNN head as a feature decoder and a Feature Pyramid Network (FPN) feature encoder, forming a joint learning system for future instance segmentation prediction. We experimentally show that the proposed method can achieve state-of-the-art performance on three video-based instance segmentation benchmarks for future instance segmentation prediction. Jianfang Hu, Jiangxin Sun, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Deep Graph Metric Learning for Weakly Supervised Person Re-IdentificationabstractIn conventional person re-identification (re-id), the images used for model training in the training probe set and training gallery set are all assumed to be instance-level samples that are manually labeled from raw surveillance video (likely with the assistance of detection) in a frame-by-frame manner. This labeling across multiple non-overlapping camera views from raw video surveillance is expensive and time consuming. To overcome these issues, we consider a weakly supervised person re-id modeling that aims to find the raw video clips where a given target person appears. In our weakly supervised setting, during training, given a sample of a person captured in one camera view, our weakly supervised approach aims to train a re-id model without further instance-level labeling for this person in another camera view. The weak setting refers to matching a target person with an untrimmed gallery video where we only know that the identity appears in the video without the requirement of annotating the identity in any frame of the video during the training procedure. The weakly supervised person re-id is challenging since it not only suffers from the difficulties occurring in conventional person re-id (e.g., visual ambiguity and appearance variations caused by occlusions, pose variations, background clutter, etc.), but more importantly, is also challenged by weakly supervised information because the instance-level labels and the ground-truth locations for person instances (i.e., the ground-truth bounding boxes of person instances) are absent. To solve the weakly supervised person re-id problem, we develop deep graph metric learning (DGML). On the one hand, DGML measures the consistency between intra-video spatial graphs of consecutive frames, where the spatial graph captures neighborhood relationship about the detected person instances in each frame. On the other hand, DGML distinguishes the inter-video spatial graphs captured from different camera views at different sites simultaneously. To further explicitly embed weak supervision into the DGML and solve the weakly supervised person re-id problem, we introduce weakly supervised regularization (WSR), which utilizes multiple weak video-level labels to learn discriminative features by means of a weak identity loss and a cross-video alignment loss. We conduct extensive experiments to demonstrate the feasibility of the weakly supervised person re-id approach and its special cases (e.g., its bag-to-bag extension) and show that the proposed DGML is effective. Jingke Meng, Wei-Shi Zheng 0001, Jian-Huang Lai, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Adaptive Action AssessmentabstractAction assessment, the process of evaluating how well an action is performed, is an important task in human action analysis. Action assessment has experienced considerable development based on visual cues; however, existing methods neglect to adaptively learn different architectures for varied types of actions and are therefore limited in achieving high-performance assessment for each type of action. In fact, every type of action has specific evaluation criteria, and human experts are trained for years to correctly evaluate a single type of action. Therefore, it is difficult for a single assessment architecture to achieve high performance for all types of actions. However, manually designing an assessment architecture for each specific type of action is very difficult and impracticable. This work addresses this problem by adaptively designing different assessment architectures for different types of actions, and the proposed approach is therefore called the adaptive action assessment. In order to facilitate our adaptive action assessment by exploiting the specific joint interactions for each type of action, a set of graph-based joint relations is learned for each type of action by means of trainable joint relation graphs built according to the human skeleton structure, and the learned joint relation graphs can visually interpret the assessment process. In addition, we introduce using a normalized mean squared error loss (N-MSE loss) and a Pearson loss that perform automatic score normalization to operate adaptive assessment training. The experiments on four benchmarks for action assessment demonstrate the effectiveness and feasibility of the proposed method. We also demonstrate the visual interpretability of our model by visualizing the details of the assessment process. Jibin Gao, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Unsupervised Intrinsic Image Decomposition Using Internal Self-Similarity CuesabstractRecent learning-based intrinsic image decomposition methods have achieved remarkable progress. However, they usually require massive ground truth intrinsic images for supervised learning, which limits their applicability on real-world images since obtaining ground truth intrinsic decomposition for natural images is very challenging. In this paper, we present an unsupervised framework that is able to learn the decomposition effectively from a single natural image by training solely with the image itself. Our approach is built upon the observations that the reflectance of a natural image typically has high internal self-similarity of patches, and a convolutional generation network tends to boost the self-similarity of an image when trained for image reconstruction. Based on the observations, an unsupervised intrinsic decomposition network (UIDNet) consisting of two fully convolutional encoder-decoder sub-networks, i.e., reflectance prediction network (RPN) and shading prediction network (SPN), is devised to decompose an image into reflectance and shading by promoting the internal self-similarity of the reflectance component, in a way that jointly trains RPN and SPN to reproduce the given image. A novel loss function is also designed to make effective the training for intrinsic decomposition. Experimental results on three benchmark real-world datasets demonstrate the superiority of the proposed method. Qing Zhang 0006, Lei Zhu 0003, Wei Sun 0007, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Learning multimodal relationship interaction for visual relationship detection
Zhixuan Liu, Wei-Shi Zheng 0001 |
Pattern Recognit. | 2 |
| 2022 | Semi-Supervised Action Quality Assessment With Self-Supervised Segment Feature RecoveryabstractAction Quality Assessment aims to evaluate how well an action performs. Existing methods have achieved remarkable progress on fully-supervised action assessment. However, in real-world applications, with expert’s experience, it is not always feasible to manually label all samples. Therefore, it is important to study the problem of semi-supervised action assessment with only a small amount of samples annotated. A major challenge for semi-supervised action assessment is how to exploit the temporal pattern from unlabeled videos. Inspired by the temporal dependencies of the action execution, we propose a self-supervised learning on the unlabeled videos by recovering the feature of a masked segment of an unlabeled video. Furthermore, we leverage adversarial learning to align the representation distribution of the labeled and the unlabeled samples to close their gap in the sample space since unlabeled samples always come from unseen actions. Finally, we propose an adversarial self-supervised framework for semi-supervised action quality assessment. The extensive experimental results on the MTL-AQA and the Rhythmic Gymnastics datasets will demonstrate the effectiveness of our framework, achieving the state-of-the-art performances of semi-supervised action quality assessment. Jibin Gao, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Spectral-Spatial Transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search FrameworkabstractNeural networks have dominated the research of hyperspectral image classification, attributing to the feature learning capacity of convolution operations. However, the fixed geometric structure of convolution kernels hinders long-range interaction between features from distant locations. In this article, we propose a novel spectral–spatial transformer network (SSTN), which consists of spatial attention and spectral association modules, to overcome the constraints of convolution kernels. Also, we design a factorized architecture search (FAS) framework that involves two independent subprocedures to determine the layer-level operation choices and block-level orders of SSTN. Unlike conventional neural architecture search (NAS) that requires a bilevel optimization of both network parameters and architecture settings, the FAS focuses only on finding out optimal architecture settings to enable a stable and fast architecture search. Extensive experiments conducted on five popular HSI benchmarks demonstrate the versatility of SSTNs over other state-of-the-art (SOTA) methods and justify the FAS strategy. On the University of Houston dataset, SSTN obtains comparable overall accuracy to SOTA methods with a small fraction (1.2%) of multiply-and-accumulate operations compared to a strong baseline spectral–spatial residual network (SSRN). Most importantly, SSTNs outperform other SOTA networks using only 1.2% or fewer MACs of SSRNs on the Indian Pines, the Kennedy Space Center, the University of Pavia, and the Pavia Center datasets. Zilong Zhong, Ying Li 0036, Lingfei Ma, Jonathan Li 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Probabilistic Temporal Modeling for Unintentional Action LocalizationabstractHumans have the inherent advantage of understanding action intention, while it is an enormous challenge to train the machine to localize unintentional action in videos due to the lack of reliable annotations for stable training. The annotations of unintentional action are unreliable since different annotators are affected by their subjective appraisals and intrinsic ambiguity, which brings heavy difficulties for the training. To address this issue, we propose a probabilistic framework for unintentional action localization by modeling the uncertainty of annotations. Our framework consists of two main components, including Temporal Label Aggregation (TLA) and Dense Probabilistic Localization (DPL). We first formulate each annotated failure moment as a temporal label distribution. Then we propose a TLA component to aggregate temporal label distributions of different failure moments in an online manner and generate dense probabilistic supervision. Based on TLA, We further develop a DPL component to jointly train three heads (i.e., probabilistic dense classification, probabilistic temporal detection, and probabilistic regression) with different supervision granularities and make them highly collaborative. We evaluate our approach on the largest unintentional action dataset OOPS and demonstrate that our approach can achieve significant improvement over the baseline and state-of-the-art methods. Jinglin Xu, Guangyi Chen 0002, Nuoxing Zhou, Wei-Shi Zheng 0001, Jiwen Lu |
IEEE Trans. Image Process. | 4 |
| 2022 | Conditional Feature Embedding by Visual Clue Correspondence Graph for Person Re-IdentificationabstractAlthough Person Re-Identification has made impressive progress, difficult cases like occlusion, change of view-point, and similar clothing still bring great challenges. In order to tackle these challenges, extracting discriminative feature representation is crucial. Most of the existing methods focus on extracting ReID features from individual images separately. However, when matching two images, we propose that the ReID features of a query image should be dynamically adjusted based on the contextual information from the gallery image it matches. We call this type of ReID features conditional feature embedding. In this paper, we propose a novel ReID framework that extracts conditional feature embedding based on the aligned visual clues between image pairs, called Clue Alignment based Conditional Embedding (CACE-Net). CACE-Net applies an attention module to build a detailed correspondence graph between crucial visual clues in image pairs and uses discrepancy-based GCN to embed the obtained complex correspondence information into the conditional features. The experiments show that CACE-Net achieves state-of-the-art performance on three public datasets. Fufu Yu, Xinyang Jiang, Yifei Gong, Wei-Shi Zheng 0001, Feng Zheng 0001, Xing Sun 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Deep Shape-Aware Person Re-Identification for Overcoming Moderate Clothing ChangesabstractAlthough person re-identification (person re-id) has advanced substantially in recent years, most methods are based on the assumption that the identities would not change clothes. This assumption might not hold in practice considering criminals who intentionally change clothes. In this work, we attempt to solve person re-id under moderate clothing change. Since the human body shape is considered as relatively more invariant under moderate clothing changes, we propose to learn a reliable shape-aware feature representation by mutually learning both colorful images and contour images. Instead of directly extracting shape features from contour images, we utilize contour feature learning as regularization and excavate more effective shape-aware feature representations from colorful images. We propose a multi-scale appearance and contour deep infomax (MAC-DIM) to maximize mutual information between colorful appearance features and contour shape features, and in this way, the extracted appearance features are constrained to be shape-aware in terms of both low-level visual properties and high-level semantics. To better model the long-range human body shape and explicitly capture contour segment relations, we introduce hierarchical graph modeling as aggregation headers, propagating structural context through graph convolutional networks (GCNs). The extensive results on benchmarks under clothing changes demonstrate the effectiveness of our shape-aware feature learning scheme. Wei-Shi Zheng 0001, Qize Yang, Jingke Meng, Richang Hong, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | A Blind Color Separation Model for Faithful Palette-Based Image RecoloringabstractPalette-based image recoloring provides a simple yet effective way for color adjustment, which allows users to interactively manipulate the color of an image by editing a compact color palette. While remarkable progress has been made by previous methods, they have the common limitations that may produce unfaithful image recoloring results i.e., the obtained result does not respond faithfully to the palette adjustment, and tend to induce visual artifacts such as color bleeding and distortion. To address these limitations, we in this paper present a novel color separation model for palette-based recoloring. Akin to previous methods, our color separation model is built upon the assumption that color of each pixel in an image can be formulated as a linear combination of a small set of same basis colors. However, different from previous palette-based recoloring methods which typically rely on heuristic rules to build the color separation model, we experimentally reveal the underlying relationship between the color separation and the palette-based recoloring, and summarize three specialized color separation priors that allow more faithful palette-based recoloring. Based on these priors, we devise a blind color separation model that not only does not require known palette as input as done in previous methods, but also enables more effective palette-based recoloring with much less visual artifacts. Experiments on two datasets demonstrate that our method outperforms the state-of-the-art palette-based recoloring methods. In addition, we show some applications enabled by the proposed color separation model, including automatic pattern coloring generation, green screen keying and region-controllable color transfer. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Semisupervised Feature Learning by Deep Entropy-Sparsity Subspace ClusteringabstractWhile feature learning by deep neural networks is currently widely used, it is still very challenging to perform this task, given the very limited quantity of labeled data. To solve this problem, we propose to unite subspace clustering with deep semisupervised feature learning to form a unified learning framework to pursue feature learning by subspace clustering. More specifically, we develop a deep entropy-sparsity subspace clustering (deep ESSC) model, which forces a deep neural network to learn features using subspace clustering constrained by our designed entropy-sparsity scheme. The model can inherently harmonize deep semisupervised feature learning and subspace clustering simultaneously by the proposed self-similarity preserving strategy. To optimize the deep ESSC model, we introduce two unconstrained variables to eliminate the two constraints via softmax functions. We provide a general algebraic-treatment scheme for solving the proposed deep ESSC model. Extensive experiments with comprehensive analysis substantiate that our deep ESSC model is more effective than the related methods. Wei-Shi Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | One for More: Selecting Generalizable Samples for Generalizable ReID ModelabstractCurrent training objectives of existing person Re-IDentification (ReID) models only ensure that the loss of the model decreases on selected training batch, with no regards to the performance on samples outside the batch. It will inevitably cause the model to over-fit the data in the dominant position (e.g., head data in imbalanced class, easy samples or noisy samples). The latest resampling methods address the issue by designing specific criterion to select specific samples that trains the model generalize more on certain type of data (e.g., hard samples, tail data), which is not adaptive to the inconsistent real world ReID data distributions. Therefore, instead of simply presuming on what samples are generalizable, this paper proposes a one-for-more training objective that directly takes the generalization ability of selected samples as a loss function and learn a sampler to automatically select generalizable samples. More importantly, our proposed one-for-more based sampler can be seamlessly integrated into the ReID training framework which is able to simultaneously train ReID models and the sampler in an end-to-end fashion. The experimental results show that our method can effectively improve the ReID model training and boost the performance of ReID models. Enwei Zhang, Xinyang Jiang, Hao Cheng 0012, Ancong Wu, Fufu Yu, Ke Li 0015, Feng Zheng 0001, Wei-Shi Zheng 0001, Xing Sun 0001 |
AAAI | 9 |
| 2021 | Learning 3D Shape Feature for Texture-Insensitive Person Re-IdentificationabstractIt is well acknowledged that person re-identification (person ReID) highly relies on visual texture information like clothing. Despite significant progress has been made in recent years, texture-confusing situations like clothing changing and persons wearing the same clothes receive little attention from most existing ReID methods. In this paper, rather than relying on texture based information, we propose to improve the robustness of person ReID against clothing texture by exploiting the information of a person’s 3D shape. Existing shape learning schemas for person ReID either ignore the 3D information of a person, or require extra physical devices to collect 3D source data. Differently, we propose a novel ReID learning framework that directly extracts a texture-insensitive 3D shape embedding from a 2D image by adding 3D body reconstruction as an auxiliary task and regularization, called 3D Shape Learning (3DSL). The 3D reconstruction based regularization forces the ReID model to decouple the 3D shape information from the visual texture, and acquire discriminative 3D shape ReID features. To solve the problem of lacking 3D ground truth, we design an adversarial self-supervised projection (ASSP) model, performing 3D reconstruction without ground truth. Extensive experiments on common ReID datasets and texture-confusing datasets validate the effectiveness of our model. Xinyang Jiang, Fudong Wang 0001, Jun Zhang 0018, Feng Zheng 0001, Xing Sun 0001, Wei-Shi Zheng 0001 |
CVPR | 7 |
| 2021 | MIST: Multiple Instance Self-Training Framework for Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WS-VAD) is to distinguish anomalies from normal events based on discriminative representations. Most existing works are limited in insufficient video representations. In this work, we develop a multiple instance self-training framework (MIST) to efficiently refine task-specific discriminative representations with only video-level annotations. In particular, MIST is composed of 1) a multiple instance pseudo label generator, which adapts a sparse continuous sampling strategy to produce more reliable clip-level pseudo labels, and 2) a self-guided attention boosted feature encoder that aims to automatically focus on anomalous regions in frames while extracting task-specific representations. Moreover, we adopt a self-training scheme to optimize both components and finally obtain a task-specific feature encoder. Extensive experiments on two public datasets demonstrate the efficacy of our method, and our method performs comparably to or even better than existing supervised and weakly supervised methods, specifically obtaining a frame-level AUC 94.83% on ShanghaiTech. Jia-Chang Feng, Fa-Ting Hong, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2021 | Fine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-IdentificationabstractRecently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is challenging since pedestrian images with clothes change exhibit large intra-class variation and small inter-class variation. Some significant features for identification are embedded in unobvious body shape differences across pedestrians. To explore such body shape cues for cloth-changing Re-ID, we propose a Fine-grained Shape-Appearance Mutual learning framework (FSAM), a two-stream framework that learns fine-grained discriminative body shape knowledge in a shape stream and transfers it to an appearance stream to complement the cloth-unrelated knowledge in the appearance features. Specifically, in the shape stream, FSAM learns fine-grained discriminative mask with the guidance of identities and extracts fine-grained body shape features by a pose-specific multi-branch network. To complement cloth-unrelated shape knowledge in the appearance stream, dense interactive mutual learning is performed across low-level and high-level features to transfer knowledge from shape stream to appearance stream, which enables the appearance stream to be deployed independently without extra computation for mask estimation. We evaluated our method on benchmark cloth-changing Re-ID datasets and achieved the start-of-the-art performance. Peixian Hong, Ancong Wu, Xintong Han, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2021 | Combined Depth Space Based Architecture Search for Person Re-IdentificationabstractMost works on person re-identification (ReID) take advantage of large backbone networks such as ResNet, which are designed for image classification instead of ReID, for feature extraction. However, these backbones may not be computationally efficient or the most suitable architectures for ReID. In this work, we aim to design a lightweight and suitable network for ReID. We propose a novel search space called Combined Depth Space (CDS), based on which we search for an efficient network architecture, which we call CDNet, via a differentiable architecture search algorithm. Through the use of the combined basic building blocks in CDS, CDNet tends to focus on combined pattern information that is typically found in images of pedestrians. We then propose a low-cost search strategy named the Top-k Sample Search strategy to make full use of the search space and avoid trapping in local optimal result. Furthermore, an effective Fine-grained Balance Neck (FBLNeck), which is removable at the inference time, is presented to balance the effects of triplet loss and softmax loss during the training process. Extensive experiments show that our CDNet (~1.8 M parameters) has comparable performance with state-of-the-art lightweight networks. Hanjun Li 0004, Gaojie Wu, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2021 | Graph-Based High-Order Relation Modeling for Long-Term Action RecognitionabstractLong-term actions involve many important visual concepts, e.g., objects, motions, and sub-actions, and there are various relations among these concepts, which we call basic relations. These basic relations will jointly affect each other during the temporal evolution of long-term actions, which forms the high-order relations that are essential for long-term action recognition. In this paper, we propose a Graph-based High-order Relation Modeling (GHRM) module to exploit the high-order relations in the long-term actions for long-term action recognition. In GHRM, each basic relation in the long-term actions will be modeled by a graph, where each node represents a segment in a long video. Moreover, when modeling each basic relation, the information from all the other basic relations will be incorporated by GHRM, and thus the high-order relations in the long-term actions can be well exploited. To better exploit the high-order relations along the time dimension, we design a GHRM-layer consisting of a Temporal-GHRM branch and a Semantic-GHRM branch, which aims to model the local temporal high-order relations and global semantic high-order relations. The experimental results on three long-term action recognition datasets, namely, Breakfast, Charades, and MultiThumos, demonstrate the effectiveness of our model. Kun-Yu Lin, Haoxin Li, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2021 | Predictive Feature Learning for Future Segmentation PredictionabstractFuture segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative (with rich details) and are always of high resolution/dimension. Hence, the complicated spatiotemporal variations of these features are difficult to predict, which motivates us to learn a more predictive representation. In this work, we develop a novel framework called Predictive Feature Autoencoder. In the proposed framework, we construct an autoencoder which serves as a bridge between the segmentation features and the predictor. In the latent feature learned by the autoencoder, global structures are enhanced and local details are suppressed so that it is more predictive. In order to reduce the risk of vanishing the suppressed details during recurrent feature prediction, we further introduce a reconstruction constraint in the prediction module. Extensive experiments show the effectiveness of the proposed approach and our method outperforms state-of-the-arts by a considerable margin. Zihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu, Jian-Huang Lai, Wei-Shi Zheng 0001 |
ICCV | 6 |
| 2021 | Learning to Know Where to See: A Visibility-Aware Approach for Occluded Person Re-identificationabstractPerson re-identification (ReID) has gained an impressive progress in recent years. However, the occlusion is still a common and challenging problem for recent ReID methods. Several mainstream methods utilize extra cues (e.g., human pose information) to distinguish human parts from obstacles to alleviate the occlusion problem. Although achieving inspiring progress, these methods severely rely on the fine-grained extra cues, and are sensitive to the estimation error in the extra cues. In this paper, we show that existing methods may degrade if the extra information is sparse or noisy. Thus we propose a simple yet effective method that is robust to sparse and noisy pose information. This is achieved by discretizing pose information to the visibility label of body parts, so as to suppress the influence of occluded regions. We show in our experiments that leveraging pose information in this way is more effective and robust. Besides, our method can be embedded into most person ReID models easily. Extensive experiments validate the effectiveness of our model on common occluded person ReID datasets. Jinrui Yang, Fufu Yu, Xinyang Jiang, Mengdan Zhang, Xing Sun 0001, Ying-Cong Chen, Wei-Shi Zheng 0001 |
ICCV | 8 |
| 2021 | Weakly Supervised Text-based Person Re-IdentificationabstractThe conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting called weakly supervised text-based person re-identification, where only the text-image pairs are available without the requirement of annotating identities during the training phase. To this end, we propose a Cross-Modal Mutual Training (CMMT) framework. Specifically, to alleviate the intra-class variations, a clustering method is utilized to generate pseudo labels for both visual and textual instances. To further re-fine the clustering results, CMMT provides a Mutual Pseudo Label Refinement module, which leverages the clustering results in one modality to refine that in the other modality constrained by the text-image pairwise relationship. Mean-while, CMMT introduces a Text-IoU Guided Cross-Modal Projection Matching loss to resolve the cross-modal matching ambiguity problem. A Text-IoU Guided Hard Sample Mining method is also proposed for learning discriminative textual-visual joint embeddings. We conduct extensive experiments to demonstrate the effectiveness of the proposed CMMT, and the results show that CMMT performs favorably against existing text-based person re-identification methods. Our code will be available at https://github.com/X-BrainLab/WS_Text-ReID. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Wei-Shi Zheng 0001, Nong Sang |
ICCV | 4 |
| 2021 | Cross-Scene Person Trajectory Anomaly Detection Based on Re-IdentificationabstractIn this work, we consider the cross-scene person trajectory anomaly detection problem, which detects the anomalous trajectories across multiple nonoverlapping scenes. This problem is highly significant for public security, but it is still underexplored. Since the trajectory is not continuous across nonoverlapping camera views, we take use of person reidentification (re-ID) to associate the same pedestrian in different scenes while mitigating its inaccuracy by a directional probabilistic graph. To better distinguishing normal samples from anomalies, We formulate a maximized margin graph autoencoder (MMGAE) model, and the reconstruction error of the MMGAE is regarded as an anomaly indicator for the sample. To verify the effectiveness of our approach, we collected and labeled a new dataset. we also explore the impact of the re-ID performance on the anomaly detection problem and the effect of an inaccurately constructed graph on the MMGAE. Yuanxun Li, Ancong Wu, Wei-Shi Zheng 0001 |
ICME | 3 |
| 2021 | Temporal Label Aggregation for Unintentional Action LocalizationabstractHumans can easily understand whether a person’s action is intentional or not. However, it is very challenging to teach a machine to recognize this due to the lack of referable comparisons and reliable annotations. Given a video with unintentional action, the annotations are usually unreliable due to the intrinsic ambiguity from multiple annotators and the subjective appraisals. To address this problem, we propose a new framework which online aggregates multiple probabilistic labels for unintentional action localization. Specifically, we first model the uncertainty of annotations with a temporal probability distribution, and then develop a label attention model to aggregate the reliable annotations in an online manner. We evaluate our method on the public OOPS dataset where each video contains multiple annotations of unintentional action and our experimental results show that mining reliable supervision information from multiple unreliable annotations achieves significant improvements over the baseline methods. Nuoxing Zhou, Guangyi Chen 0002, Jinglin Xu, Wei-Shi Zheng 0001, Jiwen Lu |
ICME | 4 |
| 2021 | Alleviating Data Imbalance Issue with Perturbed Input During Inference
Kanghao Chen, Huijuan Lu, Chenghua Zeng, Wei-Shi Zheng 0001 |
MICCAI (5) | 6 |
| 2021 | Data Augmentation in Logit Space for Medical Image Classification with Limited Training Data
Yangwen Hu, Zhehao Zhong, Hongmei Liu 0001, Zhijun Tan, Wei-Shi Zheng 0001 |
MICCAI (5) | 6 |
| 2021 | Continual Learning with Bayesian Model Based on a Fixed Pre-trained Feature Extractor
Zhiying Cui, Changhong Zhong 0001, Wei-Shi Zheng 0001 |
MICCAI (5) | 6 |
| 2021 | Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant ScenesabstractPerson re-identification (Re-ID) aims to match person images across non-overlapping camera views. The majority of Re-ID methods focus on small-scale surveillance systems in which each pedestrian is captured in different camera views of adjacent scenes. However, in large-scale surveillance systems that cover larger areas, it is required to track a pedestrian of interest across distant scenes (e.g., a criminal suspect escapes from one city to another). Since most pedestrians appear in limited local areas, it is difficult to collect training data with cross-camera pairs of the same person. In this work, we study intra-camera supervised person re-identification across distant scenes (ICS-DS Re-ID), which uses cross-camera unpaired data with intra-camera identity labels for training. It is challenging as cross-camera paired data plays a crucial role for learning camera-invariant features in most existing Re-ID methods. To learn camera-invariant representation from cross-camera unpaired training data, we propose a cross-camera feature prediction method to mine cross-camera self supervision information from camera-specific feature distribution by transforming fake cross-camera positive feature pairs and minimize the distances of the fake pairs. Furthermore, we automatically localize and extract local-level feature by a transformer. Joint learning of global-level and local-level features forms a global-local cross-camera feature prediction scheme for mining fine-grained cross-camera self supervision information. Finally, cross-camera self supervision and intra-camera supervision are aggregated in a framework. The experiments are conducted in the ICS-DS setting on Market-SCT, Duke-SCT and MSMT17-SCT datasets. The evaluation results demonstrate the superiority of our method, which gains significant improvements of 15.4 Rank-1 and 22.3 mAP on Market-SCT as compared to the second best method. Our code is available at https://github.com/g3956/CCFP. Wenhang Ge, Chunyan Pan, Ancong Wu, Hongwei Zheng 0002, Wei-Shi Zheng 0001 |
ACM Multimedia | 5 |
| 2021 | Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WS-TAL) is a challenging task that aims to localize action instances in the given video with video-level categorical supervision. Previous works use the appearance and motion features extracted from pre-trained feature encoder directly,e.g., feature concatenation or score-level fusion. In this work, we argue that the features extracted from the pre-trained extractors,e.g., I3D, which are trained for trimmed video action classification, but not specific for WS-TAL task, leading to inevitable redundancy and sub-optimization. Therefore, the feature re-calibration is needed for reducing the task-irrelevant information redundancy. Here, we propose a cross-modal consensus network(CO2-Net) to tackle this problem. In CO2-Net, we mainly introduce two identical proposed cross-modal consensus modules (CCM) that design a cross-modal attention mechanism to filter out the task-irrelevant information redundancy using the global information from the main modality and the cross-modal local information from the auxiliary modality. Moreover, we further explore inter-modality consistency, where we treat the attention weights derived from each CCM as the pseudo targets of the attention weights derived from another CCM to maintain the consistency between the predictions derived from two CCMs, forming a mutual learning manner. Finally, we conduct extensive experiments on two commonly used temporal action localization datasets, THUMOS14 and ActivityNet1.2, to verify our method, which we achieve state-of-the-art results. The experimental results show that our proposed cross-modal consensus module can produce more representative features for temporal action localization. Fa-Ting Hong, Jia-Chang Feng, Dan Xu 0002, Ying Shan, Wei-Shi Zheng 0001 |
ACM Multimedia | 5 |
| 2021 | Discriminator-free Generative Adversarial AttackabstractThe Deep Neural Networks are vulnerable to adversarial examples (Figure 1), making the DNNs-based systems collapsed by adding the inconspicuous perturbations to the images. Most of the existing works for adversarial attack are gradient-based and suffer from the latency efficiencies and the load on GPU memory. The generative-based adversarial attacks can get rid of this limitation, and some relative works propose the approaches based on GAN. However, suffering from the difficulty of the convergence of training a GAN, the adversarial examples have either bad attack ability or bad visual quality. In this work, we find that the discriminator could be not necessary for generative-based adversarial attack, and propose the Symmetric Saliency-based Auto-Encoder (SSAE) to generate the perturbations, which is composed of the saliency map module and the angle-norm disentanglement of the features module. The advantage of our proposed method lies in that it is not depending on discriminator, and uses the generative saliency map to pay more attention to label-relevant regions. The extensive experiments among the various tasks, datasets, and models demonstrate that the adversarial examples generated by SSAE not only make the widely-used models collapse, but also achieves good visual quality. The code is available at: https://github.com/BravoLu/SSAE. Shaohao Lu, Yuqiao Xian, Xing Sun 0001, Feiyue Huang, Wei-Shi Zheng 0001 |
ACM Multimedia | 8 |
| 2021 | Action-guided 3D Human Motion PredictionabstractThe ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human motion is closely related to the action being performed, we propose to explore action context to guide motion prediction. Specifically, we construct an action-specific memory bank to store representative motion dynamics for each action category, and design a query-read process to retrieve some motion dynamics from the memory bank. The retrieved dynamics are consistent with the action depicted in the observed video frames and serve as a strong prior knowledge to guide motion prediction. We further formulate an action constraint loss to ensure the global semantic consistency of the predicted motion. Extensive experiments demonstrate the effectiveness of the proposed approach, and we achieve state-of-the-art performance on 3D human motion prediction. Jiangxin Sun, Zihang Lin, Xintong Han, Jianfang Hu, Jia Xu 0011, Wei-Shi Zheng 0001 |
NeurIPS | 6 |
| 2021 | Joint regression and learning from pairwise rankings for personalized image aesthetic assessmentabstractRecent image aesthetic assessment methods have achieved remarkable progress due to the emergence of deep convolutional neural networks (CNNs). However, these methods focus primarily on predicting generally perceived preference of an image, making them usually have limited practicability, since each user may have completely different preferences for the same image. To address this problem, this paper presents a novel approach for predicting personalized image aesthetics that fit an individual user’s personal taste. We achieve this in a coarse to fine manner, by joint regression and learning from pairwise rankings. Specifically, we first collect a small subset of personal images from a user and invite him/her to rank the preference of some randomly sampled image pairs. We then search for the K -nearest neighbors of the personal images within a large-scale dataset labeled with average human aesthetic scores, and use these images as well as the associated scores to train a generic aesthetic assessment model by CNN-based regression. Next, we fine-tune the generic model to accommodate the personal preference by training over the rankings with a pairwise hinge loss. Experiments demonstrate that our method can effectively learn personalized image aesthetic preferences, clearly outperforming state-of-the-art methods. Moreover, we show that the learned personalized image aesthetic benefits a wide variety of applications. Qing Zhang 0006, Jian-Hao Fan, Wei Sun 0007, Wei-Shi Zheng 0001 |
Comput. Vis. Media | 5 |
| 2021 | Letter-Level Online Writer Identification
Zelin Chen, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Identity-constrained noise modeling with metric learning for face anti-spoofing
Yaowen Xu, Lifang Wu, Meng Jian, Wei-Shi Zheng 0001, Zhuming Wang |
Neurocomputing | 4 |
| 2021 | Person Re-Identification by Contour Sketch Under Moderate Clothing ChangeabstractPerson re-identification (re-id), the process of matching pedestrian images across different camera views, is an important task in visual surveillance. Substantial development of re-id has recently been observed, and the majority of existing models are largely dependent on color appearance and assume that pedestrians do not change their clothes across camera views. This limitation, however, can be an issue for re-id when tracking a person at different places and at different time if that person (e.g., a criminal suspect) changes his/her clothes, causing most existing methods to fail, since they are heavily relying on color appearance, and thus, they are inclined to match a person to another person wearing similar clothes. In this work, we call the person re-id under clothing change the "cross-clothes person re-id." In particular, we consider the case when a person only changes his clothes moderately as a first attempt at solving this problem based on visible light images; that is, we assume that a person wears clothes of a similar thickness, and thus the shape of a person would not change significantly when the weather does not change substantially within a short period of time. We perform cross-clothes person re-id based on a contour sketch of person image to take advantage of the shape of the human body instead of color information for extracting features that are robust to moderate clothing change. To select/sample more reliable and discriminative curve patterns on a body contour sketch, we introduce a learning-based spatial polar transformation (SPT) layer in the deep neural network to transform contour sketch images for extracting reliable and discriminant convolutional neural network (CNN) features in a polar coordinate space. An angle-specific extractor (ASE) is applied in the following layers to extract more fine-grained discriminant angle-specific features. By varying the sampling range of the SPT, we develop a multistream network for aggregating multi-granularity features to better identify a person. Due to the lack of a large-scale dataset for cross-clothes person re-id, we contribute a new dataset that consists of 33,698 images from 221 identities. Our experiments illustrate the challenges of cross-clothes person re-id and demonstrate the effectiveness of our proposed method. Qize Yang, Ancong Wu, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Weakly-supervised semantic segmentation with saliency and incremental supervision updating
Wenfeng Luo, Meng Yang 0001, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2021 | Joint adaptive manifold and embedding learning for unsupervised feature selection
Jian-Sheng Wu, Meng-Xiao Song, Weidong Min, Jian-Huang Lai, Wei-Shi Zheng 0001 |
Pattern Recognit. | 5 |
| 2021 | Online deep transferable dictionary learning
Ancong Wu, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2021 | Arbitrary-View Human Action Recognition: A Varying-View RGB-D Action DatasetabstractCurrent researches of action recognition which focus on single-view and multi-view recognition can hardly satisfy the requirements of human-robot interaction (HRI) applications for recognizing human actions from arbitrary views. Arbitrary-view recognition is still a challenging issue due to view changes and visual occlusions. In addition, the lack of datasets also sets up barriers. To provide data for arbitrary-view action recognition, we collect a new large-scale RGB-D action dataset for arbitrary-view action analysis, including RGB videos, depth and skeleton sequences. The dataset includes action samples captured in 8 fixed viewpoints and varying-view sequences which cover the entire 360° view angles. In total, 118 persons are invited to act 40 action categories. Our dataset involves more participants, more viewpoints and a large number of samples. More importantly, it is the first dataset containing the entire 360° varying-view sequences. The dataset provides sufficient data for multi-view, cross-view and arbitrary-view action analysis. Besides, we propose a View-guided Skeleton CNN (VS-CNN) to tackle the problem of arbitrary-view action recognition. Experiment results show that the VS-CNN achieves superior performance, and our dataset provides valuable but challenging data for the evaluation of arbitrary-view recognition. Yanli Ji, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Guest Editorial Introduction to the Special Issue on Large-Scale Visual Sensor Networks: Architectures and ApplicationsabstractLarge–scale visual sensor networks have become progressively an essential part of our daily lives underpinning many technological, financial, and social advancements today, with applications in smart cities, traffic monitoring, environmental pollution control, public safety, and crime prevention. Paolo Spagnolo, Hamid K. Aghajan, George Bebis, Shaogang Gong, Amy Loutfi, Leonid Sigal, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Enhancing Underexposed Photos Using Perceptually Bidirectional SimilarityabstractAlthough remarkable progress has been made, existing methods for enhancing underexposed photos tend to produce visually unpleasing results due to the existence of visual artifacts (e.g., color distortion, loss of details and uneven exposure). We observed that this is because they fail to ensure the perceptual consistency of visual information between the source underexposed image and its enhanced output. To obtain high-quality results free of these artifacts, we present a novel underexposed photo enhancement approach that is able to maintain the perceptual consistency. We achieve this by proposing an effective criterion, referred to as perceptually bidirectional similarity, which explicitly describes how to ensure the perceptual consistency. Particularly, we adopt the Retinex theory and cast the enhancement problem as a constrained illumination estimation optimization, where we formulate perceptually bidirectional similarity as constraints on illumination and solve for the illumination which can recover the desired artifact-free enhancement results. In addition, we describe a video enhancement framework that adopts the presented illumination estimation for handling underexposed videos. To this end, a probabilistic approach is introduced to propagate illuminations of sampled keyframes to the entire video by tackling a Bayesian Maximum A Posteriori problem. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time AspectabstractRecently, the research interest of person re-identification (ReID) has gradually turned to video-based methods, which acquire a person representation by aggregating frame features of an entire video. However, existing video-based ReID methods do not consider the semantic difference brought by the outputs of different network stages, which potentially compromises the information richness of the person features. Furthermore, traditional methods ignore important relationship among frames, which causes information redundancy in fusion along the time axis. To address these issues, we propose a novel general temporal fusion framework to aggregate frame features on both semantic aspect and time aspect. As for the semantic aspect, a multi-stage fusion network is explored to fuse richer frame features at multiple semantic levels, which can effectively reduce the information loss caused by the traditional single-stage fusion. While, for the time axis, the existing intra-frame attention method is improved by adding a novel inter-frame attention module, which effectively reduces the information redundancy in temporal fusion by taking the relationship among frames into consideration. The experimental results show that our approach can effectively improve the video-based re-identification accuracy, achieving the state-of-the-art performance. Xinyang Jiang, Yifei Gong, Qize Yang, Feiyue Huang, Wei-Shi Zheng 0001, Feng Zheng 0001, Xing Sun 0001 |
AAAI | 6 |
| 2020 | Deep Camouflage ImagesabstractThis paper addresses the problem of creating camouflage images. Such images typically contain one or more hidden objects embedded into a background image, so that viewers are required to consciously focus to discover them. Previous methods basically rely on hand-crafted features and texture synthesis to create camouflage images. However, due to lack of reliable understanding of what essentially makes an object recognizable, they typically result in either complete standout or complete invisible hidden objects. Moreover, they may fail to produce seamless and natural images because of the sensitivity to appearance differences. To overcome these limitations, we present a novel neural style transfer approach that adopts the visual perception mechanism to create camouflage images, which allows us to hide objects more effectively while producing natural-looking results. In particular, we design an attention-aware camouflage loss to adaptively mask out information that make the hidden objects visually standout, and also leave subtle yet enough feature clues for viewers to perceive the hidden objects. To remove the appearance discontinuities between the hidden objects and the background, we formulate a naturalness regularization to constrain the hidden objects to maintain the manifold structure of the covered background. Extensive experiments show the advantages of our approach over existing camouflage methods and state-of-the-art neural style transfer algorithms. Qing Zhang 0006, Gelin Yin, Yongwei Nie, Wei-Shi Zheng 0001 |
AAAI | 4 |
| 2020 | Viewpoint-Aware Loss with Angular Regularization for Person Re-IdentificationabstractAlthough great progress in supervised person re-identification (Re-ID) has been made recently, due to the viewpoint variation of a person, Re-ID remains a massive visual challenge. Most existing viewpoint-based person Re-ID methods project images from each viewpoint into separated and unrelated sub-feature spaces. They only model the identity-level distribution inside an individual viewpoint but ignore the underlying relationship between different viewpoints. To address this problem, we propose a novel approach, called Viewpoint-Aware Loss with Angular Regularization (VA-reID). Instead of one subspace for each viewpoint, our method projects the feature from different viewpoints into a unified hypersphere and effectively models the feature distribution on both the identity-level and the viewpoint-level. In addition, rather than modeling different viewpoints as hard labels used for conventional viewpoint classification, we introduce viewpoint-aware adaptive label smoothing regularization (VALSR) that assigns the adaptive soft label to feature representation. VALSR can effectively solve the ambiguity of the viewpoint cluster label assignment. Extensive experiments on the Market1501 and DukeMTMC-reID datasets demonstrated that our method outperforms the state-of-the-art supervised Re-ID methods. Zhihui Zhu, Xinyang Jiang, Feng Zheng 0001, Feiyue Huang, Xing Sun 0001, Wei-Shi Zheng 0001 |
AAAI | 7 |
| 2020 | Anomaly Detection on Electroencephalography with Self-supervised LearningabstractEpilepsy is one of the most common neurological diseases in humans, and electroencephalography (EEG) is the most widely used method for clinicians to detect epileptic seizures. However, it is error-prone to detect epileptic seizures by manually observing EEG, and labeling epilepsy data is an expensive and time-consuming process. In this study, without requiring any epileptic EEG data and only based on normal EEGs, a new selfsupervised learning method is proposed for anomaly detection on EEG signals. In particular, a series of scaling transformations are performed on the original EEG data to generated self-labeled scaled EEG data, where different labels correspond to different scaling transformations. Then using the self-labeled normal EEG dataset, a multi-class classifier can be trained to accurately predict the scaling transformations on new normal EEG data, but not accurately on abnormal (epileptic) EEGs. The inconsistency between the predicted scaling transformations and the groundtruth scaling transformations can then be used to measure the degree of abnormality in a new EEG data. Comprehensive experimental evaluations demonstrate that the proposed self-supervised method outperforms classic anomaly detection methods including one-class support vector machine (SVM) and autoencoders. The robustness of the proposed method also has been empirically proved with different classifier structures and by varying relevant hyper-parameters. Yaojia Zheng, Wei-Shi Zheng 0001 |
BIBM | 5 |
| 2020 | Learning to Detect Important People in Unlabelled Images for Semi-Supervised Important People DetectionabstractImportant people detection is to automatically detect the individuals who play the most important roles in a social event image, which requires the designed model to understand a high-level pattern. However, existing methods rely heavily on supervised learning using large quantities of annotated image samples, which are more costly to collect for important people detection than for individual entity recognition (i.e., object recognition). To overcome this problem, we propose learning important people detection on partially annotated images. Our approach iteratively learns to assign pseudo-labels to individuals in un-annotated images and learns to update the important people detection model based on data with both labels and pseudo-labels. To alleviate the pseudo-labelling imbalance problem, we introduce a ranking strategy for pseudo-label estimation, and also introduce two weighting strategies: one for weighting the confidence that individuals are important people to strengthen the learning on important people and the other for neglecting noisy unlabelled images (i.e., images without any important people). We have collected two large-scale datasets for evaluation. The extensive experimental results clearly confirm the efficacy of our method attained by leveraging unlabelled images for improving the performance of important people detection. Fa-Ting Hong, Wei-Hong Li 0001, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2020 | Adaptive Interaction Modeling via Graph Operations SearchabstractInteraction modeling is important for video action analysis. Recently, several works design specific structures to model interactions in videos. However, their structures are manually designed and non-adaptive, which require structures design efforts and more importantly could not model interactions adaptively. In this paper, we automate the process of structures design to learn adaptive structures for interaction modeling. We propose to search the network structures with differentiable architecture search mechanism, which learns to construct adaptive structures for different videos to facilitate adaptive interaction modeling. To this end, we first design the search space with several basic graph operations that explicitly capture different relations in videos. We experimentally demonstrate that our architecture search framework learns to construct adaptive interaction modeling structures, which provides more understanding about the relations between the structures and some interaction characteristics, and also releases the requirement of structures design efforts. Additionally, we show that the designed basic graph operations in the search space are able to model different interactions in videos. The experiments on two interaction datasets show that our method achieves competitive performance with state-of-the-arts. Haoxin Li, Wei-Shi Zheng 0001, Haifeng Hu 0001, Jian-Huang Lai |
CVPR | 2 |
| 2020 | Spatial-Temporal Graph Convolutional Network for Video-Based Person Re-IdentificationabstractWhile video-based person re-identification (Re-ID) has drawn increasing attention and made great progress in recent years, it is still very challenging to effectively overcome the occlusion problem and the visual ambiguity problem for visually similar negative samples. On the other hand, we observe that different frames of a video can provide complementary information for each other, and the structural information of pedestrians can provide extra discriminative cues for appearance features. Thus, modeling the temporal relations of different frames and the spatial relations within a frame has the potential for solving the above problems. In this work, we propose a novel Spatial-Temporal Graph Convolutional Network (STGCN) to solve these problems. The STGCN includes two GCN branches, a spatial one and a temporal one. The spatial branch extracts structural information of a human body. The temporal branch mines discriminative cues from adjacent frames. By jointly optimizing these branches, our model extracts robust spatial-temporal information that is complementary with appearance information. As shown in the experiments, our model achieves state-of-the-art results on MARS and DukeMTMC-VideoReID datasets. Jinrui Yang, Wei-Shi Zheng 0001, Qize Yang, Ying-Cong Chen, Qi Tian 0001 |
CVPR | 2 |
| 2020 | Weakly Supervised Discriminative Feature Learning With State Information for Person IdentificationabstractUnsupervised learning of identity-discriminative visual feature is appealing in real-world tasks where manual labelling is costly. However, the images of an identity can be visually discrepant when images are taken under different states, e.g. different camera views and poses. This visual discrepancy leads to great difficulty in unsupervised discriminative learning. Fortunately, in real-world tasks we could often know the states without human annotation, e.g. we can easily have the camera view labels in person re-identification and facial pose labels in face recognition. In this work we propose utilizing the state information as weak supervision to address the visual discrepancy caused by different states. We formulate a simple pseudo label model and utilize the state information in an attempt to refine the assigned pseudo labels by the weakly supervised decision boundary rectification and weakly supervised feature drift regularization. We evaluate our model on unsupervised person re-identification and pose-invariant face recognition. Despite the simplicity of our method, it could outperform the state-of-the-art results on Duke-reID, MultiPIE and CFP datasets with a standard ResNet-50 backbone. We also find our model could perform comparably with the standard supervised fine-tuning results on the three datasets. Code is available at https://github.com/KovenYu/state-information}. Hong-Xing Yu, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2020 | Squeeze-and-Attention Networks for Semantic SegmentationabstractThe recent integration of attention mechanisms into segmentation networks improves their representational capabilities through a great emphasis on more informative features. However, these attention mechanisms ignore an implicit sub-task of semantic segmentation and are constrained by the grid structure of convolution kernels. In this paper, we propose a novel squeeze-and-attention network (SANet) architecture that leverages an effective squeeze-and-attention (SA) module to account for two distinctive characteristics of segmentation: i) pixel-group attention, and ii) pixel-wise prediction. Specifically, the proposed SA modules impose pixel-group attention on conventional convolution by introducing an 'attention' convolutional channel, thus taking into account spatial-channel inter-dependencies in an efficient manner. The final segmentation results are produced by merging outputs from four hierarchical stages of a SANet to integrate multi-scale contexts for obtaining an enhanced pixel-wise prediction. Empirical experiments on two challenging public datasets validate the effectiveness of the proposed SANets, which achieves 83.2 % mIoU (without COCO pre-training) on PASCAL VOC and a state-of-the-art mIoU of 54.4 % on PASCAL Context. Zilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu, Ibrahim Ben Daya, Wei-Shi Zheng 0001, Jonathan Li 0001, Alexander Wong |
CVPR | 7 |
| 2020 | An Asymmetric Modeling for Action Assessment
Jibin Gao, Wei-Shi Zheng 0001, Chengying Gao, Yaowei Wang 0001, Wei Zeng 0006, Jian-Huang Lai |
ECCV (30) | 2 |
| 2020 | MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li 0001, Wei-Shi Zheng 0001 |
ECCV (13) | 4 |
| 2020 | Contextual Heterogeneous Graph Network for Human-Object Interaction Detection
Wei-Shi Zheng 0001, Yingbiao Ling |
ECCV (17) | 2 |
| 2020 | Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
Shizhen Zhao, Changxin Gao, Jun Zhang 0018, Hao Cheng 0012, Chuchu Han, Xinyang Jiang, Wei-Shi Zheng 0001, Nong Sang, Xing Sun 0001 |
ECCV (6) | 8 |
| 2020 | Semi-supervised Person Re-identification by Attribute Similarity GuidanceabstractAlthough supervised person re-identification (RE-ID) has achieved great progress with deep learning, it requires time-consuming annotation of a large number of pedestrian identities. To reduce labeling cost, we attempt to reduce cross-camera identity annotations and exploit pedestrian attribute annotations as auxiliary information instead. The pedestrian attributes, such as outfit styles, contain coarse semantic knowledge. Although pedestrian attributes are annotated without exhaustive searching in a camera network, which is much easier than cross-camera identity annotation, ambiguity exists in attributes when different persons have similar outfits. To solve this problem, we propose an Attribute Similarity Guidance loss (ASG) to guide appearance feature learning for RE-ID by selective attribute similarity preservation to avoid the impact of such ambiguity. Finally, we develop an attribute-guided self training framework to jointly utilize attribute annotations, unlabeled data and limited labeled data for semi-supervised learning. Extensive experiments on Market-1501 and DukeMTMC-ReID show the superiority of our method for semi-supervised RE-ID. Peixian Hong, Ancong Wu, Wei-Shi Zheng 0001 |
ICPR | 3 |
| 2020 | Cross-View Relation Networks for Mammogram Mass DetectionabstractIn medical image analysis, multi-view modeling is crucial for pathology detection when the target lesion is presented in different views, e.g. mass lesions in breasts. Currently mammogram is the most effective imaging modality for mass lesion detection of breast cancer at the early stage. The pathological information from the two paired views (i.e., medio-Iateral oblique and cranio-caudal) are highly relational and complementary, which is crucial for diagnosis in clinical practice. Existing mass detection methods do not consider learning synergistic features from the two relational views. For the first time, we propose a novel mass detection framework to capture the latent relation information from the two paired views of a same mass in mammogram. We evaluate our model on a public mammogram dataset and a large-scale private dataset, demonstrating that the proposed method outperforms existing feature fusion approaches and state-of-the-art mass detection methods. We further analyze the performance gains from the relation modeling. Our quantitative and qualitative results suggest that jointly learning cross- view features boosts the detection performance of existing models, which is a promising avenue for mass detection task in mammogram. Jiechao Ma, Xiang Li 0032, Hongwei Li 0004, Bjoern Menze, Wei-Shi Zheng 0001 |
ICPR | 6 |
| 2020 | Neighbor Combinatorial Attention for Critical Structure MiningabstractGraph convolutional networks (GCNs) have been widely used to process graph-structured data. However, existing GNN methods do not explicitly extract critical structures, which reflect the intrinsic property of a graph. In this work, we propose a novel GCN module named Neighbor Combinatorial ATtention (NCAT) to find critical structure in graph-structured data. NCAT attempts to match combinatorial neighbors with learnable patterns and assigns different weights to each combination based on the matching degree between the patterns and combinations. By stacking several NCAT modules, we can extract hierarchical structures that is helpful for down-stream tasks. Our experimental results show that NCAT achieves state-of-the-art performance on several benchmark graph classification datasets. In addition, we interpret what kind of features our model learned by visualizing the extracted critical structures. Tanli Zuo, Yukun Qiu, Wei-Shi Zheng 0001 |
IJCAI | 3 |
| 2020 | A Block Decomposition Algorithm for Sparse OptimizationabstractSparse optimization is a central problem in machine learning and computer vision. However, this problem is inherently NP-hard and thus difficult to solve in general. Combinatorial search methods find the global optimal solution but are confined to small-sized problems, while coordinate descent methods are efficient but often suffer from poor local minima. This paper considers a new block decomposition algorithm that combines the effectiveness of combinatorial search methods and the efficiency of coordinate descent methods. Specifically, we consider a random strategy or/and a greedy strategy to select a subset of coordinates as the working set, and then perform a global combinatorial search over the working set based on the original objective function. We show that our method finds stronger stationary points than Amir Beck et al.'s coordinate-wise optimization method. In addition, we establish the convergence rate of our algorithm. Our experiments on solving sparse regularized and sparsity constrained least squares optimization problems demonstrate that our method achieves state-of-the-art performance in terms of accuracy. For example, our method generally outperforms the well-known greedy pursuit method. Ganzhao Yuan, Li Shen 0008, Wei-Shi Zheng 0001 |
KDD | 3 |
| 2020 | Continual Learning of New Diseases with Dual Distillation and Ensemble Strategy
Zhuoyun Li, Changhong Zhong 0001, Wei-Shi Zheng 0001 |
MICCAI (1) | 4 |
| 2020 | Abnormality Detection in Chest X-Ray Images Using Uncertainty Prediction Autoencoders
Feifei Xue, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Hongmei Liu 0001 |
MICCAI (6) | 5 |
| 2020 | Deep kNN for Medical Image Classification
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
MICCAI (1) | 5 |
| 2020 | Transductive Multi-Object Tracking in Complex Events by Interactive Self-TrainingabstractRecently, multi-object tracking (MOT) for estimating trajectories of pedestrians has undergone fast development and played an important role in human-centric video analysis. However, video analysis in complex events (e.g. scenes in HiEve dataset) is still under-explored. In complex real-world scenarios, domain gap in unseen testing scenes and severe occlusion problem that disconnects tracks are challenging for existing online MOT methods without domain adaptation. To alleviate domain gap, we study the problem in a transductive learning setting, which assumes that unlabeled testing data is available for learning offline tracking. We propose a transductive interactive self-training method to adapt the tracking model to unseen crowded scenes with unlabeled testing data by means of teacher-student interative learning. To reduce prediction variance in an unseen domain, we train two different models and teach one model with pseudo labels of unlabeled data predicted by the other model interactively. To improve robustness against occlusions during self-training, we exploit disconnected track interpolation (DTI) to refine the predicted pseudo labels. Our method achieved MOTA of 60.23 on HiEve dataset and won the first place of Multi-person Motion Tracking in Complex Events (with Private Detection) in the ACM MM Grand Challenge on Large-scale Human-centric Video Analysis in Complex Events. Ancong Wu, Chengzhi Lin, Bogao Chen, Weihao Huang, Wei-Shi Zheng 0001 |
ACM Multimedia | 6 |
| 2020 | Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long VideosabstractThe objective of action quality assessment is to score sports videos. However, most existing works focus only on video dynamic information (i.e., motion information) but ignore the specific postures that an athlete is performing in a video, which is important for action assessment in long videos. In this work, we present a novel hybrid dynAmic-static Context-aware attenTION NETwork (ACTION-NET) for action assessment in long videos. To learn more discriminative representations for videos, we not only learn the video dynamic information but also focus on the static postures of the detected athletes in specific frames, which represent the action quality at certain moments, along with the help of the proposed hybrid dynamic-static architecture. Moreover, we leverage a context-aware attention module consisting of a temporal instance-wise graph convolutional network unit and an attention unit for both streams to extract more robust stream features, where the former is for exploring the relations between instances and the latter for assigning a proper weight to each instance. Finally, we combine the features of the two streams to regress the final video score, supervised by ground-truth scores given by experts. Additionally, we have collected and annotated the new Rhythmic Gymnastics dataset, which contains videos of four different types of gymnastics routines, for evaluation of action quality assessment in long videos. Extensive experimental results validate the efficacy of our proposed method, which outperforms related approaches. Ling-An Zeng, Fa-Ting Hong, Wei-Shi Zheng 0001, Qi-Zhi Yu, Wei Zeng 0006, Yaowei Wang 0001, Jian-Huang Lai |
ACM Multimedia | 3 |
| 2020 | Aggregating Spatio-temporal Context for Video Object Segmentation
Jianfang Hu, Wei-Shi Zheng 0001 |
PRCV (1) | 3 |
| 2020 | S3D-CNN: skeleton-based 3D consecutive-low-pooling neural network for fall detection
Xin Xiong 0016, Weidong Min, Wei-Shi Zheng 0001, Pin Liao, Hao Yang 0027 |
Appl. Intell. | 3 |
| 2020 | RGB-IR Person Re-identification by Cross-Modality Similarity Preservation
Ancong Wu, Wei-Shi Zheng 0001, Shaogang Gong, Jian-Huang Lai |
Int. J. Comput. Vis. | 2 |
| 2020 | Fine-Grained Person Re-identification
Jiahang Yin, Ancong Wu, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | On large appearance change in visual tracking
Yun Liang 0003, Meihua Wang, Yanwen Guo 0001, Wei-Shi Zheng 0001 |
Neural Comput. Appl. | 4 |
| 2020 | Unsupervised Person Re-Identification by Deep Asymmetric Metric EmbeddingabstractPerson re-identification (Re-ID) aims to match identities across non-overlapping camera views. Researchers have proposed many supervised Re-ID models which require quantities of cross-view pairwise labelled data. This limits their scalabilities to many applications where a large amount of data from multiple disjoint camera views is available but unlabelled. Although some unsupervised Re-ID models have been proposed to address the scalability problem, they often suffer from the view-specific bias problem which is caused by dramatic variances across different camera views, e.g., different illumination, viewpoints and occlusion. The dramatic variances induce specific feature distortions in different camera views, which can be very disturbing in finding cross-view discriminative information for Re-ID in the unsupervised scenarios, since no label information is available to help alleviate the bias. We propose to explicitly address this problem by learning an unsupervised asymmetric distance metric based on cross-view clustering. The asymmetric distance metric allows specific feature transformations for each camera view to tackle the specific feature distortions. We then design a novel unsupervised loss function to embed the asymmetric metric into a deep neural network, and therefore develop a novel unsupervised deep framework named the DEep Clustering-based Asymmetric MEtric Learning (DECAMEL). In such a way, DECAMEL jointly learns the feature representation and the unsupervised asymmetric metric. DECAMEL learns a compact cross-view cluster structure of Re-ID data, and thus help alleviate the view-specific bias and facilitate mining the potential cross-view discriminative information for unsupervised Re-ID. Extensive experiments on seven benchmark datasets whose sizes span several orders show the effectiveness of our framework. Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Joint graph regularized dictionary learning and sparse ranking for multi-modal multi-shot person re-identification
Aihua Zheng, Bo Jiang 0002, Wei-Shi Zheng 0001, Bin Luo 0001 |
Pattern Recognit. | 4 |
| 2020 | Fast Collective Activity Recognition Under Weak SupervisionabstractCollective activity recognition, which tells what activity a group of people is performing, is a cutting-edge research topic in computer vision. Different from action performed by individuals, collective activity needs to consider the complex interactions among different people. However, most previous works require exhaustive annotations such as accurate label information of individual actions, pairwise interactions, and poses, which could not be easily available in practice. Moreover, most of them treat human detection as a decoupled task before collective activity recognition and leverage all detected persons. This not only ignores the mutual relation between the two tasks, which makes it hard for filtering out irrelevant people, but also probably increases the computation burden when reasoning the collective activities. In this paper, we propose a fast weakly supervised deep learning architecture for collective activity recognition. For fast inference, we propose to make the actor detection and weakly supervised collective activity reasoning collaborate in an end-to-end framework by sharing convolutional layers between them. The joint learning makes the two tasks united and reinforced each other, so that it is more effective to filter out the outliers who are not involved in the activity. For the weakly supervised learning, we propose a latent embedding scheme for mining person-group interactive relationship to get rid of the use of any pairwise relation between people and the individual action labels as well. The experimental results show that the proposed framework achieves comparable or even better performance as compared to the state-of-the-art on three datasets. Our joint modelling reasons collective activities at the speed of 22.65 fps, which is the fastest ever known and substantially makes collective activity recognition more towards real-time applications. Peizhen Zhang, Yongyi Tang, Jianfang Hu, Wei-Shi Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Action Knowledge Transfer for Action Prediction with Partial VideosabstractPredicting action class from partially observed videos, which is known as action prediction, is an important task in computer vision field with many applications. The challenge for action prediction mainly lies in the lack of discriminative action information for the partially observed videos. To tackle this challenge, in this work, we propose to transfer action knowledge learned from fully observed videos for improving the prediction of partially observed videos. Specifically, we develop a two-stage learning framework for action knowledge transfer. At the first stage, we learn feature embeddings and discriminative action classifier from full videos. The knowledge in the learned embeddings and classifier is then transferred to the partial videos at the second stage. Our experiments on the UCF-101 and HMDB-51 datasets show that the proposed action knowledge transfer method can significantly improve the performance of action prediction, especially for the actions with small observation ratios (e.g., 10%). We also experimentally illustrate that our method outperforms all the state-of-the-art action prediction systems. Yijun Cai, Haoxin Li, Jianfang Hu, Wei-Shi Zheng 0001 |
AAAI | 4 |
| 2019 | Deep Dual Relation Modeling for Egocentric Interaction RecognitionabstractEgocentric interaction recognition aims to recognize the camera wearer's interactions with the interactor who faces the camera wearer in egocentric videos. In such a human-human interaction analysis problem, it is crucial to explore the relations between the camera wearer and the interactor. However, most existing works directly model the interactions as a whole and lack modeling the relations between the two interacting persons. To exploit the strong relations for egocentric interaction recognition, we introduce a dual relation modeling framework which learns to model the relations between the camera wearer and the interactor based on the individual action representations of the two persons. Specifically, we develop a novel interactive LSTM module, the key component of our framework, to explicitly model the relations between the two interacting persons based on their individual action representations, which are collaboratively learned with an interactor attention module and a global-local motion module. Experimental results on three egocentric interaction datasets show the effectiveness of our method and advantage over state-of-the-arts. Haoxin Li, Yijun Cai, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2019 | Learning to Learn Relation for Important People Detection in Still ImagesabstractHumans can easily recognize the importance of people in social event images, and they always focus on the most important individuals. However, learning to learn the relation between people in an image, and inferring the most important person based on this relation, remains undeveloped. In this work, we propose a deep imPOrtance relatIon NeTwork (POINT) that combines both relation modeling and feature learning. In particular, we infer two types of interaction modules: the person-person interaction module that learns the interaction between people and the event-person interaction module that learns to describe how a person is involved in the event occurring in an image. We then estimate the importance relations among people from both interactions and encode the relation feature from the importance relations. In this way, POINT automatically learns several types of relation features in parallel, and we aggregate these relation features and the person's feature to form the importance feature for important people classification. Extensive experimental results show that our method is effective for important people detection and verify the efficacy of learning to learn relations for important people detection. Wei-Hong Li 0001, Fa-Ting Hong, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2019 | Weakly Supervised Person Re-IdentificationabstractIn the conventional person re-id setting, it is assumed that the labeled images are the person images within the bounding box for each individual; this labeling across multiple nonoverlapping camera views from raw video surveillance is costly and time-consuming. To overcome this difficulty, we consider weakly supervised person re-id modeling. The weak setting refers to matching a target person with an untrimmed gallery video where we only know that the identity appears in the video without the requirement of annotating the identity in any frame of the video during the training procedure. Hence, for a video, there could be multiple video-level labels. We cast this weakly supervised person re-id challenge into a multi-instance multi-label learning (MIML) problem. In particular, we develop a Cross-View MIML (CV-MIML) method that is able to explore potential intraclass person images from all the camera views by incorporating the intra-bag alignment and the cross-view bag alignment. Finally, the CV-MIML method is embedded into an existing deep neural network for developing the Deep Cross-View MIML (Deep CV-MIML) model. We have performed extensive experiments to show the feasibility of the proposed weakly supervised setting and verify the effectiveness of our method compared to related methods on four weakly labeled datasets. Jingke Meng, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2019 | Weakly Supervised Open-Set Domain Adaptation by Dual-Domain CollaborationabstractIn conventional domain adaptation, a critical assumption is that there exists a fully labeled domain (source) that contains the same label space as another unlabeled or scarcely labeled domain (target). However, in the real world, there often exist application scenarios in which both domains are partially labeled and not all classes are shared between these two domains. Thus, it is meaningful to let partially labeled domains learn from each other to classify all the unlabeled samples in each domain under an open-set setting. We consider this problem as weakly supervised open-set domain adaptation. To address this practical setting, we propose the Collaborative Distribution Alignment (CDA) method, which performs knowledge transfer bilaterally and works collaboratively to classify unlabeled data and identify outlier samples. Extensive experiments on the Office benchmark and an application on person reidentification show that our method achieves state-of-the-art performance. Shuhan Tan, Jiening Jiao, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2019 | Progressive Teacher-Student Learning for Early Action PredictionabstractThe goal of early action prediction is to recognize actions from partially observed videos with incomplete action executions, which is quite different from action recognition. Predicting early actions is very challenging since the partially observed videos do not contain enough action information for recognition. In this paper, we aim at improving early action prediction by proposing a novel teacherstudent learning framework. Our framework involves a teacher model for recognizing actions from full videos, a student model for predicting early actions from partial videos, and a teacher-student learning block for distilling progressive knowledge from teacher to student, crossing different tasks. Extensive experiments on three public action datasets show that the proposed progressive teacher-student learning framework can consistently improve performance of early action prediction model. We have also reported the state-of-the-art performances for early action prediction on all of these sets. Xionghui Wang, Jianfang Hu, Jian-Huang Lai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
CVPR | 5 |
| 2019 | Underexposed Photo Enhancement Using Deep Illumination EstimationabstractThis paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability to learn complex photographic adjustment from expert-retouched input/output image pairs. Based on this model, we formulate a loss function that adopts constraints and priors on the illumination, prepare a new dataset of 3,000 underexposed image pairs, and train the network to effectively learn a rich variety of adjustment for diverse lighting conditions. By these means, our network is able to recover clear details, distinct contrast, and natural color in the enhancement results. We perform extensive experiments on the benchmark MIT-Adobe FiveK dataset and our new dataset, and show that our network is effective to deal with previously challenging images. Ruixing Wang, Qing Zhang 0006, Chi-Wing Fu, Xiaoyong Shen, Wei-Shi Zheng 0001, Jiaya Jia |
CVPR | 5 |
| 2019 | Distilled Person Re-Identification: Towards a More Scalable SystemabstractPerson re-identification (Re-ID), for matching pedestrians across non-overlapping camera views, has made great progress in supervised learning with abundant labelled data. However, the scalability problem is the bottleneck for applications in large-scale systems. We consider the scalability problem of Re-ID from three aspects: (1) low labelling cost by reducing label amount, (2) low extension cost by reusing existing knowledge and (3) low testing computation cost by using lightweight models. The requirements render scalable Re-ID a challenging problem. To solve these problems in a unified system, we propose a Multi-teacher Adaptive Similarity Distillation Framework, which requires only a few labelled identities of target domain to transfer knowledge from multiple teacher models to a user-specified lightweight student model without accessing source domain data. We propose the Log-Euclidean Similarity Distillation Loss for Re-ID and further integrate the Adaptive Knowledge Aggregator to select effective teacher models to transfer target-adaptive knowledge. Extensive evaluations show that our method can extend with high scalability and the performance is comparable to the state-of-the-art unsupervised and semi-supervised Re-ID methods. Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai |
CVPR | 2 |
| 2019 | Patch-Based Discriminative Feature Learning for Unsupervised Person Re-IdentificationabstractWhile discriminative local features have been shown effective in solving the person re-identification problem, they are limited to be trained on fully pairwise labelled data which is expensive to obtain. In this work, we overcome this problem by proposing a patch-based unsupervised learning framework in order to learn discriminative feature from patches instead of the whole images. The patch-based learning leverages similarity between patches to learn a discriminative model. Specifically, we develop a PatchNet to select patches from the feature map and learn discriminative features for these patches. To provide effective guidance for the PatchNet to learn discriminative patch feature on unlabeled datasets, we propose an unsupervised patch-based discriminative feature learning loss. In addition, we design an image-level feature learning loss to leverage all the patch features of the same image to serve as an image-level guidance for the PatchNet. Extensive experiments validate the superiority of our method for unsupervised person re-id. Our code is available at https://github.com/QizeYang/PAUL. Qize Yang, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2019 | Unsupervised Person Re-Identification by Soft Multilabel LearningabstractAlthough unsupervised person re-identification (RE-ID) has drawn increasing research attentions due to its potential to address the scalability problem of supervised RE-ID models, it is very challenging to learn discriminative information in the absence of pairwise labels across disjoint camera views. To overcome this problem, we propose a deep model for the soft multilabel learning for unsupervised RE-ID. The idea is to learn a soft multilabel (real-valued label likelihood vector) for each unlabeled person by comparing the unlabeled person with a set of known reference persons from an auxiliary domain. We propose the soft multilabel-guided hard negative mining to learn a discriminative embedding for the unlabeled target domain by exploring the similarity consistency of the visual features and the soft multilabels of unlabeled target pairs. Since most target pairs are cross-view pairs, we develop the cross-view consistent soft multilabel learning to achieve the learning goal that the soft multilabels are consistently good across different camera views. To enable effecient soft multilabel learning, we introduce the reference agent learning to represent each reference person by a reference agent in a joint embedding. We evaluate our unified deep model on Market-1501 and DukeMTMC-reID. Our model outperforms the state-of-the-art unsupervised RE-ID methods by clear margins. Code is available at https://github.com/KovenYu/MAR. Hong-Xing Yu, Wei-Shi Zheng 0001, Ancong Wu, Shaogang Gong, Jian-Huang Lai |
CVPR | 2 |
| 2019 | A Decomposition Algorithm for the Sparse Generalized Eigenvalue ProblemabstractThe sparse generalized eigenvalue problem arises in a number of standard and modern statistical learning models, including sparse principal component analysis, sparse Fisher discriminant analysis, and sparse canonical correlation analysis. However, this problem is difficult to solve since it is NP-hard. In this paper, we consider a new effective decomposition method to tackle this problem. Specifically, we use random or/and swapping strategies to find a working set and perform global combinatorial search over the small subset of variables. We consider a bisection search method and a coordinate descent method for solving the quadratic fractional programming subproblem. In addition, we provide some theoretical analysis for the proposed method. Our experiments on synthetic data and real-world data have shown that our method significantly and consistently outperforms existing solutions in term of accuracy. Ganzhao Yuan, Li Shen 0008, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2019 | Action Assessment by Joint Relation GraphsabstractWe present a new model to assess the performance of actions from videos, through graph-based joint relation modelling. Previous works mainly focused on the whole scene including the performer's body and background, yet they ignored the detailed joint interactions. This is insufficient for fine-grained, accurate action assessment, because the action quality of each joint is dependent of its neighbouring joints. Therefore, we propose to learn the detailed joint motion based on the joint relations. We build trainable Joint Relation Graphs, and analyze joint motion on them. We propose two novel modules, the Joint Commonality Module and the Joint Difference Module, for joint motion learning. The Joint Commonality Module models the general motion for certain body parts, and the Joint Difference Module models the motion differences within body parts. We evaluate our method on six public Olympic actions for performance assessment. Our method outperforms previous approaches (+0.0912) and the whole-scene analysis (+0.0623) in the Spearman's Rank Correlation. We also demonstrate our model's ability to interpret the action assessment process. Jibin Gao, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2019 | Unsupervised Person Re-Identification by Camera-Aware Similarity Consistency LearningabstractFor matching pedestrians across disjoint camera views in surveillance, person re-identification (Re-ID) has made great progress in supervised learning. However, it is infeasible to label data in a number of new scenes when extending a Re-ID system. Thus, studying unsupervised learning for Re-ID is important for saving labelling cost. Yet, cross-camera scene variation is a key challenge for unsupervised Re-ID, such as illumination, background and viewpoint variations, which cause domain shift in the feature space and result in inconsistent pairwise similarity distributions that degrade matching performance. To alleviate the effect of cross-camera scene variation, we propose a Camera-Aware Similarity Consistency Loss to learn consistent pairwise similarity distributions for intra-camera matching and cross-camera matching. To avoid learning ineffective knowledge in consistency learning, we preserve the prior common knowledge of intra-camera matching in the pretrained model as reliable guiding information, which does not suffer from cross-camera scene variation as cross-camera matching. To learn similarity consistency more effectively, we further develop a coarse-to-fine consistency learning scheme to learn consistency globally and locally in two steps. Experiments show that our method outperformed the state-of-the-art unsupervised Re-ID methods. Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai |
ICCV | 2 |
| 2019 | Towards Photo-Realistic Visible Watermark Removal with Conditional Generative Adversarial Networks
Xiang Li 0032, Chan Lu, Danni Cheng, Wei-Hong Li 0001, Mei Cao, Jiechao Ma, Wei-Shi Zheng 0001 |
ICIG (1) | 8 |
| 2019 | Learning Prediction of Emotional Change on BehaviorsabstractWe are interested in the topic about the prediction of emotional change caused by other people's behaviors. Although information of other people's behaviors is intuitive and important for arousing emotional change, existing works on prediction of emotional change rarely take it into consideration. In this work, we propose a multimodal deep network, called Emotional Change Prediction Network (ECPNet), for predicting emotional changes caused by behaviors of other people. ECPNet can learn useful cues from visual and acoustic information and extract both local and global spatiotemporal multimodal features of behaviors. Furthermore, the proposed network can learn the temporal relations of the behaviors in videos which help us to explore the behaviors' cumulative effects on other people's emotions. In order to quantify our model, we establish the first dataset on emotional change prediction by behaviors. Experiment results show that our approach can effectively model emotional change and perform better than other baselines. Wei-Shi Zheng 0001 |
ICIP | 2 |
| 2019 | Unsupervised Learning for Optical Flow Estimation Using Pyramid Convolution LSTMabstractMost of current Convolution Neural Network (CNN) based methods for optical flow estimation focus on learning optical flow on synthetic datasets with groundtruth, which is not practical. In this paper, we propose an unsupervised optical flow estimation framework named PCLNet. It uses pyramid Convolution LSTM (ConvLSTM) with the constraint of adjacent frame reconstruction, which allows flexibly estimating multi-frame optical flows from any video clip. Besides, by decoupling motion feature learning and optical flow representation, our method avoids complex short-cut connections used in existing frameworks while improving accuracy of optical flow estimation. Moreover, different from those methods using specialized CNN architectures for capturing motion, our framework directly learns optical flow from the features of generic CNNs and thus can be easily embedded in any CNN based frameworks for other tasks. Extensive experiments have verified that our method not only estimates optical flow effectively and accurately, but also obtains comparable performance on action recognition. Shuosen Guan, Haoxin Li, Wei-Shi Zheng 0001 |
ICME | 3 |
| 2019 | Affective Video Content Analyses by Using Cross-Modal Embedding Learning FeaturesabstractMost existing methods on affective video content analyses are dedicated to single media, either visual content or audio content and few attempts for combined analysis of the two media signals are made. In this paper, we employ a cross-modal embedding learning approach to learn the compact feature representations of different modalities that are discriminative for analyzing the emotion attributes of the video. Specifically, we introduce inter-modal similarity constraints and intra-modal similarity constraints to promote the joint embedding learning procedure for obtaining the robust features. In order to capture cues in different grains, global and local features are extracted from both visual and audio signals, thereafter a unified framework consisting with global and local features embedding networks is built for affective video content analyses. Experiments show that our proposed approach significantly outperforms the state-of-the-art methods and demonstrate the effectiveness of our approach. Benchao Li, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2019 | Fast Person Search PipelineabstractIn this work, we study the person search problem, which incorporates both person detection and person re-identification. The high time consumption of person search system obstructs the development of this field in practice. For the purpose of real-time person search, we analyze the bottleneck of the person search task and therefore develop a more efficient method called the Fast Person Search Pipeline (FPSP), consisting of a multi-scale person detector as well as a part-based person re-identification descriptor. Extensive experiments show that our proposed method FPSP can accomplish the person search task approximately FIVE times faster than the state-of-the-art methods while achieving a competitive accuracy. Jianheng Li, Fuhang Liang, Yuanxun Li, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2019 | Pedestrian re-Identification Based on Tree Branch Network with Local and Global LearningabstractDeep part-based methods in recent literature have revealed the great potential of learning local part-level representation for pedestrian image in the task of person re-identification. However, global features that capture discriminative holistic information of human body are usually ignored or not well exploited. This motivates us to investigate joint learning global and local features from pedestrian images. Specifically, in this work, we propose a novel framework termed tree branch network (TBN) for person re-identification. Given a pedestrain image, the feature maps generated by the backbone CNN, are partitioned recursively into several pieces, each of which is followed by a bottleneck structure that learns finer-grained features for each level in the hierarchical tree-like framework. In this way, representations are learned in a coarse-to-fine manner and finally assembled to produce more discriminative image descriptions. Experimental results demonstrate the effectiveness of the global and local feature learning method in the proposed TBN framework. We also show significant improvement in performance over state-of-the-art methods on three public benchmarks: Market-1501, CUHK-03 and DukeMTMC. Meng Yang 0001, Zhihui Lai 0001, Wei-Shi Zheng 0001, Zitong Yu |
ICME | 4 |
| 2019 | Deep Semi-Supervised Person Re-Identification with External MemoryabstractTo overcome the scalability problem of supervised person re-identification (Re-ID), we consider the semi-supervised person Re-ID problem of learning from a limited number of labeled images of a few identities and a large number of unlabeled images. To this end, we propose an external-memory-based deep semi-supervised person Re-ID model (EDS). Based on the external memory, two loss functions are designed so as to effectively cope with the relation between labeled and unlabeled data for overcoming the limitation of batch size in each epoch in deep learning. Therefore, an effective deep semi-supervised learning method can be performed. Extensive experiments validate the superiority of the proposed method for semi-supervised person Re-ID. Qize Yang, Ancong Wu, Wei-Shi Zheng 0001 |
ICME | 3 |
| 2019 | Particle Swarm Loss for Lightweight Object DetectionabstractCurrently in object detection, deep learning based detectors are gaining their momentum. However, the supervision involved in the widely-used anchor paradigm within the detection pipeline is inadequate. Traditional object detectors opt for densely picking anchors to increase the training samples for faster convergence and better detection quality. However, dense anchor scheme requires extra computational budget which renders it infeasible for lightweight detectors. To address the problem, inspired by the cognitive consistency, we propose a novel Particle Swarm Loss for lightweight object detection. Experiments upon the MS-COCO challenge show that detectors compensated by PS loss can not only converge faster but also acquire better detection quality than their vanilla versions (YOLOv3 and SSD improves 2.0% and 2.5% on the harsh AP50respectively) without extra computational overhead. In addition, we propose a dapper backbone with high cost-efficiency for the resource-limited scenarios. Peizhen Zhang, Feng Zheng 0001, Junlong Du, Jun Zhang 0018, Wei-Shi Zheng 0001 |
ICME | 6 |
| 2019 | Part-Based Convolutional Network for Imbalanced Age EstimationabstractAge estimation based on unconstrained face images remains a challenging problem in computer vision and pattern recognition. We address this by ensembling part-based features and designing a cost sensitive loss to overcome the general imbalance data in age estimation. Specifically, we treat age estimation as a multi-class classification problem and mainly make two contributions: (i) We present a Part-based Convolutional Network(PCN) for age estimation to extract regional features instead of the holistic ones, which can preserve more discriminative features in facial regions like forehead, canthus, cheek, jaw, etc; (ii) A balanced loss for multi-class classification is designed to handle the extremely imbalance age distribution. The loss pays more attention to the hard examples, and can automatically adjust the weights of samples according to their contribution. State-of-the-art performance was achieved on FG-NET, MORPH and CACD databases, validating the effectiveness of the proposed approach. Jun-Yong Zhu, Wei-Shi Zheng 0001 |
ICME | 3 |
| 2019 | DBDNet: Learning Bi-directional Dynamics for Early Action PredictionabstractPredicting future actions from observed partial videos is very challenging as the missing future is uncertain and sometimes has multiple possibilities. To obtain a reliable future estimation, a novel encoder-decoder architecture is proposed for integrating the tasks of synthesizing future motions from observed videos and reconstructing observed motions from synthesized future motions in an unified framework, which can capture the bi-directional dynamics depicted in partial videos along the temporal (past-to-future) direction and reverse chronological (future-back-to-past) direction. We then employ a bi-directional long short-term memory (Bi-LSTM) architecture to exploit the learned bi-directional dynamics for predicting early actions. Our experiments on two benchmark action datasets show that learning bi-directional dynamics benefits the early action prediction and our system clearly outperforms the state-of-the-art methods. Guoliang Pang, Xionghui Wang, Jianfang Hu, Qing Zhang 0006, Wei-Shi Zheng 0001 |
IJCAI | 5 |
| 2019 | Improving Robustness of Medical Image Diagnosis with Denoising Convolutional Neural Networks
Feifei Xue, Wei-Shi Zheng 0001 |
MICCAI (6) | 5 |
| 2019 | Biomarker Localization by Combining CNN Classifier and Generative Adversarial Network
Shuhan Tan, Siyamalan Manivannan, Haotian Lin 0001, Wei-Shi Zheng 0001 |
MICCAI (1) | 7 |
| 2019 | Predicting Future Instance Segmentation with Contextual Pyramid ConvLSTMsabstractDespite the remarkable progress in instance segmentation, the problem of predicting future instance segmentation remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting pyramid features to represent unobserved future frames. However, they mainly predict features for each pyramid level independently, and ignore the underlying structural relationship between features of different levels. Jiangxin Sun, Jiafeng Xie, Jianfang Hu, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001 |
ACM Multimedia | 7 |
| 2019 | Contour-Guided Person Re-identification
Qize Yang, Jingke Meng, Wei-Shi Zheng 0001, Jian-Huang Lai |
PRCV (3) | 4 |
| 2019 | Ensemble Transductive Learning for Skin Lesion Segmentation
Zhiying Cui, Longshi Wu, Wei-Shi Zheng 0001 |
PRCV (2) | 4 |
| 2019 | Dual Illumination Estimation for Robust Exposure CorrectionabstractAbstract Exposure correction is one of the fundamental tasks in image processing and computational photography. While various methods have been proposed, they either fail to produce visually pleasing results, or only work well for limited types of image (e.g., underexposed images). In this paper, we present a novel automatic exposure correction method, which is able to robustly produce high‐quality results for images of various exposure conditions (e.g., underexposed, overexposed, and partially under‐ and over‐exposed). At the core of our approach is the proposed dual illumination estimation, where we separately cast the under‐and over‐exposure correction as trivial illumination estimation of the input image and the inverted input image. By performing dual illumination estimation, we obtain two intermediate exposure correction results for the input image, with one fixes the underexposed regions and the other one restores the overexposed regions. A multi‐exposure image fusion technique is then employed to adaptively blend the visually best exposed parts in the two intermediate exposure correction images and the input image into a globally well‐exposed image. Experiments on a number of challenging images demonstrate the effectiveness of the proposed approach and its superiority over the state‐of‐the‐art methods and popular automatic exposure correction tools. Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
Comput. Graph. Forum | 3 |
| 2019 | Robust joint representation with triple local feature for face recognition with single sample per person
Xing Wang 0012, Bob Zhang 0001, Meng Yang 0001, Kangyin Ke, Wei-Shi Zheng 0001 |
Knowl. Based Syst. | 5 |
| 2019 | Early Action Prediction by Soft RegressionabstractWe propose a novel approach for predicting on-going action with the assistance of a low-cost depth camera. Our approach introduces a soft regression-based early prediction framework. In this framework, we estimate soft labels for the subsequences at different progress levels, jointly learned with an action predictor. Our formulation of soft regression framework 1) overcomes a usual assumption in existing early action prediction systems that the progress level of on-going sequence is given in the testing stage; and 2) presents a theoretical framework to better resolve the ambiguity and uncertainty of subsequences at early performing stage. The proposed soft regression framework is further enhanced in order to take the relationships among subsequences and the discrepancy of soft labels over different classes into consideration, so that a Multiple Soft labels Recurrent Neural Network (MSRNN) is finally developed. For real-time performance, we also introduce a new RGB-D feature called "local accumulative frame feature (LAFF)", which can be computed efficiently by constructing an integral feature map. Our experiments on three RGB-D benchmark datasets and an unconstrained RGB action set demonstrate that the proposed regression-based early action prediction model outperforms existing models significantly and also show that the early action prediction on RGB-D sequence is more accurate than that on RGB channel. Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | One-pass person re-identification by sketch online discriminant analysis
Wei-Hong Li 0001, Allen Z. Zhong, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2019 | Deep asymmetric video-based person re-identification
Jingke Meng, Ancong Wu, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2018 | Deep Low-Resolution Person Re-IdentificationabstractPerson images captured by public surveillance cameras often have low resolutions (LR) in addition to uncontrolled pose variations, background clutters and occlusions. This gives rise to the resolution mismatch problem when matched against the high resolution (HR) gallery images (typically available in enrolment), which adversely affects the performance of person re-identification (re-id) that aims to associate images of the same person captured at different locations and different time. Most existing re-id methods either ignore this problem or simply upscale LR images. In this work, we address this problem by developing a novel approach called Super-resolution and Identity joiNt learninG (SING) to simultaneously optimise image super-resolution and person re-id matching. This approach is instantiated by designing a hybrid deep Convolutional Neural Network for improving cross-resolution re-id performance. We further introduce an adaptive fusion algorithm for accommodating multi-resolution LR images. Extensive evaluations show the advantages of our method over related state-of-the-art re-id and super-resolution methods on cross-resolution re-id benchmarks. Jiening Jiao, Wei-Shi Zheng 0001, Ancong Wu, Xiatian Zhu, Shaogang Gong |
AAAI | 2 |
| 2018 | Improving Fast Segmentation With Teacher-Student Learning
Jiafeng Xie, Bing Shuai, Jianfang Hu, Wei-Shi Zheng 0001 |
BMVC | 5 |
| 2018 | Deep Bilinear Learning for RGB-D Action Recognition
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
ECCV (7) | 2 |
| 2018 | Adversarial Open-World Person Re-Identification
Xiang Li 0032, Ancong Wu, Wei-Shi Zheng 0001 |
ECCV (2) | 3 |
| 2018 | Letter-Level Writer IdentificationabstractWriter Identification aims to identify a certain writer from a given group of candidates by their handwriting. Although it is very significant in security systems like bank account verification systems, existing works focus on document-level or text-level writer identification. This limits their scalabilities and flexibilities in realistic scenarios as they require complete document or text. To facilitate the realistic applications of writer identification, we propose a novel technology, letter-level writer identification, which requires only a few letters as the identification cue. It is challenging due to large intra-class discrepancy and implicit identifiable writing cues. Considering these challenges, we propose a novel deep model called Multi-Branch Encoding net (Mul-BEnc). To evaluate our model and provide a benchmark for this problem, we have collected a large Letter-stroke sequence Writer identification DataBase (LetWriterDB). The experimental results validate the effectiveness of our model. Zelin Chen, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001 |
FG | 4 |
| 2018 | PersonRank: Detecting Important People in ImagesabstractAlways, some individuals in images are more important/attractive than the others in some events such as presentation, basketball game or speech. However, it is challenging to ?nd important people among all individuals in an image directly based on their spatial or appearance information due to the existence of diverse variations of pose, action, appearance of persons an various changes of occasions. We overcome this challenge by constructing a multiple HyperInteraction Graph that treats each individual in an image as a node and inferring the most active node from the interactions estimated by using various types of cues. We model a pairwise interaction between people as an edge message communicated between nodes, resulting in a bidirectional pairwise-interaction graph. To enrich the person-person interaction estimation, we further introduce a unidirectional hyper-interaction graph that models the consensus of interactions between a focal person and any person in his/her local region around. Finally, we modify the PageRank algorithm to infer the activeness of people on the multiple Hybrid-Interaction Graph (HIG), the union of the pairwise-interaction and hyper-interaction graphs, and we call our algorithm the PersonRank. In order to provide publicable datasets for evaluation, we have contributed a new dataset called Multi-scene Important People Image Dataset and gathered a NCAA Basketball Image Dataset from sports game sequences. We have demonstrated that the proposed PersonRank outperforms related methods clearly and substantially. Our code and datasets are available at https://weihonglee.github.io/Projects/PersonRank.htm. Wei-Hong Li 0001, Benchao Li, Wei-Shi Zheng 0001 |
FG | 3 |
| 2018 | GL-PAM RGB-D Gesture RecognitionabstractThe existing approaches for RGB-D gesture recognition mainly developed their systems based on the global features extracted from full sequences, which makes them unreliable for capturing some important movements. In this paper, we propose to combine the global and local context information extracted from posture, appearance, and motion sequences. Our experimental results on a large scale RGB-D gesture dataset show that the proposed global and local contexts can complement well with each other for efficiently characterizing gestures, and thus achieve the 2nd place in the ChaLearn LAP Large-scale Isolated Gesture Recognition Challenge (Round 2). Benchao Li, Wanhua Li 0001, Yongyi Tang, Jianfang Hu, Wei-Shi Zheng 0001 |
ICIP | 5 |
| 2018 | Light Person Re-Identification by Multi-Cue Tiny NetabstractNowadays, person re-identification (re-id) receives much attraction and it is for matching person images across disjoint camera views. Although many methods are developed, none of the state-of-the-art models (especially for the deep models) can be deployed by a camera's chip due to limited memory and weak computation. To address this problem, we propose a framework called Multi-Cue Tiny Net, which is a combination of tiny convolutional neural networks (CNNs) with small model size. And we call the person re-identification based on limited memory and weak computation the Light Person Re-Identification. Three different tiny nets are used to learn the complementary features and four cues of images are used for learning features with different information. All the features will be concatenated to form a fusion feature. Besides, to reduce the dimension of the features, and keep the small model size and high performance accuracy, we pre-trained the tiny nets and then retrained them after adding a fully connected layer to the global average pooling layer. Experimental results on challenging person re-identification dataset show that our approach yields promising accuracy with light memory and computation. Weiji Wu, Ancong Wu, Wei-Shi Zheng 0001 |
ICIP | 3 |
| 2018 | Unsupervised Learning for Forecasting Action RepresentationsabstractMost of previous works on future forecasting require a mass of videos with frame-level labels which would probably limit their application, since labelling video frame requires much tremendous efforts. In this paper, we present a unsupervised learning framework to anticipate the future representation by utilizing temporal historical information and train the anticipating capacity only using unlabelled videos. Compared to existing methods that predict the future representation from a static image, our proposed model presents a novel temporal context learning model for estimating the temporal evolution tendency by compacting outputs of all time steps in a LST-M. We evaluate the proposed model on two different activity datasets, TV Human Interaction dataset and THUMOS Validation and Test sets. We have demonstrated the effectiveness of our model in anticipating future representation task. Wei-Shi Zheng 0001 |
ICIP | 2 |
| 2018 | Learning Intrinsic Image Decomposition by Deep Neural Network with Perceptual LossabstractIntrinsic Image Decomposition (IID) refers to recovering the albedo and shading from images, and it plays important roles in addressing computer vision tasks such as illumination-invariant object recognition and image recoloring. IID is an ill-posed problem and lacks of actual labelled samples for learning. This paper presents a deep neural network (DNN) based method to address this problem. To facilitate the training of DNN, we synthesize an intrinsic image dataset through rendering \pmb 3D models. To make the learnt model well generalize to realworld images with better visual results, we employ the perceptual loss in model learning. The perceptual loss is constructed upon the activations of a neural network pre-trained on real images, such as the VGG network. Such a loss function implicitly introduces knowledge from real-world images and has the ability of multi-level semantic understanding on the decomposed results. Experimental results show that our model trained on synthetic single-object dataset can produce well decomposition results not only on synthetic images but also on real-world scene-level images (containing multiple objects). Guangyun Han, Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai |
ICPR | 3 |
| 2018 | Does A Body Image Tell Age?abstractAge estimation is an important task in computer vision and is widely used in applications. However, such a technology is largely affected by the resolution of face, and it would be a challenge if one has to estimate the age of a person at a distance. While body image of a person is often captured more clearly, when and how to use body-based visual cues for age estimation are largely under studied. In this work, we argue that body-based visual cues are better for estimating the age group and can assist the estimation of exact age value. For this purpose, we develop a Body-based Age Net (BAN) that unifies selective local convolution features and contextual convolution features. The network is designed based on two assumptions: 1) a person's wearing is closely related to his/her age group property; 2) some selective local parts of a body are more discriminative for age group estimation. We have contributed a large-scale and publicly available Body Age (BAG) dataset. We have quantitatively evaluated the proposed model on BAG. Baoyu Yuan, Ancong Wu, Wei-Shi Zheng 0001 |
ICPR | 3 |
| 2018 | MIXGAN: Learning Concepts from Different Domains for Mixture GenerationabstractIn this work, we present an interesting attempt on mixture generation: absorbing different image concepts (e.g., content and style) from different domains and thus generating a new domain with learned concepts. In particular, we propose a mixture generative adversarial network (MIXGAN). MIXGAN learns concepts of content and style from two domains respectively, and thus can join them for mixture generation in a new domain, i.e., generating images with content from one domain and style from another. MIXGAN overcomes the limitation of current GAN-based models which either generate new images in the same domain as they observed in training stage, or require off-the-shelf content templates for transferring or translation. Extensive experimental results demonstrate the effectiveness of MIXGAN as compared to related state-of-the-art GAN-based models. Guang-Yuan Hao, Hong-Xing Yu, Wei-Shi Zheng 0001 |
IJCAI | 3 |
| 2018 | Long-Term Human Motion Prediction by Modeling Motion Context and Enhancing Motion DynamicsabstractHuman motion prediction aims at generating future frames of human motion based on an observed sequence of skeletons. Recent methods employ the latest hidden states of a recurrent neural network (RNN) to encode the historical skeletons, which can only address short-term prediction. In this work, we propose a motion context modeling by summarizing the historical human motion with respect to the current prediction. A modified highway unit (MHU) is proposed for efficiently eliminating motionless joints and estimating next pose given the motion context. Furthermore, we enhance the motion dynamic by minimizing the gram matrix loss for long-term motion prediction. Experimental results show that the proposed model can promisingly forecast the human future movements, which yields superior performances over related state-of-the-art approaches. Moreover, specifying the motion context with the activity labels enables our model to perform human motion transfer. Yongyi Tang, Lin Ma 0002, Wei Liu 0005, Wei-Shi Zheng 0001 |
IJCAI | 4 |
| 2018 | Representation Learning for Scene Graph Completion via Jointly Structural and Visual EmbeddingabstractThis paper focuses on scene graph completion which aims at predicting new relations between two entities utilizing existing scene graphs and images. By comparing with the well-known knowledge graph, we first identify that each scene graph is associated with an image and each entity of a visual triple in a scene graph is composed of its entity type with attributes and grounded with a bounding box in its corresponding image. We then propose an end-to-end model named Representation Learning via Jointly Structural and Visual Embedding (RLSV) to take advantages of structural and visual information in scene graphs. In RLSV model, we provide a fully-convolutional module to extract the visual embeddings of a visual triple and apply hierarchical projection to combine the structural and visual embeddings of a visual triple. In experiments, we evaluate our model on two scene graph completion tasks: link prediction and visual triple classification, and further analyze by case studies. Experimental results demonstrate that our model outperforms all baselines in both tasks, which justifies the significance of combining structural and visual information for scene graph completion. Hai Wan, Yonghao Luo, Bo Peng 0041, Wei-Shi Zheng 0001 |
IJCAI | 4 |
| 2018 | Adversarial Attribute-Image Person Re-identificationabstractWhile attributes have been widely used for person re-identification (Re-ID) which aims at matching the same person images across disjoint camera views, they are used either as extra features or for performing multi-task learning to assist the image-image matching task. However, how to find a set of person images according to a given attribute description, which is very practical in many surveillance applications, remains a rarely investigated cross-modality matching problem in person Re-ID. In this work, we present this challenge and leverage adversarial learning to formulate the attribute-image cross-modality person Re-ID model. By imposing a semantic consistency constraint across modalities as a regularization, the adversarial learning enables to generate image-analogous concepts of query attributes for matching the corresponding images at both global level and semantic ID level. We conducted extensive experiments on three attribute datasets and demonstrated that the regularized adversarial modelling is so far the most effective method for the attribute-image cross-modality person Re-ID problem. Zhou Yin, Wei-Shi Zheng 0001, Ancong Wu, Hong-Xing Yu, Hai Wan, Feiyue Huang, Jian-Huang Lai |
IJCAI | 2 |
| 2018 | A Large-scale RGB-D Database for Arbitrary-view Human Action RecognitionabstractCurrent researches mainly focus on single-view and multiview human action recognition, which can hardly satisfy the requirements of human-robot interaction (HRI) applications to recognize actions from arbitrary views. The lack of databases also sets up barriers. In this paper, we newly collect a large-scale RGB-D action database for arbitrary-view action analysis, including RGB videos, depth and skeleton sequences. The database includes action samples captured in 8 fixed viewpoints and varying-view sequences which covers the entire 360 view angles. In total, 118 persons are invited to act 40 action categories, and 25,600 video samples are collected. Our database involves more articipants, more viewpoints and a large number of samples. More importantly, it is the first database containing the entire 360? varying-view sequences. The database provides sufficient data for cross-view and arbitrary-view action analysis. Besides, we propose a View-guided Skeleton CNN (VS-CNN) to tackle the problem of arbitrary-view action recognition. Experiment results show that the VS-CNN achieves superior performance. Yanli Ji, Feixiang Xu, Yang Yang 0002, Fumin Shen, Heng Tao Shen, Wei-Shi Zheng 0001 |
ACM Multimedia | 6 |
| 2018 | High-Quality Exposure Correction of Underexposed PhotosabstractWe address the problem of correcting the exposure of underexposed photos. Previous methods have tackled this problem from many different perspectives and achieved remarkable progress. However, they usually fail to produce natural-looking results due to the existence of visual artifacts such as color distortion, loss of detail, exposure inconsistency, etc. We find that the main reason why existing methods induce these artifacts is because they break a perceptually similarity between the input and output. Based on this observation, an effective criterion, termed as perceptually bidirectional similarity (PBS) is proposed. Based on this criterion and the Retinex theory, we cast the exposure correction problem as an illumination estimation optimization, where PBS is defined as three constraints for estimating illumination that can generate the desired result with even exposure, vivid color and clear textures. Qualitative and quantitative comparisons, and the user study demonstrate the superiority of our method over the state-of-the-art methods. Qing Zhang 0006, Ganzhao Yuan, Chunxia Xiao, Lei Zhu 0003, Wei-Shi Zheng 0001 |
ACM Multimedia | 5 |
| 2018 | Large-Scale Visible Watermark Detection and Removal with Deep Convolutional Networks
Danni Cheng, Xiang Li 0032, Wei-Hong Li 0001, Chan Lu, Fake Li, Wei-Shi Zheng 0001 |
PRCV (3) | 7 |
| 2018 | TW-Co-k-means: Two-level weighted collaborative k-means for multi-view clustering
Chang-Dong Wang 0001, Dong Huang 0001, Wei-Shi Zheng 0001 |
Knowl. Based Syst. | 4 |
| 2018 | Person Re-Identification by Camera Correlation Aware Feature AugmentationabstractThe challenge of person re-identification (re-id) is to match individual images of the same person captured by different non-overlapping camera views against significant and unknown cross-view feature distortion. While a large number of distance metric/subspace learning models have been developed for re-id, the cross-view transformations they learned are view-generic and thus potentially less effective in quantifying the feature distortion inherent to each camera view. Learning view-specific feature transformations for re-id (i.e., view-specific re-id), an under-studied approach, becomes an alternative resort for this problem. In this work, we formulate a novel view-specific person re-identification framework from the feature augmentation point of view, called Camera coR relation Aware Feature augmenTation (CRAFT). Specifically, CRAFT performs cross-view adaptation by automatically measuring camera correlation from cross-view visual data distribution and adaptively conducting feature augmentation to transform the original features into a new adaptive space. Through our augmentation framework, view-generic learning algorithms can be readily generalized to learn and optimize view-specific sub-models whilst simultaneously modelling view-generic discrimination information. Therefore, our framework not only inherits the strength of view-generic model learning but also provides an effective way to take into account view specific characteristics. Our CRAFT framework can be extended to jointly learn view-specific feature transformations for person re-id across a large network with more than two cameras, a largely under-investigated but realistic re-id setting. Additionally, we present a domain-generic deep person appearance representation which is designed particularly to be towards view invariant for facilitating cross-view adaptation by CRAFT. We conducted extensively comparative experiments to validate the superiority and advantages of our proposed framework over state-of-the-art competitors on contemporary challenging person re-id datasets. Ying-Cong Chen, Xiatian Zhu, Wei-Shi Zheng 0001, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Maximal granularity structure and generalized multi-view discriminant analysis for person re-identification
Cairong Zhao, Xuekuan Wang, Duoqian Miao 0001, Hanli Wang, Wei-Shi Zheng 0001, Yong Xu 0001, David Zhang 0001 |
Pattern Recognit. | 5 |
| 2018 | Learning an Intrinsic Image Decomposer Using Synthesized RGB-D DatasetabstractIntrinsic image decomposition refers to recover the albedo and shading from images, which is an ill-posed problem in signal processing. As realistic labeled data are severely lacking, it is difficult to apply learning methods in this issue. In this letter, we propose using a synthesized dataset to facilitate the solving of this problem. A physically based renderer is used to generate color images and their underlying ground-truth albedo and shading from three-dimensional models. Additionally, we render a Kinect-like noisy depth map for each instance. We utilize this synthetic dataset to train a deep neural network for intrinsic image decomposition and further fine-tune it for real-world images. Our model supports both RGB and RGB-D as input, and it employs both high-level and low-level features to avoid blurry outputs. Experimental results verify the effectiveness of our model on realistic images. Guangyun Han, Xiaohua Xie, Jian-Huang Lai, Wei-Shi Zheng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Euler Clustering on Large-Scale DatasetabstractOur concern is nonlinear clustering on large-scale dataset. While existing popular kernels (RBF, Polynomials, Spatial Pyramid, etc.) are popularly used for implicitly mapping data into a high-dimensional or infinite dimensional space in order to generalise linear clustering methods, using these kernels cannot make kernel clustering approaches directly applicable for large scale dataset, since large scale kernel matrix or similarity matrix consumes a lot of memory (e.g., 7,450 GB memory over 1 million samples of data). To solve this problem, we introduce an Euler clustering approach. Euler clustering employs Euler kernels in order to intrinsically map the input data onto a complex space of the same dimension as the input or twice, so that Euler clustering can get rid of kernel trick and does not need to rely on any approximation or random sampling on kernel function/matrix, whilst performing a more robust nonlinear clustering against noise and outliers. Moreover, since the original Euler kernel cannot generate a non-negative similarity matrix and thus is inapplicable to spectral clustering, we introduce a positive Euler kernel, and more importantly we have proved when it can generate a non-negative similarity matrix. We apply Euler kernel and the proposed positive Euler kernel to kernel k-means and spectral clustering so as to develop Euler k-means and Euler spectral clustering, respectively. An efficient Stiefel-manifold-based gradient method and an equivalent weighted positive Euler k-means are derived for fast computation of Euler spectral clustering and further alleviating the impact of discretization of the cluster membership indicators in Euler spectral clustering. The results show that the proposed Euler clustering approach achieves overall better clustering performance compared to using popular Mercer kernels and approximation models, whilst keeping the computational complexity of the same magnitude as the most popular linear clustering method k-means. Jian-Sheng Wu, Wei-Shi Zheng 0001, Jian-Huang Lai, Ching Y. Suen |
IEEE Trans. Big Data | 2 |
| 2018 | Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection MatrixabstractPerson re-identification plays an important role in video surveillance and forensics applications. In many cases, person re-identification needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of the suspect. We call re-identification in this scenario as image to video person reidentification (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature projection matrix and heterogeneous dictionary pair learning (PHDL) approach for IVPR. Specifically, the PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence caused by the variations within each video on the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted by using the coding coefficients. Furthermore, to ensure that the obtained coding coefficients own favorable discriminability, the PHDL designs a point-to-set coefficient discriminant term. To make better use of the complementary spatial-temporal and visual appearance information contained in pedestrian video data, we further propose a multi-view PHDL approach, which can fuse different video information effectively in the dictionary learning process. Experiments on four publicly available person sequence data sets demonstrate the effectiveness of the proposed approaches. Xiaoke Zhu, Xiaoyuan Jing, Xinge You, Wangmeng Zuo, Shiguang Shan, Wei-Shi Zheng 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2018 | Global-Local Temporal Saliency Action PredictionabstractAction prediction on a partially observed action sequence is a very challenging task. To address this challenge, we first design a global-local distance model, where a global-temporal distance compares subsequences as a whole and local-temporal distance focuses on individual segment. Our distance model introduces temporal saliency for each segment to adapt its contribution. Finally, a global-local temporal action prediction model is formulated in order to jointly learn and fuse these two types of distances. Such a prediction model is capable of recognizing action of: 1) an on-going sequence and 2) a sequence with arbitrarily frames missing between the beginning and end (known as gap-filling). Our proposed model is tested and compared with related action prediction models on BIT, UCF11, and HMDB data sets. The results demonstrated the effectiveness of our proposal. In particular, we showed the benefit of our proposed model on predicting unseen action types and the advantage on addressing the gapfilling problem as compared with recently developed action prediction models. Shaofan Lai, Wei-Shi Zheng 0001, Jianfang Hu, Jianguo Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Fast Open-World Person Re-IdentificationabstractExisting person re-identification (re-id) methods typically assume that: 1) any probe person is guaranteed to appear in the gallery target population during deployment (i.e., closed-world) and 2) the probe set contains only a limited number of people (i.e., small search scale). Both assumptions are artificial and breached in real-world applications, since the probe population in target people search can be extremely vast in practice due to the ambiguity of probe search space boundary. Therefore, it is unrealistic that any probe person is assumed as one target people, and a large-scale search in person images is inherently demanded. In this paper, we introduce a new person re-id search setting, called large scale open-world (LSOW) re-id, characterized by huge size probe images and open person population in search thus more close to practical deployments. Under LSOW, the under-studied problem of person re-id efficiency is essential in addition to that of commonly studied re-id accuracy. We, therefore, develop a novel fast person re-id method, called Cross-view Identity Correlation and vErification (X-ICE) hashing, for joint learning of cross-view identity representation binarisation and discrimination in a unified manner. Extensive comparative experiments on three large-scale benchmarks have been conducted to validate the superiority and advantages of the proposed X-ICE method over a wide range of the state-of-the-art hashing models, person re-id methods, and their combinations. Xiatian Zhu, Botong Wu, Dongcheng Huang, Wei-Shi Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Online HashingabstractAlthough hash function learning algorithms have achieved great success in recent years, most existing hash models are off-line, which are not suitable for processing sequential or online data. To address this problem, this paper proposes an online hash model to accommodate data coming in stream for online learning. Specifically, a new loss function is proposed to measure the similarity loss between a pair of data samples in hamming space. Then, a structured hash model is derived and optimized in a passive-aggressive way. Theoretical analysis on the upper bound of the cumulative loss for the proposed online hash model is provided. Furthermore, we extend our online hashing (OH) from a single model to a multimodel OH that trains multiple models so as to retain diverse OH models in order to avoid biased update. The competitive efficiency and effectiveness of the proposed online hash models are verified through extensive experiments on several large-scale data sets as compared with related hashing methods. Long-Kai Huang, Qiang Yang 0010, Wei-Shi Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Learning Heterogeneous Dictionary Pair with Feature Projection Matrix for Pedestrian Video Retrieval via Single Query ImageabstractPerson re-identification (re-id) plays an important role in video surveillance and forensics applications. In many cases, person re-id needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of him. We call re-id in this scenario as image to video person re-id (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature projection matrix and heterogeneous dictionary pair learning (PHDL) approach for IVPR. Specifically, PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence of variations within each video to the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted using coding coefficients. Furthermore, to ensure that the obtained coding coefficients have favorable discriminability, PHDL designs a point-to-set coefficient discriminant term. Experiments on the public iLIDS-VID and PRID 2011 datasets demonstrate the effectiveness of the proposed approach. Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Yunhong Wang 0001, Wangmeng Zuo, Wei-Shi Zheng 0001 |
AAAI | 6 |
| 2017 | Latent embeddings for collective activity recognitionabstractRather than simply recognizing the action of a person individually, collective activity recognition aims to find out what a group of people is acting in a collective scene. Previous state-of-the-art methods using hand-crafted potentials in conventional graphical model which can only define a limited range of relations. Thus, the complex structural dependencies among individuals involved in a collective scenario cannot be fully modeled. In this paper, we overcome these limitations by embedding latent variables into feature space and learning the feature mapping functions in a deep learning framework. The embeddings of latent variables build a global relation containing person-group interactions and richer contextual information by jointly modeling broader range of individuals. Besides, we assemble attention mechanism during embedding for achieving more compact representations. We evaluate our method on three collective activity datasets, where we contribute a much larger dataset in this work. The proposed model has achieved clearly better performance as compared to the state-of-the-art methods in our experiments. Yongyi Tang, Peizhen Zhang, Jianfang Hu, Wei-Shi Zheng 0001 |
AVSS | 4 |
| 2017 | A Matrix Splitting Method for Composite Function MinimizationabstractComposite function minimization captures a wide spectrum of applications in both computer vision and machine learning. It includes bound constrained optimization and cardinality regularized optimization as special cases. This paper proposes and analyzes a new Matrix Splitting Method (MSM) for minimizing composite functions. It can be viewed as a generalization of the classical Gauss-Seidel method and the Successive Over-Relaxation method for solving linear systems in the literature. Incorporating a new Gaussian elimination procedure, the matrix splitting method achieves state-of-the-art performance. For convex problems, we establish the global convergence, convergence rate, and iteration complexity of MSM, while for non-convex problems, we prove its global convergence. Finally, we validate the performance of our matrix splitting method on two particular applications: nonnegative matrix factorization and cardinality regularized sparse coding. Extensive experiments show that our method outperforms existing composite function minimization techniques in term of both efficiency and efficacy. Ganzhao Yuan, Wei-Shi Zheng 0001, Bernard Ghanem |
CVPR | 2 |
| 2017 | RGB-Infrared Cross-Modality Person Re-identificationabstractPerson re-identification (Re-ID) is an important problem in video surveillance, aiming to match pedestrian images across camera views. Currently, most works focus on RGB-based Re-ID. However, in some applications, RGB images are not suitable, e.g. in a dark environment or at night. Infrared (IR) imaging becomes necessary in many visual systems. To that end, matching RGB images with infrared images is required, which are heterogeneous with very different visual characteristics. For person Re-ID, this is a very challenging cross-modality problem that has not been studied so far. In this work, we address the RGB-IR cross-modality Re-ID problem and contribute a new multiple modality Re-ID dataset named SYSU-MM01, including RGB and IR images of 491 identities from 6 cameras, giving in total 287,628 RGB images and 15,792 IR images. To explore the RGB-IR Re-ID problem, we evaluate existing popular cross-domain models, including three commonly used neural network structures (one-stream, two-stream and asymmetric FC layer) and analyse the relation between them. We further propose deep zero-padding for training one-stream network towards automatically evolving domain-specific nodes in the network for cross-modality matching. Our experiments show that RGB-IR cross-modality matching is very challenging but still feasible using the proposed model with deep zero-padding, giving the best performance. Our dataset is available at http:// isee.sysu.edu.cn/project/RGBIRReID.htm. Ancong Wu, Wei-Shi Zheng 0001, Hong-Xing Yu, Shaogang Gong, Jian-Huang Lai |
ICCV | 2 |
| 2017 | Cross-View Asymmetric Metric Learning for Unsupervised Person Re-IdentificationabstractWhile metric learning is important for Person reidentification (RE-ID), a significant problem in visual surveillance for cross-view pedestrian matching, existing metric models for RE-ID are mostly based on supervised learning that requires quantities of labeled samples in all pairs of camera views for training. However, this limits their scalabilities to realistic applications, in which a large amount of data over multiple disjoint camera views is available but not labelled. To overcome the problem, we propose unsupervised asymmetric metric learning for unsupervised RE-ID. Our model aims to learn an asymmetric metric, i.e., specific projection for each view, based on asymmetric clustering on cross-view person images. Our model finds a shared space where view-specific bias is alleviated and thus better matching performance can be achieved. Extensive experiments have been conducted on a baseline and five large-scale RE-ID datasets to demonstrate the effectiveness of the proposed model. Through the comparison, we show that our model works much more suitable for unsupervised RE-ID compared to classical unsupervised metric learning models. We also compare with existing unsupervised REID methods, and our model outperforms them with notable margins. Specifically, we report the results on large-scale unlabelled RE-ID dataset, which is important but unfortunately less concerned in literatures. Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2017 | Correlation Based Identity Filter: An Efficient Framework for Person Search
Wei-Hong Li 0001, Yafang Mao, Ancong Wu, Wei-Shi Zheng 0001 |
ICIG (1) | 4 |
| 2017 | Efficient symmetry-driven fully convolutional network for multimodal brain tumor segmentationabstractIn this paper, we present a novel and efficient method for brain tumor (and sub regions) segmentation in multimodal MR images based on a fully convolutional network (FCN) that enables end-to-end training and fast inference. Our structure consists of a downsampling path and three upsampling paths, which extract multi-level contextual information by concatenating hierarchical feature representation from each upsam-pling path. Meanwhile, we introduce a symmetry-driven FCN by the proposal of using symmetry difference images. The model was evaluated on Brain Tumor Image Segmentation Benchmark (BRATS) 2013 challenge dataset and achieved the state-of-the-art results while the computational cost is less than competitors. Haocheng Shen, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
ICIP | 3 |
| 2017 | Multi-view collaborative locally adaptive clustering with Minkowski metric
Chang-Dong Wang 0001, Dong Huang 0001, Wei-Shi Zheng 0001 |
Expert Syst. Appl. | 4 |
| 2017 | Corrigendum to "How many clusters? A robust PSO-based local density model" [Neurocomputing 207 (2016) 264-275]
Hui-Liang Ling, Jian-Sheng Wu, Yi Zhou 0005, Wei-Shi Zheng 0001 |
Neurocomputing | 4 |
| 2017 | Special issue on 2016 International Conference on Intelligence Science and Big Data Engineering
Wei-Shi Zheng 0001, Xinbo Gao 0001, Yuanqing Li 0003, Jian-Huang Lai |
Neurocomputing | 1 |
| 2017 | Jointly Learning Heterogeneous Features for RGB-D Activity RecognitionabstractIn this paper, we focus on heterogeneous features learning for RGB-D activity recognition. We find that features from different channels (RGB, depth) could share some similar hidden structures, and then propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogeneous multi-task learning. The proposed model formed in a unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to exploit latent shared features across different feature channels, 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces, and 3) transferring feature-specific intermediate transforms (i-transforms) for learning fusion of heterogeneous features across datasets. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by a simple inference model. Extensive experimental results on four activity datasets have demonstrated the efficacy of the proposed method. A new RGB-D activity dataset focusing on human-object interaction is further contributed, which presents more challenges for RGB-D activity benchmarking. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Multi-task mid-level feature learning for micro-expression recognition
Jiachi He, Jianfang Hu, Wei-Shi Zheng 0001 |
Pattern Recognit. | 4 |
| 2017 | Sparse transfer for facial shape-from-shading
Jianfang Hu, Wei-Shi Zheng 0001, Xiaohua Xie, Jian-Huang Lai |
Pattern Recognit. | 2 |
| 2017 | Multiple metric learning based on bar-shape descriptor for person re-identification
Cairong Zhao, Xuekuan Wang, Wai Keung Wong, Wei-Shi Zheng 0001, Jian Yang 0003, Duoqian Miao 0001 |
Pattern Recognit. | 4 |
| 2017 | Illumination invariant single face image recognition under heterogeneous lighting condition
Jun-Yong Zhu, Wei-Shi Zheng 0001, Jian-Huang Lai |
Pattern Recognit. | 2 |
| 2017 | An Asymmetric Distance Model for Cross-View Feature Mapping in Person ReidentificationabstractPerson reidentification, which matches person images of the same identity across nonoverlapping camera views, becomes an important component for cross-camera-view activity analysis. Most (if not all) person reidentification algorithms are designed based on appearance features. However, appearance features are not stable across nonoverlapping camera views under dramatic lighting change, and those algorithms assume that two cross-view images of the same person can be well represented either by exploring robust and invariant features or by learning matching distance. Such an assumption ignores the nature that images are captured under different camera views with different camera characteristics and environments, and thus, mostly there exists large discrepancy between the extracted features under different views. To solve this problem, we formulate an asymmetric distance model for learning camera-specific projections to transform the unmatched features of each view into a common space where discriminative features across view space are extracted. A cross-view consistency regularization is further introduced to model the correlation between view-specific feature transformations of different camera views, which reflects their nature relations and plays a significant role in avoiding overfitting. A kernel cross-view discriminant component analysis is also presented. Extensive experiments have been conducted to show that asymmetric distance modeling is important for person reidentification, which matches the concerns on cross-disjoint-view matching, reporting superior performance compared with related distance learning methods on six publically available data sets. Ying-Cong Chen, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Robust Depth-Based Person Re-IdentificationabstractPerson re-identification (re-id) aims to match people across non-overlapping camera views. So far the RGB-based appearance is widely used in most existing works. However, when people appeared in extreme illumination or changed clothes, the RGB appearance-based re-id methods tended to fail. To overcome this problem, we propose to exploit depth information to provide more invariant body shape and skeleton information regardless of illumination and color change. More specifically, we exploit depth voxel covariance descriptor and further propose a locally rotation invariant depth shape descriptor called Eigen-depth feature to describe pedestrian body shape. We prove that the distance between any two covariance matrices on the Riemannian manifold is equivalent to the Euclidean distance between the corresponding Eigen-depth features. Furthermore, we propose a kernelized implicit feature transfer scheme to estimate Eigen-depth feature implicitly from RGB image when depth information is not available. We find that combining the estimated depth features with RGB-based appearance features can sometimes help to better reduce visual ambiguities of appearance features caused by illumination and similar clothes. The effectiveness of our models was validated on publicly available depth pedestrian datasets as compared to related methods for re-id. Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai |
IEEE Trans. Image Process. | 2 |
| 2017 | Semi-Supervised Multi-View Discrete Hashing for Fast Image SearchabstractHashing is an important method for fast neighbor search on large scale dataset in Hamming space. While most research on hash models are focusing on single-view data, recently the multi-view approaches with a majority of unsupervised multi-view hash models have been considered. Despite of existence of millions of unlabeled data samples, it is believed that labeling a handful of data will remarkably improve the searching performance. In this paper, we propose a semi-supervised multi-view hash model. Besides incorporating a portion of label information into the model, the proposed multi-view model differs from existing multi-view hash models in three-fold: 1) a composite discrete hash learning modeling that is able to minimize the loss jointly on multi-view features when using relaxation on learning hashing codes; 2) exploring statistically uncorrelated multi-view features for generating hash codes; and 3) a composite locality preserving modeling for locally compact coding. Extensive experiments have been conducted to show the effectiveness of the proposed semi-supervised multi-view hash model as compared with related multi-view hash models and semi-supervised hash models. Wei-Shi Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Top-Push Video-Based Person Re-identificationabstractMost existing person re-identification (re-id) models focus on matching still person images across disjoint camera views. Since only limited information can be exploited from still images, it is hard (if not impossible) to overcome the occlusion, pose and camera-view change, and lighting variation problems. In comparison, video-based re-id methods can utilize extra space-time information, which contains much more rich cues for matching to overcome the mentioned problems. However, we find that when using video-based representation, some inter-class difference can be much more obscure than the one when using still-image-based representation, because different people could not only have similar appearance but also have similar motions and actions which are hard to align. To solve this problem, we propose a top-push distance learning model (TDL), in which we integrate a top-push constrain for matching video features of persons. The top-push constraint enforces the optimization on top-rank matching in re-id, so as to make the matching model more effective towards selecting more discriminative features to distinguish different persons. Our experiments show that the proposed video-based reid framework outperforms the state-of-the-art video-based re-id methods. Jinjie You, Ancong Wu, Xiang Li 0032, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2016 | Real-Time RGB-D Activity Prediction by Soft Regression
Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai |
ECCV (1) | 2 |
| 2016 | Embedding Deep Metric for Person Re-identification: A Study Against Large Variations
Hailin Shi, Yang Yang 0062, Xiangyu Zhu 0001, Shengcai Liao, Zhen Lei 0001, Wei-Shi Zheng 0001, Stan Z. Li |
ECCV (1) | 6 |
| 2016 | Distance learning by treating negative samples differently and exploiting impostors with symmetric triplet constraint for person re-identificationabstractDistance learning (DL) is an effective technique for person reidentification (PR-ID). DL based methods learn the distance metric by exploiting the discriminative information contained in samples. In PR-ID, different types of negative samples own different amounts of discriminative information, and impostor samples usually own more than other well separable negative samples (WSN-samples). Therefore, how to make full use of the different discriminative information conveyed by all negative samples in the DL process is a critical issue to be investigated. In this paper, we propose a novel DL approach for PR-ID. Specifically, for each target sample, we divide its negative samples into impostors and WSN-samples. Then we learn the distance metric by utilizing impostors and WSN-samples differently. For impostors, we design a symmetric triplet constraint, which requires the impostor to be far away from both samples of its corresponding positive sample pair simultaneously; for WSN-samples, we require them to keep their favorable separability. Experimental results on three benchmark datasets demonstrate the effectiveness and efficiency of our approach. Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Wei-Shi Zheng 0001, Ruimin Hu, Chunxia Xiao, Chao Liang 0001 |
ICME | 4 |
| 2016 | HEp-2 specimen classification via deep CNNs and pattern histogramabstractAutomatic classification of Human Epithelial Type-2 (HEp-2) specimen patterns is an important yet challenging problem in medical image analysis. Most prior works have primarily focused on cells images classification problem which is one of the early essential steps in the system pipeline, while less attention has been paid to the classification of whole-specimen ones. In this work, a specimen pattern recognition system combining convolutional neural networks (CNNs) and pattern histogram was proposed. The pattern histograms were obtained based on the prediction of each single cell inside the specimens. Two strategies were designed to predicted the pattern of a whole specimen: 1) the most dominant cell pattern in pattern histogram was represented as the specimen pattern, 2) the pattern histograms were employed as bags of patterns and then were trained and predicted separately by a SVM classifier. Experimental results show that the proposed system is effective and achieves high classification accuracy on public benchmark datasets. We further evaluate the robustness of the proposed framework by testing trained CNNs on another different dataset, demonstrating that the system is robust to inter-lab data. Hongwei Li 0004, Wei-Shi Zheng 0001, Xiaohua Xie, Jianguo Zhang 0001 |
ICPR | 3 |
| 2016 | Facial skin beautification via sparse representation over learned layer dictionaryabstractIn this paper, we propose a facial skin beautification framework to remove facial spots based on layer dictionary learning and sparse representation. More precisely, we first decompose the face image into three layers: lighting layer, detail layer and color layer. The corresponding detail layer dictionary are learned by using 60 thousands beauty images collected from the Internet. Thereafter, the detail layer of the image is reconstructed by using sparse representation. Moreover, a binary mask obtained from the learned layer is used to transform detail information from original detail layer to the learned one. The experiment results demonstrate that the proposed method is more effective in eliminating moles, flaws and wrinkles in face image compared with representative commercial systems like PicTreat, Portrait+, Portraitrue and MeituPic. Xiaobin Chang, Xiaohua Xie, Jianfang Hu, Wei-Shi Zheng 0001 |
IJCNN | 5 |
| 2016 | Sketch metric learningabstractThe main theme of this paper is to develop a systematic framework to learn a Mahalanobis distance metric based on matrix sketching. Within this framework, we present a novel sketch metric learning algorithm which sequentially sketches the received samples from training dataset and formulates a new kind of constraint for metric learning. This is in contrast to the traditional constraints that are only consisted of data from training dataset. In this paper, one training instance in the constraint is replaced by a pseudo center, which is generated during the sketching stage. Due to this change, our learning algorithm can focus on pushing every received sample to its corresponding similar pseudo center closer and pulling it far away from the dissimilar one. In addition, it can further achieve better performance of some kinds of time-varying process (e.g. on object tracking) than the compared related competitors. We demonstrate how to implement other methods in our algorithm framework and experiment to show that our method outperforms the competitors and relevant baselines on multiple datasets. Yuting Mai, Wei-Hong Li 0001, Yongyi Tang, Xixi Bi, Wei-Shi Zheng 0001 |
IJCNN | 5 |
| 2016 | Online visual tracking via correlation filter with convolutional networksabstractRobust online visual tracking is a challenging task of the computer vision due to its violent variation within the video sequences. To approach these issues, deep networks have been applied in order to improve accuracy and correlation filter based trackers perform excellent efficiency and adaptation to scale. In this paper, we present a novel method with convolutional networks and correlation filter. A simple two-layer convolutional network is constructed to learn robust representations, which encode the inner geometric layout and local structural information, and the tracking framework resorts to learning discriminative correlation filters based on them. For our method satisfies both veracity and efficiency by finding a compromise between two theories, it performs favorably against several state-of-the-art methods with 50 public challenging videos. Juan Zha, Chang-Dong Wang 0001, Wei-Shi Zheng 0001 |
VCIP | 5 |
| 2016 | An enhanced deep feature representation for person re-identificationabstractFeature representation and metric learning are two critical components in person re-identification models. In this paper, we focus on the feature representation and claim that hand-crafted histogram features can be complementary to Convolutional Neural Network (CNN) features. We propose a novel feature extraction model called Feature Fusion Net (FFN) for pedestrian image representation. In FFN, back propagation makes CNN features constrained by the handcrafted features. Utilizing color histogram features (RGB, HSV, YCbCr, Lab and YIQ) and texture features (multi-scale and multi-orientation Gabor features), we get a new deep feature representation that is more discriminative and compact. Experiments on three challenging datasets (VIPeR, CUHK01, PRID450s) validates the effectiveness of our proposal. Shangxuan Wu, Ying-Cong Chen, Xiang Li 0032, Ancong Wu, Jinjie You, Wei-Shi Zheng 0001 |
WACV | 6 |
| 2016 | Learning object-specific DAGs for multi-label material recognition
Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001 |
Comput. Vis. Image Underst. | 3 |
| 2016 | How many clusters? A robust PSO-based local density model
Hui-Liang Ling, Jian-Sheng Wu, Yi Zhou 0005, Wei-Shi Zheng 0001 |
Neurocomputing | 4 |
| 2016 | Towards Open-World Person Re-Identification by One-Shot Group-Based VerificationabstractSolving the problem of matching people across non-overlapping multi-camera views, known as person re-identification (re-id), has received increasing interests in computer vision. In a real-world application scenario, a watch-list (gallery set) of a handful of known target people are provided with very few (in many cases only a single) image(s) (shots) per target. Existing re-id methods are largely unsuitable to address this open-world re-id challenge because they are designed for (1) a closed-world scenario where the gallery and probe sets are assumed to contain exactly the same people, (2) person-wise identification whereby the model attempts to verify exhaustively against each individual in the gallery set, and (3) learning a matching model using multi-shots. In this paper, a novel transfer local relative distance comparison (t-LRDC) model is formulated to address the open-world person re-identification problem by one-shot group-based verification. The model is designed to mine and transfer useful information from a labelled open-world non-target dataset. Extensive experiments demonstrate that the proposed approach outperforms both non-transfer learning and existing transfer learning based re-id methods. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | One-pass online learning: A local approach
Zhaoze Zhou, Wei-Shi Zheng 0001, Jianfang Hu, Yong Xu 0001, Jane You |
Pattern Recognit. | 2 |
| 2016 | Exemplar-Based Recognition of Human-Object InteractionsabstractHuman action can be recognized from a single still image by modeling human-object interactions (HOIs), which infers the mutual spatial structure information between human and the manipulated object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object and estimation of human pose; they are thus sensitive to large variations of human poses, occlusion, and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar-based approach is proposed in this paper. Our approach learns a set of spatial pose-object interaction exemplars, which are probabilistic density functions describing spatially how a person is interacting with a manipulated object for different activities. Specifically, a new framework consisting of an exemplar-based HOI descriptor and an associated matching model is formulated for robust human action recognition in still images. In addition, the framework is extended to perform HOI recognition in videos, where the proposed exemplar representation is used for implicit frame selection to negate irrelevant or noisy frames by temporal structured HOI modeling. Extensive experiments are carried out on two image action datasets and two video action datasets. The results demonstrate the effectiveness of our proposed methods and show that our approach is able to achieve state-of-the-art performance, compared with several recently proposed competitors. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Cross-Scenario Transfer Person ReidentificationabstractPerson reidentification (Re-ID) matches images of the same person captured in disjoint camera views and at different times. To obtain a reliable similarity measurement between images, manually annotating a large amount of pairwise cross-camera-view person images is deemed necessary. However, this kind of annotation is both costly and impractical for efficiently deploying a Re-ID system to a completely new scenario, a new setting of nonoverlapping camera views between which person images are to be matched. To solve this problem, we consider utilizing other existing person images captured in other scenarios to help the Re-ID system in a target (new) scenario, provided that a few samples are captured under the new scenario. More specifically, we tackle this problem by jointly learning the similarity measurements for Re-ID in different scenarios in an asymmetric way. To model the joint learning, we consider that the Re-ID models share certain component across tasks. A distinct consideration in our multitask modeling is to extract the discriminant shared component that reduces the cross-task data overlap in the shared latent space during the joint learning, so as to enhance the target inter-class separation in the shared latent space. For this purpose, we propose to maximize the cross-task data discrepancy on the shared component during asymmetric multitask learning (MTL), along with maximizing the local inter-class variation and minimizing local intra-class variation on all tasks. We call our proposed method the constrained asymmetric multitask discriminant component analysis (cAMT-DCA). We show that cAMT-DCA can be solved by a simple eigen decomposition with a closed form, getting rid of any iterative learning used in most conventional MTL analyses. The experimental results show that the proposed transfer model gains a clear improvement against the related nontransfer and general multitask person Re-ID models. Wei-Shi Zheng 0001, Xiang Li 0032, Jianguo Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Jointly learning heterogeneous features for RGB-D activity recognitionabstractIn this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogenous multi-task learning. The proposed model in an unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to enable the multi-task classifier learning, and 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by two inference models. Extensive results on three activity datasets have demonstrated the efficacy of the proposed method. In addition, a novel RGB-D activity dataset focusing on human-object interaction is collected for evaluating the proposed method, which will be made available to the community for RGB-D activity benchmarking and analysis. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
CVPR | 2 |
| 2015 | Multi-Scale Learning for Low-Resolution Person Re-IdentificationabstractIn real world person re-identification (re-id), images of people captured at very different resolutions from different locations need be matched. Existing re-id models typically normalise all person images to the same size. However, a low-resolution (LR) image contains much less information about a person, and direct image scaling and simple size normalisation as done in conventional re-id methods cannot compensate for the loss of information. To solve this LR person re-id problem, we propose a novel joint multi-scale learning framework, termed joint multi-scale discriminant component analysis (JUDEA). The key component of this framework is a heterogeneous class mean discrepancy (HCMD) criterion for cross-scale image domain alignment, which is optimised simultaneously with discriminant modelling across multiple scales in the joint learning framework. Our experiments show that the proposed JUDEA framework outperforms existing representative re-id methods as well as other related LR visual matching models applied for the LR person re-id problem. Xiang Li 0032, Wei-Shi Zheng 0001, Tao Xiang 0002, Shaogang Gong |
ICCV | 2 |
| 2015 | Partial Person Re-IdentificationabstractWe address a new partial person re-identification (re-id) problem, where only a partial observation of a person is available for matching across different non-overlapping camera views. This differs significantly from the conventional person re-id setting where it is assumed that the full body of a person is detected and aligned. To solve this more challenging and realistic re-id problem without the implicit assumption of manual body-parts alignment, we propose a matching framework consisting of 1) a local patch-level matching model based on a novel sparse representation classification formulation with explicit patch ambiguity modelling, and 2) a global part-based matching model providing complementary spatial layout information. Our framework is evaluated on a new partial person re-id dataset as well as two existing datasets modified to include partial person images. The results show that the proposed method outperforms significantly existing re-id methods as well as other partial visual matching methods. Wei-Shi Zheng 0001, Xiang Li 0032, Tao Xiang 0002, Shengcai Liao, Jian-Huang Lai, Shaogang Gong |
ICCV | 1 |
| 2015 | Mirror Representation for Modeling View-Specific Transform in Person Re-Identification
Ying-Cong Chen, Wei-Shi Zheng 0001, Jian-Huang Lai |
IJCAI | 2 |
| 2015 | Quantized Correlation Hashing for Fast Cross-Modal Search
Botong Wu, Qiang Yang 0010, Wei-Shi Zheng 0001, Yizhou Wang 0001, Jingdong Wang 0001 |
IJCAI | 3 |
| 2015 | Approximate kernel competitive learning
Jian-Sheng Wu, Wei-Shi Zheng 0001, Jian-Huang Lai |
Neural Networks | 2 |
| 2015 | Learning Person-Person Interaction in Collective Activity RecognitionabstractCollective activity is a collection of atomic activities (individual person's activity) and can hardly be distinguished by an atomic activity in isolation. The interactions among people are important cues for recognizing collective activity. In this paper, we concentrate on modeling the person-person interactions for collective activity recognition. Rather than relying on hand-craft description of the person-person interaction, we propose a novel learning-based approach that is capable of computing the class-specific person-person interaction patterns. In particular, we model each class of collective activity by an interaction matrix, which is designed to measure the connection between any pair of atomic activities in a collective activity instance. We then formulate an interaction response (IR) model by assembling all these measurements and make the IR class specific and distinct from each other. A multitask IR is further proposed to jointly learn different person-person interaction patterns simultaneously in order to learn the relation between different person-person interactions and keep more distinct activity-specific factor for each interaction at the same time. Our model is able to exploit discriminative low-rank representation of person-person interaction. Experimental results on two challenging data sets demonstrate our proposed model is comparable with the state-of-the-art models and show that learning person-person interactions plays a critical role in collective activity recognition. Xiaobin Chang, Wei-Shi Zheng 0001, Jianguo Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Half-Quadratic-Based Iterative Minimization for Robust Sparse RepresentationabstractRobust sparse representation has shown significant potential in solving challenging problems in computer vision such as biometrics and visual surveillance. Although several robust sparse models have been proposed and promising results have been obtained, they are either for error correction or for error detection, and learning a general framework that systematically unifies these two aspects and explores their relation is still an open problem. In this paper, we develop a half-quadratic (HQ) framework to solve the robust sparse representation problem. By defining different kinds of half-quadratic functions, the proposed HQ framework is applicable to performing both error correction and error detection. More specifically, by using the additive form of HQ, we propose an ℓ1-regularized error correction method by iteratively recovering corrupted data from errors incurred by noises and outliers; by using the multiplicative form of HQ, we propose an ℓ1-regularized error detection method by learning from uncorrupted data iteratively. We also show that the ℓ1-regularization solved by soft-thresholding function has a dual relationship to Huber M-estimator, which theoretically guarantees the performance of robust sparse representation in terms of M-estimation. Experiments on robust face recognition under severe occlusion and corruption validate our framework and findings. Ran He 0001, Wei-Shi Zheng 0001, Tieniu Tan, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Matching NIR Face to VIS Face Using TransductionabstractVisual versus near infrared (VIS-NIR) face image matching uses an NIR face image as the probe and conventional VIS face images as enrollment. It takes advantage of the NIR face technology in tackling illumination changes and low-light condition and can cater for more applications where the enrollment is done using VIS face images such as ID card photos. Existing VIS-NIR techniques assume that during classifier learning, the VIS images of each target people have their NIR counterparts. However, since corresponding VIS-NIR image pairs of the same people are not always available, which is often the case, so those methods cannot be applied. To address this problem, we propose a transductive method named transductive heterogeneous face matching (THFM) to adapt the VIS-NIR matching learned from training with available image pairs to all people in the target set. In addition, we propose a simple feature representation for effective VIS-NIR matching, which can be computed in three steps, namely Log-DoG filtering, local encoding, and uniform feature normalization, to reduce heterogeneities between VIS and NIR images. The transduction approach can reduce the domain difference due to heterogeneous data and learn the discriminative model for target people simultaneously. To the best of our knowledge, it is the first attempt to formulate the VIS-NIR matching using transduction to address the generalization problem for matching. Experimental results validate the effectiveness of our proposed method on the heterogeneous face biometric databases. Jun-Yong Zhu, Wei-Shi Zheng 0001, Jian-Huang Lai, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Recognising Human-Object Interaction via Exemplar Based ModellingabstractHuman action can be recognised from a single still image by modelling Human-object interaction (HOI), which infers the mutual spatial structure information between human and object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object, and estimation of human pose. They are thus sensitive to large variations of human poses, occlusion and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar based approach is proposed in this work. Our approach learns a set of spatial pose-object interaction exemplars, which are density functions describing how a person is interacting with a manipulated object for different activities spatially in a probabilistic way. A representation based on our HOI exemplar thus has great potential for being robust to the errors in human/object detection and pose estimation. A new framework consists of a proposed exemplar based HOI descriptor and an activity specific matching model that learns the parameters is formulated for robust human activity recognition. Experiments on two benchmark activity datasets demonstrate that the proposed approach obtains state-of-the-art performance. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002 |
ICCV | 2 |
| 2013 | Human Re-identification by Matching Compositional Template with Cluster SamplingabstractThis paper aims at a newly raising task in visual surveillance: re-identifying people at a distance by matching body information, given several reference examples. Most of existing works solve this task by matching a reference template with the target individual, but often suffer from large human appearance variability (e.g. different poses/views, illumination) and high false positives in matching caused by conjunctions, occlusions or surrounding clutters. Addressing these problems, we construct a simple yet expressive template from a few reference images of a certain individual, which represents the body as an articulated assembly of compositional and alternative parts, and propose an effective matching algorithm with cluster sampling. This algorithm is designed within a candidacy graph whose vertices are matching candidates (i.e. a pair of source and target body parts), and iterates in two steps for convergence. (i) It generates possible partial matches based on compatible and competitive relations among body parts. (ii) It confirms the partial matches to generate a new matching solution, which is accepted by the Markov Chain Monte Carlo (MCMC) mechanism. In the experiments, we demonstrate the superior performance of our approach on three public databases compared to existing methods. Yuanlu Xu, Liang Lin 0004, Wei-Shi Zheng 0001, Xiaobai Liu |
ICCV | 3 |
| 2013 | Online Hashing
Long-Kai Huang, Qiang Yang 0010, Wei-Shi Zheng 0001 |
IJCAI | 3 |
| 2013 | Euler Clustering
Jian-Sheng Wu, Wei-Shi Zheng 0001, Jian-Huang Lai |
IJCAI | 2 |
| 2013 | Smart Hashing Update for Fast Response
Qiang Yang 0010, Long-Kai Huang, Wei-Shi Zheng 0001, Yingbiao Ling |
IJCAI | 3 |
| 2013 | Learning from Partially Annotated OPT Images by Contextual Relevance Ranking
Wenqi Li 0001, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Maria Coats, Frank A. Carey, Stephen J. McKenna |
MICCAI (3) | 3 |
| 2013 | Robust spectral regression for face recognition
Yanqing Guo, Ran He 0001, Wei-Shi Zheng 0001, Xiangwei Kong 0001, Zhaofeng He 0001 |
Neurocomputing | 3 |
| 2013 | A fast convex conjugated algorithm for sparse recovery
Ran He 0001, Xiao-Tong Yuan, Wei-Shi Zheng 0001 |
Neurocomputing | 3 |
| 2013 | Discriminant subspace learning constrained by locally statistical uncorrelation for face recognition
Wei-Shi Zheng 0001, Xiaohong Xu, Jian-Huang Lai |
Neural Networks | 2 |
| 2013 | Reidentification by Relative Distance ComparisonabstractMatching people across nonoverlapping camera views at different locations and different times, known as person reidentification, is both a hard and important problem for associating behavior of people observed in a large distributed space over a prolonged period of time. Person reidentification is fundamentally challenging because of the large visual appearance changes caused by variations in view angle, lighting, background clutter, and occlusion. To address these challenges, most previous approaches aim to model and extract distinctive and reliable visual features. However, seeking an optimal and robust similarity measure that quantifies a wide range of features against realistic viewing conditions from a distance is still an open and unsolved problem for person reidentification. In this paper, we formulate person reidentification as a relative distance comparison (RDC) learning problem in order to learn the optimal similarity measure between a pair of person images. This approach avoids treating all features indiscriminately and does not assume the existence of some universally distinctive and reliable features. To that end, a novel relative distance comparison model is introduced. The model is formulated to maximize the likelihood of a pair of true matches having a relatively smaller distance than that of a wrong match pair in a soft discriminant manner. Moreover, in order to maintain the tractability of the model in large scale learning, we further develop an ensemble RDC model. Extensive experiments on three publicly available benchmarking datasets are carried out to demonstrate the clear superiority of the proposed RDC models over related popular person reidentification techniques. The results also show that the new RDC models are more robust against visual appearance changes and less susceptible to model overfitting compared to other related existing models. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | SVStream: A Support Vector-Based Algorithm for Clustering Data StreamsabstractIn this paper, we propose a novel data stream clustering algorithm, termed SVStream, which is based on support vector domain description and support vector clustering. In the proposed algorithm, the data elements of a stream are mapped into a kernel space, and the support vectors are used as the summary information of the historical elements to construct cluster boundaries of arbitrary shape. To adapt to both dramatic and gradual changes, multiple spheres are dynamically maintained, each describing the corresponding data domain presented in the data stream. By allowing for bounded support vectors (BSVs), the proposed SVStream algorithm is capable of identifying overlapping clusters. A BSV decaying mechanism is designed to automatically detect and remove outliers (noise). We perform experiments over synthetic and real data streams, with the overlapping, evolving, and noise situations taken into consideration. Comparison results with state-of-the-art data stream clustering methods demonstrate the effectiveness and efficiency of the proposed method. Chang-Dong Wang 0001, Jian-Huang Lai, Dong Huang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Two-Stage Nonnegative Sparse Representation for Large-Scale Face RecognitionabstractThis paper proposes a novel nonnegative sparse representation approach, called two-stage sparse representation (TSR), for robust face recognition on a large-scale database. Based on the divide and conquer strategy, TSR decomposes the procedure of robust face recognition into outlier detection stage and recognition stage. In the first stage, we propose a general multisubspace framework to learn a robust metric in which noise and outliers in image pixels are detected. Potential loss functions, including L1 , L2,1, and correntropy are studied. In the second stage, based on the learned metric and collaborative representation, we propose an efficient nonnegative sparse representation algorithm to find an approximation solution of sparse representation. According to the L1 ball theory in sparse representation, the approximated solution is unique and can be optimized efficiently. Then a filtering strategy is developed to avoid the computation of the sparse representation on the whole large-scale dataset. Moreover, theoretical analysis also gives the necessary condition for nonnegative least squares technique to find a sparse solution. Extensive experiments on several public databases have demonstrated that the proposed TSR approach, in general, achieves better classification accuracy than the state-of-the-art sparse representation methods. More importantly, a significant reduction of computational costs is reached in comparison with sparse representation classifier; this enables the TSR to be more suitable for robust face recognition on a large-scale dataset. Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | l2, 1 Regularized correntropy for robust feature selectionabstractIn this paper, we study the problem of robust feature extraction based on l2,1regularized correntropy in both theoretical and algorithmic manner. In theoretical part, we point out that an l2,1-norm minimization can be justified from the viewpoint of half-quadratic (HQ) optimization, which facilitates convergence study and algorithmic development. In particular, a general formulation is accordingly proposed to unify l1-norm and l2,1-norm minimization within a common framework. In algorithmic part, we propose an l2,1regularized correntropy algorithm to extract informative features meanwhile to remove outliers from training data. A new alternate minimization algorithm is also developed to optimize the non-convex correntropy objective. In terms of face recognition, we apply the proposed method to obtain an appearance-based model, called Sparse-Fisherfaces. Extensive experiments show that our method can select robust and sparse features, and outperforms several state-of-the-art subspace methods on largescale and open face recognition datasets. Ran He 0001, Tieniu Tan, Liang Wang 0001, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2012 | Transfer re-identification: From person to set-based verificationabstractSolving the person re-identification problem has become important for understanding people's behaviours in a multicamera network of non-overlapping views. In this work, we address the problem of re-identification from a set-based verification perspective. More specifically, we have a small set of target people on a watch list (a set) and we aim to verify whether a query image of a person is on this watch list. This differs from the existing person re-identification problem in that the probe is verified against a small set of known people but requires much higher degree of verification accuracy with very limited sampling data for each candidate in the set. That is, rather than recognising everybody in the scene, we consider identifying a small set of target people against non-target people when there is only a limited number of target training samples and a large number of unlabelled (unknown) non-target samples available. To this end, we formulate a transfer learning framework for mining discriminant information from non-target people data to solve the watch list set verification problem. Based on the proposed approach, we introduce the concepts of multi-shot and one-shot verifications. We also design new criteria for evaluating the performance of the proposed transfer learning method against the i-LIDS and ETHZ data sets. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
CVPR | 1 |
| 2012 | A Novel Approach for Stroke Extraction of Off-Line Chinese Handwritten Characters Based on Optimum PathsabstractIn recognition of Off-line handwritten characters and signatures, stroke extraction is often a crucial step. Given the large number of Chinese handwritten characters, pattern matching based on structural decomposition and analysis is useful and essential to Off-line Chinese recognition to reduce ambiguity. Two challenging problems for stroke extraction are: 1) how to extract primary strokes and 2) how to solve the segmentation ambiguities at intersection points. In this paper, we introduce a novel approach based on Optimum Paths(AOP) to solve this problem. Optimum Paths(AOP) are derived from the degree information and continuation property, we use them to tackle these two problems. Compared with other methods, the proposed approach has extracted strokes from Off-line Chinese handwritten characters with better performance. Jun Tan 0001, Jian-Huang Lai, Wei-Shi Zheng 0001, Ching Y. Suen |
ICFHR | 3 |
| 2012 | Transductive VIS-NIR face matchingabstractThe visual-near infrared (VIS-NIR) face matching, sharing the illumination-invariant property of NIR face image and remaining the use of existing VIS face images as enrollment, has been a popular issue in recent years. However, existing techniques assume that there are sufficient pairwise VIS and NIR images for each person during training, which is not realistic in VIS-NIR matching problem, as no NIR images are available for people who have already been registered in the existing face recognition system and only a handful of pairwise VIS and NIR face images captured from new people are available. To address this problem, we formulate the VIS-NIR matching as a transductive learning problem, which is a first attempt to our best knowledge. Moreover, we propose a transductive method named Transductive Heterogeneous Face Matching (THFM) by alleviating the domains difference and learning the discriminative model for target simultaneously, making it possible to take the query/probe NIR images into account in a transductive way. Experimental results validate the effectiveness of our approach on the heterogeneous face biometric database. Jun-Yong Zhu, Wei-Shi Zheng 0001, Jian-Huang Lai |
ICIP | 2 |
| 2012 | A facial sparse descriptor for single image based face recognition
Na Liu 0017, Jian-Huang Lai, Wei-Shi Zheng 0001 |
Neurocomputing | 3 |
| 2012 | Radical Extraction Using Affine Sparse Matrix Factorization for Printed Chinese Characters RecognitionabstractEach Chinese character is comprised of radicals, where a single character (compound character) contains one (or more than one) radicals. For human cognitive perspective, a Chinese character can be recognized by identifying its radicals and their spatial relationship. This human cognitive law may be followed in computer recognition. However, extracting Chinese character radicals automatically by computer is still an unsolved problem. In this paper, we propose using an improved sparse matrix factorization which integrates affine transformation, namely affine sparse matrix factorization (ASMF), for automatically extracting radicals from Chinese characters. Here the affine transformation is vitally important because it can address the poor-alignment problem of characters that may be caused by internal diversity of radicals and image segmentation. Consequently we develop a radical-based Chinese character recognition model. Because the number of radicals is much less than the number of Chinese characters, the radical-based recognition performs a far smaller category classification than the whole character-based recognition, resulting in a more robust recognition system. The experiments on standard Chinese character datasets show that the proposed method gets higher recognition rates than related Chinese character recognition methods. Jun Tan 0001, Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | Robust large margin discriminant tangent analysis for face recognition
Nanhai Yang, Ran He 0001, Wei-Shi Zheng 0001, Xiukun Wang |
Neural Comput. Appl. | 3 |
| 2012 | Quantifying and Transferring Contextual Information in Object DetectionabstractContext is critical for reducing the uncertainty in object detection. However, context modeling is challenging because there are often many different types of contextual information coexisting with different degrees of relevance to the detection of target object(s) in different images. It is therefore crucial to devise a context model to automatically quantify and select the most effective contextual information for assisting in detecting the target object. Nevertheless, the diversity of contextual information means that learning a robust context model requires a larger training set than learning the target object appearance model, which may not be available in practice. In this work, a novel context modeling framework is proposed without the need for any prior scene segmentation or context annotation. We formulate a polar geometric context descriptor for representing multiple types of contextual information. In order to quantify context, we propose a new maximum margin context (MMC) model to evaluate and measure the usefulness of contextual information directly and explicitly through a discriminant context inference method. Furthermore, to address the problem of context learning with limited data, we exploit the idea of transfer learning based on the observation that although two categories of objects can have very different visual appearance, there can be similarity in their context and/or the way contextual information helps to distinguish target objects from nontarget objects. To that end, two novel context transfer learning models are proposed which utilize training samples from source object classes to improve the learning of the context model for a target object class based on a joint maximum margin learning framework. Experiments are carried out on PASCAL VOC2005 and VOC2007 data sets, a luggage detection data set extracted from the i-LIDS data set, and a vehicle detection data set extracted from outdoor surveillance footage. Our results validate the effectiveness of the proposed models for quantifying and transferring contextual information, and demonstrate that they outperform related alternative context models. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Extracting non-negative basis images using pixel dispersion penalty
Wei-Shi Zheng 0001, Jian-Huang Lai, Shengcai Liao, Ran He 0001 |
Pattern Recognit. | 1 |
| 2011 | Recovery of corrupted low-rank matrices via half-quadratic based nonconvex minimizationabstractRecovering arbitrarily corrupted low-rank matrices arises in computer vision applications, including bioinformatic data analysis and visual tracking. The methods used involve minimizing a combination of nuclear norm and l1norm. We show that by replacing the l1norm on error items with nonconvex M-estimators, exact recovery of densely corrupted low-rank matrices is possible. The robustness of the proposed method is guaranteed by the M-estimator theory. The multiplicative form of half-quadratic optimization is used to simplify the nonconvex optimization problem so that it can be efficiently solved by iterative regularization scheme. Simulation results corroborate our claims and demonstrate the efficiency of our proposed method under tough conditions. Ran He 0001, Zhenan Sun, Tieniu Tan, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2011 | Nonnegative sparse coding for discriminative semi-supervised learningabstractAn informative and discriminative graph plays an important role in the graph-based semi-supervised learning methods. This paper introduces a nonnegative sparse algorithm and its approximated algorithm based on the l0-l1equivalence theory to compute the nonnegative sparse weights of a graph. Hence, the sparse probability graph (SPG) is termed for representing the proposed method. The nonnegative sparse weights in the graph naturally serve as clustering indicators, benefiting for semi-supervised learning. More important, our approximation algorithm speeds up the computation of the nonnegative sparse coding, which is still a bottle-neck for any previous attempts of sparse non-negative graph learning. And it is much more efficient than using l1-norm sparsity technique for learning large scale sparse graph. Finally, for discriminative semi-supervised learning, an adaptive label propagation algorithm is also proposed to iteratively predict the labels of data on the SPG. Promising experimental results show that the nonnegative sparse coding is efficient and effective for discriminative semi-supervised learning. Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001 |
CVPR | 2 |
| 2011 | Person re-identification by probabilistic relative distance comparisonabstractMatching people across non-overlapping camera views, known as person re-identification, is challenging due to the lack of spatial and temporal constraints and large visual appearance changes caused by variations in view angle, lighting, background clutter and occlusion. To address these challenges, most previous approaches aim to extract visual features that are both distinctive and stable under appearance changes. However, most visual features and their combinations under realistic conditions are neither stable nor distinctive thus should not be used indiscriminately. In this paper, we propose to formulate person re-identification as a distance learning problem, which aims to learn the optimal distance that can maximises matching accuracy regardless the choice of representation. To that end, we introduce a novel Probabilistic Relative Distance Comparison (PRDC) model, which differs from most existing distance learning methods in that, rather than minimising intra-class variation whilst maximising intra-class variation, it aims to maximise the probability of a pair of true match having a smaller distance than that of a wrong match pair. This makes our model more tolerant to appearance changes and less susceptible to model over-fitting. Extensive experiments are carried out to demonstrate that 1) by formulating the person re-identification problem as a distance learning problem, notable improvement on matching accuracy can be obtained against conventional person re-identification techniques, which is particularly significant when the training sample size is small; and 2) our PRDC outperforms not only existing distance learning methods but also alternative learning methods based on boosting and learning to rank. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
CVPR | 1 |
| 2011 | A Regularized Correntropy Framework for Robust Pattern RecognitionabstractThis letter proposes a new multiple linear regression model using regularized correntropy for robust pattern recognition. First, we motivate the use of correntropy to improve the robustness of the classical mean square error (MSE) criterion that is sensitive to outliers. Then an l1regularization scheme is imposed on the correntropy to learn robust and sparse representations. Based on the half-quadratic optimization technique, we propose a novel algorithm to solve the nonlinear optimization problem. Second, we develop a new correntropy-based classifier based on the learned regularization scheme for robust object recognition. Extensive experiments over several applications confirm that the correntropy-based l1regularization can improve recognition accuracy and receiver operator characteristic curves under noise corruption and occlusion. Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001 |
Neural Comput. | 2 |
| 2011 | Maximum Correntropy Criterion for Robust Face RecognitionabstractIn this paper, we present a sparse correntropy framework for computing robust sparse representations of face images for recognition. Compared with the state-of-the-art l(1)norm-based sparse representation classifier (SRC), which assumes that noise also has a sparse representation, our sparse algorithm is developed based on the maximum correntropy criterion, which is much more insensitive to outliers. In order to develop a more tractable and practical approach, we in particular impose nonnegativity constraint on the variables in the maximum correntropy criterion and develop a half-quadratic optimization technique to approximately maximize the objective function in an alternating way so that the complex optimization problem is reduced to learning a sparse representation through a weighted linear least squares problem with nonnegativity constraint at each iteration. Our extensive experiments demonstrate that the proposed method is more robust and efficient in dealing with the occlusion and corruption problems in face recognition as compared to the related state-of-the-art methods. In particular, it shows that the proposed method can improve both recognition accuracy and receiver operator characteristic (ROC) curves, while the computational cost is much lower than the SRC algorithms. Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Spatial-temporal consistent labeling of tracked pedestrians across non-overlapping camera views
Guoyun Lian, Jian-Huang Lai, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2011 | Non-ideal class non-point light source quotient image for face relighting
Xiaohua Xie, Jian-Huang Lai, Ching Y. Suen, Wei-Shi Zheng 0001 |
Signal Process. | 4 |
| 2011 | Robust Principal Component Analysis Based on Maximum Correntropy CriterionabstractPrincipal component analysis (PCA) minimizes the mean square error (MSE) and is sensitive to outliers. In this paper, we present a new rotational-invariant PCA based on maximum correntropy criterion (MCC). A half-quadratic optimization algorithm is adopted to compute the correntropy objective. At each iteration, the complex optimization problem is reduced to a quadratic problem that can be efficiently solved by a standard optimization method. The proposed method exhibits the following benefits: 1) it is robust to outliers through the mechanism of MCC which can be more theoretically solid than a heuristic rule based on MSE; 2) it requires no assumption about the zero-mean of data for processing and can estimate data mean during optimization; and 3) its optimal solution consists of principal eigenvectors of a robust covariance matrix corresponding to the largest eigenvalues. In addition, kernel techniques are further introduced in the proposed method to deal with nonlinearly distributed data. Numerical results demonstrate that the proposed method can outperform robust rotational-invariant PCAs based on L(1) norm when outliers occur. Ran He 0001, Bao-Gang Hu, Wei-Shi Zheng 0001, Xiangwei Kong 0001 |
IEEE Trans. Image Process. | 3 |
| 2011 | Normalization of Face Illumination Based on Large-and Small-Scale FeaturesabstractA face image can be represented by a combination of large-and small-scale features. It is well-known that the variations of illumination mainly affect the large-scale features (low-frequency components), and not so much the small-scale features. Therefore, in relevant existing methods only the small-scale features are extracted as illumination-invariant features for face recognition, while the large-scale intrinsic features are always ignored. In this paper, we argue that both large-and small-scale features of a face image are important for face restoration and recognition. Moreover, we suggest that illumination normalization should be performed mainly on the large-scale features of a face image rather than on the original face image. A novel method of normalizing both the Small-and Large-scale (S&L) features of a face image is proposed. In this method, a single face image is first decomposed into large-and small-scale features. After that, illumination normalization is mainly performed on the large-scale features, and only a minor correction is made on the small-scale features. Finally, a normalized face image is generated by combining the processed large-and small-scale features. In addition, an optional visual compensation step is suggested for improving the visual quality of the normalized image. Experiments on CMU-PIE, Extended Yale B, and FRGC 2.0 face databases show that by using the proposed method significantly better recognition performance and visual results can be obtained as compared to related state-of-the-art methods. Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen, Ching Y. Suen |
IEEE Trans. Image Process. | 2 |
| 2010 | Two-Stage Sparse Representation for Robust Recognition on Large-Scale DatabaseabstractThis paper proposes a novel robust sparse representation method, called the two-stage sparse representation (TSR), for robust recognition on a large-scale database. Based on the divide and conquer strategy, TSR divides the procedure of robust recognition into outlier detection stage and recognition stage. In the first stage, a weighted linear regression is used to learn a metric in which noise and outliers in image pixels are detected. In the second stage, based on the learnt metric, the large-scale dataset is firstly filtered into a small set according to the nearest neighbor criterion. Then a sparse representation is computed by the non-negative least squares technique. The sparse solution is unique and can be optimized efficiently. The extensive numerical experiments on several public databases demonstrate that the proposed TSR approach generally obtains better classification accuracy than the state of the art Sparse Representation Classification (SRC). At the same time, by using the TSR, a significant reduction of computational cost is reached by over fifty times in comparison with the SRC, which enables the TSR to be deployed more suitably for large-scale dataset. Ran He 0001, Bao-Gang Hu, Wei-Shi Zheng 0001, Yanqing Guo |
AAAI | 3 |
| 2010 | Unsupervised Selective Transfer Learning for Object Recognition
Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
ACCV (2) | 1 |
| 2010 | Person Re-Identification by Support Vector RankingabstractSolving the person re-identification problem involves matching observations of individuals across disjoint camera views. The problem becomes particularly hard in a busy public scene as the number of possible matches is very high. This is further compounded by significant appearance changes due to varying lighting conditions, viewing angles and body poses across camera views. To address this problem, existing approaches focus on extracting or learning discriminative features followed by template matching using a distance measure. The novelty of this work is that we reformulate the person reidentification problem as a ranking problem and learn a subspace where the potential true match is given highest ranking rather than any direct distance measure. By doing so, we convert the person re-identification problem from an absolute scoring problem to a relative ranking problem. We further develop an novel Ensemble RankSVM to overcome the scalability limitation problem suffered by existing SVM-based ranking methods. This new model reduces significantly memory usage therefore is much more scalable, whilst maintaining high-level performance. We present extensive experiments to demonstrate the performance gain of the proposed ranking approach over existing template matching and classification models. 1 Bryan James Prosser, Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |