EDBT 2026 Demo / reviewers in the wild / expert
Chong Wang 0001
dblp:72/1334-1
· DBLP profile ↗
55ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0001-6016-6545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Systems, architecture and hardware · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning A Bank of Transferable Prompts for Vision-Language Models
Zhongwei Huang, Chong Wang 0001, Endai Huang, Ran Zhou 0002, Haitao Gan, Yingying Zhu 0001, Xiaoyu Shen 0001 |
ICMR | 3 |
| 2026 | Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Boosting representation diversity in video transformers via segmented contrastive masked autoencoders
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Neurocomputing | 6 |
| 2026 | Semantic Boosting via Knowledge Sharing and Feedback for Video Anomaly DetectionabstractVision-language models have the potential to enrich purely visual tasks by utilizing the combined representation of images/videos and corresponding textual descriptions. Recent advances in video anomaly detection have also integrated textual information to enhance the understanding of abnormal events. However, existing approaches often merge visual and textual modalities in a straightforward, bottom-up manner, failing to fully explore their interconnections. Moreover, textual captions themselves do not inherently convey “abnormal” attributes. Consequently, these joint representations tend to highlight all salient input features without adequately focusing on high-level tasks such as video anomaly detection. To direct the model’s attention towards anomalies more effectively, we propose incorporating a top-down mechanism into weakly supervised video anomaly detection tasks. A new Knowledge Sharing and Feedback (KSF) framework is designed to unify the representation of anomalies across both video and text. Specifically, we develop a category pattern sharing module that performs knowledge matching, acting as an alignment bridge between abnormal events and their corresponding descriptions. This ensures consistent representations for identical anomalies while maintaining distinct representations for different ones. Following this alignment process, matched high-level semantic priors are fed back into the forward path to enhance differentiation between abnormal and normal patterns. Comprehensive experiments on three benchmark datasets demonstrate the superiority of our proposed method in learning the implicit definition of anomaly patterns. The code is available at https://github.com/XJ-Cai/KSF. Xiaojie Cai, Yucheng Qian, Chong Wang 0001, Xiaohao Peng, Yuanbin Qian, Jiafei Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Learning Suspected Anomalies from Event Prompts for Video Anomaly DetectionabstractMost models for Weakly Supervised Video Anomaly Detection (WS-VAD) rely on multiple instance learning, aiming to distinguish normal and abnormal snippets without specifying the type of anomaly. However, the ambiguous nature of anomaly definitions across contexts may introduce inaccuracy in discriminating abnormal and normal events. To show the model what is anomalous, a novel framework is proposed to guide the learning of suspected anomalies from event prompts. Given a textual prompt dictionary of potential anomaly events and the captions generated from anomaly videos, the semantic anomaly similarity between them could be calculated to identify the suspected events for each video snippet. It enables a new multi-prompt learning process to constrain the visual-semantic features across all videos, as well as provides a new way to label pseudo anomalies for self-training. To demonstrate its effectiveness, comprehensive experiments and detailed ablation studies are conducted on four datasets, namely XD-Violence, UCF-Crime, TAD, and ShanghaiTech. Our proposed model outperforms most state-of-the-art methods in terms of AP or AUC (86.5%, 90.4%, 94.4%, and 97.4%). Furthermore, it shows promising performance in open-set and cross-dataset cases. The data, code, and models can be found at: https://github.com/shiwoaz/lap . Chenchen Tao, Xiaohao Peng, Chong Wang 0001, Jiafei Wu, Puning Zhao, Jun Wang 0071, Jiangbo Qian |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksabstractVideo anomaly detection plays a significant role in intelligent surveillance systems. To enhance model's anomaly recognition ability, previous works have typically involved RGB, optical flow, and text features. Recently, dynamic vision sensors (DVS) have emerged as a promising technology, which capture visual information as discrete events with a very high dynamic range and temporal resolution. It reduces data redundancy and enhances the capture capacity of moving objects compared to conventional camera. To introduce this rich dynamic information into the surveillance field, we created the first DVS video anomaly detection benchmark, namely UCF-Crime-DVS. To fully utilize this new data modality, a multi-scale spiking fusion network (MSF) is designed based on spiking neural networks (SNNs). This work explores the potential application of dynamic information from event data in video anomaly detection. Our experiments demonstrate the effectiveness of our framework on UCF-Crime-DVS and its superior performance compared to other models, establishing a new baseline for SNN-based weakly supervised video anomaly detection. Yuanbin Qian, Shuhan Ye, Chong Wang 0001, Xiaojie Cai, Jiangbo Qian, Jiafei Wu |
AAAI | 3 |
| 2025 | Differential Private Stochastic Optimization with Heavy-tailed Data: Towards Optimal RatesabstractWe study convex optimization problems under differential privacy (DP). With heavy-tailed gradients, existing works achieve suboptimal rates. The main obstacle is that existing gradient estimators have suboptimal tail property, resulting in a superfluous factor of d in the union bound. In this paper, we explore algorithms achieving optimal rates of DP optimization with heavy-tailed gradients. Our first method is a simple clipping approach. Under bounded p-th order moments of gradients, with n samples, it achieves minimax optimal population risk with epsilon less than 1/d. We then propose an iterative updating method, which is more complex but achieves this rate for all epsilon smaller than 1. The results significantly improve over existing methods. Such improvement relies on a careful treatment of the tail behavior of gradient estimators. Our results match the minimax lower bound, indicating that the theoretical limit of stochastic convex optimization under DP is achievable. Puning Zhao, Jiafei Wu, Zhe Liu 0001, Chong Wang 0001, Rongfei Fan, Qingming Li |
AAAI | 4 |
| 2025 | Cross Knowledge Distillation between Artificial and Spiking Neural NetworksabstractRecently, Spiking Neural Networks (SNNs) have demonstrated rich potential in computer vision domain due to their high biological plausibility, event-driven characteristic and energy-saving efficiency. Still, limited annotated event-based datasets and immature SNN architectures result in their performance inferior to that of Artificial Neural Networks (ANNs). To enhance the performance of SNNs on their optimal data format, DVS data, we explore using RGB data and well-performing ANNs to implement knowledge distillation. In this case, solving cross-modality and cross-architecture challenges is necessary. In this paper, we propose cross knowledge distillation (CKD), which not only leverages semantic similarity and sliding replacement to mitigate the cross-modality challenge, but also uses an indirect phased knowledge distillation to mitigate the cross-architecture challenge. We validated our method on main-stream neuromorphic datasets, including N-Caltech101 and CEP-DVS. The experimental results show that our method outperforms current State-of-the-Art methods. The code will be available at https://github.com/ShawnYE618/CKD. Shuhan Ye, Yuanbin Qian, Chong Wang 0001, Sunqi Lin, Jiangbo Qian |
ICME | 3 |
| 2025 | Improving crowdsourced label quality by peer-to-peer federated learning
Xiangming Lu, Jiangbo Qian, Chong Wang 0001, Diqun Yan, Youhui Zhang |
Appl. Intell. | 3 |
| 2025 | Robust Federated Learning Under Realistic Corruption: An Iterative Filtering ApproachabstractRobustness is one of the critical concerns in federated learning. Existing research focuses primarily on the worst case, typically modeled as the Byzantine attack, which alters the gradients in an optimal way. However, in practice, the corruption usually happens randomly, and is much weaker than the Byzantine attack. Therefore, existing methods overestimate the power of corruption, resulting in unnecessary sacrifice of performance. In this article, we build practical algorithms that can withstand realistic corruption, which is weaker than the Byzantine attack, in a better way. Toward this goal, we propose a new iterative filtering approach. In each iteration, it calculates the geometric median of all gradient vectors uploaded from clients and remove the gradients that are far away from the geometric median. A theoretical analysis is then provided, showing that under suitable parameter regimes, gradient vectors from corrupted clients are filtered if the noise is large, while those from benign clients are never filtered throughout the training process. For realistic gradient noise, our approach significantly outperforms existing methods, while the performance under the worst-case attack (i.e., the Byzantine attack) remains nearly the same. Experiments on both synthesized and real data validate our theoretical results, as well as the practical performance of our approach. In particular, we have achieved 3%–10% increase in MNIST and CIFAR10 datasets. Jiafei Wu, Puning Zhao, Chong Wang 0001, Zhe Liu 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Enhancing open-vocabulary object detection through region-word and region-vision matching
Yi Chen 0001, Chong Wang 0001, Sunqi Lin, Jinhui Xiang, Jiangbo Qian |
Multim. Syst. | 2 |
| 2025 | Event-Based Video Reconstruction Via Spatial-Temporal Heterogeneous Spiking Neural NetworkabstractEvent cameras detect per-pixel brightness changes and output asynchronous event streams with high temporal resolution, high dynamic range, and low latency. However, the unstructured nature of event streams means that humans cannot analyze and interpret them in the same way as natural images. Event-based video reconstruction is a widely used method aimed at reconstructing intuitive videos from event streams. Most reconstruction methods based on traditional artificial neural networks (ANNs) have high energy consumption, which counteracts the low-power advantage of event cameras. Spiking neural networks (SNNs) are a new generation of event-driven neural networks that encode information via discrete spikes, which leads to greater computational efficiency. Previous methods based on SNNs overlooked the asynchronous nature of event streams, leading to reconstructions that suffer from artifacts, flickering, low contrast, etc. In this work, we analyze event streams and spiking neurons and explain poor reconstruction quality. We specifically propose a novel spatial-temporal heterogeneous (STH) spiking neuron suitable for reconstructing asynchronous event streams. The STH neuron adjusts the membrane decay coefficient adaptively and has better spatiotemporal perception. In addition, we propose a temporal-frequency calibration module (TFCM) based on the Fourier transform to improve the contrast of the reconstructions. On the basis of the above proposed neuron and module, we construct two SNN-based models, referred to as the STHSNN and TFCSNN. The goal of the former is to reduce the artifacts and flickering in reconstructions, whereas the latter focuses on enhancing the contrast. The experimental results demonstrate that our models can yield reconstructions in various scenarios, achieving better quality and lower energy consumption than previous SNNs. Specifically, the TFCSNN and STHSNN achieve top-2 performance among the SNN-based models, with energy consumption reductions of 3.48 times and 12.40 times, respectively. Lijun Guo, Chong Wang 0001, Guoqi Li 0002, Jiangbo Qian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SpikeHCD: Spiking Transformer With Parallel Neurons and Memory-Enhanced Attention for Hyperspectral Change DetectionabstractHyperspectral change detection is a critical technology in remote sensing, widely applied in urban planning, environmental monitoring, and disaster detection. However, hyperspectral data exhibits higher spectral dimensionality compared to conventional RGB data, making existing methods struggle to balance high accuracy and low energy consumption. As the third generation of neural networks, spiking neural networks (SNNs) demonstrate the advantage of low energy efficiency, but the iterative computation process in spiking neurons significantly increases training and inference burdens when applied to hyperspectral change detection. To address these challenges, we propose a novel spiking Transformer with parallel neurons and memory-enhanced attention for hyperspectral change detection named SpikeHCD, the first SNNs specifically designed for hyperspectral change detection. SpikeHCD not only maintains low-energy advantage but also employs a probability-driven parallel spiking neurons (PPSN) to improve computational efficiency, enabling more effective application in remote sensing tasks. We further design a memory-enhanced spiking attention (MSA) module to enhance temporal modeling capability, and thoroughly extract spatial-spectral features. Additionally, a spiking difference module (SDM) is introduced to capture change features across different timesteps. Experimental results demonstrate that SpikeHCD can achieve several state-of-the-art (SOTA) results on multiple hyperspectral datasets, with faster detection, lower energy consumption, and fewer number of parameters. The codes are available at https://github.com/mzhcode/HCD_snn. Zihao Mei, Chong Wang 0001, Lijun Guo, Jiangbo Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Enhancing Home Exercise Experiences with Video Motion-Tracking for Automatic Display Height AdjustmentabstractThe increasing demand for home fitness solutions underscores the need for interactive displays that enhance user experiences. This study introduces a technology that autonomously adjusts display height using the skeletal information of demonstrators from videos, catering to home fitness needs. A user study involving thirty participants compared fixed height, manual adjustment, and automatic adjustment conditions. Head flexion angles and NASA-TLX survey responses were used for evaluation. Results showed a significant reduction in head flexion angles with automatic adjustment, promoting proper spinal alignment. NASA-TLX responses indicated lower mental, effort, and frustration ratings, along with improved performance and perceived support in the automatic adjustment condition compared to other conditions. These findings confirm that motion-based height adjustment improves posture and enhances the overall interactive experience. This research demonstrates the feasibility of integrating responsive ergonomics into interactive displays and suggests the importance of further personalization, conducting diverse user studies, and refining algorithms to fully leverage the potential of this technology. Chong Wang 0001, Pinyan Tang |
CHI | 5 |
| 2024 | SPEC-NERF: Multi-Spectral Neural Radiance FieldsabstractWe propose Multi-spectral Neural Radiance Fields(Spec-NeRF) for jointly reconstructing a multispectral radiance field and spectral sensitivity functions(SSFs) of the camera from a set of color images filtered by different filters. The proposed method focuses on modeling the physical imaging process, and applies the estimated SSFs and radiance field to synthesize novel views of multispectral scenes. In this method, the data acquisition requires only a low-cost trichromatic camera and several off-the-shelf color filters, making it more practical than using specialized 3D scanning and spectral imaging equipment. Our experiments on both synthetic and real scenario datasets demonstrate that utilizing filtered RGB images with learnable NeRF and SSFs can achieve high fidelity and promising spectral reconstruction while retaining the inherent capability of NeRF to comprehend geometric structures. Code is available at https://github.com/CPREgroup/SpecNeRF-v2. Ciliang Sun, Chong Wang 0001, Jinhui Xiang |
ICASSP | 4 |
| 2024 | Zero-Shot Object Detection with Partitioned Contrastive Feature AlignmentabstractHow to properly align the extracted visual features with certain semantic embeddings of unseen objects is crucial to the problem of Zero-Shot Object Detection (ZSD). To give a better guess of those unseen visual features, a partitioned contrast strategy is proposed in this paper to train the visual and attribute feature alignment networks. To be specific, four types of contrast are considered, including the visual-to-visual, visual-to-attribute, attribute-to-visual and attribute-to-attribute contrasts. Combining with two cross-batch memory banks of the visual features and unseen attribute features, it is effective to adjust the alignment rules for unseen visual features. Experimental results on the MS-COCO dataset show the superiority of the proposed model. Our code is available at: https://github.com/lihh1023/PCFA-ZSD. Haohe Li, Chong Wang 0001, Shenghao Yu, Zheng Huo, Jiangbo Qian |
ICASSP | 2 |
| 2024 | Distill Vision Transformers to CNNs via Teacher CollaborationabstractThe vision transformer (ViT) has recently emerged as a leading approach in various domains, outperforming other methods. Therefore, it is logical to explore the possibility of transferring the superior knowledge from ViT to more compact and cost-effective convolutional neural networks (CNNs). However, due to substantial architectural disparities in representation and logits between these models, conventional knowledge distillation methods have proven ineffective in this context. To address this issue, a novel cross-architecture knowledge distillation scheme based on teacher collaboration is proposed to alleviate the architecture gap. Two different teachers, i.e. one ViT and one CNN, are utilized to simultaneously distill the student by feature reaggregation and logit correction. The experiments show that the proposed scheme outperforms conventional methods on CIFAR-100 dataset. The code is available at https://github.com/SunkiLin/RCD. Sunqi Lin, Chong Wang 0001, Chenchen Tao, Xinmiao Dai |
ICASSP | 2 |
| 2024 | Swin transformer-based traffic video text tracking
Jinyao Yu, Jiangbo Qian, Chong Wang 0001, Yihong Dong |
Appl. Intell. | 4 |
| 2024 | Animation line art colorization based on the optical flow methodabstractAbstract Coloring an animation sketch sequence is a challenging task in computer vision since the information contained in line sketches is too sparse, and the colors need to be uniform between continuous frames. Many the existing colorization algorithms can only be applied to one image and can be considered color filling algorithms. Such algorithms only provide a color result that fits within a reasonable range and can not be applied to the coloring of frame sequences. This paper proposes an end‐to‐end two‐stage optical flow colorization network to solve the animation frame sequence colorization problem. The first stage of the network finds the direction of the color pixel flow from the detail change between a given reference frame and the next frame of line artwork and then completes the initial coloring process. The second stage of the network performs color correction and clarifies the output of the first stage. Since our algorithm does not directly colorize the image but finds the path of the color change to colorize it, it ensures a consistent color space for the sequence frames after colorization. We conduct experiments on an animation dataset, and the results show that our algorithm is effective. The code is available at https://github.com/silenye/Colorization . Jiangbo Qian, Chong Wang 0001, Yihong Dong, Baisong Liu |
Comput. Animat. Virtual Worlds | 3 |
| 2024 | A Vision Enhancement and Feature Fusion Multiscale Detection NetworkabstractAbstract In the field of object detection, there is often a high level of occlusion in real scenes, which can very easily interfere with the accuracy of the detector. Currently, most detectors use a convolutional neural network (CNN) as a backbone network, but the robustness of CNNs for detection under cover is poor, and the absence of object pixels makes conventional convolution ineffective in extracting features, leading to a decrease in detection accuracy. To address these two problems, we propose VFN (A Vision Enhancement and Feature Fusion Multiscale Detection Network), which first builds a multiscale backbone network using different stages of the Swin Transformer, and then utilizes a vision enhancement module using dilated convolution to enhance the vision of feature points at different scales and address the problem of missing pixels. Finally, the feature guidance module enables features at each scale to be enhanced by fusing with each other. The total accuracy demonstrated by VFN on both the PASCAL VOC dataset and the CrowdHuman dataset is better than that of other methods, and its ability to find occluded objects is also better, demonstrating the effectiveness of our method.The code is available at https://github.com/qcw666/vfn . Chengwu Qian, Jiangbo Qian, Chong Wang 0001, Xulun Ye, Caiming Zhong |
Neural Process. Lett. | 3 |
| 2024 | Dynamic Sensing and Correlation Loss Detector for Small Object Detection in Remote Sensing ImagesabstractRecently, significant object detection achievements have been emerged for optical remote sensing images. However, the performance and efficiency of small object detection are still highly unsatisfactory because of the scale diversity between the objects; furthermore, small objects always have small amounts of effective information that are difficult to locate. To address this problem, we propose a novel dynamic sensing and correlation loss detector (DCDet) for performing object detection in remote sensing images. The detector consists of two modules: a small-object dynamic sensing (SODS) module and a simple but effective correlation loss function (CrLoss). SODS is utilized to capture the information of small objects in a scale sequence. We consider the feature pyramid as a set of video frames when the camera is zoomed in on the image and use the object focusing module in dynamic sensing to always focus on the small objects in each video frame. The detection performance achieved for small objects is improved by shifting the detector’s attention from the entire image to small objects within the frame to provide a multiscale feature representation of the small objects and their contextual information. The CrLoss is a special correlation loss for remote sensing image object detection tasks and directly optimizes the correlation coefficient to improve the performance of a detector. Extensive experiments conducted on the publicly available DOTA, DIOR-R and HRSC2016 datasets show that our DCDet outperforms the existing state-of-the-art remote sensing object detection methods in terms of many evaluation metrics. Chongchong Shen, Jiangbo Qian, Chong Wang 0001, Diqun Yan, Caiming Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Restructuring the Teacher and Student in Self-DistillationabstractKnowledge distillation aims to achieve model compression by transferring knowledge from complex teacher models to lightweight student models. To reduce reliance on pre-trained teacher models, self-distillation methods utilize knowledge from the model itself as additional supervision. However, their performance is limited by the same or similar network architecture between the teacher and student. In order to increase architecture variety, we propose a new self-distillation framework called restructured self-distillation (RSD), which involves restructuring both the teacher and student networks. The self-distilled model is expanded into a multi-branch topology to create a more powerful teacher. During training, diverse student sub-networks are generated by randomly discarding the teacher's branches. Additionally, the teacher and student models are linked by a randomly inserted feature mixture block, introducing additional knowledge distillation in the mixed feature space. To avoid extra inference costs, the branches of the teacher model are then converted back to its original structure equivalently. Comprehensive experiments have demonstrated the effectiveness of our proposed framework for most architectures on CIFAR-10/100 and ImageNet datasets. Code is available at https://github.com/YujieZheng99/RSD. Chong Wang 0001, Chenchen Tao, Sunqi Lin, Jiangbo Qian, Jiafei Wu |
IEEE Trans. Image Process. | 2 |
| 2024 | Feature Reconstruction With Disruption for Unsupervised Video Anomaly DetectionabstractUnsupervised video anomaly detection (UVAD) has gained significant attention due to its label-free nature. Typically, UVAD methods can be categorized into two branches, i.e. the one-class classification (OCC) methods and fully UVAD ones. However, the former may suffer from data imbalance and high false alarm rates, while the latter relies heavily on feature representation and pseudo-labels. In this paper, a novel feature reconstruction and disruption model (FRD-UVAD) is proposed for effective feature refinement and better pseudo-label generation in fully UVAD, based on cascade cross-attention transformers, a latent anomaly memory bank and an auxiliary scorer. The clip features are reconstructed using the space-time intra-clip information, as well as cross-inter-clip knowledge. Moreover, instead of blindly reconstructing all training features as OCC methods, a new disruption process is proposed to cooperate with the feature reconstruction simultaneously. Using the collected pseudo anomaly samples, it is able to emphasize the feature differences between normal and abnormal events. Additionally, a pre-trained UVAD scorer is utilized as a different criteria for anomaly prediction, which further refines the pseudo-labels. To demonstrate its effectiveness, comprehensive experiments and detailed ablation studies are conducted on three video benchmarks, namely CUHK Avenue, ShanghaiTech and UCF-Crime. Our proposed model (FRD-UVAD) achieves the best AUC performance (91.23%, 80.14%, and 82.12%) on all three datasets, surpassing other state-of-the-art OCC and fully UVAD methods. Furthermore, it obtains the lowest false alarm rate with a lower scene dependency, compared with other OCC methods. The code is available athttps://github.com/tcc-power/FRD-unsupervised-video-anomaly-detection. Chenchen Tao, Chong Wang 0001, Sunqi Lin, Suhang Cai, Jiangbo Qian |
IEEE Trans. Multim. | 2 |
| 2023 | CaSE-NeRF: Camera Settings Editing of Neural Radiance Fields
Ciliang Sun, Chong Wang 0001, Xinmiao Dai |
CGI | 4 |
| 2023 | Animal Re-Identification Algorithm for Posture DiversityabstractRe-identification (Re-ID) technology is important for wildlife conservation and intelligent farm management. With the development of deep learning, the performance of animal ReID based on computer vision has been improved. However, variations in animal pose push a negative impact on recognition performance. In this paper, a Multi-pose Feature Fusion Network (MPFNet) is proposed to improve the performance of the Re-ID. First, we construct three pose modules for the three postures, that is, standing, sitting, and lying, respectively. In each pose module, there are two parallel branches, one is a global branch for extracting global features, and the other is a local branch for extracting local features. In addition, to obtain more effective feature representations, we weighted fusion for the global branching of the three pose modules. We validate the efficiency of MPFNet on both the self-built MPDD dog dataset and the public ATRW Amur Tiger dataset. Experimental results show that MPFNet can obtain better recognition performance than other state-of-the-art Re-ID methods. The source of code will be public available at https://github.com/hezhimin7028/MPFNet. Jiangbo Qian, Diqun Yan, Chong Wang 0001 |
ICASSP | 4 |
| 2023 | Enlightening the Student in Knowledge DistillationabstractKnowledge distillation is a common method of model compression, which uses large models (teacher networks) to guide the training of small models (student networks). However, the student may find a hard time absorbing the knowledge from a sophisticated teacher due to the capacity and confidence gaps between them. To address this issue, a new knowledge distillation and refinement (KDrefine) framework is proposed to enlighten the student by expending and refining its network structure. In addition, a confidence refinement strategy is utilized to generate adaptive soften logits for efficient distillation. The experiments show that the proposed framework outperforms state-of-the-art methods on both CIFAR-100 and Tiny-ImageNet datasets. The code is available at https://github.com/YujieZheng99/KDrefine. Chong Wang 0001, Yi Chen 0001, Jiangbo Qian, Jun Wang 0071, Jiafei Wu |
ICASSP | 2 |
| 2023 | Synthetic Feature Assessment for Zero-Shot Object DetectionabstractZero-shot object detection aims to simultaneously identify and localize classes that were not presented during training. Many generative model-based methods have shown promising performance by synthesizing the visual features of unseen classes from semantic embeddings. However, these synthetic features are inevitably of varied quality, which may be far from the ground truth. It degrades the performance of trained unseen classifier. Instead of tweaking the generative model, a new idea of feature quality assessment is proposed to utilize both the good and bad features to optimize the classifier in the right direction. Moreover, contrastive learning is also introduced to enhance the feature uniqueness between unseen and seen classes, which helps the feature assessment implicitly. To demonstrate the effectiveness of the proposed algorithm, comprehensive experiments are conducted on the MS COCO dataset and PASCAL VOC dataset, the state-of-the-art performance is achieved. Our code is available at: https://github.com/Dai1029/SFA-ZSD. Xinmiao Dai, Chong Wang 0001, Haohe Li, Sunqi Lin, Li Dong 0006, Jiafei Wu, Jun Wang 0071 |
ICME | 2 |
| 2023 | Pseudo-label Diversity Exploitation for Few-Shot Object Detection
Chong Wang 0001, Zhengjie Ye, Jiacheng Deng 0001 |
MMM (2) | 2 |
| 2023 | Zero-shot object detection with contrastive semantic association network
Haohe Li, Chong Wang 0001, Yiling Gong, Xinmiao Dai |
Appl. Intell. | 2 |
| 2023 | Swin transformer-based supervised hashing
Liangkang Peng, Jiangbo Qian, Chong Wang 0001, Baisong Liu, Yihong Dong |
Appl. Intell. | 3 |
| 2023 | Harmonized Portrait-Background Image CompositionabstractAbstract Portrait‐background image composition is a widely used operation in selfie editing, video meeting, and other portrait applications. To guarantee the realism of the composited images, the appearance of the foreground portraits needs to be adjusted to fit the new background images. Existing image harmonization approaches are proposed to handle general foreground objects, thus lack the special ability to adjust portrait foregrounds. In this paper, we present a novel end‐to‐end network architecture to learn both the content features and style features for portrait‐background composition. The method adjusts the appearance of portraits to make them compatible with backgrounds, while the generation of the composited images satisfies the prior of a style‐based generator. We also propose a pipeline to generate high‐quality and high‐variety synthesized image datasets for training and evaluation. The proposed method outperforms other state‐of‐the‐art methods both on the synthesized dataset and the real composited images and shows robust performance in video applications. Yijiang Wang, Chong Wang 0001, Xulun Ye |
Comput. Graph. Forum | 3 |
| 2023 | Feature Differentiation Reconstruction Network for Weakly-Supervised Video Anomaly DetectionabstractRecent research into video anomaly detection under weakly supervised settings has made significant progress in identifying anomalies with only coarse-grained annotations. Mainstream weakly supervised methods improve detection performance by generating high-quality pseudo labels for video segments. However, these pseudo-label-based methods have been ordinarily hindered by manually-set constraint rules as the bottleneck. In this paper, we propose the Feature Differentiation Reconstruction Network (FDR-Net), which no longer relies on pseudo labels and instead uses a differential reconstruction strategy to improve the discriminability of the representation. Concretely, video features are first randomly masked out and then reconstructed with distinct targets for normal and abnormal videos during the differential reconstruction process. Besides, we also introduce a dense transformer-based encoder to refine spatial-temporal relationships among video segments. Comprehensive experiments on ShanghaiTech demonstrate the superior performance of our model. Yiling Gong, Sihui Luo 0001, Chong Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Balanced Stripe-Wise Pruning In The FilterabstractNeural network pruning offers a promising prospect to compress and accelerate modern deep convolution networks. The stripe-wise pruning method with a finer granularity than traditional methods has become the focus of research. Inspired by the previous work, a new balanced stripe-wise pruning, including the balanced pruning strategy and dynamic pruning threshold, is proposed to achieve higher performance. Specifically, the survived inter-filter stripes and the intra-filter stripes are redistributed by a balanced pruning strategy. Meanwhile the dynamic pruning threshold method makes survival rates will be further balanced across all layers. Comprehensive experiments are conducted on two public datasets (CIFAR-10 and TinyImageNet-200) for different models (ResNet and VGG). The experimental results show that the proposed model is capable of reducing the most parameters, yet achieving the highest accuracy. Our code is available at: https://github.com/ajdt1111/BSWP. Zheng Huo, Chong Wang 0001, Jun Wang 0071, Jiafei Wu |
ICASSP | 2 |
| 2022 | Novel Instance Mining with Pseudo-Margin Evaluation for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) enables the detector to recognize novel objects only using limited training samples, which could greatly alleviate model’s dependency on data. Most existing methods include two training stages, namely base training and fine-tuning. However, the unlabeled novel instances in the base set were untouched in previous works, which can be re-used to enhance the FSOD performance. Thus, a new instance mining model is proposed in this paper to excavate the novel samples from the base set. The detector is thus fine-tuned again by these additional free novel instances. Meanwhile, a novel pseudo-margin evaluation algorithm is designed to address the quality problem of pseudo-labels brought by those new novel instances. The experimental results on MS-COCO dataset show the effectiveness of the proposed model, which does not require any additional training samples or parameters. Our code is available at: https://github.com/liuweijie19980216/NimPme. Chong Wang 0001, Shenghao Yu, Chenchen Tao, Jun Wang 0071, Jiafei Wu |
ICASSP | 2 |
| 2022 | Pruning Dynamic Group Convolution with Static SubstituteabstractDeep learning networks are gradually deployed in edge applications, such as phones and cameras, which has a restriction of the computational resources. Thus, to improve the computational efficiency, numerous types of group convolution-based frameworks have been studied, including the dynamic group convolution (DGC). However, it is worth noting that sometimes the same channels are selected in different dynamic heads in DGC, which violates its original intention. It indicates the number of dynamic heads exceeds what is really needed. In this paper, a novel model by pruning dynamic group convolution with static substitute is proposed. Specifically, those channels pruned in the dynamic head can be reselected in the static head of the group convolution in the proposed model. In addition, a new polarization regularization is introduced to prune more useless channels with less accuracy loss. The experiment results on two image classification benchmarks (CIFAR-I00 and TinylmageNet-200) show promising performance. Jieyong Che, Chong Wang 0001, Xinmiao Dai, Jun Wang 0071, Jiafei Wu |
ICME | 2 |
| 2022 | Multi-Scale Continuity-Aware Refinement Network for Weakly Supervised Video Anomaly DetectionabstractIn many previous work, weakly supervised video anomaly detection is formulated as a multiple instance learning (MIL) problem, which represents the video as a bag of multiple instances. However, most MIL-based frameworks only focused on identifying anomalous events from the given instances, without considering the event continuity. Motivated by the fact that abnormal events tend to be more continuous in real-world videos, a Multi-scale Continuity-aware Refinement Network (MCR) is proposed in this paper. It utilizes the property of multi-scale continuity to refine anomaly scores by introducing differential contextual information of instances. At the same time, multi-scale attention is designed to produce a video-level weights in order to select the proper scale and fuse all scores at different scales. Experimental results of MCR show noticeable improvement on two public datasets, specifically obtaining a frame-level AUC 94.92% on ShanghaiTech dataset. Yiling Gong, Chong Wang 0001, Xinmiao Dai, Shenghao Yu, Lehong Xiang, Jiafei Wu |
ICME | 2 |
| 2022 | TCA-VAD: Temporal Context Alignment Network for Weakly Supervised Video Anomly DetectionabstractVideo Anomaly detection (VAD) with weakly supervised is usually formulated as a multiple instance learning (MIL) problem. Although the current MIL-based methods have achieved promising detection performance, the temporal dependencies in videos are not well exploited. There may multiple abnormal clips in a given anomaly video, while the previous work only focused on the most abnormal one. To address above issues, a temporal context alignment (TCA) network for video anomaly detection is proposed in this work. Its merits are three-fold, 1) a sparse continuous sampling strategy is proposed to adapt the varying length of untrimmed videos; 2) a multi-scale attention module is used to establish the video temporal dependencies; 3) a top-k loss strategy is used to enlarge the distance between the top-k normal and abnormal clips. Extensive experiments demonstrate the noticeable anomaly discriminability of the proposed network on two public datasets (ShanghaiTech and UCF-Crime). Shenghao Yu, Chong Wang 0001, Lehong Xiang, Jiafei Wu |
ICME | 2 |
| 2022 | Multi-view k-proximal plane clustering
Feixiang Sun, Xijiong Xie, Jiangbo Qian, Chong Wang 0001, Guoqing Chao |
Appl. Intell. | 6 |
| 2022 | Band Selection for HSI Classification Using Binary Constrained OptimizationabstractHyperspectral images (HSIs) containing tens to hundreds of bands can be used in various image classification tasks. However, due to the high data redundancy of the spectral information, the acquiring and analysis of HSIs are usually relatively time-consuming and wasteful of storage space, and therefore limit the practical application of HSIs. Selecting a subset of bands without sacrificing classification accuracy is a strategy to relieve such problems. In this letter, we present an optimization-based method, which can jointly optimize the band selection (BS) and the classification network parameters for HSIs. The proposed method regards the discrete selection problem as a continuous constrained optimization problem and adaptively selects the informative band subsets for classification. Besides, the experimental results on three public datasets show that our BS method outperforms the state-of-the-art methods in terms of classification accuracy. Xueyan Tian, Chong Wang 0001, Lijun Guo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Reweighted Dynamic Group ConvolutionabstractComputational efficiency of modern deep convolutional networks is always desired in real world applications. Various types of group convolutions have been widely used to reduce the complexity of the networks. Inspired by the previous work, a new reweighted dynamic group convolution (RDGC) structure, including a reweighted pruning module and a survival loss, is proposed in this work for more precise channel pruning. Specifically, a layer-wise pruning rate adjustment strategy is employed with an intra-cluster loss to guide the learning of self-attention implicitly. The proposed model is more capable to select the suitable channels to construct the filter groups. Its effectiveness is demonstrated by extensive experiments on four public datasets, namely CIFAR-10, CIFAR-100 and two fine-grained image datasets (Stanford Dogs and Stanford Cars). Chong Wang 0001, Zheng Huo, Linlin Gao |
ICASSP | 2 |
| 2021 | Cross-Epoch Learning for Weakly Supervised Anomaly Detection in Surveillance VideosabstractWeakly Supervised Anomaly Detection (WSAD) in surveillance videos is a complex task since usually only video-level annotations are available. Previous work treated it as a regression problem by giving different scores on normal and anomaly events. However, the widely used mini-batch training strategy may suffer from the data imbalance between these two types of events, which limits the model's performance. In this work, a cross-epoch learning (XEL) strategy associated with a hard instance bank (HIB) is proposed to introduce additional information from previous training epochs. Two new losses are proposed for XEL to achieve a higher detection rate as well as a lower false alarm rate of anomaly events. Moreover, the proposed XEL can be directly integrated into any existing WSAD framework. Experimental results of three XEL embedded models have shown promising AUC improvement (3%~7%) on two public datasets, surpassing the state-of-the-art methods. Our code is available at: https://github.com/sdjsngs/XEL-WSAD. Shenghao Yu, Chong Wang 0001, Qiao-mei Ma, Jiafei Wu |
IEEE Signal Process. Lett. | 2 |
| 2018 | An Improved Guided Filtering Algorithm for Image EnhancementabstractGuided image filter (GIF) is popular in image processing and computer vision for the properties of edge-preserving and low computational complexity, but GIF may suffer from over-smoothing (halo artifacts) near sharp edges and under-smoothing at flat regions. There is a tradeoff between them in the original cost function of GIF. In this paper, an improved guided filter (IGIF) is proposed by incorporating an adaptive structure aware constraint. The adaptive structure aware constraint can well preserve edges and smooth details through assigning different weights to different local structure. Simultaneously, thanks to the L1penalty, the proposed IGIF can exactly remove small details at the flat regions. To illustrate the effectiveness of the proposed IGIF, we apply it to image enhancement. Experimental results show that the proposed filter can produce enhanced images with better visual quality as well as quantitative performance. Jiafei Wu, Chong Wang 0001, Yongze Xu |
ICME | 2 |
| 2018 | Practical Radiometric Compensation for Projection Display on Textured Surfaces using a Multidimensional ModelabstractAbstract Radiometric compensation methods remove the effect of the underlying spatially varying surface reflectance of the texture when projecting on textured surfaces. All prior work sample the surface reflectance dependent radiometric transfer function from the projector to the camera at every pixel that requires the camera to observe tens or hundreds of images projected by the projector. In this paper, we cast the radiometric compensation problem as a sampling and reconstruction of multi‐dimensional radiometric transfer function that models the color transfer function from the projector to an observing camera and the surface reflectance in a unified manner. Such a multi‐dimensional representation makes no assumption about linearity of the projector to camera color transfer function and can therefore handle projectors with non‐linear color transfer functions(e.g. DLP, LCOS, LED‐based or laser‐based). We show that with a well‐curated sampling of this multi‐dimensional function, achieved by exploiting the following key properties, is adequate for its accurate representation: (a) the spectral reflectance of most real‐world materials are smooth and can be well‐represented using a lower‐dimension function; (b) the reflectance properties of the underlying texture have strong redundancies – for example, multiple pixels or even regions can have similar surface reflectance; (c) the color transfer function from the projector to camera have strong input coherence. The proposed sampling allows us to reduce the number of projected images that needs to be observed by a camera by up to two orders of magnitude, the minimum being only two. We then present a new multi‐dimensional scattered data interpolation technique to reconstruct the radiometric transfer function at a high spatial density (i.e. at every pixel) to compute the compensation image. We show that the accuracy of our interpolation technique is higher than any existing methods. Aditi Majumder, Meenakshisundaram Gopi, Chong Wang 0001 |
Comput. Graph. Forum | 4 |
| 2018 | Locally Linear Embedded Sparse Coding for Spectral Reconstruction From RGB ImagesabstractTraining-based spectral reconstruction is an efficient, inexpensive technique to recover spectral images from the RGB images captured by trichromatic cameras. Existing methods handle training samples individually without any consideration of local spatial and spectral correlations between samples, which results in high metamerism and inaccurate reconstruction. In this letter, we exploit for the first time the concept of spectral image reconstruction from RGB images with both chromatic and texture priors. We reduce redundancy of the sample set by applying a volume maximization based selection strategy. Taking advantage of the local linearity and sparsity of spectra in dictionary learning, we propose a locally linear embedded sparse reconstruction method taking into account both RGB values of pixels and the features of patch texture. Experimental results show that our method is significantly more accurate than the state-of-the-art methods. Chong Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2018 | Efficient spectral reconstruction using a trichromatic camera via sample optimization
Chong Wang 0001, Qingshu Yuan |
Vis. Comput. | 2 |
| 2018 | Superpixel-based color-depth restoration and dynamic environment modeling for Kinect-assisted image-based rendering systems
Chong Wang 0001, S. C. Chan 0001, Li Zhang 0041, Harry Shum |
Vis. Comput. | 1 |
| 2017 | A hand gesture recognition system based on canonical superpixel-graph
Chong Wang 0001, Zhong Liu 0004, Minfeng Zhu 0002, S. C. Chan 0001 |
Signal Process. Image Commun. | 1 |
| 2016 | Hand gesture recognition based on canonical formed superpixel earth mover's distanceabstractThis paper presents a new hand gesture recognition algorithm based on canonical formed superpixel earth mover's distance (CF-SP-EMD). SP-EMD is a recently proposed distance metric designed for depth based hand gesture recognition, which shows promising performance. However, in real life, people may have their own habits while performing certain hand gestures. This will yield a variety of hand shapes with different finger poses as compared with the standard templates. Such variety may affect the accuracy of SP-EMD and hence will degrade its performance. In this paper, we propose a new distance metric CF-SP-EMD to alleviate the problem. We organize superpixels in canonical forms that can factor out nonstandard finger poses, resulting a well-structured fingerpose-neutral shape representation for hand gestures. Experimental results using public gesture datasets show that the proposed CF-SP-EMD can achieve better performance for hand gesture recognition, compared with the state-of-art algorithms. Chong Wang 0001, Zhong Liu 0004 |
ICME | 1 |
| 2015 | Multi-view articulated human body tracking with textured deformable mesh modelabstractThis paper proposes a multi-view articulated human motion tracking approach with textured deformable mesh model. Firstly, a subject-specific mesh model is initialized by using linear blend skinning method. The model is then textured according to the multi-view image observations. We introduce a segmentation-based method to refine the appearance of the subject. With the textured mesh model, a color-based likelihood (CbL) is also proposed for human body tracking with Annealed Particle Filter (APF). Experiments in the paper show that the performance can be considerately improved by using CbL as the measurement for pose tracking. Zhong Liu 0004, S. C. Chan 0001, Chong Wang 0001, Shuai Zhang 0004 |
ISCAS | 3 |
| 2015 | Depth map restoration and upsampling for kinect v2 based on IR-depth consistency and joint adaptive kernel regressionabstractThis paper presents a depth map restoration scheme for both the raw and projected depth map from Kinect v2 sensor. Based on IR-depth consistency, erroneous depth readings around foreground objects are removed by an edge aware consistency correction method. Moreover, a joint adaptive kernel regression algorithm is designed to upsample the sparse depth map after the projection from Kinect v2 sensor's depth camera to its full HD video camera. The structural information in the high resolution color image is implicitly utilized to guide the upsampling of depth map. The effectiveness of the proposed upsampling algorithm is illustrated by experimental results and comparisons on both real Kinect v2 data and Middlebury dataset. Chong Wang 0001, Zhouchi Lin, S. C. Chan 0001 |
ISCAS | 1 |
| 2015 | Superpixel-Based Hand Gesture Recognition With Kinect Depth CameraabstractThis paper presents a new superpixel-based hand gesture recognition system based on a novel superpixel earth mover's distance metric, together with Kinect depth camera. The depth and skeleton information from Kinect are effectively utilized to produce markerless hand extraction. The hand shapes, corresponding textures and depths are represented in the form of superpixels, which effectively retain the overall shapes and color of the gestures to be recognized. Based on this representation, a novel distance metric, superpixel earth mover's distance (SP-EMD), is proposed to measure the dissimilarity between the hand gestures. This measurement is not only robust to distortion and articulation, but also invariant to scaling, translation and rotation with proper preprocessing. The effectiveness of the proposed distance metric and recognition algorithm are illustrated by extensive experiments with our own gesture dataset as well as two other public datasets. Simulation results show that the proposed system is able to achieve high mean accuracy and fast recognition speed. Its superiority is further demonstrated by comparisons with other conventional techniques and two real-life applications. Chong Wang 0001, Zhong Liu 0004, S. C. Chan 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | A new bandwidth adaptive non-local kernel regression algorithm for image/video restoration and its GPU realizationabstractThis paper presents a new bandwidth adaptive nonlocal kernel regression (BA-NLKR) algorithm for image and video restoration. NLKR is a recent approach for improving the performance of conventional steering kernel regression (SKR) and local polynomial regression (LPR) in image/video processing. Its bandwidth, which controls the amount of smoothing, however is chosen empirically. The proposed algorithm incorporates the intersecting confidence intervals (ICI) bandwidth selection method into the framework of NLKR to facilitate automatic bandwidth selection so as to achieve better performance. A parallel implementation of the proposed algorithm is also introduced to reduce significantly its computation time. The effectiveness of the proposed algorithm is illustrated by experimental results on both single image and videos super resolution and denoising. Chong Wang 0001, S. C. Chan 0001 |
ISCAS | 1 |
| 2012 | A Multi-Camera Approach to Image-Based Rendering and 3-D/Multiview Display of Ancient Chinese ArtifactsabstractThis paper proposes an image-based approach for the capturing, rendering and display of ancient Chinese artifacts for cultural heritage preservation. A multiple-camera circular array is proposed to record images of the artifacts, which forms a simplified circular light field (SCLF). A systematic image-based approach and associate algorithms such as segmentation, depth estimation and shape morphing are developed for rendering new views of the Chinese artifacts. An object-based compression scheme is also proposed to reduce the data size for storage and transmission of the texture, depth maps and alpha maps associated with the object-based circular light field. Spatial redundancies among the various images are exploited to improve the coding performance, while avoiding excessive complexity in selective decoding of the light field to support fast rendering speed. To allow the Chinese artifacts to be viewed over the internet, scalable prioritized transmission and rendering schemes of the SCLF with low latency were also developed. The multiple views so synthesized enable the ancient artifacts to be displayed in 3-D/multi-view displays. Several collections from the University Museum and Art Gallery at The University of Hong Kong were captured and excellent rendering results are obtained. King To Ng, Chong Wang 0001, S. C. Chan 0001, Harry Shum |
IEEE Trans. Multim. | 3 |
| 2011 | Realistic and interactive image-based rendering of ancient chinese artifacts using a multiple camera arrayabstractThis paper proposes a system for photorealistic interactive rendering of ancient Chinese artifacts for cultural heritage preservation using multiview images captured by a circular multiple-camera array. It employs 3D reconstruction and precomputed shadow field techniques to enable real-time relighting and object interaction. Moreover, Gabor features are employed to improve the robustness of line matching along epipolar lines and robust radial basis function modeling is employed to suppress possible outliers arising from false matching. Using the 3D model reconstructed, the precomputed shadow field is employed to provide real-time rendering/relighting and object movement, after acceleration on a graphic processing unit (GPU). Excellent rendering results are obtained and the ancient Chinese artifacts can be displayed in modern multi-view displays and conventional stereo systems. Chong Wang 0001, S. C. Chan 0001, Harry Shum |
ISCAS | 1 |
| 2010 | Cooperative co-evolutionary algorithm in satellite imaging scheduling of cooperative multiple centersabstractIn this paper, satellite imaging scheduling of cooperative multiple centers, an example of multi-agent systems, is discussed with a proposal of a novel Cooperative CoEvolutionary Planning Algorithm (CCEPA). Considering the numbers of the multi-center and the characteristics of targets to be observed, CCEPA decomposes the tasks into smaller components and evolves multiple solutions in the form of cooperative subpopulations. At the same time, for such evolutionary algorithm based on scheduling, a novel fixed-length binary encoding mechanism for tasks assigned to each center is also proposed. Incorporated with various features like archiving, dynamic sharing, the CCEPA is capable of maintaining archive diversity in the evolution and distributing the solutions uniformly along the Pareto front. Simulation and analysis show that the proposed algorithm can solve the problem effectively. Chong Wang 0001, Ning Jing, Jun Li 0020, Jun Wang 0071, Hao Chen 0046 |
IEEE Congress on Evolutionary Computation | 1 |