Shuang Wu 0001

dblp:85/3231-1 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0003-1237-2656ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 10 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Instruct Where the Model Fails: Generative Data Augmentation via Guided Self-contrastive Fine-tuning
abstract
Data augmentation is expected to bring about unseen features of training set, enhancing the model’s ability to generalize in situations where data is limited. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with stereotypes and imperceptible bias when used to augment training data, owing to dataset misalignment and the generator’s ignorance of the downstream model. We improve downstream task awareness in generated images by proposing a task-aware fine-tuning strategy that actively detects failures of downstream task in the target model to fine-tune the generation process between epochs. The dynamic fine-tuning strategy is achieved by (1) inspecting misalignment between generated data and original data via VLM captioners and (2) adjusts both prompts and diffusion model so that the strategy dynamically guides the generator by focusing on the detected bias of VLM. This is done via re-captioning the overfitted data as well as finetuning the diffusion trajectory in a contrastive manner. To co-operate with the VLM captioner, the contrastive fine-tuning process dynamically adjusts different parts of the diffusion trajectory based on detected misalignment, thus shifting the the generated distribution away from making the downstream model overfit. Our experiments on few-shot class incremental learning show that our instruction-guided finetuning strategy consistently assists the downstream model with higher classification accuracy compared to generative data augmentation baselines such as Stable Diffusion and GPT-4o, and state-of-the-art non-generative strategies.
Weijian Ma, Ruoxin Chen, Ke-Yue Zhang, Shuang Wu 0001, Shouhong Ding
AAAI4
2025 Dual Data Alignment Makes AI-Generated Image Detector Easier Generalizable
abstract
The rapid increase in AI-generated images (AIGIs) underscores the need for detection methods. Existing detectors are often trained on biased datasets, leading to overfitting on spurious correlations between non-causal image attributes and real/synthetic labels. While these biased features enhance performance on the training data, they result in substantial performance degradation when tested on unbiased datasets. A common solution is to perform data alignment through generative reconstruction, matching the content between real and synthetic images. However, we find that pixel-level alignment alone is inadequate, as the reconstructed images still suffer from frequency-level misalignment, perpetuating spurious correlations. To illustrate, we observe that reconstruction models restore the high-frequency details lost in real images, inadvertently creating a frequency-level misalignment, where synthetic images appear to have richer high-frequency content than real ones. This misalignment leads to models associating high-frequency features with synthetic labels, further reinforcing biased cues. To resolve this, we propose Dual Data Alignment (DDA), which aligns both the pixel and frequency domains. DDA generates synthetic images that closely resemble real ones by fusing real and synthetic image pairs in both domains, enhancing the detector's ability to identify forgeries without relying on biased features. Moreover, we introduce two new test sets: DDA-COCO, containing DDA-aligned synthetic images, and EvalGEN, featuring the latest generative models. Our extensive evaluations demonstrate that a detector trained exclusively on DDA-aligned MSCOCO improves across diverse benchmarks. Code is available at https://github.com/roy-ch/Dual-Data-Alignment.
Ruoxin Chen, Junwei Xi, Zhiyuan Yan 0002, Ke-Yue Zhang, Shuang Wu 0001, Isabel Guan, Taiping Yao, Shouhong Ding
NeurIPS5
2025 EnfoMax: Domain entropy and mutual information maximization for domain generalized face anti-spoofing
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
Neurocomputing3
2025 Exploring the adversarial robustness of face forgery detection with decision-based black-box attacks
Zhaoyu Chen 0001, Bo Li 0115, Kaixun Jiang, Shuang Wu 0001, Shouhong Ding
Knowl. Based Syst.4
2024 Re-Thinking Data Availability Attacks Against Deep Neural Networks
abstract
The unauthorized use of personal data for commercial purposes and the covert acquisition of private data for training machine learning models continue to raise concerns. To address these issues, researchers have proposed availability attacks that aim to render data unexploitable. However, many availability attack methods can be easily disrupted by adversarial training. Although some robust methods can resist adversarial training, their protective effects are limited. In this paper, we re-examine the existing availability attack methods and propose a novel two-stage min-max-min optimization paradigm to generate robust unlearnable noise. The inner min stage is utilized to generate unlearnable noise, while the outer min-max stage simulates the training process of the poisoned model. Additionally, we formulate the attack effects and use it to constrain the optimization objective. Comprehensive experiments have revealed that the noise generated by our method can lead to a decline in test accuracy for adversarially trained poisoned models by up to approximately 30%, in comparison to SOTA methods.11Code is available at EuterpeK/Rethinking-Data-Availability-Attacks
Bin Fang 0009, Bo Li 0115, Shuang Wu 0001, Shouhong Ding, Ran Yi 0002, Lizhuang Ma
CVPR3
2024 MFAE: Masked Frequency Autoencoders for Domain Generalization Face Anti-Spoofing
abstract
The generalizable face anti-spoofing (FAS) has attracted much attention recently. Even though many existing methods perform well under intra-domain settings, the model’s performance in the unseen domain is not satisfying. In this paper, we shift our attention to the frequency domain to seek a solution. Specifically, we examine the characteristics of different frequency band components of FAS images and observe that the model’s cross-domain performance is very sensitive to low-frequency features. To alleviate this sensitivity and improve the model’s performance in FAS cross-domain tasks, we propose a new approach called Masked Frequency Autoencoders (MFAE). MFAE randomly masks a portion of frequencies on the low-frequency spectrum of the image and then reconstructs the image from the resulting embedding. This innovative Masked Image Modeling (MIM) strategy can be used as a self-supervised task for pre-training vision transformers (ViTs), which can reduce the ViT encoder’s sensitivity to domain shifts. Additionally, we add an auxiliary content-regularization decoder in our MFAE to encourage the encoder to be insensitive to low-frequency features. The results show that the model insensitive to low-frequency features performs well on extensive public datasets and outperforms other state-of-the-art methods in cross-domain FAS tasks.
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
IEEE Trans. Inf. Forensics Secur.3
2023 Delving into the Adversarial Robustness of Federated Learning
abstract
In Federated Learning (FL), models are as fragile as centrally trained models against adversarial examples. However, the adversarial robustness of federated learning remains largely unexplored. This paper casts light on the challenge of adversarial robustness of federated learning. To facilitate a better understanding of the adversarial vulnerability of the existing FL methods, we conduct comprehensive robustness evaluations on various attacks and adversarial training methods. Moreover, we reveal the negative impacts induced by directly adopting adversarial training in FL, which seriously hurts the test accuracy, especially in non-IID settings. In this work, we propose a novel algorithm called Decision Boundary based Federated Adversarial Training (DBFAT), which consists of two components (local re-weighting and global regularization) to improve both accuracy and robustness of FL systems. Extensive experiments on multiple datasets demonstrate that DBFAT consistently outperforms other baselines under both IID and non-IID settings.
Jie Zhang 0081, Bo Li 0115, Chen Chen 0043, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001
AAAI5
2023 Attack Can Benefit: An Adversarial Approach to Recognizing Facial Expressions under Noisy Annotations
abstract
The real-world Facial Expression Recognition (FER) datasets usually exhibit complex scenarios with coupled noise annotations and imbalanced classes distribution, which undoubtedly impede the development of FER methods. To address the aforementioned issues, in this paper, we propose a novel and flexible method to spot noisy labels by leveraging adversarial attack, termed as Geometry Aware Adversarial Vulnerability Estimation (GAAVE). Different from existing state-of-the-art methods of noisy label learning (NLL), our method has no reliance on additional information and is thus easy to generalize to the large-scale real-world FER datasets. Besides, the combination of Dataset Splitting module and Subset Refactoring module mitigates the impact of class imbalance, and the Self-Annotator module facilitates the sufficient use of all training data. Extensive experiments on RAF-DB, FERPlus, AffectNet, and CIFAR-10 datasets validate the effectiveness of our method. The stabilized enhancement based on different methods demonstrates the flexibility of our proposed GAAVE.
Jiawen Zheng, Bo Li 0115, Shengchuan Zhang, Shuang Wu 0001, Liujuan Cao, Shouhong Ding
AAAI4
2023 Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition
abstract
Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the imbalance of short- and long-term temporal relationships in DFER. Therefore, we introduce the Multi-3D Dynamic Facial Expression Learning (M3DFEL) framework, which utilizes Multi-Instance Learning (MIL) to handle inexact labels. M3DFEL generates 3D-instances to model the strong short-term temporal relationship and utilizes 3DCNNs for feature extraction. The Dynamic Long-term Instance Aggregation Module (DLIAM) is then utilized to learn the long-term temporal relationships and dynamically aggregate the instances. Our experiments on DFEW and FERV39K datasets show that M3DFEL outperforms existing state-of-the-art approaches with a vanilla R3D18 backbone. The source code is available at https://github.com/faceeyes/M3DFEL.
Hanyang Wang 0001, Bo Li 0115, Shuang Wu 0001, Feng Liu 0039, Shouhong Ding, Aimin Zhou
CVPR3
2023 Content-based Unrestricted Adversarial Attack
abstract
Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. However, current works usually sacrifice unrestricted degrees and subjectively select some image content to guarantee the photorealism of unrestricted adversarial examples, which limits its attack performance. To ensure the photorealism of adversarial examples and boost attack performance, we propose a novel unrestricted attack framework called Content-based Unrestricted Adversarial Attack. By leveraging a low-dimensional manifold that represents natural images, we map the images onto the manifold and optimize them along its adversarial direction. Therefore, within this framework, we implement Adversarial Content Attack (ACA) based on Stable Diffusion and can generate high transferable unrestricted adversarial examples with various adversarial contents. Extensive experimentation and visualization demonstrate the efficacy of ACA, particularly in surpassing state-of-the-art attacks by an average of 13.3-50.4\% and 16.8-48.0\% in normally trained models and defense methods, respectively.
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Kaixun Jiang, Shouhong Ding
NeurIPS3
2023 Query-Efficient Decision-Based Black-Box Patch Attack
abstract
Deep neural networks (DNNs) have been showed to be highly vulnerable to imperceptible adversarial perturbations. As a complementary type of adversary, patch attacks that introduce perceptible perturbations to the images have attracted the interest of researchers. Existing patch attacks rely on the architecture of the model or the probabilities of predictions and perform poorly in the decision-based setting, which can still construct a perturbation with the minimal information exposed – the top-1 predicted label. In this work, we first explore the decision-based patch attack. To enhance the attack efficiency, we model the patches using paired key-points and use targeted images as the initialization of patches, and parameter optimizations are all performed on the integer domain. Then, we propose a differential evolutionary algorithm named DevoPatch for query-efficient decision-based patch attacks. Experiments demonstrate that DevoPatch outperforms the state-of-the-art black-box patch attacks in terms of patch area and attack success rate within a given query budget on image classification and face verification. Additionally, we conduct the vulnerability evaluation of ViT and MLP on image classification in the decision-based patch attack setting for the first time. Using DevoPatch, we can evaluate the robustness of models to black-box patch attacks. We believe this method could inspire the design and deployment of robust vision models based on various DNN architectures in the future.
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Shouhong Ding
IEEE Trans. Inf. Forensics Secur.3
2022 Towards Efficient Data Free Blackbox Adversarial Attack
abstract
Classic black-box adversarial attacks can take advantage of transferable adversarial examples generated by a similar substitute model to successfully fool the target model. However, these substitute models need to be trained by target models' training data, which is hard to acquire due to privacy or transmission reasons. Recognizing the limited availability of real data for adversarial queries, recent works proposed to train substitute models in a data-free black-box scenario. However, their generative adversarial networks (GANs) based framework suffers from the convergence failure and the model collapse, resulting in low efficiency. In this paper, by rethinking the collaborative relationship between the generator and the substitute model, we design a novel black-box attack framework. The proposed method can efficiently imitate the target model through a small number of queries and achieve high attack success rate. The comprehensive experiments over six datasets demonstrate the effectiveness of our method against the state-of-the-art attacks. Especially, we conduct both label-only and probability-only attacks on the Microsoft Azure online model, and achieve a 100% attack success rate with only 0.46% query budget of the SOTA method [49].
Jie Zhang 0081, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Lei Zhang 0197, Chao Wu 0001
CVPR4
2022 Towards Practical Certifiable Patch Defense with Vision Transformer
abstract
Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks. Existing certifiable patch defenses sacrifice the clean accuracy of classifiers and only obtain a low certified accuracy on toy datasets. Furthermore, the clean and certified accuracy of these methods is still significantly lower than the accuracy of normal classification networks, which limits their application in practice. To move towards a practical certifiable patch defense, we introduce Vision Transformer (ViT) into the framework of Derandomized Smoothing (DS). Specifically, we propose a progressive smoothed image modeling task to train Vision Transformer, which can capture the more discriminable local context of an image while preserving the global semantic information. For efficient inference and deployment in the real world, we innovatively reconstruct the global self-attention structure of the original ViT into isolated band unit self-attention. On ImageNet, under 2% area patch attacks our method achieves 41.70% certified accuracy, a nearly 1-fold increase over the previous best method (26.00%). Simultaneously, our method achieves 78.58% clean accuracy, which is quite close to the normal ResNet-101 accuracy. Extensive experiments show that our method obtains state-of-the-art clean and certified accuracy with inferring efficiently on CIFAR-10 and ImageNet.
Zhaoyu Chen 0001, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding
CVPR4
2022 Detecting Camouflaged Object in Frequency Domain
abstract
Camouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception ability of human eyes. Hence, we claim that the goal of COD task is not just to mimic the human visual ability in a single RGB domain, but to go beyond the human biological vision. We then introduce the frequency domain as an additional clue to better detect camouflaged objects from backgrounds. To well involve the frequency clues into the CNN models, we present a powerful network with two special components. We first design a novel frequency enhancement module (FEM) to dig clues of camouflaged objects in the frequency domain. It contains the offline discrete cosine transform followed by the learnable enhancement. Then we use a feature alignment to fuse the features from RGB domain and frequency domain. Moreover, to further make full use of the frequency information, we propose the high-order relation module (HOR) to handle the rich fusion feature. Comprehensive experiments on three widely-used COD datasets show the proposed method significantly outperforms other state-of-the-art methods by a large margin.
Yijie Zhong 0001, Bo Li 0115, Lv Tang, Senyun Kuang, Shuang Wu 0001, Shouhong Ding
CVPR5
2022 Shape Matters: Deformable Patch Attack
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Jianghe Xu, Shouhong Ding
ECCV (4)3
2022 Federated Learning with Label Distribution Skew via Logits Calibration
abstract
Traditional federated optimization methods perform poorly with heterogeneous data (i.e. , accuracy reduction), especially for highly skewed data. In this paper, we investigate the label distribution skew in FL, where the distribution of labels varies across clients. First, we investigate the label distribution skew from a statistical view. We demonstrate both theoretically and empirically that previous methods based on softmax cross-entropy are not suitable, which can result in local models heavily overfitting to minority classes and missing classes. Additionally, we theoretically introduce a deviation bound to measure the deviation of the gradient after local update. At last, we propose FedLC (\textbf{Fed}erated learning via \textbf{L}ogits \textbf{C}alibration), which calibrates the logits before softmax cross-entropy according to the probability of occurrence of each class. FedLC applies a fine-grained calibrated cross-entropy loss to local update by adding a pairwise label margin. Extensive experiments on federated datasets and real-world datasets demonstrate that FedLC leads to a more accurate global model and much improved performance. Furthermore, integrating other FL methods into our approach can further enhance the performance of the global model.
Jie Zhang 0081, Zhiqi Li 0004, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001
ICML5
2022 DENSE: Data-Free One-Shot Federated Learning
abstract
One-shot Federated Learning (FL) has recently emerged as a promising approach, which allows the central server to learn a model in a single communication round. Despite the low communication cost, existing one-shot FL methods are mostly impractical or face inherent limitations, \eg a public dataset is required, clients' models are homogeneous, and additional data/model information need to be uploaded. To overcome these issues, we propose a novel two-stage \textbf{D}ata-fre\textbf{E} o\textbf{N}e-\textbf{S}hot federated l\textbf{E}arning (DENSE) framework, which trains the global model by a data generation stage and a model distillation stage. DENSE is a practical one-shot FL method that can be applied in reality due to the following advantages:(1) DENSE requires no additional information compared with other methods (except the model parameters) to be transferred between clients and the server;(2) DENSE does not require any auxiliary dataset for training;(3) DENSE considers model heterogeneity in FL, \ie different clients can have different model architectures.Experiments on a variety of real-world datasets demonstrate the superiority of our method.For example, DENSE outperforms the best baseline method Fed-ADI by 5.08\% on CIFAR10 dataset.
Jie Zhang 0081, Chen Chen 0043, Bo Li 0115, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chunhua Shen, Chao Wu 0001
NeurIPS5
2021 Robust and Efficient Graph Correspondence Transfer for Person Re-Identification
abstract
Spatial misalignment caused by variations in poses and viewpoints is one of the most critical issues that hinder the performance improvement in existing person re-identification (Re-ID) algorithms. Although it is straightforward to explore correspondence learning algorithms for alignment, online learning is intractable for negative pairs due to the intrinsic visual difference between negative pairs and efficiency concern. To address this problem, in this paper, we present a robust and efficient graph correspondence transfer (REGCT) approach for explicit spatial alignment in Re-ID. Specifically, we propose the off-line correspondence learning and on-line correspondence transfer framework. During training, patch-wise correspondences between positive training pairs are established via graph matching. By exploiting both spatial and visual contexts of human appearance in graph matching, meaningful semantic correspondences can be obtained. During testing, the off-line learned patch-wise correspondence templates are transferred to test pairs with similar pose-pair configurations for local feature distance calculation. To enhance the robustness of correspondence transfer, we design a novel pose context descriptor to accurately model human body configurations, and present an approach to measure the similarity between a pair of pose context descriptors. Meanwhile, to improve testing efficiency, we propose a correspondence template ensemble method using the voting mechanism, which significantly reduces the amount of patch-wise matchings involved in distance calculation. With the aforementioned strategies, the REGCT model can effectively and efficiently handle the spatial misalignment problem in Re-ID. Extensive experiments on five challenging benchmarks, including VIPeR, Road, PRID450S, 3DPES, and CUHK01, evidence the superior performance of REGCT over other state-of-the-art approaches.
Qin Zhou 0002, Heng Fan 0001, Hua Yang 0001, Hang Su 0006, Shibao Zheng, Shuang Wu 0001, Haibin Ling
IEEE Trans. Image Process.6
2018 Graph Correspondence Transfer for Person Re-Identification
abstract
In this paper, we propose a graph correspondence transfer (GCT) approach for person re-identification. Unlike existing methods, the GCT model formulates person re-identification as an off-line graph matching and on-line correspondence transferring problem. In specific, during training, the GCT model aims to learn off-line a set of correspondence templates from positive training pairs with various pose-pair configurations via patch-wise graph matching. During testing, for each pair of test samples, we select a few training pairs with the most similar pose-pair configurations as references, and transfer the correspondences of these references to test pair for feature distance calculation. The matching score is derived by aggregating distances from different references. For each probe image, the gallery image with the highest matching score is the re-identifying result. Compared to existing algorithms, our GCT can handle spatial misalignment caused by large variations in view angles and human poses owing to the benefits of patch-wise graph matching. Extensive experiments on five benchmarks including VIPeR, Road, PRID450S, 3DPES and CUHK01 evidence the superior performance of GCT model over other state-of-the-art methods.
Qin Zhou 0002, Heng Fan 0001, Shibao Zheng, Hang Su 0006, Xinzhe Li 0002, Shuang Wu 0001, Haibin Ling
AAAI6
2017 Data Generation for Improving Person Re-identification
abstract
In this paper, we explore ways to address the challenges such as data bias caused by the lack of data on person re-identification problem. We propose a data generation framework from both intra- and inter-view aspects for data augmentation to advance the performance of the existing person re-identification algorithms. Specifically, for intra-view data generation, the proposed method generates useful predicted sequences within a camera view for certain person data expansion. The generated sequences well preserve the movement information of the camera and objects, which expands the original data with longer sequence length to tackle the problem caused by insufficient data from the root. For more challenging datasets which suffer from background clutters, we propose an inter-view image generation with automatic end-to-end background substitution to eliminate the influence by the background and increase the diversity of the training data as well, which makes the recognition system learn to focus on the regions of objects and image features related to identity. We then propose a flexible data augmentation method based on our data generation approaches to improve the performance of the person re-identification and analyze the advantages and applicability of these approaches respectively. Evaluated on the challenging re-id datasets, our method outperforms existing state-of-the-art approaches without any network structure modification on the baseline neural network. Cross-datasets evaluation results show that our method has favorable generalization ability and is potentially helpful for solving similar recognition tasks due to the common issue of insufficient data.
Lin Chen 0019, Hua Yang 0001, Shuang Wu 0001
ACM Multimedia3
2017 Crowd Behavior Analysis via Curl and Divergence of Motion Trajectories
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Yawen Fan, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.1
2017 Bilinear dynamics for crowd video analysis
Shuang Wu 0001, Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Qin Zhou 0002
J. Vis. Commun. Image Represent.1
2017 Motion sketch based crowd video retrieval
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Qin Zhou 0002
Multim. Tools Appl.1
2017 Joint dictionary and metric learning for person re-identification
Qin Zhou 0002, Shibao Zheng, Haibin Ling, Hang Su 0006, Shuang Wu 0001
Pattern Recognit.5
2016 Crowd semantic segmentation based on spatial-temporal dynamics
abstract
Crowd semantic segmentation is supposed to not only accurately segment the crowd into groups but also describe them by semantic properties. We define a group as a set of members sharing common spatial-temporal dynamics, i.e., motion consistency and distribution homogeneity. This paper proposes a novel crowd semantic segmentation method, termed as joint spatial-temporal semantic segmentation, which leverages the temporal motion characteristics and spatial distribution information of crowd. We first conduct temporal motion grouping and spatial distribution grouping according to motion consistency and distribution homogeneity respectively. Then, a a joint semantic segmentation algorithm is employed to combine the motion and distribution groups into semantic groups. States of these groups are described in terms of motion pattern and density level. Experiments show that our proposed method is effective to obtain favorable segmentation with semantic descriptions.
Jijia Li, Hua Yang 0001, Shuang Wu 0001
AVSS3
2016 Motion sketch based crowd video retrieval via motion structure coding
abstract
Crowd video retrieval is an important problem in surveillance video management in the era of big data, e.g., video indexing and browsing. In this paper, we address this issue from the motion-level perspective by using hand-drawn sketches as queries. Motion sketch based crowd video retrieval naturally suffers from challenges in motion-level video indexing and sketch representation. We tackle them by leveraging the motion structure coding algorithm to extract robust structure-preserved motion descriptors. For video indexing, we use motion decomposition to separate the sub-motion vector fields with typical patterns from a set of optical flows. Then, the motion-level descriptors of the vector fields are computed and stored in the index database. To represent sketch queries, we propose a sketch vectorization algorithm followed by motion structure coding. In the retrieval stage, given a new query, the retrieval function learned by the Ranking SVM algorithm predicts the ranking score of each motion pattern in the index database. Extensive experiments are conducted on the publicly available crowd datasets, which demonstrate the robustness and effectiveness of the proposed sketch based crowd video retrieval system.
Shuang Wu 0001, Hang Su 0006, Shibao Zheng, Hua Yang 0001, Qin Zhou 0002
ICIP1
2015 Kernelized View Adaptive Subspace Learning for Person Re-identification
Qin Zhou 0002, Shibao Zheng, Hang Su 0006, Hua Yang 0001, Shuang Wu 0001
BMVC6
2015 Towards active annotation for detection of numerous and scattered objects
abstract
Object detection is an active study area in the field of computer vision and image understanding. In this paper, we propose an active annotation algorithm by addressing the detection of numerous and scattered objects in a view, e.g., hundreds of cells in microscopy images. In particular, object detection is implemented by classifying pixels into specific classes with graph-based semi-supervised learning and grouping neighboring pixels with the same label. Sample or seed selection is conducted based on a novel annotation criterion that minimizes the expected prediction error. The most informative samples are therefore annotated actively, which are subsequently propagated to the unlabeled samples via a pairwise affinity graph. Experimental results conducted on two real world datasets validate that our proposed scheme quickly reaches high quality results and reduces human efforts significantly.
Hang Su 0006, Hua Yang 0001, Shibao Zheng, Sha Wei, Shuang Wu 0001
ICME6
2015 Hierarchical video summarization with loitering indication
abstract
In this paper, a hierarchical and informative summarization framework is proposed, which facilitates rapid video browsing. Moreover, a method for loitering detection is exploited to indicate potential abnormal behaviors. The hierarchical framework includes two levels: a holistic-level and an object-level. The holistic-level summarization provides viewers with a comprehensive and compact representation of the original video, while the object-level summarization extracts the narrative information of each object, including trajectory, direction, time, changes of appearance and indication of the loitering behavior. The two summarizations are formulated as two different energy minimization problems, which are solved by the proposed heuristic algorithms. Our framework is evaluated on two publicly available datasets. Experimental results demonstrate that the proposed method performs favourably in providing holistic- and object-level information, fast browsing, and loitering detection.
Ruipeng Lu, Hua Yang 0001, Ji Zhu 0002, Shuang Wu 0001, Jia Wang 0004, David Bull 0001
VCIP4