Shibao Zheng

dblp:20/1917 · DBLP profile ↗
← Back
62ranked-venue papers
0as first author
28since 2021 · last 2025
0000-0002-5060-0210ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 since 2021Artificial intelligence and machine learning · 27 · 18 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Assessing Robustness of Multi-Modal Large Language Models in Image Classification through Hierarchical WordNet-Based Evaluation
abstract
The advancement of multi-modal large language models (MLLMs) has significantly enhanced their capability to process and understand diverse data types, integrating text, images, and other modalities. Despite their impressive performance, evaluating the robustness of these models remains challenging due to the difficulty of aligning their text-based responses with image classification labels. Traditional approaches rely on CLIP scores or other large language models as judges, but these methods lack scientific rigor and fail to capture robustness across different semantic levels. In this paper, we propose a novel evaluation metric that systematically assesses the robustness of MLLMs in image classification using WordNet’s hierarchical structure. Specifically, we parse the text descriptions generated by MLLMs to extract all nouns, then calculate their distances to the groundtruth label in WordNet as a similarity metric. The minimum distance among all nouns is used as the similarity score between the text description and the label. By using WordNet, we can also evaluate classification performance at different semantic levels. Through extensive experiments, we demonstrate that our WordNet-based evaluation metric offers a deeper understanding of MLLMs’ robustness, paving the way for more resilient and reliable models in real-world applications.
Chang Liu 0077, Hai Chen, Shibao Zheng
ICASSP4
2025 Multi-Style Facial Sketch Synthesis through Masked Generative Modeling
abstract
The facial sketch synthesis (FSS) model, capable of generating sketch portraits from given facial photographs, holds profound implications across multiple domains, encompassing cross-modal face recognition, entertainment, art, media, among others. However, the production of high-quality sketches remains a formidable task, primarily due to the challenges and flaws associated with three key factors: (1) the scarcity of artist-drawn data, (2) the constraints imposed by limited style types, and (3) the deficiencies of processing input information in existing models. To address these difficulties, we propose a lightweight end-to-end synthesis model that efficiently converts images to corresponding multi-stylized sketches, obviating the necessity for any supplementary inputs (e.g., 3D geometry). In this study, we overcome the issue of data insufficiency by incorporating semi-supervised learning into the training process. Additionally, we employ a feature extraction module and style embeddings to proficiently steer the generative transformer during the iterative prediction of masked image tokens, thus achieving a continuous stylized output that retains facial features accurately in sketches. The extensive experiments demonstrate that our method consistently outperforms previous algorithms across multiple benchmarks, exhibiting a discernible disparity.
Guo Lu, Shibao Zheng
ICASSP3
2025 A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking
Chang Liu 0077, Yinpeng Dong, Wenzhao Xiang 0001, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001, Yuefeng Chen, Yuan He 0011, Hui Xue 0001, Shibao Zheng
Int. J. Comput. Vis.10
2025 Improving model generalization by on-manifold adversarial augmentation in the frequency domain
abstract
Deep Neural Networks (DNNs) often suffer from performance drops when training and test data distributions differ. Ensuring model generalization for Out-Of-Distribution (OOD) data is crucial, but current models still struggle with accuracy on such data. Recent studies have shown that regular or off-manifold adversarial examples as data augmentation improve OOD generalization. Building on this, we provide theoretical validation that on-manifold adversarial examples can enhance OOD generalization even more. However, generating these examples is challenging due to the complexity of real manifolds. To address this, we propose AdvWavAug, an on-manifold adversarial data augmentation method using a Wavelet module. This approach, based on the AdvProp training framework, leverages wavelet transformation to project an image into the wavelet domain and modifies it within the estimated data manifold. Experiments on various models and datasets, including ImageNet and its distorted versions, show that our method significantly improves model generalization, especially for OOD data.
Chang Liu 0077, Wenzhao Xiang 0001, Yuan He 0011, Hui Xue 0001, Shibao Zheng, Hang Su 0006
J. Vis. Commun. Image Represent.5
2025 RobustPrompt: Learning to defend against adversarial attacks with adaptive visual prompts
Chang Liu 0077, Wenzhao Xiang 0001, Yinpeng Dong, Xingxing Zhang 0001, Ranjie Duan, Shibao Zheng, Hang Su 0006
Pattern Recognit. Lett.7
2025 CamoEnv: Transferable and environment-consistent adversarial camouflage in autonomous driving
Xiao Yang 0028, Hang Su 0006, Shibao Zheng
Pattern Recognit. Lett.4
2025 ImAdv: Transferable Implicit Adversarial Attack for 3D Object Detectors in Autonomous Driving
abstract
3D adversarial attacks have garnered significant attention in the realm of autonomous driving security due to their high feasibility and multi-view effectiveness. However, existing 3D attacks have limited transferability, primarily due to their overfitting to surrogate models. To address this limitation, we introduce a novel 3D adversarial attack method based on implicit texture modeling, termed ImAdv, against 3D object detection models. Specifically, ImAdv utilizes a positional encoder and a MLP to map the 3D coordinates of an object's surface to the RGB color space, thereby reformulating the object's texture within an implicit framework. This method significantly reduces the parameter number for color modeling, thus mitigating overfitting and improving transferability. Furthermore, we propose two innovative techniques to enhance the transferability, Random Texture Reset (RandReset) and Texture Model Averaging. RandReset randomly restores portions of the adversarial texture, increasing the training set diversity and mitigating overfitting. Texture Model Averaging employs self-ensembling of multiple texture checkpoints during the training phase to reduce overfitting in the final texture model. Comprehensive experiments demonstrate the superiority of our methods, which outperform previous methods by 17.18% in average black-box attack success rate. Additionally, our method shows strong transferability and practicality in zero-shot cross-task attacks and physical attacks.
Xiao Yang 0028, Hang Su 0006, Shu Zhao 0005, Shibao Zheng
IEEE Trans. Big Data5
2025 DiFace: Cross-Modal Face Recognition through Controlled Diffusion
abstract
Diffusion probabilistic models (DPMs) have exhibited exceptional proficiency in generating visual media of outstanding quality and realism. Nonetheless, their potential in non-generative domains, such as face recognition (FR), has yet to be thoroughly investigated. Meanwhile, despite the extensive development of multi-modal FR methods, their emphasis has predominantly centered on visual modalities. In this context, FR through textual description presents a unique and promising solution that not only transcends the limitations from application scenarios but also expands the potential for research in the field of cross-modal FR. It is regrettable that this avenue remains underutilized, a consequence from the challenges mainly associated with three aspects: 1) the intrinsic imprecision of verbal descriptions; 2) the significant gaps between texts and images; and 3) the immense hurdle posed by insufficient databases. To tackle this problem, we present DiFace, an end-to-end solution that effectively achieves FR via text through a controllable diffusion process, by establishing its theoretical connection with probability transport. Our approach not only unleashes the potential of DPMs across a wide range of tasks but also showcases a remarkable improvement in accuracy for text-image FR, as evidenced by our experiments on verification and identification.
Guo Lu, Shibao Zheng
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
abstract
Although pre-trained models such as Contrastive Language-Image Pre-Training (CLIP) show impressive generalization results, their robustness is still limited under Out-of-Distribution (OOD) scenarios. Instead of undesirably leveraging human annotation as commonly done, it is possible to leverage the visual understanding power of Multi-modal Large Language Models (MLLMs). However, MLLMs struggle with vision problems due to task incompatibility, thus hindering their effectiveness. In this paper, we propose to effectively leverage MLLMs via Machine Vision Therapy which aims to rectify erroneous predictions of specific vision models. By supervising vision models using MLLM predictions, visual robustness can be boosted in a nearly unsupervised manner. Moreover, we propose a Denoising In-Context Learning (DICL) strategy to solve the incompatibility issue. Concretely, by examining the noise probability of each example through a transition matrix, we construct an instruction containing a correct exemplar and a probable erroneous one, which enables MLLMs to detect and rectify the incorrect predictions of vision models. Under mild assumptions, we theoretically show that our DICL method is guaranteed to find the ground truth. Through extensive experiments on various OOD datasets, our method demonstrates powerful capabilities for enhancing visual robustness under many OOD scenarios.
Chang Liu 0077, Yinpeng Dong, Hang Su 0006, Shibao Zheng, Tongliang Liu
ICML5
2024 FreeFlow: A Unified Viewpoint on Diffusion Probabilistic Models via Optimal Transport and Fluid Mechanics
Guo Lu, Shibao Zheng
ICONIP (3)3
2024 Black-box attacks on face recognition via affine-invariant training
Hang Su 0006, Shibao Zheng
Neural Comput. Appl.3
2023 Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving
abstract
3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular benchmarks. However, there still lacks a systematic understanding of the robustness of these vision-dependent BEV models, which is closely related to the safety of autonomous driving systems. In this paper, we evaluate the natural and adversarial robustness of various representative models under extensive settings, to fully understand their behaviors influenced by explicit BEV features compared with those without BEV. In addition to the classic settings, we propose a 3D consistent patch attack by applying adversarial patches in the 3D space to guarantee the spatiotemporal consistency, which is more realistic for the scenario of autonomous driving. With substantial experiments, we draw several findings: 1) BEV models tend to be more stable than previous methods under different natural conditions and common corruptions due to the expressive spatial representations; 2) BEV models are more vulnerable to adversarial noises, mainly caused by the redundant BEV features; 3) Camera-LiDARfusion models have superior performance under different settings with multi-modal inputs, but BEV fusion model is still vulnerable to adversarial noises of both point cloud and image. These findings alert the safety issue in the applications of BEV detectors and could facilitate the development of more robust models.
Yichi Zhang 0012, Hai Chen, Yinpeng Dong, Shu Zhao 0005, Wenbo Ding 0004, Jiachen Zhong, Shibao Zheng
CVPR8
2023 Hierarchical Semantic Perceptual Listener Head Video Generation: A High-performance Pipeline
abstract
In dyadic speaker-listener interactions, the listener's head reactions, together with the speaker's head movements, form an important non-verbal semantic expression. The listener Head generation task aims to synthesize the responsive listener's head videos based on audios of the speaker and reference images of the listener. Compared to the talking-head generation, it is more challenging to capture the correlation cues from the speaker's audio and visual information. Following the ViCo baseline scheme, we propose a high-performance solution by enhancing the hierarchical semantic extraction capability of the audio encoder module and improving the decoder part, renderer and post-processing modules. Our solution gets the first place on the official leaderboard for the track of listening head generation. This paper is a technical report of ViCo@2023 Conversational Head Generation Challenge at the ACM Multimedia 2023 conference.
Zhigang Chang, Weitai Hu, Qing Yang 0033, Shibao Zheng
ACM Multimedia4
2023 Improving the robustness of adversarial attacks using an affine-invariant gradient estimator
Wenzhao Xiang 0001, Hang Su 0006, Chang Liu 0077, Yandong Guo, Shibao Zheng
Comput. Vis. Image Underst.5
2023 To make yourself invisible with Adversarial Semantic Contours
Yichi Zhang 0012, Hang Su 0006, Jun Zhu 0001, Shibao Zheng, Yuan He 0011, Hui Xue 0001
Comput. Vis. Image Underst.5
2023 Learning comprehensive global features in person re-identification: Ensuring discriminativeness of more local regions
Jiali Xi, Jianqiang Huang 0001, Shibao Zheng, Qin Zhou 0002, Bernt Schiele, Xian-Sheng Hua 0001, Qianru Sun
Pattern Recognit.3
2022 Subtask-dominated Supervised Pretraining Transfer Learning for Person Search
Hua Yang 0001, Shibao Zheng
BMVC3
2022 A Morphological Fingerprint Minutiae Annotation Algorithm for Deep Learning Datasets
abstract
Minutiae extraction is vital for automatic finger-print identification system. Fingerprint datasets (databases) for minutiae extraction are beneficial to improve the performance of subsequent indexing and matching tasks, particularly combined with advanced deep learning technique. However, the learning-based methods are restricted by the limitation of public labeled datasets due to legitimate privacy concerns. To thoroughly address these issues (composed of lacking of the available databases and general annotation technique), a novel fingerprint image labeling approach is first focused on in this paper. Specifically, we devise a knowledge-oriented minutiae annotation method, which consists of image segmentation, normalization, orientation and frequency estimations, enhancement, binarization, thining, minutiae extraction, deleting redundancy and then revising the wrong annotations manually. Next, a minutiae database is created via the approach, and then experiments verify its effectiveness for training deep learning model. Comprehensive experimental results on two datasets manifest the ability of the constructed database to improve the prediction accuracy of CNN (achieving precision and recall up to 89.97% and 88.38%, respectively) compared with public software and other similar deep learning extraction model which in turn reveals the effectiveness of our method.
Hongtian Zhao, Shibao Zheng
ISCAS2
2022 Modeling context appearance changes for person re-identification via IPES-GCN
Hua Yang 0001, Ji Zhu 0002, Qin Zhou 0002, Shibao Zheng
Neurocomputing5
2022 Momentum source-proxy guided initialization for unsupervised domain adaptive person re-identification
Jiali Xi, Qin Zhou 0002, Xinzhe Li 0002, Shibao Zheng
Neurocomputing4
2022 Making person search enjoy the merits of person re-identification
Hua Yang 0001, Qin Zhou 0002, Shibao Zheng
Pattern Recognit.4
2021 Mask Region-oriented Diabetic Retinopathies Detection in Ophthalmic Medical Images via Non-local Attention
abstract
Accurate lesions detection on retinopathy images is crucial for the diagnosis of diabetes. However, it is always hampered by various characteristics of lesions such as shape, color, texture, and similarities. Most advanced algorithms still cannot automatically detect common lesions, e.g. exudate, hemorrhage, and cotton-wool spots, being used for comprehensive analysis of disease state. To this end, we present a multi-functional detection model for diabetic retinopathies and further analyze disease mechanisms overall. Specifically, this paper attempts to implement a multi-lesion detector via modified Mask region-oriented CNN, which can be used for the aforementioned retinopathies. Meanwhile, a non-local attention module is introduced to the detector as a spatial attention mechanism for handling the global information missing problem. In addition, to boost the effectiveness of the detector, the dilated operation is implemented for dataset preprocessing. Improvement is achieved both algorithmically and architecturally, via investigating thoroughly the most probable lesion category with a novel ensemble learning framework. Extensive experiments on standard datasets for three different tasks evidence the superior performance of the proposed method over state-of-the-art methods.
Hongtian Zhao, Haoyue Peng, Zhongpai Gao, Shibao Zheng
IJCNN4
2021 Seq-Masks: Bridging the gap between appearance and gait modeling for video-based person re-identification
abstract
Video-based person re-identification (Re-ID) aims to match person images in video sequences captured by disjoint surveillance cameras. Traditional video-based person Re-ID methods focus on exploring appearance information, thus, vulnerable against illumination changes, scene noises, camera parameters, and especially clothes/carrying variations. Gait recognition provides an implicit biometric solution to alleviate the above headache. Nonetheless, it experiences severe performance degeneration as camera view varies. In an attempt to address these problems, in this paper, we propose a framework that utilizes the sequence masks (SeqMasks) in the video to integrate appearance information and gait modeling in a close fashion. Specifically, to sufficiently validate the effectiveness of our method, we build a novel dataset named MaskMARS based on MARS. Comprehensive experiments on our proposed large wild video Re-ID dataset MaskMARS evidenced our extraordinary performance and generalization capability. Validations on the gait recognition metric CASIA-B dataset further demonstrated the capability of our hybrid model. Our codes and dataset MaskMARS will be open-sourced as a strong baseline.
Zhigang Chang, Yongbiao Chen, Shibao Zheng
VCIP5
2021 Learning to teach and learn for semi-supervised few-shot image classification
Xinzhe Li 0002, Jianqiang Huang 0001, Yaoyao Liu 0001, Qin Zhou 0002, Shibao Zheng, Bernt Schiele, Qianru Sun
Comput. Vis. Image Underst.5
2021 Graph similarity rectification for person search
Hua Yang 0001, Ji Zhu 0002, Xinzhe Li 0002, Zhigang Chang, Shibao Zheng
Neurocomputing6
2021 DeepSDM: Boundary-aware pneumothorax segmentation in chest X-ray images
Xueqing Peng, Lili Shi, Shibao Zheng, Weiya Shi
Neurocomputing6
2021 Gradient-based conditional generative adversarial network for non-uniform blind deblurring via DenseResNet
Hongtian Zhao, Hang Su 0006, Shibao Zheng
J. Vis. Commun. Image Represent.4
2021 Robust and Efficient Graph Correspondence Transfer for Person Re-Identification
abstract
Spatial misalignment caused by variations in poses and viewpoints is one of the most critical issues that hinder the performance improvement in existing person re-identification (Re-ID) algorithms. Although it is straightforward to explore correspondence learning algorithms for alignment, online learning is intractable for negative pairs due to the intrinsic visual difference between negative pairs and efficiency concern. To address this problem, in this paper, we present a robust and efficient graph correspondence transfer (REGCT) approach for explicit spatial alignment in Re-ID. Specifically, we propose the off-line correspondence learning and on-line correspondence transfer framework. During training, patch-wise correspondences between positive training pairs are established via graph matching. By exploiting both spatial and visual contexts of human appearance in graph matching, meaningful semantic correspondences can be obtained. During testing, the off-line learned patch-wise correspondence templates are transferred to test pairs with similar pose-pair configurations for local feature distance calculation. To enhance the robustness of correspondence transfer, we design a novel pose context descriptor to accurately model human body configurations, and present an approach to measure the similarity between a pair of pose context descriptors. Meanwhile, to improve testing efficiency, we propose a correspondence template ensemble method using the voting mechanism, which significantly reduces the amount of patch-wise matchings involved in distance calculation. With the aforementioned strategies, the REGCT model can effectively and efficiently handle the spatial misalignment problem in Re-ID. Extensive experiments on five challenging benchmarks, including VIPeR, Road, PRID450S, 3DPES, and CUHK01, evidence the superior performance of REGCT over other state-of-the-art approaches.
Qin Zhou 0002, Heng Fan 0001, Hua Yang 0001, Hang Su 0006, Shibao Zheng, Shuang Wu 0001, Haibin Ling
IEEE Trans. Image Process.5
2020 Weighted bilinear coding over salient body parts for person re-identification
Zhigang Chang, Heng Fan 0001, Hang Su 0006, Hua Yang 0001, Shibao Zheng, Haibin Ling
Neurocomputing6
2020 Convolutional neural network with adaptive inferential framework for skeleton-based action recognition
Hong'en Huang, Hang Su 0006, Zhigang Chang, Mingyang Yu 0005, Jialin Gao, Xinzhe Li 0002, Shibao Zheng
J. Vis. Commun. Image Represent.7
2019 Face Anti-Spoofing: Model Matters, so Does Data
abstract
Face anti-spoofing is an important task in full-stack face applications including face detection, verification, and recognition. Previous approaches build models on datasets which do not simulate the real-world data well (e.g., small scale, insignificant variance, etc.). Existing models may rely on auxiliary information, which prevents these anti-spoofing solutions from generalizing well in practice. In this paper, we present a data collection solution along with a data synthesis technique to simulate digital medium-based face spoofing attacks, which can easily help us obtain a large amount of training data well reflecting the real-world scenarios. Through exploiting a novel Spatio-Temporal Anti-Spoof Network (STASN), we are able to push the performance on public face anti-spoofing datasets over state-of-the-art methods by a large margin. Since the proposed model can automatically attend to discriminative regions, it makes analyzing the behaviors of the network possible.We conduct extensive experiments and show that the proposed model can distinguish spoof faces by extracting features from a variety of regions to seek out subtle evidences such as borders, moire patterns, reflection artifacts, etc.
Wenhan Luo, Linchao Bao, Yuan Gao 0015, Dihong Gong, Shibao Zheng, Zhifeng Li 0001, Wei Liu 0005
CVPR6
2019 Reference-oriented Loss for Person Re-identification
abstract
Deep metric learning methods are quite effective in exploring discriminative feature embeddings, among which triplet loss and its variants are widely utilized. However, in existing methods, the tightness information for intra-class samples is ignored, leading to large intra-class divergence and severe inter-class overlapping problem. To address this issue, a novel loss function called reference-oriented triplet loss is proposed in this paper. The proposed method introduces several reference images to guide training. More specifically, distances between the reference image and images of the same identity are required to be as similar as possible. By introducing reference images, images from the same class become much closer with each other and the inter-class overlapping problem is alleviated. Comparing to baseline batch hard triplet loss, the mAP accuracy increases by 3.75%/5.69% on person re-ID datasets Market1501 and DukeMTMC-Reid. Comparison results with state-of-the-art algorithms also demonstrate effectiveness of the proposed algorithm.
Zhigang Chang, Shibao Zheng, Tai-Pang Wu
IJCNN4
2019 Learning to Self-Train for Semi-Supervised Few-Shot Classification
abstract
Few-shot classification (FSC) is challenging due to the scarcity of labeled training data (e.g. only one labeled data point per class). Meta-learning has shown to achieve promising results by learning to initialize a classification model for FSC. In this paper we propose a novel semi-supervised meta-learning method called learning to self-train (LST) that leverages unlabeled data and specifically meta-learns how to cherry-pick and label such unsupervised data to further improve performance. To this end, we train the LST model through a large number of semi-supervised few-shot tasks. On each task, we train a few-shot model to predict pseudo labels for unlabeled data, and then iterate the self-training steps on labeled and pseudo-labeled data with each step followed by fine-tuning. We additionally learn a soft weighting network (SWN) to optimize the self-training weights of pseudo labels so that better ones can contribute more to gradient descent optimization. We evaluate our LST method on two ImageNet benchmarks for semi-supervised few-shot classification and achieve large improvements over the state-of-the-art.
Xinzhe Li 0002, Qianru Sun, Yaoyao Liu 0001, Qin Zhou 0002, Shibao Zheng, Tat-Seng Chua, Bernt Schiele
NeurIPS5
2019 Distribution Context Aware Loss for Person Re-identification
abstract
To learn the optimal similarity function between probe and gallery images in Person re-identification, effective deep metric learning methods have been extensively explored to obtain discriminative feature embedding. However, existing metric loss like triplet loss and its variants always emphasize pair-wise relations but ignore the distribution context in feature space, leading to inconsistency and sub-optimal. In fact, the similarity of one pair not only decides the match of this pair, but also has potential impacts on other sample pairs. In this paper, we propose a novel Distribution Context Aware (DCA) loss based on triplet loss to combine both numerical similarity and relation similarity in feature space for better clustering. Extensive experiments on three benchmarks including Market-1501, DukeMTMC-reID and MSMT17, evidence the favorable performance of our method against the corresponding baseline and other state-of-the-art methods.
Zhigang Chang, Qin Zhou 0002, Shibao Zheng, Hua Yang 0001, Tai-Pang Wu
VCIP4
2018 Graph Correspondence Transfer for Person Re-Identification
abstract
In this paper, we propose a graph correspondence transfer (GCT) approach for person re-identification. Unlike existing methods, the GCT model formulates person re-identification as an off-line graph matching and on-line correspondence transferring problem. In specific, during training, the GCT model aims to learn off-line a set of correspondence templates from positive training pairs with various pose-pair configurations via patch-wise graph matching. During testing, for each pair of test samples, we select a few training pairs with the most similar pose-pair configurations as references, and transfer the correspondences of these references to test pair for feature distance calculation. The matching score is derived by aggregating distances from different references. For each probe image, the gallery image with the highest matching score is the re-identifying result. Compared to existing algorithms, our GCT can handle spatial misalignment caused by large variations in view angles and human poses owing to the benefits of patch-wise graph matching. Extensive experiments on five benchmarks including VIPeR, Road, PRID450S, 3DPES and CUHK01 evidence the superior performance of GCT model over other state-of-the-art methods.
Qin Zhou 0002, Heng Fan 0001, Shibao Zheng, Hang Su 0006, Xinzhe Li 0002, Shuang Wu 0001, Haibin Ling
AAAI3
2018 A Mobile Health Solution for Medication Adherence Intervention and its Real World Evidence
abstract
This paper proposes a mobile health solution to improve patients' medication adherence. It leverages both mobile and cloud technologies with proprietary algorithms to provide automatic, personalized and contextual adherence interventions. After a nine-month real-world trial, the evidence of medication adherence demonstrates the effectiveness of our solution.
Shibao Zheng, Youren Yang, Chu Feng
HealthCom4
2018 Recognizing Minimal Facial Sketch by Generating Photorealistic Faces With the Guidance of Descriptive Attributes
abstract
Cross-modal sketch-photo recognition is of vital importance in law enforcement and public security. Most existing methods are dedicated to bridging the gap between the low-level visual features of sketches and photo images, which is limited due to intrinsic differences in pixel values. In this paper, based on the intuition that sketches and photo images are highly correlated in the semantic domain, we propose to jointly utilize the low-level visual features and high-level facial attributes to enhance the representation ability of sketches. More specifically, a Multi-Modal Conditional GAN (MMC-GAN) is proposed to generate face images for further face recognition based on the generated images. During training, an identity-preserving constraint is further introduced to improve the discriminative ability of the synthetic images. Extensive experiments demonstrate that the effectiveness of attribute-aided face synthesis and recognition.
Xiao Yang 0028, Hang Su 0006, Qin Zhou 0002, Xinzhe Li 0002, Shibao Zheng
ICASSP5
2017 Crowd Behavior Analysis via Curl and Divergence of Motion Trajectories
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Yawen Fan, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2017 Bilinear dynamics for crowd video analysis
Shuang Wu 0001, Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Qin Zhou 0002
J. Vis. Commun. Image Represent.4
2017 Motion sketch based crowd video retrieval
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Qin Zhou 0002
Multim. Tools Appl.3
2017 Joint dictionary and metric learning for person re-identification
Qin Zhou 0002, Shibao Zheng, Haibin Ling, Hang Su 0006, Shuang Wu 0001
Pattern Recognit.2
2016 Joint instance and feature importance re-weighting for person reidentification
abstract
Person reidentification refers to the task of recognizing the same person under different non-overlapping camera views. Presently, person reidentification based on metric learning is proved to be effective among various techniques, which exploits the labeled data to learn a subspace that maximizes the inter-person divergence while minimizes the intra-person divergence. However, these methods fail to take the different impacts of various instances and local features into account. To address this issue, we propose to learn a projection matrix such that the importance of different instances and local features are re-weighted jointly. We also come up with a simplified formulation of the proposed algorithm, thus it can be solved by the efficient UDFS optimization algorithm. Extensive experiments on the VIPeR and iLIDS datasets demonstrate the effectiveness and efficiency of our algorithm.
Qin Zhou 0002, Shibao Zheng, Hua Yang 0001, Hang Su 0006
ICASSP2
2016 Motion sketch based crowd video retrieval via motion structure coding
abstract
Crowd video retrieval is an important problem in surveillance video management in the era of big data, e.g., video indexing and browsing. In this paper, we address this issue from the motion-level perspective by using hand-drawn sketches as queries. Motion sketch based crowd video retrieval naturally suffers from challenges in motion-level video indexing and sketch representation. We tackle them by leveraging the motion structure coding algorithm to extract robust structure-preserved motion descriptors. For video indexing, we use motion decomposition to separate the sub-motion vector fields with typical patterns from a set of optical flows. Then, the motion-level descriptors of the vector fields are computed and stored in the index database. To represent sketch queries, we propose a sketch vectorization algorithm followed by motion structure coding. In the retrieval stage, given a new query, the retrieval function learned by the Ranking SVM algorithm predicts the ranking score of each motion pattern in the index database. Extensive experiments are conducted on the publicly available crowd datasets, which demonstrate the robustness and effectiveness of the proposed sketch based crowd video retrieval system.
Shuang Wu 0001, Hang Su 0006, Shibao Zheng, Hua Yang 0001, Qin Zhou 0002
ICIP3
2015 Kernelized View Adaptive Subspace Learning for Person Re-identification
Qin Zhou 0002, Shibao Zheng, Hang Su 0006, Hua Yang 0001, Shuang Wu 0001
BMVC2
2015 Towards active annotation for detection of numerous and scattered objects
abstract
Object detection is an active study area in the field of computer vision and image understanding. In this paper, we propose an active annotation algorithm by addressing the detection of numerous and scattered objects in a view, e.g., hundreds of cells in microscopy images. In particular, object detection is implemented by classifying pixels into specific classes with graph-based semi-supervised learning and grouping neighboring pixels with the same label. Sample or seed selection is conducted based on a novel annotation criterion that minimizes the expected prediction error. The most informative samples are therefore annotated actively, which are subsequently propagated to the unlabeled samples via a pairwise affinity graph. Experimental results conducted on two real world datasets validate that our proposed scheme quickly reaches high quality results and reduces human efforts significantly.
Hang Su 0006, Hua Yang 0001, Shibao Zheng, Sha Wei, Shuang Wu 0001
ICME3
2014 The large-scale crowd analysis based on sparse spatial-temporal local binary pattern
Hua Yang 0001, Yihua Cao, Hang Su 0006, Yawen Fan, Shibao Zheng
Multim. Tools Appl.5
2013 Vehicle logo recognition based on Bag-of-Words
abstract
The recognition of vehicle manufacturer logo is a crucial and very challenging problem, which is still an area with few published effective methods. This paper proposes a new fast and reliable system for Vehicle Logo Recognition (VLR) based on Bag-of-Words (BoW). In our system, vehicle logo images are represented as histograms of visual words and classified by SVM in three steps: firstly, extract dense-SIFT features; secondly, quantize features into visual words by `Soft-assignment' thirdly, build histograms of visual words with spatial information. Compared with traditional VLR methods, experiment results show that our proposed system achieves higher recognition accuracy with less processing time. The proposed system is evaluated on a dataset of 840 low-resolution vehicle logo images with about 30×30 pixels, which verifies that our system is practical and effective.
Shuyuan Yu, Shibao Zheng, Hua Yang 0001, Longfei Liang
AVSS2
2013 H.264/Advanced Video Control Perceptual Optimization Coding Based on JND-Directed Coefficient Suppression
abstract
The field of video coding has been exploring the compact representation of video data, where perceptual redundancies in addition to signal redundancies are removed for higher compression. Many research efforts have been dedicated to modeling the human visual system's characteristics. The resulting models have been integrated into video coding frameworks in different ways. Among them, coding enhancements with the just noticeable distortion (JND) model have drawn much attention in recent years due to its significant gains. A common application of the JND model is the adjustment of quantization by a multiplying factor corresponding to the JND threshold. In this paper, we propose an alternative perceptual video coding method to improve upon the current H.264/advanced video control (AVC) framework based on an independent JND-directed suppression tool. This new tool is capable of finely tuning the quantization using a JND-normalized error model. To make full use of this new rate distortion adjustment component the Lagrange multiplier for rate distortion optimization is derived in terms of the equivalent distortion. Because the H.264/AVC integer discrete cosine transform (DCT) is different from classic DCT, on which state-of-the-art JND models are computed, we analytically derive a JND mapping formula between the integer DCT domain and the classic DCT domain which permits us to reuse the JND models in a more natural way. In addition, the JND threshold can be refined by adopting a saliency algorithm in the coding framework and we reduce the complexity of the JND computation by reusing the motion estimation of the encoder. Another benefit of the proposed scheme is that it remains fully compliant with the existing H.264/AVC standard. Subjective experimental results show that significant bit saving can be obtained using our method while maintaining a similar visual quality to the traditional H.264/AVC coded video.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.3
2013 The Large-Scale Crowd Behavior Perception Based on Spatio-Temporal Viscous Fluid Field
abstract
Over the past decades, a wide attention has been paid to crowd control and management in the intelligent video surveillance area. Among the tasks for automatic surveillance video analysis, crowd motion modeling lays a crucial foundation for numerous subsequent analysis but encounters many unsolved challenges due to occlusions among pedestrians, complicated motion patterns in crowded scenarios, etc. Addressing the unsolved challenges, the authors propose a novel spatio-temporal viscous fluid field to model crowd motion patterns by exploring both appearance of crowd behaviors and interaction among pedestrians. Large-scale crowd events are hereby recognized based on characteristics of the fluid field. First, a spatio-temporal variation matrix is proposed to measure the local fluctuation of video signals in both spatial and temporal domains. After that, eigenvalue analysis is applied on the matrix to extract the principal fluctuations resulting in an abstract fluid field. Interaction force is then explored based on shear force in viscous fluid, incorporating with the fluctuations to characterize motion properties of a crowd. The authors then construct a codebook by clustering neighboring pixels with similar spatio-temporal features, and consequently, crowd behaviors are recognized using the latent Dirichlet allocation model. The convincing results obtained from the experiments on published datasets demonstrate that the proposed method obtains high-quality results for large-scale crowd behavior perception in terms of both robustness and effectiveness.
Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Sha Wei
IEEE Trans. Inf. Forensics Secur.3
2013 Raptor Codes Based Unequal Protection for Compressed Video According to Packet Priority
abstract
Raptor codes are state-of-the-art forward error correction (FEC) solutions for multimedia transmission, which have been applied to unequal error protection (UEP) of multi-layered media such as scalable video coding. In this paper, we address the problem of UEP for single-layered video over packet erasure channels. By exploiting the different priorities of video packets inside a group of pictures (GOP) and making full use of the good characteristics of standardized Raptor codes at large block length, we propose an optimized UEP framework for single-layered video and develop an efficient algorithm to solve it. Simulation results show that significant gains can be obtained by our method in case of packet losses.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
IEEE Trans. Multim.3
2012 Crowd Event Perception Based on Spatio-temporal Viscous Fluid Field
abstract
Over the past decades, a wide attention has been paid to crowd control and management in intelligent video surveillance area. In this paper, the authors propose a novel spatiotemporal viscous fluid field to recognize large-scale crowd event with respect to both appearance and driven factor of crowd behavior. Firstly, a spatiotemporal variation matrix is proposed to exploit motion property of a crowd. In particular, the paper exploits characteristics of the matrix with eigenvalue decomposition algorithm and constructs an abstract fluid field to model the crowd motion pattern, which is denoted by spatiotemporal fluid field. Secondly, the paper proposes a spatiotemporal force field to exploit the interaction force between the pedestrians. Furthermore, the fluid and force field constructs a spatiotemporal viscous fluid field. Thirdly, after generating feature with bag of word model, the authors utilize latent Dirichlet allocation model to recognize crowd behavior. The experiments on PETS2009 and UMN datasets show that the proposed method has a better performance for large-scale crowd behavior perception in both robustness and effectiveness comparing with the conventional methods.
Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Sha Wei
AVSS3
2012 Optimized nested protection for video Region of Interest with Raptor codes
abstract
Due to the best effort feature of many existing transmission channels, video streams often suffer from inevitable transmission errors. In this paper, we propose a scheme of robust video transmission based on the state-of-the-art Raptor codes, whose applications are in full swing now. And considering Region of Interest (ROI) often draws much attention in images, the scheme adopts a nested protection framework to show partialities to ROI areas for better protection. Different from many existing Raptor codes based UEP methods, our scheme is developed based on the easy-to-use standardized Raptor codes. Experimental results show that significant robustness can be obtained for the video streams, especially for the ROI areas.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
VCIP3
2011 Geometric Motion Flow (GMF): A New Feature for Traffic Surveillance
abstract
Motion analysis is still a challenging task in many computer vision applications. This paper proposes a new low-level motion feature based on geometric regularity information, particularly suited for traffic surveillance. Firstly, a novel concept of temporal geometry consistency constraint (TGCC) is introduced, which exploits the fact that the geometric structure of a rigid object remains consistent across consecutive frames. Furthermore, the spatial geometric flow is adopted to characterize image structure. Finally, the video motion is represented as a set of geometric flow that moves in the temporal direction. In this case, the method yields a promising illumination robust moderately dense geometric motion flow (GMF) and has more explicit motion boundaries. The GMF could also be used for higher level motion modeling and structural inference tasks, as an effective low-level feature. Extensive experiment results on real video demonstrate the effectiveness and robustness of the proposed method for vehicle motion analysis.
Yawen Fan, Hua Yang 0001, Shibao Zheng, Hang Su 0006
ICIG3
2011 The large-scale crowd density estimation based on sparse spatiotemporal local binary pattern
abstract
Over the past decade, a wide attention has been paid to the crowd control and management in intelligent video surveillance area. This paper proposes a sparse spatiotemporal local binary pattern (SST-LBP) descriptor to extract the dynamic texture of the walking crowd with the application to crowd density estimation. Firstly, the sparse selected location is extracted, which is notably variant in temporal domain and scale invariant in spatial domain. Afterwards, considering the spatial and temporal symmetry, the authors propose a sparse spatiotemporal local binary pattern algorithm and utilize its statistical property to describe the crowd feature. Finally, the crowd features are classified into a range of density levels by adopting support vector machine. The experiments on real video show that the proposed SST-LBP method is effective and robust on the large-scale crowd density estimation. Compared with the other methods, the proposed method does not base on the premise that the background should be extracted perfectly, which is too complicated to implement in real surveillance.
Hua Yang 0001, Hang Su 0006, Shibao Zheng, Sha Wei, Yawen Fan
ICME3
2010 The Large-Scale Crowd Density Estimation Based on Effective Region Feature Extraction Method
Hang Su 0006, Hua Yang 0001, Shibao Zheng
ACCV (3)3
2010 Improving H.264/AVC video coding with adaptive coefficient suppression
abstract
Video coding has been widely adopted to achieve pleasant video quality at constrained bitrate. In this paper, adaptive frequency coefficient suppression directed by Human Visual System (HVS) is presented for H.264 video coding. Firstly, starting from Just Noticeable Distortion (JND) models for the classic DCT domain, we deduce a JND threshold for the H.264 transform domain with decent adaptation. Then the resultant threshold is used to adaptively suppress the transform coefficients of prediction residuals. It should be noted that our scheme is fully compatible with the H.264 standard. And experimental results show that compared to normal methods, significant bitrate reduction can be obtained by our scheme at similar subjective quality.
Shibao Zheng
ISCAS3
2009 Offset based leaky prediction for error resilient ROI coding
abstract
During the period of transmission, video data usually suffer from transmission errors inevitably. Intra update is a common approach to stop error propagation. However, damaged images cannot recover until next update in case of errors, which often leads to annoying effect. In this paper, we propose an enhanced leaky prediction approach that enables the region-of-interest (ROI) of images to recover gently from the immediate succeeding frame of erroneous ones in favor of better human perception. Moreover, an optimized offset compensation technique is designed to improve coding performance. Experimental results show that the proposed scheme can achieve better image quality for ROI and the fluctuation of bit rate is greatly reduced, compared to the intra update method.
Shibao Zheng
ICME3
2009 Distortion-Minimized Video Slicing for Unequal Loss Protection
abstract
This letter proposes a distortion-minimized slicing scheme for the unequal loss protection of compressed video bitstreams transported over packet networks. Unlike most existing slicing methods where each slice includes nearly equal number of video macroblocks, the proposed scheme reorders the macroblocks in one video frame according to their importance, and then divides them into two slices with an unequal proportion. According to given channel conditions and a novel unequal loss protection scheme, the more appropriate macroblock division ratio can be found out to achieve minimized end-to-end distortion. Simulation results show that the proposed slicing scheme outperforms state-of-the-art approaches with fixed macroblock division ratio.
Hua Yang 0001, Shibao Zheng
IEEE Signal Process. Lett.4
2008 A two-stage rho-domain rate control scheme for H.264 encoder
abstract
In this paper, we present a two-stage rate control scheme for H.264/AVC encoder. In frame-layer, a novel bit allocation method is introduced, which estimate the target frame bits based on the coding statistics of the first stage. In MB-layer, we adopt the rho-domain rate control scheme to get the refine QP of each MB for residual signal recoding. Experimental results show that our algorithm can efficiently smooth the visual quality fluctuation between frames, on the basis of meeting bit rate constraints accurately.
Guixu Lin, Shibao Zheng, Jianling Hu
ICME2
2008 Variable block size selection for a transcoder based on MB movement information
abstract
To improve coding performance, the new video compression standard H.264 employs seven variable block sizes for one Macro Block (MB) to conduct motion estimation and compensation. MPEG-2 only has one size 16×16. This paper presents a novel fast variable block size selection method for an inter-MB in a video transcoder from MPEG-2 to H.264 with downscaling by a factor two in each dimension based on the MB motion information. Without conducting motion re-estimation, an optimal block size is decided. Experiment results show that this method saves the transcoder complexity dramatically with little compression performance degradation.
Jun Sun 0005, Rong Xie 0004, Shibao Zheng, Songyu Yu
ICME4
2007 Bit Allocation for Fine-Granular SNR Scalability Coding with Hierarchical B Pictures
abstract
Hierarchical B pictures are devised to achieve temporal scalability in the scalable extension of H.264/AVC (SVC) which is under standardization. The fine-granular SNR scalability (FGS) can be provided by progressive refinement (PR) slices in SVC. In this paper, we firstly investigate error propagation in the case of discarding PR slices and obtain a rate difference distortion optimization criterion to improve coding efficiency of base layer. Then we consider the full rate case and propose a rate distortion slope criterion to enhance FGS coding efficiency at high rate. Finally the criterion to boost coding efficiency in the whole range of FGS rate is derived by combing the criterions derived previously. The proposed method is compared to the approach in SVC test model and up to 0.3dB coding gains are achieved.
Jun Xu 0040, Li Song 0001, Shibao Zheng, Xiaokang Yang 0001, Rong Xie 0004
ICME3
2007 A Time and Storage Optimized Hardware Design for Context-Based Adaptive Binary Arithmetic Decoding in H.264/AVC
abstract
This paper proposes a hardware architecture for Context-based Adaptive Binary Arithmetic Code (CABAC) decoding in H.264/AVC. The proposed architecture takes both bin decoding efficiency and control efficiency into account. This architecture improves time and storage efficiency by taking full use of the new found characters of syntax elements (SEs). In this architecture, two controllers are designed to decode SEs. One is the main controller, and the other is the sub controller, which controls the decoding of residual block SEs. The parallel working mode of the two controller improves the time-consuming performance of the system. Experimental result shows that our design is quite rich for main profile CIF video stream at 30fps.
Shibao Zheng, Zhonghua Huang
ICME2