Hua Zhang 0008

dblp:69/2745-8 · DBLP profile ↗
← Back
46ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-7627-4142ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Emotional Baby Is Truly Deadly: Does Your Multimodal Large Reasoning Model Have Emotional Flattery Towards Humans?
abstract
Multimodal large reasoning models (MLRMs) have advanced visual-textual integration, enabling sophisticated human-AI interaction. While prior work has exposed MLRMs to visual jailbreaks, it remains underexplored how their reasoning capabilities reshape the security landscape under adversarial inputs. To fill this gap, we conduct a systematic security assessment of MLRMs and uncover a security-reasoning paradox: although deeper reasoning boosts cross‑modal risk recognition, it also creates cognitive blind spots that adversaries can exploit. We observe that MLRMs oriented toward human-centric service are highly susceptible to users' emotional cues during the deep-thinking stage, often overriding safety protocols or built‑in safety checks under high emotional intensity. Inspired by this key insight, we propose EmoAgent, an autonomous adversarial emotion-agent that orchestrates exaggerated affective prompts to hijack reasoning pathways. Even when visual risks are correctly identified, models can still produce harmful completions through emotional misalignment. We further identify persistent high-risk failure modes in transparent deep-thinking scenarios, such as MLRMs generating harmful reasoning masked behind seemingly safe responses. These failures expose misalignments between internal inference and surface-level behavior, eluding existing content-based safeguards. To quantify these risks, we introduce three metrics: (1) Risk-Reasoning Stealth Score (RRSS) for harmful reasoning beneath benign outputs; (2) Risk-Visual Neglect Rate (RVNR) for unsafe completions despite visual risk recognition; and (3) Refusal Attitude Inconsistency (RAIC) for evaluating refusal unstability under prompt variants. Extensive experiments on advanced MLRMs demonstrate the effectiveness of EmoAgent and reveal deeper emotional cognitive misalignments in model safety.
Yuan Xun, Xiaojun Jia, Simeng Qin, Hua Zhang 0008
AAAI5
2026 Diagnosing Hidden Instabilities in Model Editing via Uncertainty Quantification
abstract
Zihan Gu, TianYi Zhang, Xinyan Zhang, Zhiyuan Wang, Han Zhang, Yuhao Wei, Jiacheng Lu, Tianyi Ma, Xingsheng Zhang, Hua Zhang, Yue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zihan Gu, Han Zhang 0008, Yuhao Wei, Xingsheng Zhang, Hua Zhang 0008
ACL (1)10
2026 PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
abstract
Diffusion models (DMs) have made significant advances in high-quality text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyright concerns. Malicious actors can misuse these models to generate unauthorized content, such as realistic portraits or artistic style replicas. Existing proactive defenses primarily rely on applying adversarial perturbations to user-provided reference images to disrupt training. However, these approaches face some limitations in real-world scenarios: they rely on the assumption that all training images are pre-perturbed, and they are prone to failure when datasets contain unperturbed images or undergo minor data transformations. In this paper, we introduce PersGuard, a novel backdoor-based framework designed to prevent unauthorized personalization of pre-trained T2I diffusion models. Unlike perturbation-based methods, we assume protectors can control and embed protective backdoors into the models before their release. This mechanism ensures that if a downstream user fine-tunes the model on protected images, the model retains the backdoor and generates predefined protective outputs; conversely, for unprotected images, the backdoor is effectively removed during fine-tuning to ensure normal model utility. We formulate the backdoor injection as a unified optimization problem incorporating three objectives: a backdoor behavior loss to activate protection, a prior preservation loss to maintain standard generation capabilities, and a novel backdoor retention loss. The retention loss is specifically designed to mirror personalization loss, ensuring the backdoor remains robust during the downstream fine-tuning process. Extensive experiments across various scenarios, including gray-box and black-box settings, multi-object protection, and facial identity protection, demonstrate that PersGuard provides superior privacy protection compared to existing perturbation-based methods.
Xiaojun Jia, Yuan Xun, Hua Zhang 0008, Xiaochun Cao
IEEE Trans. Dependable Secur. Comput.4
2025 Interpreting Object-level Foundation Models via Visual Precision Search
abstract
Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models’ decisions has grown increasingly challenging. Existing interpretable attribution methods for object-level task interpretation have notable limitations: (1) gradient-based methods lack precise localization due to visual-textual fusion in foundation models, and (2) perturbation-based methods produce noisy saliency maps, limiting fine-grained interpretability. To address these, we propose a Visual Precision Search method that generates accurate attribution maps with fewer regions. Our method bypasses internal model parameters to overcome attribution issues from multimodal fusion, dividing inputs into sparse sub-regions and using consistency and collaboration scores to accurately identify critical decision-making regions. We also conducted a theoretical analysis of the boundary guarantees and scope of applicability of our method. Experiments on RefCOCO, MS COCO, and LVIS show our approach enhances object-level task interpretability over SOTA for Grounding DINO and Florence-2 across various evaluation metrics, with faithfulness gains of 23.7%, 31.6%, and 20.1% on MS COCO, LVIS, and RefCOCO for Grounding DINO, and 50.7% and 66.9% on MS COCO and RefCOCO for Florence-2. Additionally, our method can interpret failures in visual grounding and object detection tasks, surpassing existing methods across multiple evaluation metrics. The code is released at https://github.com/RuoyuChen10/VPS.
Ruoyu Chen 0001, Siyuan Liang 0004, Jingzhi Li 0002, Shiming Liu, Maosen Li, Zhen Huang 0006, Hua Zhang 0008, Xiaochun Cao
CVPR7
2025 FaceInsight: A Multimodal Large Language Model for Face Perception
abstract
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, a versatile face perception MLLM that provides fine-grained information. Our approach introduces visual textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments show that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings.
Jingzhi Li 0002, Changjiang Luo, Ruoyu Chen 0001, Hua Zhang 0008, Wenqi Ren, Jianhou Gan, Xiaochun Cao
ACM Multimedia4
2025 Anti-Fake Vaccine: Safeguarding Privacy Against Face Swapping via Visual-Semantic Dual Degradation
Jingzhi Li 0002, Changjiang Luo, Hua Zhang 0008, Yang Cao 0011, Xiaochun Cao
Int. J. Comput. Vis.3
2025 Robust Label Propagation and Graph Embedding for Cross-Domain Image Classification
abstract
Cross-domain label propagation (LP) faces two main challenges: 1) learning domain-invariant and 2) discriminative feature representations and obtaining high-confidence predicted labels. The distribution differences between domains can make labels difficult to propagate across domains. Low-quality labels can distort the modeling process associated with label-induced loss, resulting in decreased performance. We propose a novel cross-domain image classification method, namely, robust LP and graph embedding (RLPGE). We introduce a nuclear norm maximization constraint in order to make the predicted labels more diverse in categories while preserving their discriminability. The graph embedding process brings two nearby same-class samples close in the embedding subspace, ensuring domain invariance and local discriminability of the embedded features. For optimal graph learning, we simultaneously optimize the cross-domain graph and two intradomain graphs using both features and labels, enhancing their local discriminability and robustness to feature noise. We conducted comprehensive experiments on four cross-domain image classification datasets. The results demonstrate that our proposed RLPGE method outperforming some state-of-the-art approaches
Chengjin Yu, Wuchang Liang, Wei Wang 0335, Yuan-Ting Yan, Hua Zhang 0008
IEEE Internet Things J.7
2025 An unsupervised medical image registration network for intelligent medical education
Jie Mu, Jing Zhang 0037, Tiantian Yan, Wei Wang 0335, Hua Zhang 0008, Wenqi Ren
Neural Comput. Appl.7
2025 Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
abstract
The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature space for novel categories with limited data on the basis of the base category space, which could adapt the learned detection model to unknown scenarios. Most existing fine-tuning-based approaches tackle the challenge via pre-training a feature extractor based on the base categories and then fine-tuning the detector through the novel categories. However, limited by insufficient samples for novel categories, two issues still exist: (1) the features of the novel category are easily implicitly represented by the features of the base category, leading to inseparable classifier boundaries, (2) novel categories with fewer data are not enough to fully represent the distribution, where the model fine-tuning is prone to overfitting. To address these issues, we introduce the side information to alleviate the negative influences derived from the feature space and sample viewpoints and formulate a novel generalized feature representation learning method for FSOD. Specifically, we first utilize embedding side information to construct a knowledge matrix to quantify the semantic relationship between the base and novel categories. Then, to strengthen the discrimination between semantically similar categories, we further develop contextual semantic supervised contrastive learning which embeds side information. Furthermore, to prevent overfitting problems caused by sparse samples, a side-information guided region-aware masked module is introduced to augment the diversity of samples, which finds and abandons biased information that discriminates between similar categories via counterfactual explanation, and refines the discriminative representation space further. Finally, we theoretically analyze the generalization bound for introducing our proposed module and demonstrate that our proposed model can effectively reduce the upper bound of the generalization error. Extensive experiments using ResNet and ViT backbones on PASCAL VOC, MS COCO, LVIS V1, FSOD-1 K, and FSVOD-500 benchmarks demonstrate that our model outperforms the previous state-of-the-art methods, significantly improving the ability of FSOD in most shots/splits. The code is released athttps://github.com/RuoyuChen10/CCL-FSOD.
Ruoyu Chen 0001, Hua Zhang 0008, Jingzhi Li 0002, Li Liu 0002, Zhen Huang 0006, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Less is More: Fewer Interpretable Region via Submodular Subset Selection
abstract
Image attribution algorithms aim to identify important regions that are highly relevant to model decisions. Although existing attribution solutions can effectively assign importance to target elements, they still face the following challenges: 1) existing attribution methods generate inaccurate small regions thus misleading the direction of correct attribution, and 2) the model cannot produce good attribution results for samples with wrong predictions. To address the above challenges, this paper re-models the above image attribution problem as a submodular subset selection problem, aiming to enhance model interpretability using fewer regions. To address the lack of attention to local regions, we construct a novel submodular function to discover more accurate small interpretation regions. To enhance the attribution effect for all samples, we also impose four different constraints on the selection of sub-regions, i.e., confidence, effectiveness, consistency, and collaboration scores, to assess the importance of various subsets. Moreover, our theoretical analysis substantiates that the proposed function is in fact submodular. Extensive experiments show that the proposed method outperforms SOTA methods on two face datasets (Celeb-A and VGG-Face2) and one fine-grained dataset (CUB-200-2011). For correctly predicted samples, the proposed method improves the Deletion and Insertion scores with an average of 4.9\% and 2.5\% gain relative to HSIC-Attribution. For incorrectly predicted samples, our method achieves gains of 81.0\% and 18.4\% compared to the HSIC-Attribution algorithm in the average highest confidence and Insertion score respectively. The code is released at https://github.com/RuoyuChen10/SMDL-Attribution.
Ruoyu Chen 0001, Hua Zhang 0008, Siyuan Liang 0004, Jingzhi Li 0002, Xiaochun Cao
ICLR2
2024 Multiple Adverse Weather Conditions Adaptation for Object Detection via Causal Intervention
abstract
Most state-of-the-art object detection methods have achieved impressive perfomrace on several public benchmarks, which are trained with high definition images. However, existing detectors are often sensitive to the visual variations and out-of-distribution data due to the domain gap caused by various confounders, e.g. the adverse weathre conditions. To bridge the gap, previous methods have been mainly exploring domain alignment, which requires to collect an amount of domain-specific training samples. In this paper, we introduce a novel domain adaptation model to discover a weather condition invariant feature representation. Specifically, we first employ a memory network to develop a confounder dictionary, which stores prototypes of object features under various scenarios. To guarantee the representativeness of each prototype in the dictionary, a dynamic item extraction strategy is used to update the memory dictionary. After that, we introduce a causal intervention reasoning module to explore the invariant representation of a specific object under different weather conditions. Finally, a categorical consistency regularization is used to constrain the similarities between categories in order to automatically search for the aligned instances among distinct domains. Experiments are conducted on several public benchmarks (RTTS, Foggy-Cityscapes, RID, and BDD 100K) with state-of-the-art performance achieved under multiple weather conditions.
Hua Zhang 0008, Xiaohong Li 0001, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Toward Generalized Few-Shot Open-Set Object Detection
abstract
Open-set object detection (OSOD) aims to detect the known categories and reject unknown objects in a dynamic world, which has achieved significant attention. However, previous approaches only consider this problem in data-abundant conditions, while neglecting the few-shot scenes. In this paper, we seek a solution for the generalized few-shot open-set object detection (G-FOOD), which aims to avoid detecting unknown classes as known classes with a high confidence score while maintaining the performance of few-shot detection. The main challenge for this task is that few training samples induce the model to overfit on the known classes, resulting in a poor open-set performance. We propose a new G-FOOD algorithm to tackle this issue, named Few-shOt Open-set Detector (FOOD), which contains a novel class weight sparsification classifier (CWSC) and a novel unknown decoupling learner (UDL). To prevent over-fitting, CWSC randomly sparses parts of the normalized weights for the logit prediction of all classes, and then decreases the co-adaptability between the class and its neighbors. Alongside, UDL decouples training the unknown class and enables the model to form a compact unknown decision boundary. Thus, the unknown objects can be identified with a confidence probability without any threshold, prototype, or generation. We compare our method with several state-of-the-art OSOD methods in few-shot scenes and observe that our method improves the F-score of unknown classes by 4.80%-9.08% across all shots in VOC-COCO dataset settings.
Binyi Su, Hua Zhang 0008, Jingzhi Li 0002, Zhong Zhou
IEEE Trans. Image Process.2
2023 HSIC-based Moving Weight Averaging for Few-Shot Open-Set Object Detection
abstract
We study the problem of few-shot open-set object detection (FOOD), whose goal is to quickly adapt a model to a small set of labeled samples and reject unknown class samples. Recent works usually use the weight sparsification for unknown rejection, but due to the lack of tailored considerations for data-scarce scenarios, the performance is not satisfactory. In this work, we solve the challenging few-shot open-set object detection problems from three aspects. First, different from previous pseudo-unknown sample mining methods, we employ the evidential uncertainty estimated by the Dirichlet distribution of probability to mine the pseudo-unknown samples from the foreground and background proposal space. Second, based on the statistical analysis between the number of pseudo-unknown samples and the Intersection over Union (IoU), we propose an IoU-aware unknown objective, which sharps the unknown decision boundary by considering the localization quality. Third, to suppress the over-fitting problem and improve the model's generalization ability for unknown rejection, we propose the HSIC-based (Hilbert-Schmidt Independence Criterion) moving weight averaging to update the weights of classification and regression heads, which considers the degree of independence between the current weights and previous weights stored in the long-term memory banks. We compare our method with several state-of-the-art methods and observe that our method improves the mean recall of unknown classes by 12.87% across all shots in the VOC-COCO dataset settings. Our code is available at https://github.com/binyisu/food.
Binyi Su, Hua Zhang 0008, Zhong Zhou
ACM Multimedia2
2023 Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
abstract
Recently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in edge devices. The core challenge of developing compact and efficient DNNs lies in how to balance the competing goals of achieving high accuracy and high efficiency. In this paper we propose two novel types of convolutions, dubbed Pixel Difference Convolution (PDC) and Binary PDC (Bi-PDC) which enjoy the following benefits: capturing higher-order local differential information, computationally efficient, and able to be integrated with existing DNNs. With PDC and Bi-PDC, we further present two lightweight deep networks named Pixel Difference Networks (PiDiNet) and Binary PiDiNet (Bi-PiDiNet) respectively to learn highly efficient yet more accurate representations for visual tasks including edge detection and object recognition. Extensive experiments on popular datasets (BSDS500, ImageNet, LFW, YTF, etc.) show that PiDiNet and Bi-PiDiNet achieve the best accuracy-efficiency trade-off. For edge detection, PiDiNet is the first network that can be trained without ImageNet, and can achieve the human-level performance on BSDS500 at 100 FPS and with 1 M parameters. For object recognition, among existing Binary DNNs, Bi-PiDiNet achieves the best accuracy and a nearly 2× reduction of computational cost on ResNet18.
Zhuo Su 0002, Longguang Wang, Hua Zhang 0008, Zhen Liu 0004, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Privacy-Enhancing Face Obfuscation Guided by Semantic-Aware Attribution Maps
abstract
Face recognition technology is increasingly being integrated into our daily life, e.g. Face ID. With the advancement of machine learning algorithms, the personal information such as age, gender, and race can be easily deduced from the recorded face images in these applications. This poses a serious privacy threat to individuals who do not want to be profiled, as face images are collected for biometric purposes. Existing methods mostly focus on adding the invisible adversarial perturbations into the images to make automatic inference infeasible. However, the application scenarios of these methods are limited due to the perturbations depending on the specific model. In this paper, we introduce a novel face privacy-enhancing framework by obfuscating the stored faces, which could maintain the data utility (face identity) while protecting the privacy of users (facial attributes). Specifically, we first develop a feature attribution module to discover the identity-related facial parts. Within this module, we introduce a pixel importance estimation model based on Shapley value to obtain a pixel-level attribution map, and then each pixel on the attribution map is aggregated into semantic facial parts, which are used to quantify the importance of different facial parts. Next, we design a privacy-enhancing module to generate the high-quality obfuscated images, which can modify the privacy semantic content and preserve the identity-related information. Using the proposed method, users can choose the single or multiple attributes to be obfuscated without affecting identity matching. Extensive experiments conducted on CelebA-HQ and VGGFace2-HQ benchmarks demonstrate the effectiveness and generalization ability of our method.
Jingzhi Li 0002, Hua Zhang 0008, Siyuan Liang 0004, Pengwen Dai, Xiaochun Cao
IEEE Trans. Inf. Forensics Secur.2
2023 Prediction With Visual Evidence: Sketch Classification Explanation via Stroke-Level Attributions
abstract
Sketch classification models have been extensively investigated by designing a task-driven deep neural network. Despite their successful performances, few works have attempted to explain the prediction of sketch classifiers. To explain the prediction of classifiers, an intuitive way is to visualize the activation maps via computing the gradients. However, visualization based explanations are constrained by several factors when directly applying them to interpret the sketch classifiers: (i) low-semantic visualization regions for human understanding. and (ii) neglecting of the inter-class correlations among distinct categories. To address these issues, we introduce a novel explanation method to interpret the decision of sketch classifiers with stroke-level evidences. Specifically, to achieve stroke-level semantic regions, we first develop a sketch parser that parses the sketch into strokes while preserving their geometric structures. Then, we design a counterfactual map generator to discover the stroke-level principal components for a specific category. Finally, based on the counterfactual feature maps, our model could explain the question of "why the sketch is classified as X" by providing positive and negative semantic explanation evidences. Experiments conducted on two public sketch benchmarks, Sketchy-COCO and TU-Berlin, demonstrate the effectiveness of our proposed model. Furthermore, our model could provide more discriminative and human understandable explanations compared with these existing works.
Sixuan Liu, Jingzhi Li 0002, Hua Zhang 0008, Long Xu 0001, Xiaochun Cao
IEEE Trans. Image Process.3
2023 New Wine Old Bottles: Feistel Structure Revised
abstract
This paper mainly investigates the iterative structures whose decryption is similar to the encryption. Firstly, we unify many well-known structures which share similar procedures between the decryption and the encryption, and give a sufficient and necessary condition for this structure to be bijective, which reveals many new insights into the Feistel structure as well as the Lai-Massey structure. Secondly, we analyze the security of the unified structure against the known cryptanalysis. By extending the dual structure from a Feistel structure to the unified structure, we prove that a differential of the unified structure is impossible if and only if it is a zero-correlation linear hull of its dual structure, which presents a generalized link between the impossible differential and zero-correlation linear cryptanalysis shown in CRYPTO 2015. Significantly, several constraints on the linear components of the cipher and the permutation on the branches of the cipher are specified to make the structure resilient to differential and linear cryptanalysis. Furthermore, in the case that the order of the permutation equals the number of the branches$n$, we prove that there always exist a$(3n-1)$-round impossible differential and a$(3n-1)$-round zero-correlation linear hull of the structure, and also present an algorithm to construct these distinguishers. Finally, we propose some novel structures which might be used in future block cipher designs.
Bing Sun 0001, Li Liu 0002, Hua Zhang 0008, Chao Li 0002
IEEE Trans. Inf. Theory6
2023 Sim2Word: Explaining Similarity with Representative Attribute Words via Counterfactual Explanations
abstract
Recently, we have witnessed substantial success using the deep neural network in many tasks. Although there still exist concerns about the explainability of decision making, it is beneficial for users to discern the defects in the deployed deep models. Existing explainable models either provide the image-level visualization of attention weights or generate textual descriptions as post hoc justifications. Different from existing models, in this article we propose a new interpretation method that explains the image similarity models by salience maps and attribute words. Our interpretation model contains visual salience maps generation and the counterfactual explanation generation. The former has two branches: global identity relevant region discovery and multi-attribute semantic region discovery. The first branch aims to capture the visual evidence supporting the similarity score, which is achieved by computing counterfactual feature maps. The second branch aims to discover semantic regions supporting different attributes, which helps to understand which attributes in an image might change the similarity score. Then, by fusing visual evidence from two branches, we can obtain the salience maps indicating important response evidence. The latter will generate the attribute words that best explain the similarity using the proposed erasing model. The effectiveness of our model is evaluated on the classical face verification task. Experiments conducted on two benchmarks—VGGFace2 and Celeb-A—demonstrate that our model can provide convincing interpretable explanations for the similarity. Moreover, our algorithm can be applied to evidential learning cases, such as finding the most characteristic attributes in a set of face images, and we verify its effectiveness on the VGGFace2 dataset.
Ruoyu Chen 0001, Jingzhi Li 0002, Hua Zhang 0008, Changchong Sheng, Li Liu 0002, Xiaochun Cao
ACM Trans. Multim. Comput. Commun. Appl.3
2022 FSRDD: An Efficient Few-Shot Detector for Rare City Road Damage Detection
abstract
Road damage detection (RDD) is indispensable for safe autonomous driving. Existing RDD models focus on designing feature representations following expert knowledge. However, collecting and labeling all types of samples is time-consuming and leads to insufficient training data. To alleviate the adverse effect of few training samples, a novel few-shot road damage detector (FSRDD) is proposed in this paper to detect rare road damages. The proposed FSRDD includes three stages. First, fully annotated abundant base classes are leveraged to train a base detector, where ghost attention (GA) and proposal feature metric (PFM) modules are developed to eliminate the redundant information and measure the proposal features, respectively. Second, the recognition branch of the detector is fine-tuned using a few samples of all classes. Finally, the test set is inferred with the help of an offline scale-aware prototypical calibration block (SPCB). Extensive experiments show that our FSRDD achieves 10-shot rare road damage detection with 33.4% and 12.9% mAP50 on RDD and CNRDD datasets, respectively, significantly outperforming state-of-the-art methods.
Binyi Su, Hua Zhang 0008, Zhaohui Wu 0005, Zhong Zhou
IEEE Trans. Intell. Transp. Syst.2
2022 Accurate Scene Text Detection Via Scale-Aware Data Augmentation and Shape Similarity Constraint
abstract
Scene text detection has attracted increasing concerns with the rapid development of deep neural networks in recent years. However, existing scene text detectors may overfit on the public datasets due to the limited training data, or generate inaccurate localization for arbitrary-shape scene texts. This paper presents an arbitrary-shape scene text detection method that can achieve better generalization ability and more accurate localization. We first propose a Scale-Aware Data Augmentation (SADA) technique to increase the diversity of training samples. SADA considers the scale variations and local visual variations of scene texts, which can effectively relieve the dilemma of limited training data. At the same time, SADA can enrich the training minibatch, which contributes to accelerating the training process. Furthermore, a Shape Similarity Constraint (SSC) technique is exploited to model the global shape structure of arbitrary-shape scene texts and backgrounds from the perspective of the loss function. SSC encourages the segmentation of text or non-text in the candidate boxes to be similar to the corresponding ground truth, which is helpful to localize more accurate boundaries for arbitrary-shape scene texts. Extensive experiments have demonstrated the effectiveness of the proposed techniques, and state-of-the-art performances are achieved over public arbitrary-shape scene text benchmarks (e.g.,CTW1500,Total-TextandArT).
Pengwen Dai, Yang Li 0093, Hua Zhang 0008, Jingzhi Li 0002, Xiaochun Cao
IEEE Trans. Multim.3
2021 Progressive Contour Regression for Arbitrary-Shape Scene Text Detection
abstract
State-of-the-art scene text detection methods usually model the text instance with local pixels or components from the bottom-up perspective and, therefore, are sensitive to noises and dependent on the complicated heuristic post-processing especially for arbitrary-shape texts. To relieve these two issues, instead, we propose to progressively evolve the initial text proposal to arbitrarily shaped text contours in a top-down manner. The initial horizontal text proposals are generated by estimating the center and size of texts. To reduce the range of regression, the first stage of the evolution predicts the corner points of oriented text proposals from the initial horizontal ones. In the second stage, the contours of the oriented text proposals are iteratively regressed to arbitrarily shaped ones. In the last iteration of this stage, we rescore the confidence of the final localized text by utilizing the cues from multiple contour points, rather than the single cue from the initial horizontal proposal center that may be out of arbitrary-shape text regions. Moreover, to facilitate the progressive contour evolution, we design a contour information aggregation mechanism to enrich the feature representation on text contours by considering both the circular topology and semantic context. Experiments conducted on CTW1500, Total-Text, ArT, and TD500 have demonstrated that the proposed method especially excels in line-level arbitrary-shape texts. Code is available at https://github.com/dpengwen/PCR.
Pengwen Dai, Sanyi Zhang, Hua Zhang 0008, Xiaochun Cao
CVPR3
2021 Identity-Preserving Face Anonymization via Adaptively Facial Attributes Obfuscation
abstract
With the popularity of using computer vision technology in monitoring system, there is an increasing societal concern on intruding people's privacy as the captured images/videos may contain identity-related information e.g. people's face. Existing methods on protecting such privacy focus on removing the identity-related information from faces. However, this would weaken the utility of current monitoring system. In this paper, we develop a face anonymization framework that could obfuscate visual appearance while preserving the identity discriminability. The framework is composed of two parts: an identity-aware region discovery module and an identity-aware face confusion module. The former adaptively locates the identity-independent attributes on human faces, and the latter generates the privacy-preserving faces using original faces and discovered facial attributes. To optimize the face generator, we employ a multi-task based loss function, which consists of discriminator loss, identify preserving loss, and reconstruction loss functions. Our model can achieve a balance between recognition utility and appearance anonymizing by modifying different numbers of facial attributes according to pratical demands, and provide a variety of results. Extensive experiments conducted on two public benchmarks Celeb-A and VGG-Face2 demonstrate the effectiveness of our model under distinct face recognition scenarios.
Jingzhi Li 0002, Lutong Han, Ruoyu Chen 0001, Hua Zhang 0008, Lili Wang 0006, Xiaochun Cao
ACM Multimedia4
2021 Kernel Stability for Model Selection in Kernel-Based Algorithms
abstract
Model selection is one of the fundamental problems in kernel-based algorithms, which is commonly done by minimizing an estimation of generalization error. The notion of stability and cross-validation (CV) error of learning machines consists of two widely used tools for analyzing the generalization performance. However, there are some disadvantages to both tools when applied for model selection: 1) the stability of learning machines is not practical due to the difficulty of the estimation of its specific value and 2) the CV-based estimate of generalization error usually has a relatively high variance, so it is prone to overfitting. To overcome these two limitations, we present a novel notion of kernel stability (KS) for deriving the generalization error bounds and variance bounds of CV and provide an effective approach to the application of KS for practical model selection. Unlike the existing notions of stability of the learning machine, KS is defined on the kernel matrix; hence, it can avoid the difficulty of the estimation of its value. We manifest the relationship between the KS and the popular uniform stability of the learning algorithm, and further propose several KS-based generalization error bounds and variance bounds of CV. By minimizing the proposed bounds, we present two novel KS-based criteria that can ensure good performance. Finally, we empirically analyze the performance of the proposed criteria on many benchmark data, which demonstrates that our KS-based criteria are sound and effective.
Yong Liu 0018, Shizhong Liao, Hua Zhang 0008, Wenqi Ren, Weiping Wang 0005
IEEE Trans. Cybern.3
2021 SLOAN: Scale-Adaptive Orientation Attention Network for Scene Text Recognition
abstract
Scene text recognition, the final step of the scene text reading system, has made impressive progress based on deep neural networks. However, existing recognition methods devote to dealing with the geometrically regular or irregular scene text. They are limited to the semantically arbitrary-orientation scene text. Meanwhile, previous scene text recognizers usually learn the single-scale feature representations for various-scale characters, which cannot model effective contexts for different characters. In this paper, we propose a novel scale-adaptive orientation attention network for arbitrary-orientation scene text recognition, which consists of a dynamic log-polar transformer and a sequence recognition network. Specifically, the dynamic log-polar transformer learns the log-polar origin to adaptively convert the arbitrary rotations and scales of scene texts into the shifts in the log-polar space, which is helpful to generate the rotation-aware and scale-aware visual representation. Next, the sequence recognition network is an encoder-decoder model, which incorporates a novel character-level receptive field attention module to encode more valid contexts for various-scale characters. The whole architecture can be trained in an end-to-end manner, only requiring the word image and its corresponding ground-truth text. Extensive experiments on several public datasets have demonstrated the effectiveness and superiority of our proposed method.
Pengwen Dai, Hua Zhang 0008, Xiaochun Cao
IEEE Trans. Image Process.2
2021 Learning Deep Lucas-Kanade Siamese Network for Visual Tracking
abstract
In most recent years, Siamese trackers have drawn great attention because of their well-balanced accuracy and efficiency. Although these approaches have achieved great success, the discriminative power of the conventional Siamese trackers is still limited by the insufficient template-candidate representation. Most of the existing approaches take non-aligned features to learn a similarity function for template-candidate matching, while the target object's geometrical transformation is seldom explored. To address this problem, we propose a novel Siamese tracking framework, which enables to dynamically transform the template-candidate features to a more discriminative viewpoint for similarity matching. Specifically, we reformulate the template-candidate matching problem of the conventional Siamese tracker from the perspective of Lucas-Kanade (LK) image alignment approach. A Lucas-Kanade network (LKNet) is proposed and incorporated to the Siamese architecture to learn aligned feature representations in data-driven trainable manner, which is able to enhance the model adaptability in challenging scenarios. Within this framework, we propose two Siamese trackers named LK-Siam and LK-SiamRPN to validate the effectiveness. Extensive experiments conducted on the prevalent datasets show that the proposed method is more competitive over a number of state-of-the-art methods.
Siyuan Yao, Xiaoguang Han 0001, Hua Zhang 0008, Xiao Wang 0017, Xiaochun Cao
IEEE Trans. Image Process.3
2021 Robust Online Tracking via Contrastive Spatio-Temporal Aware Network
abstract
Existing tracking-by-detection approaches using deep features have achieved promising results in recent years. However, these methods mainly exploit feature representations learned from individual static frames, thus paying little attention to the temporal smoothness between frames. This easily leads trackers to drift in the presence of large appearance variations and occlusions. To address this issue, we propose a two-stream network to learn discriminative spatio-temporal feature representations to represent the target objects. The proposed network consists of a Spatial ConvNet module and a Temporal ConvNet module. Specifically, the Spatial ConvNet adopts 2D convolutions to encode the target-specific appearance in static frames, while the Temporal ConvNet models the temporal appearance variations using 3D convolutions and learns consistent temporal patterns in a short video clip. Then we propose a proposal refinement module to adjust the predicted bounding box, which can make the target localizing outputs to be more consistent in video sequences. In addition, to improve the model adaptation during online update, we propose a contrastive online hard example mining (OHEM) strategy, which selects hard negative samples and enforces them to be embedded in a more discriminative feature space. Extensive experiments conducted on the OTB, Temple Color and VOT benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Siyuan Yao, Hua Zhang 0008, Wenqi Ren, Chao Ma 0004, Xiaoguang Han 0001, Xiaochun Cao
IEEE Trans. Image Process.2
2020 Learning Disentangled Representations for Identity Preserving Surveillance Face Camouflage
abstract
In this paper, we focus on protecting the facial privacy for people under the surveillance scenarios, by changing some visual appearances of the faces while keeping them recognizable by the current face recognition systems. This is a challenging problem because we need to retain the most important structures of the captured facial images, while modify the salient facial regions to protect personal privacy. To address this problem, we introduce a novel individual face protection model, which can camouflage the face appearance from the perspective of human visual perception and preserve the identity features of faces used for face authentication. To that end, we develop an encoder-decoder network architecture which can separately disentangle the facial feature representation into an appearance code and an identification code. Specifically, we first randomly divide the input face image into two groups, the source and target sets, where the identity and appearance codes can be correspondingly extracted. Then, we recombine the identity and appearance codes to synthesize a new face, which has the same identity as the source subject. Finally, the synthesized faces are employed to replace the original face to protect the individual privacy. Note that our model is end-to-end with a multi-task loss function, which can better preserve the identity and stabilize the training process. Experiments conducted on Cross-Age Celebrity dataset demonstrate the effectiveness of our model and validate our superiority in terms of visual quality and scalability.
Jingzhi Li 0002, Lutong Han, Hua Zhang 0008, Xiaoguang Han 0001, Jingguo Ge, Xiaochun Cao
ICPR3
2020 Single Image Dehazing via Multi-scale Convolutional Neural Networks with Holistic Edges
Wenqi Ren, Jinshan Pan, Hua Zhang 0008, Xiaochun Cao, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2020 Task-Aware Attention Model for Clothing Attribute Prediction
abstract
Clothing attribute recognition, especially in unconstrained street images, is a challenging task for multimedia. Existing methods for multi-task clothing attribute prediction often ignore the relation between specific attributes and positions. However, the attribute response is always location-sensitive, i.e., different spatial locations have various contributions to attributes. Inspired by the locality of clothing attributes, in this paper, we introduce the attention mechanism to incorporate the impact of positions for clothing attribute prediction with only image-level annotations. However, the performance improvement is limited if we directly use the traditional spatial attention model for each task since it does not take the influence from other tasks into account. Instead, we propose a novel task-aware attention mechanism, which estimates the importance of each position across different tasks. We first evaluate a task attention network with an end-to-end multi-task clothing attribute learning architecture on the shop domain. And then, we employ curriculum learning strategy, which transfers the well-trained shop domain attribute knowledge to the street domain attribute prediction. Experiments are conducted on three clothing benchmarks, i.e., cross-domain clothing attribute dataset, woman clothing dataset, and man clothing dataset. The performance of attribute prediction demonstrates the superiority of the proposed task-aware attention mechanism over several state-of-the-art methods both in shop and street domains.
Sanyi Zhang, Zhanjie Song, Xiaochun Cao, Hua Zhang 0008, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Deep Multi-Scale Context Aware Feature Aggregation for Curved Scene Text Detection
abstract
Scene text plays a significant role in image and video understanding, which has made great progress in recent years. Most existing models on text detection in the wild have the assumption that all the texts are surrounded by a rotated rectangle or quadrangle. While there also exist lots of curved texts in the wild, which would not be bounded by a regular bounding box. In this paper, we develop a novel architecture to localize the text regions, which can deal with curved-shape scene texts. Specifically, we first design a text-related feature enhancement module by incorporating the prior knowledge of the text shape to enhance the feature representations. After that, based on the enhanced features, we employ a region proposal network to generate the candidate boxes of scene texts. For each text candidate, a pyramid region-of-interest pooling attention module is utilized to extract the fixed-size features. Finally, we exploit the box-aware context-based text segmentation module and box refinement network to obtain the location of scene text. Experiments are conducted on four challenging benchmarks CTW1500, totalTEXT, ICDAR-2015 and MLT, and the experimental results have demonstrated the superiority of our model.
Pengwen Dai, Hua Zhang 0008, Xiaochun Cao
IEEE Trans. Multim.2
2019 Learning Structural Representations via Dynamic Object Landmarks Discovery for Sketch Recognition and Retrieval
abstract
State-of-the-art methods on sketch classification and retrieval are based on deep convolutional neural network to learn representations. Although deep neural networks have the ability to model images with hierarchical representations by convolution kernels, they can not automatically extract the structural representations of object categories in a human-perceptible way. Furthermore, sketch images usually have large scale visual variations caused by the styles of drawing or viewpoints, which make it difficult to develop generalized representations using the fixed computational mode of convolutional kernel. In this paper, our aim is to address the problem of fixed computational mode in feature extraction process without extra supervision. We propose a novel architecture to dynamically discover the object landmarks and learn the discriminative structural representations. Our model is composed of two components: a representative landmark discovering module that localizes the key points on the object, and a category-aware representation learning module that develops the category-specific features. Specifically, we develop a structure-aware offset layer to dynamically localize the representative landmarks, which is optimized based on the category labels without extra supervision. After that, a diversity branch is introduced to extract the global discriminative features for each category. Finally, we employ a multi-task loss function to develop an end-to-end trainable architecture. At testing time, we fuse all the predictions with different number of landmarks to achieve the final results. Through extensive experiments, we compare our model with several state-of-the-art methods on two challenging datasets TU-Berlin and Sketchy for sketch classification and retrieval, and the experimental results demonstrate the effectiveness of our proposed model.
Hua Zhang 0008, Peng She, Yong Liu 0018, Jianhou Gan, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Image Process.1
2018 Audio Visual Attribute Discovery for Fine-Grained Object Recognition
abstract
Current progresses on fine-grained recognition are mainly focus on learning the discriminative feature representation via introducing the visual supervisions e.g. part labels. However, it is time-consuming and needs the professional knowledge to obtain the accuracy annotations. Different from these existing methods based on the visual supervisions, in this paper, we introduce a novel feature named audio visual attributes via discovering the correlations between the visual and audio representations. Specifically, our unified framework is training with video-level category label, which consists of two important modules, the encoder module and the attribute discovery module, to encode the image and audio into vectors and learn the correlations between audio and images, respectively. On the encoder module, we present two types of feed forward convolutional neural network for the image and audio modalities. While an attention driven framework based on recurrent neural network is developed to generate the audio visual attribute representation. Thus, our proposed architecture can be implemented end-to-end in the step of inference. We exploit our models for the problem of fine-grained bird recognition on the CUB200-211 benchmark. The experimental results demonstrate that with the help of audio visual attribute, we achieve the superior or comparable performance to that of strongly supervised approaches on the bird recognition.
Hua Zhang 0008, Xiaochun Cao, Rui Wang 0032
AAAI1
2018 Multi-Class Learning: From Theory to Algorithm
abstract
In this paper, we study the generalization performance of multi-class classification and obtain a shaper data-dependent generalization error bound with fast convergence rate, substantially improving the state-of-art bounds in the existing data-dependent generalization analysis. The theoretical analysis motivates us to devise two effective multi-class kernel learning algorithms with statistical guarantees. Experimental results show that our proposed methods can significantly outperform the existing multi-class classification methods.
Jian Li 0040, Yong Liu 0018, Rong Yin 0001, Hua Zhang 0008, Lizhong Ding 0001, Weiping Wang 0005
NeurIPS4
2017 Sketch based image retrieval via image-aided cross domain learning
abstract
Existing methods on sketch based image retrieval (SBIR) are usually based on the hand-crafted features whose ability of representation is limited. In this paper, we propose a sketch based image retrieval method via image-aided cross domain learning. First, the deep learning model is introduced to learn the discriminative features. However, it needs a large number of images to train the deep model, which is not suitable for the sketch images. Thus, we propose to extend the sketch training images via introducing the real images. Specifically, we initialize the deep models with extra image data, and then extract the generalized boundary from real images as the sketch approximation. The using of generalized boundary is under the assumption that their domain is similar with sketch domain. Finally, the neural network is fine-tuned with the sketch approximation data. Experimental results on Flicker15 show that the proposed method has a strong ability to link the associated image-sketch pairs and the results outperform state-of-the-arts methods.
Jianjun Lei 0001, Kaifu Zheng, Hua Zhang 0008, Xiaochun Cao, Nam Ling, Yonghong Hou
ICIP3
2017 LEAF: Latent Extended Attribute Features Discovery for Visual Classification
abstract
To improve the discrimination of attribute representation, in this paper, we propose to extend the traditional attribute representations via embedding the latent high-order structure between attributes. Specifically, our aim is to construct the Latent Extended Attribute Features (LEAF) for visual classification. Since there only exist weak label for each attribute, we firstly propose a feature selection method to explore the common feature structures across categories. After that, the attribute classifiers are trained based on the selected features. Then, the category specific graph is introduced, which is composed of single attributes and their co-occurrence attribute pairs. This attribute graph is used as the initialized representation of each image. Considering our aim, we should discover the discriminative latent structure between attributes and train the robust category classifiers. To that end, we develop a joint learning objective function which is composed of the high-order representation mining term and the classifier training term. The mining term can both preserve category-specific information and discover the common structure between categories. Based on the discovery representation, the robust visual classifiers could be trained by the classifier term. Finally, an alternating optimization method is designed to seek the optimal solution of our objective function. Experimental results on the challenging datasets demonstrate the advantages of our proposed model over existing work.
Hua Zhang 0008, Rui Wang 0032, Changqing Zhang 0002, Xiaochun Cao
ACM Multimedia1
2017 Multiple Semantic Matching on Augmented N-Partite Graph for Object Co-Segmentation
abstract
Recent methods for object co-segmentation focus on discovering single co-occurring relation of candidate regions representing the foreground of multiple images. However, region extraction based only on low and middle level information often occupies a large area of background without the help of semantic context. In addition, seeking single matching solution very likely leads to discover local parts of common objects. To cope with these deficiencies, we present a new object co-segmentation framework, which takes advantages of semantic information and globally explores multiple co-occurring matching cliques based on an N-partite graph structure. To this end, we first propose to incorporate candidate generation with semantic context. Based on the regions extracted from semantic segmentation of each image, we design a merging mechanism to hierarchically generate candidates with high semantic responses. Second, all candidates are taken into consideration to globally formulate multiple maximum weighted matching cliques, which complement the discovery of part of the common objects induced by a single clique. To facilitate the discovery of multiple matching cliques, an N-partite graph, which inherently excludes intralinks between candidates from the same image, is constructed to separate multiple cliques without additional constraints. Further, we augment the graph with an additional virtual node in each part to handle irrelevant matches when the similarity between the two candidates is too small. Finally, with the explored multiple cliques, we statistically compute pixel-wise co-occurrence map for each image. Experimental results on two benchmark data sets, i.e., iCoseg and MSRC data sets achieve desirable performance and demonstrate the effectiveness of our proposed framework.
Chuan Wang 0002, Hua Zhang 0008, Liang Yang 0002, Xiaochun Cao, Hongkai Xiong
IEEE Trans. Image Process.2
2016 SketchNet: Sketch Classification with Web Images
abstract
In this study, we present a weakly supervised approach that discovers the discriminative structures of sketch images, given pairs of sketch images and web images. In contrast to traditional approaches that use global appearance features or relay on keypoint features, our aim is to automatically learn the shared latent structures that exist between sketch images and real images, even when there are significant appearance differences across its relevant real images. To accomplish this, we propose a deep convolutional neural network, named SketchNet. We firstly develop a triplet composed of sketch, positive and negative real image as the input of our neural network. To discover the coherent visual structures between the sketch and its positive pairs, we introduce the softmax as the loss function. Then a ranking mechanism is introduced to make the positive pairs obtain a higher score comparing over negative ones to achieve robust representation. Finally, we formalize above-mentioned constrains into the unified objective function, and create an ensemble feature representation to describe the sketch images. Experiments on the TUBerlin sketch benchmark demonstrate the effectiveness of our model and show that deep feature representation brings substantial improvements over other state-of-the-art methods on sketch classification.
Hua Zhang 0008, Si Liu 0001, Changqing Zhang 0002, Wenqi Ren, Rui Wang 0032, Xiaochun Cao
CVPR1
2016 Single Image Dehazing via Multi-scale Convolutional Neural Networks
Wenqi Ren, Si Liu 0001, Hua Zhang 0008, Jinshan Pan, Xiaochun Cao, Ming-Hsuan Yang 0001
ECCV (2)3
2015 Diversity-induced Multi-view Subspace Clustering
abstract
In this paper, we focus on how to boost the multi-view clustering by exploring the complementary information among multi-view features. A multi-view clustering framework, called Diversity-induced Multi-view Subspace Clustering (DiMSC), is proposed for this task. In our method, we extend the existing subspace clustering into the multi-view domain, and utilize the Hilbert Schmidt Independence Criterion (HSIC) as a diversity term to explore the complementarity of multi-view representations, which could be solved efficiently by using the alternating minimizing optimization. Compared to other multi-view clustering methods, the enhanced complementarity reduces the redundancy between the multi-view representations, and improves the accuracy of the clustering results. Experiments on both image and video face clustering well demonstrate that the proposed method outperforms the state-of-the-art methods.
Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Hua Zhang 0008
CVPR5
2015 Deep People Counting in Extremely Dense Crowds
abstract
People counting in extremely dense crowds is an important step for video surveillance and anomaly warning. The problem becomes especially more challenging due to the lack of training samples, severe occlusions, cluttered scenes and variation of perspective. Existing methods either resort to auxiliary human and face detectors or surrogate by estimating the density of crowds. Most of them rely on hand-crafted features, such as SIFT, HOG etc, and thus are prone to fail when density grows or the training sample is scarce. In this paper we propose an end-to-end deep convolutional neural networks (CNN) regression model for counting people of images in extremely dense crowds. Our method has following characteristics. Firstly, it is a deep model built on CNN to automatically learn effective features for counting. Besides, to weaken influence of background like buildings and trees, we purposely enrich the training data with expanded negative samples whose ground truth counting is set as zero. With these negative samples, the robustness can be enhanced. Extensive experimental results show that our method achieves superior performance than the state-of-the-arts in term of the mean and variance of absolute difference.
Chuan Wang 0002, Hua Zhang 0008, Liang Yang 0002, Si Liu 0001, Xiaochun Cao
ACM Multimedia2
2015 SLED: Semantic Label Embedding Dictionary Representation for Multilabel Image Annotation
abstract
Most existing methods on weakly supervised image annotation rely on jointly unsupervised feature representation, the components of which are not directly correlated with specific labels. In practical cases, however, there is a big gap between the training and the testing data, say the label combination of the testing data is not always consistent with that of the training. To bridge the gap, this paper presents a semantic label embedding dictionary representation that not only achieves the discriminative feature representation for each label in the image, but also mines the semantic relevance between co-occurrence labels for context information. More specifically, to enhance the discriminative representation of labels, the training data is first divided into a set of overlapped groups by graph shift based on the exclusive label graph. Afterward, given a group of exclusive labels, we try to learn multiple label-specific dictionaries to explicitly decorrelate the feature representation of each label. A joint optimization approach is proposed according to the Fisher discrimination criterion for seeking its solution. Then, to discover the context information hidden in the co-occurrence labels, we explore the semantic relationship between visual words in dictionaries and labels in a multitask learning way with respect to the reconstruction coefficients of the training data. In the annotation stage, with the discriminative dictionaries and exclusive label groups as well as a group sparsity constraint, the reconstruction coefficients of a test image can be easily obtained. Finally, we introduce a label propagation scheme to compute the score of each label for the test image based on its reconstruction coefficients. Experimental results on three challenging data sets demonstrate that our proposed method leads to significant performance gains over existing methods.
Xiaochun Cao, Hua Zhang 0008, Xiaojie Guo 0001, Si Liu 0001, Dan Meng 0002
IEEE Trans. Image Process.2
2014 Image Retrieval and Ranking via Consistently Reconstructing Multi-attribute Queries
Xiaochun Cao, Hua Zhang 0008, Xiaojie Guo 0001, Si Liu 0001, Xiaowu Chen 0001
ECCV (1)2
2014 Action recognition using 3D DAISY descriptor
Xiaochun Cao, Hua Zhang 0008, Qiguang Liu
Mach. Vis. Appl.2
2013 SYM-FISH: A Symmetry-Aware Flip Invariant Sketch Histogram Shape Descriptor
abstract
Recently, studies on sketch, such as sketch retrieval and sketch classification, have received more attention in the computer vision community. One of its most fundamental and essential problems is how to more effectively describe a sketch image. Many existing descriptors, such as shape context, have achieved great success. In this paper, we propose a new descriptor, namely Symmetric-aware Flip Invariant Sketch Histogram (SYM-FISH) to refine the shape context feature. Its extraction process includes three steps. First the Flip Invariant Sketch Histogram (FISH) descriptor is extracted on the input image, which is a flip-invariant version of the shape context feature. Then we explore the symmetry character of the image by calculating the kurtosis coefficient. Finally, the SYM-FISH is generated by constructing a symmetry table. The new SYM-FISH descriptor supplements the original shape context by encoding the symmetric information, which is a pervasive characteristic of natural scene and objects. We evaluate the efficacy of the novel descriptor in two applications, i.e., sketch retrieval and sketch classification. Extensive experiments on three datasets well demonstrate the effectiveness and robustness of the proposed SYM-FISH descriptor.
Xiaochun Cao, Hua Zhang 0008, Si Liu 0001, Xiaojie Guo 0001, Liang Lin 0004
ICCV2
2012 Action recognition based on spatial-temporal pyramid sparse coding
Hua Zhang 0008, Xiaochun Cao
ICPR2
2010 Water Reflection Detection Using a Flip Invariant Shape Detector
abstract
Water reflection detection is a tough task in computer vision, since the reflection is distorted by ripples irregularly. This paper proposes an effective method to detect water reflections. We introduce a descriptor that is not only invariant to scales, rotations and affine transformations, but also tolerant to the flip transformation and even non-rigid distortions, such as ripple effects. We analyze the structure of our descriptor and show how it outperforms the existing mirror feature descriptors in the context of water reflection. The experimental results demonstrate that our method is able to detect the water reflections.
Hua Zhang 0008, Xiaojie Guo 0001, Xiaochun Cao
ICPR1