Yaoyao Zhong

dblp:231/1745 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0002-2671-9350ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 7 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark
abstract
Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks.However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition.Existing evaluation methods, which typically focus either on the final answer or on the intermediate reasoning steps, reduce mathematical reasoning to a shallow input-output mapping, overlooking its inherently multi-stage and multi-dimensional cognitive nature.Inspired by Pólya's problem-solving theory, we propose SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: Semantic Understanding, Mathematical Reasoning, Arithmetic Computation, and Reflection & Refinement, and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs.We apply SMART to 22 state-of-the-art open-and closed-source LLMs and uncover substantial discrepancies in their capabilities across dimensions.Our findings reveal genuine weaknesses in current models and motivate a new metric, the All-Pass Score, designed to better capture true problem-solving capability.
Yujie Hou, Yaoyao Zhong, Ting Zhang 0002, Xuetao Ma 0001, Hua Huang 0001
ACL (1)3
2025 VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things
abstract
Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To address the challenge, we build VIoTGPT, the framework based on LLMs to correctly interact with humans, query knowledge videos, and invoke vision models to analyze multimedia data collaboratively. To support VIoTGPT and related future works, we meticulously crafted the VIoT-Tool dataset, including the training dataset and the benchmark involving 11 representative vision models across three categories based on semi-automatic annotations. To guide LLM to act as the intelligent agent towards intelligent VIoT, we resort to ReAct instruction tuning method based on VIoT-Tool to learn the tool capability. Quantitative and qualitative experiments and analyses demonstrate the effectiveness of VIoTGPT. We believe VIoTGPT contributes to improving human-centered experiences in VIoT applications.
Yaoyao Zhong, Mengshi Qi, Yuhan Qiu, Huadong Ma
AAAI1
2025 AdvCloak: Customized adversarial cloak for privacy protection
Xuannan Liu, Yaoyao Zhong, Xing Cui, Yuhang Zhang 0016, Peipei Li 0002, Weihong Deng
Pattern Recognit.2
2025 RegPalm: Toward Large-Scale Open-Set Palmprint Recognition by Reducing Pattern Variance
abstract
Despite the recent significant progress in palmprint recognition, there are still challenges in scaling up this technology for real-world scenarios. One major challenge in developing practical, highly accurate recognition models is the shortage of comprehensive public datasets that can be used to evaluate performance at extremely low false accept rates (FAR). Furthermore, obtaining high-precision recognition models is greatly hindered by pattern variance, a notable challenge with the palmprint modality given the current technology pipeline. To address the above problems, we first collect a palmprint dataset, WebPalm, that contains the largest number of identities as well as images that have been disclosed so far. To reduce pattern variance, we propose RegPalm, a novel framework that unifies palmprint orientations (UPO) and learns pairwise spatial registration of palmprints (PPR) in an end-to-end manner. UPO harmonizes the pattern variance between left and right orientations, hence enhancing the network’s perceptual capabilities. PPR decreases both inter-class and intra-class pattern variance to improve the model’s ability to recognize hard examples. RegPalm reinforces the model by discriminating subtle palmprint features, thereby improving its performance under extremely low FAR. RegPalm not only surpasses the current state-of-the-art by 9.3 percentage points (pp) and 12.2 pp in TAR@FAR=1e-6 under the 1:1 and 1:3 open-set protocols, respectively, but also consistently achieves a 16 pp improvement in TAR@FAR=1e-9 on the WebPalm benchmark. The experimental results fully reveal the practicability and superiority of RegPalm in the real world.
Yaoyao Zhong, Weilong Chai, Huiyuan Fu, Huadong Ma
IEEE Trans. Inf. Forensics Secur.1
2024 Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient Accumulation
abstract
The blooming of social media and face recognition (FR) systems has increased people’s concern about privacy and security. A new type of adversarial privacy cloak (class-universal) can be applied to all the images of regular users, to prevent malicious FR systems from acquiring their identity information. In this work, we discover the optimization dilemma in the existing methods – the local optima problem in large-batch optimization and the gradient information elimination problem in small-batch optimization. To solve these problems, we propose Gradient Accumulation (GA) to aggregate multiple small-batch gradients into a one-step iterative gradient to enhance the gradient stability and reduce the usage of quantization operations. Experiments show that our proposed method achieves high performance on the Privacy-Commons dataset against black-box face recognition models.
Xuannan Liu, Yaoyao Zhong, Weihong Deng, Hongzhi Shi, Xingchen Cui, Yunfeng Yin, Dongchao Wen
ICASSP2
2023 Enhancing Generalization of Universal Adversarial Perturbation through Gradient Aggregation
abstract
Deep neural networks are vulnerable to universal adversarial perturbation (UAP), an instance-agnostic perturbation capable of fooling the target model for most samples. Compared to instance-specific adversarial examples, UAP is more challenging as it needs to generalize across various samples and models. In this paper, we examine the serious dilemma of UAP generation methods from a generalization perspective – the gradient vanishing problem using small-batch stochastic gradient optimization and the local optima problem using large-batch optimization. To address these problems, we propose a simple and effective method called Stochastic Gradient Aggregation (SGA), which alleviates the gradient vanishing and escapes from poor local optima at the same time. Specifically, SGA employs the small-batch training to perform multiple iterations of inner pre-search. Then, all the inner gradients are aggregated as a one-step gradient estimation to enhance the gradient stability and reduce quantization errors. Extensive experiments on the standard ImageNet dataset demonstrate that our method significantly enhances the generalization ability of UAP and outperforms other state-of-the-art methods. The code is available at https://github.com/liuxuannan/Stochastic-Gradient-Aggregation.
Xuannan Liu, Yaoyao Zhong, Yuhang Zhang 0016, Lixiong Qin, Weihong Deng
ICCV2
2023 OPOM: Customized Invisible Cloak Towards Face Privacy Protection
abstract
While convenient in daily life, face recognition technologies also raise privacy concerns for regular users on the social media since they could be used to analyze face images and videos, efficiently and surreptitiously without any security restrictions. In this paper, we investigate the face privacy protection from a technology standpoint based on a new type of customized cloak, which can be applied to all the images of a regular user, to prevent malicious face recognition systems from uncovering their identity. Specifically, we propose a new method, named one person one mask (OPOM), to generate person-specific (class-wise) universal masks by optimizing each training sample in the direction away from the feature subspace of the source identity. To make full use of the limited training images, we investigate several modeling methods, including affine hulls, class centers and convex hulls, to obtain a better description of the feature subspace of source identities. The effectiveness of the proposed method is evaluated on both common and celebrity datasets against black-box face recognition models with different loss functions and network architectures. In addition, we discuss the advantages and potential problems of the proposed method. In particular, we conduct an application study on the privacy protection of a video dataset, Sherlock, to demonstrate the potential practical usage of the proposed method.
Yaoyao Zhong, Weihong Deng
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Video Question Answering: Datasets, Algorithms and Challenges
abstract
This survey aims to organize the recent advances in video question answering (VideoQA) and point towards future directions.We firstly categorize the datasets into: 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, and 2) factoid VideoQA and inference VideoQA, according to the technical challenges in comprehending the questions and deriving the correct answers.We then summarize the VideoQA techniques, including those mainly designed for Factoid QA (such as the early spatio-temporal attention-based methods and the recent Transformer-based ones) and those targeted at explicit relation and logic inference (such as neural modular networks, neural symbolic methods, and graph-structured methods).Aside from the backbone techniques, we also delve into specific models and derive some common and useful insights either for video modeling, question answering, or for cross-modal correspondence learning.Finally, we present the research trends of studying beyond factoid VideoQA to inference VideoQA, as well as towards the robustness and interpretability.Additionally, we maintain a repository, https://github.com/VRU-NExT/ VideoQA, to keep trace of the latest VideoQA papers, datasets, and their open-source implementations if available.With these efforts, we strongly hope this survey could shed light on the follow-up VideoQA research.
Yaoyao Zhong, Wei Ji 0008, Junbin Xiao, Yicong Li 0004, Weihong Deng, Tat-Seng Chua
EMNLP1
2022 Dynamic Training Data Dropout for Robust Deep Face Recognition
abstract
Learning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images.
Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen
IEEE Trans. Multim.1
2021 Augmented Face Representation Learning via Transitive Distillation
abstract
The wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets.
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen
FG3
2021 Adaptive Label Noise Cleaning with Meta-Supervision for Deep Face Recognition
abstract
The training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning for face recognition datasets, which can learn the distribution of the data to be cleaned and make automatic adjustments based on class differences. It first learns re-liable cleaning knowledge from well-labeled noisy data, then gradually transfers it to the target data with meta-supervision to improve performance. A threshold adapter module is also proposed to address the drift problem in transfer learning methods. Extensive experiments clean two noisy in-the-wild face recognition datasets and show the effectiveness of the proposed method to reach state-of-the-art performance on the IJB-C face recognition benchmark.
Yaobin Zhang, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen
ICCV3
2021 Towards Transferable Adversarial Attack Against Deep Face Recognition
abstract
Face recognition has achieved great success in the last five years due to the development of deep learning methods. However, deep convolutional neural networks (DCNNs) have been found to be vulnerable to adversarial examples. In particular, the existence of transferable adversarial examples can severely hinder the robustness of DCNNs since this type of attacks can be applied in a fully black-box manner without queries on the target system. In this work, we first investigate the characteristics of transferable adversarial attacks in face recognition by showing the superiority of feature-level methods over label-level methods. Then, to further improve transferability of feature-level adversarial examples, we propose DFANet, a dropout-based method used in convolutional layers, which can increase the diversity of surrogate models and obtain ensemble-like effects. Extensive experiments on state-of-the-art face models with various training databases, loss functions and network architectures show that the proposed method can significantly enhance the transferability of existing attack methods. Finally, by applying DFANet to the LFW database, we generate a new set of adversarial face pairs that can successfully attack four commercial APIs without any queries. This TALFW database is available to facilitate research on the robustness and defense of deep face recognition.
Yaoyao Zhong, Weihong Deng
IEEE Trans. Inf. Forensics Secur.1
2021 SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognition
abstract
Deep face recognition has achieved great success due to large-scale training databases and rapidly developing loss functions. The existing algorithms devote to realizing an ideal idea: minimizing the intra-class distance and maximizing the inter-class distance. However, they may neglect that there are also low quality training images which should not be optimized in this strict way. Considering the imperfection of training databases, we propose that intra-class and inter-class objectives can be optimized in a moderate way to mitigate overfitting problem, and further propose a novel loss function, named sigmoid-constrained hypersphere loss (SFace). Specifically, SFace imposes intra-class and inter-class constraints on a hypersphere manifold, which are controlled by two sigmoid gradient re-scale functions respectively. The sigmoid curves precisely re-scale the intra-class and inter-class gradients so that training samples can be optimized to some degree. Therefore, SFace can make a better balance between decreasing the intra-class distances for clean examples and preventing overfitting to the label noise, and contributes more robust deep face recognition models. Extensive experiments of models trained on CASIA-WebFace, VGGFace2, and MS-Celeb-1M databases, and evaluated on several face recognition benchmarks, such as LFW, MegaFace and IJB-C databases, have demonstrated the superiority of SFace.
Yaoyao Zhong, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen
IEEE Trans. Image Process.1
2020 Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu
ECCV (15)3
2019 Unequal-Training for Deep Face Recognition With Long-Tailed Noisy Data
abstract
Large-scale face datasets usually exhibit a massive number of classes, a long-tailed distribution, and severe label noise, which undoubtedly aggravate the difficulty of training. In this paper, we propose a training strategy that treats the head data and the tail data in an unequal way, accompanying with noise-robust loss functions, to take full advantage of their respective characteristics. Specifically, the unequal-training framework provides two training data streams: the first stream applies the head data to learn discriminative face representation supervised by Noise Resistance loss; the second stream applies the tail data to learn auxiliary information by gradually mining the stable discriminative information from confusing tail classes. Consequently, both training streams offer complementary information to deep feature learning. Extensive experiments have demonstrated the effectiveness of the new unequal-training framework and loss functions. Better yet, our method could save a significant amount of GPU memory. With our method, we achieve the best result on MegaFace Challenge 2 (MF2) given a large-scale noisy training data set.
Yaoyao Zhong, Weihong Deng, Jiani Hu, Jianteng Peng, Xunqiang Tao, Yaohai Huang
CVPR1
2019 Exploring Features and Attributes in Deep Face Recognition Using Visualization Techniques
abstract
Deep convolutional neural networks (CNNs) currently have achieved state-of-the-art results on face recognition; yet, the understanding behind the success of the deep face model is still lacking. In particular, it is still unclear the inner workings of deep face model. What effective features does a deep face model learn? What do these features represent and what is the sematic meaning of them? This work explores this problem by analyzing the classic network VGGFace using deep visualization techniques. We first explore features computed by neurons, investigating characters of features like diversity, invariance, discrimination. It's worth noting that the middle layer is the least robust to transform, which contradicts the conventional view that robustness to transform increases as the network going deeper. The most significant phenomenon we find is that high level features are correspond with complex face attributes which human could not describe using a few words. We present a quantitative analysis on these face attributes perceived by deep CNNs, understanding them and the complex relationships between them. Additionally, we also focus on the significant point, the pose invariance in face recognition. Our research is the first work to understand the inner works of deep face models, elucidating some particular phenomena in deep face recognition.
Yaoyao Zhong, Weihong Deng
FG1
2019 Fair Loss: Margin-Aware Reinforcement Learning for Deep Face Recognition
abstract
Recently, large-margin softmax loss methods, such as angular softmax loss (SphereFace), large margin cosine loss (CosFace), and additive angular margin loss (ArcFace), have demonstrated impressive performance on deep face recognition. These methods incorporate a fixed additive margin to all the classes, ignoring the class imbalance problem. However, imbalanced problem widely exists in various real-world face datasets, in which samples from some classes are in a higher number than others. We argue that the number of a class would influence its demand for the additive margin. In this paper, we introduce a new margin-aware reinforcement learning based loss function, namely fair loss, in which each class will learn an appropriate adaptive margin by Deep Q-learning. Specifically, we train an agent to learn a margin adaptive strategy for each class, and make the additive margins for different classes more reasonable. Our method has better performance than present large-margin loss functions on three benchmarks, Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace, which demonstrates that our method could learn better face representation on imbalanced face datasets.
Weihong Deng, Yaoyao Zhong, Jiani Hu, Xunqiang Tao, Yaohai Huang
ICCV3
2019 Adversarial Learning With Margin-Based Triplet Embedding Regularization
abstract
The Deep neural networks (DNNs) have achieved great success on a variety of computer vision tasks, however, they are highly vulnerable to adversarial attacks. To address this problem, we propose to improve the local smoothness of the representation space, by integrating a margin-based triplet embedding regularization term into the classification objective, so that the obtained models learn to resist adversarial examples. The regularization term consists of two steps optimizations which find potential perturbations and punish them by a large margin in an iterative way. Experimental results on MNIST, CASIA-WebFace, VGGFace2 and MS-Celeb-1M reveal that our approach increases the robustness of the network against both feature and label adversarial attacks in simple object classification and deep face recognition.
Yaoyao Zhong, Weihong Deng
ICCV1
2018 Deep Difference Analysis in Similar-looking Face recognition
abstract
Deep convolutional neural networks (DCNNs) have recently demonstrated impressive performance in face recognition. However, there is no clear understanding of what difference they find between two similar-looking faces. In this paper, we propose a visualization method that gives insight into difference of similar-looking faces found by DCNNs. This method, used as an assistant role, could help human to identify people who try to invade the biometric system using a similar-looking face. We design a crowdsourcing task to evaluate our method. With assistance of our method, accuracy of participants is greatly increased by 8%, which is also better than the accuracy of network, while participants get little improvement with assistance of Deconvolutional network or Gradient Back-propagation. The experiment result suggests that our method makes a difference in human-machine cooperation.
Yaoyao Zhong, Weihong Deng
ICPR1