Dongyue Zhao

dblp:13/8105 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-0135-0614ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
YearPublicationVenuePosition
2024 UNet-: Memory-Efficient and Feature-Enhanced Network Architecture Based on U-Net with Reduced Skip-Connections
Lingxiao Yin, Wei Tao 0001, Dongyue Zhao, Tadayuki Ito, Kinya Osa, Masami Kato, Tse-Wei Chen 0001
ACCV (7)3
2022 Dynamic Training Data Dropout for Robust Deep Face Recognition
abstract
Learning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images.
Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen
IEEE Trans. Multim.5
2021 Augmented Face Representation Learning via Transitive Distillation
abstract
The wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets.
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen
FG5
2021 Adaptive Label Noise Cleaning with Meta-Supervision for Deep Face Recognition
abstract
The training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning for face recognition datasets, which can learn the distribution of the data to be cleaned and make automatic adjustments based on class differences. It first learns re-liable cleaning knowledge from well-labeled noisy data, then gradually transfers it to the target data with meta-supervision to improve performance. A threshold adapter module is also proposed to address the drift problem in transfer learning methods. Extensive experiments clean two noisy in-the-wild face recognition datasets and show the effectiveness of the proposed method to reach state-of-the-art performance on the IJB-C face recognition benchmark.
Yaobin Zhang, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen
ICCV6
2021 SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognition
abstract
Deep face recognition has achieved great success due to large-scale training databases and rapidly developing loss functions. The existing algorithms devote to realizing an ideal idea: minimizing the intra-class distance and maximizing the inter-class distance. However, they may neglect that there are also low quality training images which should not be optimized in this strict way. Considering the imperfection of training databases, we propose that intra-class and inter-class objectives can be optimized in a moderate way to mitigate overfitting problem, and further propose a novel loss function, named sigmoid-constrained hypersphere loss (SFace). Specifically, SFace imposes intra-class and inter-class constraints on a hypersphere manifold, which are controlled by two sigmoid gradient re-scale functions respectively. The sigmoid curves precisely re-scale the intra-class and inter-class gradients so that training samples can be optimized to some degree. Therefore, SFace can make a better balance between decreasing the intra-class distances for clean examples and preventing overfitting to the label noise, and contributes more robust deep face recognition models. Extensive experiments of models trained on CASIA-WebFace, VGGFace2, and MS-Celeb-1M databases, and evaluated on several face recognition benchmarks, such as LFW, MegaFace and IJB-C databases, have demonstrated the superiority of SFace.
Yaoyao Zhong, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen
IEEE Trans. Image Process.4
2020 Global-Local GCN: Large-Scale Label Noise Cleansing for Face Recognition
abstract
In the field of face recognition, large-scale web-collected datasets are essential for learning discriminative representations, but they suffer from noisy identity labels, such as outliers and label flips. It is beneficial to automatically cleanse their label noise for improving recognition accuracy. Unfortunately, existing cleansing methods cannot accurately identify noise in the wild. To solve this problem, we propose an effective automatic label noise cleansing framework for face recognition datasets, FaceGraph. Using two cascaded graph convolutional networks, FaceGraph performs global-to-local discrimination to select useful data in a noisy environment. Extensive experiments show that cleansing widely used datasets, such as CASIA-WebFace, VGGFace2, MegaFace2, and MS-Celeb-1M, using the proposed method can improve the recognition performance of state-of-the-art representation learning methods like Arcface. Further, we cleanse massive self-collected celebrity data, namely MillionCelebs, to provide 18.8M images of 636K identities. Training with the new data, Arcface surpasses state-of-the-art performance by a notable margin to reach 95.62% TPR at 1e-5 FPR on the IJB-C benchmark.
Yaobin Zhang, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen
CVPR6
2013 Occlusion cues for image scene layering
Xiaowu Chen 0001, Dongyue Zhao, Qinping Zhao
Comput. Vis. Image Underst.3
2009 Accurate semantic image labeling by fast Geodesic Propagation
abstract
Motivated by recently raised image semantic labeling problem, this paper studies a fast Geodesic Propagation (GP) algorithm that integrates recognition proposal and image compatibility into a graphical representation. Given the recognition proposal map of the image, the initial seeds are selected as confident pixels standing on local proposal peaks by Mean-shift algorithm. The geodesic distance is then defined on a hybrid manifold, combining the color and boundary features with the recognition proposal map. Based on the geodesic distance, the semantic labeling is simultaneously propagated from the initial seeds of all classes to the rest of image pixels. This inference algorithm is capable of multi-labeling an image of 2-mega pixels in one second (with a common PC). In the experiment, we test on 21 generic semantic categories (sky, road, grass ...) on MSRC dataset, and 17 categories on LHI dataset to evaluate the performance.
Xiaowu Chen 0001, Dongyue Zhao, Yibiao Zhao, Liang Lin 0004
ICIP2