EDBT 2026 Demo / reviewers in the wild / expert
Yu Zheng 0006
dblp:87/1585-6
· DBLP profile ↗
32ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0001-8336-5651ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 7 since 2021Security and privacy · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ResProto-FD: Visual-Language Residual Prototype Sets for Generalized Face Forgery DetectionabstractWith the rapid development of generative models, such as generative adversarial networks and diffusion models, the task of face forgery detection has emerged, aiming to identify forged faces in real-world scenarios. A key challenge for current face forgery detection models is improving generalization to unknown forgeries. To address this, we propose ResProto-FD, a framework that constructs residual prototype sets to capture diverse forgery cues and discriminative differences from real faces. Our novel perspective collects prototypes from the most informative residual features generated during training, enabling better representation of various forgery traces and real-vs-fake distinctions. First, we introduce a Visual-Language Residual Learning (VLRL) module based on the CLIP model. This module constructs residual features between image and text embeddings to capture inconsistencies between visual features and associated textual semantics. In doing so, it guides the model to attend to subtle visual forgery clues and enhances the discriminative power of image representations. Furthermore, we design a Gradient-aware Residual Prototypes (GRP) mechanism— a dynamic collection strategy that selectively stores uncertain residual features based on gradient signals to build the prototype sets. This enhances the model’s ability to generalize to unknown forgery types. Extensive experiments across various datasets and forgery methods demonstrate that ResProto-FD significantly improves generalization performance and consistently outperforms state-of-the-art methods. Jiuyao Jing, Yu Zheng 0006, Chunlei Peng |
AAAI | 2 |
| 2026 | Cross-color space feature fusion for anomaly detection-based face anti-spoofing
Yu Zheng 0006, Chunlei Peng |
Pattern Recognit. Lett. | 2 |
| 2026 | Contextual Masking Distillation for Network Traffic Anomaly DetectionabstractNetwork traffic anomaly detection is critical for cybersecurity but faces challenges in accurately identifying malicious activities. Recent zero-positive approaches, which use only normal training data under the reconstruction paradigm, have shown progress. However, encrypted network traffic obscures normal–anomalous distinctions, causing confused modeling. In addition, the “identical shortcut” problem, where models reconstruct any input with similar fidelity, produces suboptimal representations and indistinguishable detection. To address these limitations, this paper introduces ConMD, a novel Contextual Masking Knowledge Distillation framework. ConMD features distillation paradigm for discriminative representations and then pursues two objectives: effective contextual information modeling and a comprehensive anomaly metric. Specifically, we introduce context-aware local-global attention mechanisms for the student network's backbone, which capture both intra-packet and inter-packet dependencies. Additionally, a context-enhanced masking training strategy is designed to facilitate contextual interactions in normal flows. Given the structural characteristics of network traffic, we also present a new anomaly scoring with multi-view awareness, which perceive comprehensive traffic patterns. ConMD combines insights from both packet- and flow-level views to highlight deviations in anomalous network flows, thereby improving detection accuracy. Extensive experiments on three real-world datasets validate the effectiveness of ConMD, yielding consistent improvements over state-of-the-art baselines, achieving up to 2.8% and 5.1% AUC gains on the DataCon2020 and CIC-IDS2017 datasets, respectively. Our model code will be released at https://github.com/ikun0124/ConMD. Xinglin Lian, Yu Zheng 0006, Fan Zhou 0002, Chunlei Peng, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Cross-Domain Matrix Compression Adaptation for Zero-Shot Sketch-Based Image RetrievalabstractThis paper explores an efficient parameter fine-tuning strategy for zero-shot sketch-based image retrieval (ZS-SBIR). We highlight a key finding: through efficient parameter fine-tuning, competitive retrieval accuracy can be achieved using only 12% of the full parameters. This superior performance is attributed to low-rank matrix compression in the redundant parameter space, which extracts more discriminative effective feature dimensions. Specifically, to fully leverage the potential of pretrained models, we propose a cross-domain matrix compression adaptation method. For lower-level features, we apply a general low-rank decomposition to extract shared basic shapes or contours information across modalities. To mitigate overfitting to local similarities, we propose a domain-specific matrix compression module that guides the model in learning high-level abstractions essential for sketch retrieval. Our method is simple and effective, balancing both general semantic information and feature variations across domains within the same category. Experimental results on the ZS-SBIR benchmark dataset show that our method not only outperforms existing state-of-the-art methods, but also requires significantly fewer training parameters. Decheng Liu, Yu Zheng 0006, Chunlei Peng |
IJCNN | 3 |
| 2025 | Face Anti-spoofing based on Contour-constrained Anomaly DetectionabstractFace recognition systems are increasingly susceptible to spoofing attacks, which pose serious security risks in biometric authentication. Traditional face anti-spoofing is typically framed as a binary classification problem, but the diversity and evolving nature of spoofing techniques hinder generalization. To address this, we reformulate face anti-spoofing as an anomaly detection task, training only on normal data. Existing reconstruction-based anomaly detection methods often rely on reconstruction errors to identify anomalies, but they tend to produce blurry images that lack high-frequency details such as edges and textures. To overcome this limitation, we propose a contour-guided face anti-spoofing approach that enhances reconstruction by preserving fine-grained details. Extensive experiments demonstrate that our method outperforms state-of-the-art techniques in face anti-spoofing. Chunlei Peng, Yu Zheng 0006 |
ICMR | 4 |
| 2025 | Metal Surface Defect Detection based on Variable Mask Ratio Multi-scale ReconstructionabstractMetal surface defect detection plays a critical role in ensuring product quality in industrial manufacturing. Traditional detection methods often suffer from incomplete feature extraction and limited defective samples, hindering the performance of supervised learning. This paper explores the potential of Vision Transformer (ViT) for metal surface defect detection by addressing domain shift issues and improving training speed. We propose a semi-supervised anomaly detection method based on a Transformer network, using a variable masking ratio to integrate generative modeling with representation learning. By fusing features from a convolutional-Transformer encoder, we utilize a feature pyramid network and a frozen pre-trained hierarchical encoder to generate multi-scale features. By employing these technologies, the proposed method captures fine-grained anomalies effectively, providing a robust solution for detecting diverse defect types. The experimental results on benchmark datasets demonstrate that the proposed method outperforms state-of-the-art defect detection techniques. Yu Zheng 0006, Chunlei Peng |
ICMR | 1 |
| 2025 | Facing Anomalies Head-On: Network Traffic Anomaly Detection via Uncertainty-Inspired Inter-Sample DifferencesabstractNetwork traffic anomaly detection is pivotal in cybersecurity, especially as data volume grows and security requirement intensifies. This study addresses critical limitations in existing reconstruction-based methods, which quantify anomalies relying on intra-sample differences and struggle to detect drifted anomalies. In response, we propose a novel approach, the Uncertainty-Inspired Inter-Sample Differences (UnDiff) method, which leverages model uncertainty to enhance anomaly detection capabilities, particularly in scenarios involving anomaly drift. By employing evidential learning, the UnDiff model gathers evidence to minimize uncertainty in normal network traffic, enhancing its ability to differentiate between normal and anomalous traffic. To overcome the limitations of intra-sample difference quantification in reconstruction-based methods, we propose a novel anomaly score based on inter-sample uncertainty deviation that directly quantifies the anomaly degree. Benefiting from a concise model design and parameterized uncertainty quantification, UnDiff achieves high efficiency. Extensive experiments on three benchmarks demonstrate UnDiff's superior performance in detecting both undrifted and drifted anomalies with minimal computational overhead. Xinglin Lian, Chengtai Cao, Xovee Xu, Yu Zheng 0006, Fan Zhou 0002 |
WWW | 5 |
| 2025 | Semi-supervised anomaly traffic detection via multi-frequency reconstruction
Xinglin Lian, Yu Zheng 0006, Zhangxuan Dang, Chunlei Peng, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2025 | Posture-Aware Robust Person Re-Identification via Optimal Transport Calibration
Ruiying Lu, Yalin Sun, Chunlei Peng, Yu Zheng 0006 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Semi-Supervised Learning for Anomaly Traffic Detection via Bidirectional Normalizing FlowsabstractWith the rapid development of the Internet, various types of anomaly traffic are threatening network security. However, the difficulty of collecting and labelling anomalous traffic is a significant challenge, so this paper proposes a semi-supervised anomaly detection framework. Considering normal and abnormal traffic have different data distributions, our framework can generate pseudo anomaly samples without prior knowledge of anomalies to achieve the detection of anomaly data. The framework comprises three principal components. Firstly, a pre-trained feature extractor is employed to extract a feature representation of the network traffic. Secondly, a bidirectional normalizing flow module establishes a reversible transformation between the latent data distribution and a Gaussian space. Through this bidirectional mapping, samples first undergo transformation manipulation within the Gaussian distribution space, and are then transported through the generative direction of normalizing flows, translating mathematical transformations into semantic feature evolutions in the latent data space. Finally, a simple classifier explicitly learns the potential differences between anomaly and normal samples to facilitate better anomaly detection. During inference, our framework requires only two modules to detect anomalous samples, leading to a considerable reduction in model size. According to the experiments, our method achieves the state-of-the-art results on the common benchmarking datasets of anomaly network traffic detection. Furthermore, it exhibits good generalisation performance across datasets. Zhangxuan Dang, Yu Zheng 0006, Xinglin Lian, Chunlei Peng, Qiuyu Chen, Xinbo Gao 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Spatial-Frequency Dual-Stream Reconstruction for Deepfake Detection
Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001 |
PRCV (11) | 4 |
| 2024 | Audio-Driven Face Photo-Sketch Video Generation
Siyue Zhou, Qun Guan, Chunlei Peng, Decheng Liu, Yu Zheng 0006 |
PRICAI (3) | 5 |
| 2024 | Face Anti-spoofing based on Multi-modal Dual-stream Anomaly DetectionabstractContemporary research often addresses the face anti-spoofing challenge through a classification paradigm. However, due to the rapidly changing and diverse characteristics of spoofing faces, it is unreasonable to regard all spoofing faces as a single category. Moreover, the swift evolution of spoofing techniques can render trained detectors ineffective. Anomaly detection offers a solution to these challenges by training exclusively on normal samples, distinguishing living samples as normal and non-living samples as anomalies. This paper presents a novel face anti-spoofing approach grounded in anomaly detection. It devises an RGB-D dual-stream network architecture that integrates multi-scale features from RGB and depth modalities via intermediate fusion. In addition, it also incorporates adversarial learning to facilitate network training. Furthermore, it proposes a novel multi-stage anomaly score generation technique for face anti-spoofing. Our experiments on three public datasets demonstrate the superiority of our method over comparative approaches. Jiuyao Jing, Yu Zheng 0006, Chunlei Peng |
TrustCom | 2 |
| 2024 | Multi-view multi-label network traffic classification based on MLP-Mixer neural network
Yu Zheng 0006, Zhangxuan Dang, Xinglin Lian, Chunlei Peng, Xinbo Gao 0001 |
Comput. Networks | 1 |
| 2023 | A Semi-Supervised Anomaly Network Traffic Detection Framework via Multimodal Traffic Information FusionabstractAnomaly traffic detection is a crucial issue in the cyber-security field. Previously, many researchers regarded anomaly traffic detection as a supervised classification problem. However, in real scenarios, anomaly network traffic is unpredictable, dynamically changing and difficult to collect. To address these limitations, we employ anomaly detection setting to propose a novel semi-supervised anomaly network traffic detection framework. It only learns features of normal samples during the training phase. Our framework utilizes low-pass filtering to extract multi-scale low-frequency information from 2-D traffic image. Furthermore, we design a two-stage fusion scheme to incorporate information from original and multi-scale low-frequency traffic image modalities. We conduct experiments on two public datasets: ISCX Tor-nonTor and USTC-TFC2016. The experimental results show that our method outperforms current state-of-the-art anomaly detection methods. Yu Zheng 0006, Xinglin Lian, Zhangxuan Dang, Chunlei Peng, Chao Yang 0016, Jianfeng Ma 0001 |
CIKM | 1 |
| 2023 | FedForgery: Generalized Face Forgery Detection With Residual Federated LearningabstractWith the continuous development of deep learning in the field of image generation models, a large number of vivid forged faces have been generated and spread on the Internet. These high-authenticity artifacts could grow into a threat to society security. Existing face forgery detection methods directly utilize the obtained public shared or centralized data for training but ignore the personal privacy and security issues when personal data couldn’t be centralizedly shared in real-world scenarios. Additionally, different distributions caused by diverse artifact types would further bring adverse influences on the forgery detection task. To solve the mentioned problems, the paper proposes a novel generalized residual Federated learning for face Forgery detection (FedForgery). The designed variational autoencoder aims to learn robust discriminative residual feature maps to detect forgery faces (with diverse or even unknown artifact types). Furthermore, the general federated learning strategy is introduced to construct distributed detection model trained collaboratively with multiple local decentralized devices, which could further boost the representation generalization. Experiments conducted on publicly available face forgery detection datasets prove the superior performance of the proposed FedForgery. The designed novel generalized face forgery detection protocols and source code would be publicly available at https://github.com/GANG370/FedForgery. Decheng Liu, Zhan Dang, Chunlei Peng, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | SketchCLIP: Text-based Attribute Manipulation for Face Sketch SynthesisabstractThis paper proposes a method of modifying the face sketch with text descriptions. Face sketch is widely used in the criminal field and digital entertainment field. Forensic painters usually draw face sketches based on descriptions provided by witnesses or clients. However, drawing a face sketch often takes lots of time and effort. Existing face sketch synthesis studies have not considered text-based sketch manipulation, and we find that applying text-driven editing methods on natural images directly to face sketches causes severe distortion of generated results. Therefore, this paper proposes a novel text-based attribute manipulation method for face sketch synthesis, named SketchCLIP. Our approach adopts text-driven attribute manipulation by using the powerful Contrastive Language-Image Pre-Training (CLIP) model, which not only conforms to the current drawing process of face sketches but also does not require tedious manual operations and allows for more diverse modifications. Besides, we design an intra-modality fine-tuning module to eliminate distortion and improve the quality of the modified face sketch. Through extensive comparison experiments on public face sketch datasets, our method is demonstrated to be very excellent in the effectiveness of the face sketch processing and the quality of modified results. Mengdi Dong, Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001 |
IJCB | 4 |
| 2022 | An Object Point Set Inductive Tracker for Multi-Object Tracking and SegmentationabstractMulti-object tracking and segmentation (MOTS) is a derivative task of multi-object tracking (MOT). The new setting encourages the learning of more discriminative high-quality embeddings. In this paper, we focus on the problem of exploring the relationship between the segmenter and the tracker, and propose an efficient Object Point set Inductive Tracker (OPITrack) based on it. First, we discover that after a single attention layer, the high-dimensional, key point embedding will show feature averaging. To alleviate this phenomenon, we propose an embedding generalization training strategy for sparse training and dense testing. This strategy allows the network to increase randomness in training and encourages the tracker to learn more discriminative features. In addition, to learn the desired embedding space, we propose a general Trip-hard sample augmentation loss. The loss uses patches that are not distinguishable by the segmenter to join the feature learning and force the embedding network to learn the difference between false positives and true positives. Our method was validated on two MOTS benchmark datasets and achieved promising results. In addition, our OPITrack can achieve better performance for the raw model while costing less video memory (VRAM) at training time. Yan Gao 0025, Yu Zheng 0006, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Integrating Part of Speech Guidance for Image CaptioningabstractTo generate an image caption, firstly, the content of the image should be fully understood; and then the semantic information contained in the image should be described using a phrase or statement that conforms to certain grammatical rules. Thus, it requires techniques from both computer vision and natural language processing to connect the two different media forms together, which is highly challenging. To adaptively adjust the effect of visual information and language information on the captioning process, in this paper, the part of speech information is proposed to novelly integrate with image captioning models based on the encoder-decoder framework. First, a part of speech prediction network is proposed to analyze and model the part of speech sequences for the words in natural language sentences; then, different mechanisms are proposed to integrate the part of speech guidance information with merge-based and inject-based image captioning models, respectively; finally, according to the integrated frameworks, a multi-task learning paradigm is proposed to facilitate model training. Experiments are conducted on two widely used image captioning datasets, Flickr30 k and COCO, and the results have validated that the image captions generated by the proposed method contain more accurate visual information and comply with language habits and grammar rules better. Ji Zhang 0005, Yu Zheng 0006, Jianping Fan 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Adaptive Fractional Dilated Convolution Network for Image Aesthetics AssessmentabstractTo leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, composition-preserving and parameter-free, is developed to tackle this issue natively in convolutional kernel level. Specifically, the fractional dilated kernel is adaptively constructed according to the image aspect ratios, where the interpolation of nearest two integer dilated kernels are used to cope with the misalignment of fractional sampling. Moreover, we provide a concise formulation for mini-batch training and utilize a grouping strategy to reduce computational overhead. As a result, it can be easily implemented by common deep learning libraries and plugged into popular CNN architectures in a computation-efficient manner. Our experimental results demonstrate that our proposed method achieves state-of-the-art performance on image aesthetics assessment over the AVA dataset. Qiuyu Chen, Wei Zhang 0016, Yi Xu 0003, Yu Zheng 0006, Jianping Fan 0001 |
CVPR | 6 |
| 2020 | Hierarchical convolutional neural network via hierarchical cluster validity based visual tree learning
Yu Zheng 0006, Qiuyu Chen, Jianping Fan 0001, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2020 | Exploiting Related and Unrelated Tasks for Hierarchical Metric Learning and Image ClassificationabstractIn multi-task learning, multiple interrelated tasks are jointly learned to achieve better performance. In many cases, if we can identify which tasks are related, we can also clearly identify which tasks are unrelated. In the past, most researchers emphasized exploiting correlations among interrelated tasks while completely ignoring the unrelated tasks that may provide valuable prior knowledge for multi-task learning. In this paper, a new approach is developed to hierarchically learn a tree of multi-task metrics by leveraging prior knowledge about both the related tasks and unrelated tasks. First, a visual tree is constructed to hierarchically organize large numbers of image categories in a coarse-to-fine fashion. Over the visual tree, a multi-task metric classifier is learned for each node by exploiting both the related and unrelated tasks, where the learning tasks for training the classifiers for the sibling child nodes under the same parent node are treated as the interrelated tasks, and the others are treated as the unrelated tasks. In addition, the node-specific metric for the parent node is propagated to its sibling child nodes to control inter-level error propagation. Our experimental results demonstrate that our hierarchical metric learning algorithm achieves better results than other state-of-the-art algorithms. Yu Zheng 0006, Jianping Fan 0001, Ji Zhang 0005, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Discriminative Fast Hierarchical Learning for Multiclass Image ClassificationabstractIn this article, a discriminative fast hierarchical learning algorithm is developed for supporting multiclass image classification, where a visual tree is seamlessly integrated with multitask learning to achieve fast training of the tree classifier hierarchically (i.e., a set of structural node classifiers over the visual tree). By partitioning a large number of categories hierarchically in a coarse-to-fine fashion, a visual tree is first constructed and further used to handle data imbalance and identify the interrelated learning tasks automatically (e.g., the tasks for learning the node classifiers for the sibling child nodes under the same parent node are strongly interrelated), and a multitask SVM classifier is trained for each nonleaf node to achieve more effective separation of its sibling child nodes at the next level of the visual tree. Both the internode visual similarities and the interlevel visual correlations are utilized to train more discriminative multitask SVM classifiers and control the interlevel error propagation effectively, and a stochastic gradient descent (SGD) algorithm is developed for learning such multitask SVM classifiers with higher efficiency. Our experimental results have demonstrated that our fast hierarchical learning algorithm can achieve very competitive results on both the classification accuracy rates and the computational efficiency. Yu Zheng 0006, Jianping Fan 0001, Ji Zhang 0005, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | A novel joint dictionary framework for sparse hyperspectral unmixing incorporating spectral library
Lin Qi 0004, Jie Li 0001, Xinbo Gao 0001, Ying Wang 0007, Chongyue Zhao, Yu Zheng 0006 |
Neurocomputing | 6 |
| 2019 | Learning multi-layer coarse-to-fine representations for large-scale image classification
Ji Zhang 0005, Yu Zheng 0006, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2019 | Exploiting Mid-Level Semantics for Large-Scale Complex Video ClassificationabstractAs the amount of available video data has grown substantially, automatic video classification has become an urgent yet challenging task. Most video classification methods focus on acquiring discriminative spacial visual features and motion patterns for video representation, especially deep learning methods, which have achieved very good results on action recognition problems. However, the performance of most of these methods drastically degenerates for more generic video classification tasks where the video contents are much more complex. Thus, in this paper, the mid-level semantics of videos are exploited to bridge the semantic gap between low-level features and high-level video semantics. Inspired by the term ``frequency-inverse document frequency'', a word weighting method for the problem of text classification is introduced to the video domain. The visual objects in videos are regarded as the words in texts, and two new weighting methods are proposed to encode videos by weighting visual objects according to the characteristics of videos. In addition, the semantic similarities between video categories and visual objects are introduced from the text domain as privileged information to facilitate classifier training on the obtained semantic representations of videos. The proposed semantic encoding method (semantic stream) is then fused with the popular two-stream CNN model for the final classification results. Experiments are conducted on two large-scale complex video datasets, CCV and ActivityNet. The experimental results validate the effectiveness of the proposed methods. Ji Zhang 0005, Yu Zheng 0006, Jianping Fan 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | From Text to Video: Exploiting Mid-Level Semantics for Large-Scale Video ClassificationabstractAutomatically classifying large scale of video data is an urgent yet challenging task. To bridge the semantic gap between low-level features and high-level video semantics, we propose a method to represent videos with their mid-level semantics. Inspired by the problem of text classification, we regard the visual objects in videos as the words in documents, and adapt the TF-IDF word weighting method to encode videos by visual objects. Some extensions upon the proposed method are also made according to the characteristics of videos. We integrate the proposed semantic encoding method with the popular two-stream CNN model for video classification. Experiments are conducted on two large-scale video datasets, CCV and ActivityNet. The experimanetal results validates the effectiveness of our method. Ji Zhang 0005, Xiao Wang 0019, Yu Zheng 0006, Jianping Fan 0001 |
ICPR | 4 |
| 2018 | A Hierarchical Cluster Validity Based Visual Tree Learning for Hierarchical Classification
Yu Zheng 0006, Jianping Fan 0001, Ji Zhang 0005, Xinbo Gao 0001 |
PRCV (3) | 1 |
| 2017 | Indoor scene recognition via multi-task metric multi-kernel learning from RGB-D images
Yu Zheng 0006, Xinbo Gao 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Hierarchical learning of multi-task sparse metrics for large-scale image classification
Yu Zheng 0006, Jianping Fan 0001, Ji Zhang 0005, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2017 | HD-MTL: Hierarchical Deep Multi-Task Learning for Large-Scale Visual RecognitionabstractIn this paper, a hierarchical deep multi-task learning (HD-MTL) algorithm is developed to support large-scale visual recognition (e.g., recognizing thousands or even tens of thousands of atomic object classes automatically). First, multiple sets of multi-level deep features are extracted from different layers of deep convolutional neural networks (deep CNNs), and they are used to achieve more effective accomplishment of the coarseto- fine tasks for hierarchical visual recognition. A visual tree is then learned by assigning the visually-similar atomic object classes with similar learning complexities into the same group, which can provide a good environment for determining the interrelated learning tasks automatically. By leveraging the inter-task relatedness (inter-class similarities) to learn more discriminative group-specific deep representations, our deep multi-task learning algorithm can train more discriminative node classifiers for distinguishing the visually-similar atomic object classes effectively. Our hierarchical deep multi-task learning (HD-MTL) algorithm can integrate two discriminative regularization terms to control the inter-level error propagation effectively, and it can provide an end-to-end approach for jointly learning more representative deep CNNs (for image representation) and more discriminative tree classifier (for large-scale visual recognition) and updating them simultaneously. Our incremental deep learning algorithms can effectively adapt both the deep CNNs and the tree classifier to the new training images and the new object classes. Our experimental results have demonstrated that our HD-MTL algorithm can achieve very competitive results on improving the accuracy rates for large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Yu Zheng 0006, Ji Zhang 0005, Jun Yu 0002, Jinye Peng 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Multispectral image classification based on improved weighted MRF Bayesian
Zhaobin Cui, Ying Wang 0007, Xinbo Gao 0001, Jie Li 0001, Yu Zheng 0006 |
Neurocomputing | 5 |