Dongyu Zhang 0002

dblp:69/65-2 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-7595-0137ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
abstract
Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic understanding, particularly when dealing with complex textual descriptions.To overcome these limitations, we propose DEIG, a novel framework for fine-grained and controllable multi-instance generation. DEIG integrates an instance Detail Extractor (IDE) that transforms text encoder embeddings into compact, instance-aware representations, and a Detail Fusion Module (DFM) that applies instance-based masked attention to prevent attribute leakage across instances. These components enable DEIG to generate visually coherent multi-instance scenes that precisely match rich, localized textual descriptions. To support fine-grained supervision, we construct a high-quality dataset with detailed, compositional instance captions generated by VLMs. We also introduce DEIG-Bench, a new benchmark with region-level annotations and multi-attribute prompts for both humans and objects.Experiments demonstrate that DEIG consistently outperforms existing approaches across multiple benchmarks in spatial consistency, semantic accuracy, and compositional generalization. Moreover, DEIG functions as a plug-and-play module, making it easily integrable into standard diffusion-based pipelines.
Shiyan Du, Conghan Yue, Dongyu Zhang 0002
AAAI4
2025 LLDB: Efficient Low-Light Image Enhancement with Difffusion Bridge
abstract
This paper investigates a low-light image enhancement method based on the diffusion bridge framework. Currently, low-light image enhancement tasks still face challenges in noise reduction and detail restoration, and existing diffusion model methods are time-consuming and have unstable diffusion processes. We conducted an in-depth study of diffusion bridge theory, integrating the advantages of the end-to-end paradigm of diffusion bridge theory and employing a nonlinear activation network to further enhance the performance of low-light image enhancement tasks. Additionally, we used a Gamma correction module for fine-tuning low-light images, significantly improving performance with almost no extra computational cost. Experiments show that Low-Light image enhancement with Difffusion Bridge (LLDB) far surpasses other methods on LOLv1 and LOLv2 datasets. Code is available at https://github.com/M-Chase/LLDB
Junlong Ma, Conghan Yue, Zhengwei Peng, Dongyu Zhang 0002
ICASSP4
2025 Enhanced Control for Diffusion Bridge in Image Restoration
abstract
Image restoration refers to the process of restoring a damaged low-quality image back to its corresponding high-quality image. Recently, a special type of diffusion bridge model has achieved more advanced results in image restoration. It can transform the direct mapping from low-quality to high-quality images into a diffusion process, restoring low-quality images through a reverse process. However, the current diffusion bridge restoration models do not emphasize the idea of conditional control, which may affect performance. This paper introduces the ECDB model enhancing the control of the diffusion bridge with low-quality images as conditions. Moreover, in response to the characteristic of diffusion models having low denoising level at larger values of t, we also propose a Conditional Fusion Schedule, which more effectively handles the conditional feature information of various modules. Experimental results prove that the ECDB model has achieved state-of-the-art results in many image restoration tasks, including deraining, inpainting and super-resolution. Code is avaliable at https://github.com/Hammour-steak/ECDB.
Conghan Yue, Zhengwei Peng, Junlong Ma, Dongyu Zhang 0002
ICASSP4
2025 Inversion-Free Image Editing via Rectified Flow
abstract
Text-based image editing has advanced significantly with large-scale text-to-image models, but challenges remain: (1) inversion-based methods require substantial resources for inversion and optimization; (2) inversion-free methods struggle to balance controllability and consistency; and (3) incompatibility with faster flow-based models and Diffusion Transformers. In this paper, we propose FlowEdit, an inversion-free text-based image editing framework based on Diffusion Transformers. FlowEdit uses flow consistency sampling for precise image reconstruction, and enhances target images with robust consistency through velocity rectify and attention control mechanism. Extensive experiments demonstrate that FlowEdit excels in editing capabilities and consistency across various tasks while maintaining high efficiency (under 4 seconds on a single 3090 GPU), making it suitable for real-time applications. Our codes are available at https://github.com/XavierPeng319/FlowEdit.git
Zhengwei Peng, Conghan Yue, Tong Duan, Dongyu Zhang 0002
ICME4
2024 EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
abstract
Building scalable vision-language models to learn from diverse, multimodal data remains an open challenge. In this paper, we introduce an Efficient Vision-languagE foundation model, namely EVE, which is one unified multimodal Transformer pre-trained solely by one unified pre-training task. Specifically, EVE encodes both vision and language within a shared Transformer network integrated with modality-aware sparse Mixture-of-Experts (MoE) modules, which capture modality-specific information by selectively switching to different experts. To unify pre-training tasks of vision and language, EVE performs masked signal modeling on image-text pairs to reconstruct masked signals, i.e., image pixels and text tokens, given visible signals. This simple yet effective pre-training objective accelerates training by 4x compared to the model pre-trained with Image-Text Contrastive and Image-Text Matching losses. Owing to the combination of the unified architecture and pre-training task, EVE is easy to scale up, enabling better downstream performance with fewer resources and faster training speed. Despite its simplicity, EVE achieves state-of-the-art performance on various vision-language downstream tasks, including visual question answering, visual reasoning, and image-text retrieval.
Longteng Guo, Shuai Shao 0005, Zehuan Yuan, Liang Lin 0004, Dongyu Zhang 0002
AAAI7
2024 Image Restoration Through Generalized Ornstein-Uhlenbeck Bridge
abstract
Diffusion models exhibit powerful generative capabilities enabling noise mapping to data via reverse stochastic differential equations. However, in image restoration, the focus is on the mapping relationship from low-quality to high-quality images. Regarding this issue, we introduce the Generalized Ornstein-Uhlenbeck Bridge (GOUB) model. By leveraging the natural mean-reverting property of the generalized OU process and further eliminating the variance of its steady-state distribution through the Doob's *h*–transform, we achieve diffusion mappings from point to point enabling the recovery of high-quality images from low-quality ones. Moreover, we unravel the fundamental mathematical essence shared by various bridge models, all of which are special instances of GOUB and empirically demonstrate the optimality of our proposed models. Additionally, we present the corresponding Mean-ODE model adept at capturing both pixel-level details and structural perceptions. Experimental outcomes showcase the state-of-the-art performance achieved by both models across diverse tasks, including inpainting, deraining, and super-resolution. Code is available at https://github.com/Hammour-steak/GOUB.
Conghan Yue, Zhengwei Peng, Junlong Ma, Shiyan Du, Pengxu Wei, Dongyu Zhang 0002
ICML6
2024 Leveraging Cross-Augmentation Consensus and Conflict for Semi-supervised Semantic Segmentation
Junhao Cao, Sibo Huang, Dongyu Zhang 0002
ICPR (27)4
2024 Semantic Correlation Adaptation for Union-Set Multi-label Image Recognition
Tao Pu 0002, Dongyu Zhang 0002, Liang Lin 0004
ICPR (3)3
2024 MKCBlock: Multi Kernel Convolution with Eliminating Dimension Expansion for Real-Time Semantic Segmentation
abstract
In recent years, the integration of the transformer architecture with convolutional layers in large-kernel convolutional neural networks has showcased remarkable accomplishments in semantic segmentation. However, the inclusion of FeedForward-Network-like (FFN-like) structures within these models results in dimension expansion, significantly amplifying memory consumption and diminishing inference speed. Additionally, prevailing real-time semantic segmentation methodologies predominantly employ smaller convolutional kernels, disregarding the potential advantages of larger kernels. These methods are associated with relatively diminutive network structures, rendering them more susceptible to the influence of feature redundancy. In light of these challenges, we’ve introduced a multi-kernel convolution block (MKCBlock) that amalgamates various convolution types and kernel sizes. This innovative and streamlined approach combines the benefits of larger kernels, circumventing dimension expansion, and mitigating feature redundancy. As a consequence, applying the MKCBlock to DDRNet-23-S on the Cityscapes dataset at 131.6 FPS resulted in a 78.2% mIoU. This indicated a 0.4% improvement over the original DDRNet-23-S, which achieved 77.8% mIoU, with a reduction in inference speed of nearly 8 FPS. Similarly, integrating the MKCBlock into PIDNet-S on the Cityscapes dataset at 99.0 FPS yielded a 79.3% mIoU. This surpassed the original PIDNet-S performance of 78.8% mIoU by 0.5%, with a decrease in inference speed of nearly 3 FPS. Overall, our approach maintains a better balance between inference speed and accuracy.
Decheng Jia, Dongyu Zhang 0002
IJCNN2
2023 IRA-FSOD: Instant-Response and Accurate Few-Shot Object Detector
abstract
Aiming at recognizing and localizing objects of novel categories with just a few reference samples, few-shot object detection (FSOD) is quite a challenging task. Previous works rely heavily on the fine-tuning process to transfer their models to the novel categories. They are flawed in the real application since the fine-tuning process is time-consuming and it suffers from serious deterioration on the low-quality support set. Based on the observation, this paper proposes an instant-response and accurate few-shot object detector (IRA-FSOD) that can detect the objects from novel categories without fine-tuning. We carefully analyze the limitations of widely-used Faster R-CNN and transform it to IRA-FSOD. Specifically, we first propose a novel semi-supervised Region Proposal Network (SS-RPN) module and a switch classifier module to precisely recognize the potential foreground instances from novel categories without fine-tuning. Moreover, we introduce two explicit inference strategies into the localization module, including explicit localization score and semi-explicit box regression, to alleviate over-fitting towards the base categories. Extensive experiments demonstrates that the proposed IRA-FSOD not only accomplish few-shot object detection with the instant-response, but also reaches state-of-the-art performance under various FSOD protocols and settings.
Junying Huang, Junhao Cao, Liang Lin 0004, Dongyu Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level Strategy
abstract
Aiming at recognizing the samples from novel categories with few reference samples, few-shot learning (FSL) is a challenging problem. We found that the existing works often build their few-shot model based on the image-level feature by mixing all local-level features, which leads to the discriminative location bias and information loss in local details. To tackle the problem, this paper returns the perspective to the local-level feature and proposes a series of local-level strategies. Specifically, we present (a) a local-agnostic training strategy to avoid the discriminative location bias between the base and novel categories, (b) a novel local-level similarity measure to capture the accurate comparison between local-level features, and (c) a local-level knowledge transfer that can synthesize different knowledge transfers from the base category according to different location features. Extensive experiments justify that our proposed local-level strategies can significantly boost the performance and achieve 2.8%–7.2% improvements over the baseline across different benchmark datasets, which also achieves the state-of-the-art accuracy.
Junying Huang, Keze Wang, Liang Lin 0004, Dongyu Zhang 0002
ICASSP5
2022 Few-Shot Generation By Modeling Stereoscopic Priors
abstract
Few-shot image generation, which aims to generate images from only a few images for a new category, has attracted some research interest in recent years. However, existing few-shot generation methods only focus on 2D images, ignoring 3D information. In this work, we propose a few-shot generative network which leverages 3D priors to improve the diversity and quality of generated images. Inspired by classic graphics rendering pipelines, we unravel the image generation process into three factors: shape, viewpoint and texture. This disentangled representation enables us to make the most of both 3D and 2D information in few-shot generation. To be specific, by changing the viewpoint and extracting textures from different real images, we can generate various new images even in data-scarce settings. Extensive experiments show the effectiveness of our method.
Yuehui Wang, Qing Wang 0018, Dongyu Zhang 0002
ICASSP3
2022 Generalizing Factorization of Gans by Characterizing Convolutional Layers
abstract
Existing unsupervised disentanglement methods in latent space of the Generative Adversarial Networks (GANs) rely on the analysis and decomposition of pre-trained weight matrix. However, they only consider the weight matrix of the fully connected layers, ignoring the convolutional layers which are indispensable for image processing in modern generative models. This results in the learned latent semantics lack inter-pretability, which is unacceptable for image editing tasks. In this paper, we propose a more generalized closed-form factor-ization of latent semantics in GANs, which takes the convolutionallayers into consideration when searching for the under-lying variation factors. Our method can be applied to a wide range of deep generators with just a few lines of code. Exten-sive experiments on multiple GAN models trained on various datasets show that our approach is capable of not only finding semantically meaningful dimensions, but also maintaining the consistency and interpretability of image content.
Yuehui Wang, Qing Wang 0018, Dongyu Zhang 0002
ICME3
2022 Cross-Domain Action Recognition via Prototypical Graph Alignment
abstract
Compared with the well-explored cross-domain image recognition, cross-domain action recognition is a more challenging task because not only spatial but also temporal domain gaps exist across domains. Previous works attempt to bridge the temporal domain gap by aligning the domain-related key segments of videos from source and target domains. However, such practice overlooks the heterogeneous temporal domain gaps among different categories and presents temporal alignment strategies in a class-irrelevant manner. To address this issue, we propose to achieve class-wise temporal alignment for cross-domain action recognition via prototypical graph alignment (PGA). Concretely, we generate segment-level prototypes for the classes of both domains to capture per-class temporal dynamics. Furthermore, intra-domain and inter-domain prototypical graphs are established to mine the temporal relationships between each input video and its corresponding intra-domain and inter-domain prototypes. In this way, a discriminative and domain adaptive video representation is obtained by holistically reasoning cross-domain temporal dynamics. To class-wisely align the cross-domain video representations, each action category is equipped with a customized class-specific domain discriminator for temporal alignment via adversarial learning. Extensive experiments on three benchmarks show that PGA yeilds state-of-the-art performance on the task of cross-domain action recognition.
Junhao Zhong, Pengxu Wei, Dongyu Zhang 0002, Liang Lin 0004
ICME4
2021 Hierarchical Transformer: Unsupervised Representation Learning for Skeleton-Based Human Action Recognition
abstract
The unsupervised representation learning for skeleton-based human action can be utilized in a variety of pose analysis applications. However, previous unsupervised methods focus on modeling the temporal dependencies in sequences, but take less effort in modeling the spatial structure in human action. To this end, we propose a novel unsupervised learning frame-work named Hierarchical Transformer for skeleton-based human action recognition. The Hierarchical Transformer consists of hierarchically aggregated self-attention modules for better capturing the spatial and temporal structure in the skeleton sequences. Furthermore, we propose to predict the motion between adjacent frames as a novel pre-training task for better capturing the long-term dependencies in sequences. Experimental results show that our method outperforms prior state-of-the-art unsupervised methods on NTU RGB+D and NW-UCLA datasets. Besides, our method also achieves state-of-the-art performance when the pre-trained model is transferred to SBU dataset, which demonstrates the generalizability of learned representation.
Yi-Bin Cheng, Xipeng Chen, Pengxu Wei, Dongyu Zhang 0002, Liang Lin 0004
ICME5
2021 Jointly Super Resolution and Degradation Learning on Unpaired Real-World Images
Xuankun Chen, Dongyu Zhang 0002
PDCAT3
2021 Few-Shot Generative Learning by Modeling Stereoscopic Priors
Yuehui Wang, Qing Wang 0018, Dongyu Zhang 0002
PDCAT3
2020 Transferable, Controllable, and Inconspicuous Adversarial Attacks on Person Re-identification With Deep Mis-Ranking
abstract
The success of DNNs has driven the extensive applications of person re-identification (ReID) into a new era. However, whether ReID inherits the vulnerability of DNNs remains unexplored. To examine the robustness of ReID systems is rather important because the insecurity of ReID systems may cause severe losses, e.g., the criminals may use the adversarial perturbations to cheat the CCTV systems. In this work, we examine the insecurity of current best-performing ReID models by proposing a learning-to-mis-rank formulation to perturb the ranking of the system output. As the cross-dataset transferability is crucial in the ReID domain, we also perform a back-box attack by developing a novel multi-stage network architecture that pyramids the features of different levels to extract general and transferable features for the adversarial perturbations. Our method can control the number of malicious pixels by using differentiable multi-shot sampling. To guarantee the inconspicuousness of the attack, we also propose a new perception loss to achieve better visual quality. Extensive experiments on four of the largest ReID benchmarks (i.e., Market1501, CUHK03, DukeMTMC, and MSMT17) not only show the effectiveness of our method, but also provides directions of the future improvement in the robustness of ReID systems. For example, the accuracy of one of the best-performing ReID systems drops sharply from 91.8% to 1.4% after being attacked by our method. Some attack results are shown in Fig. 1. The code is available at: https://github.com/whj363636/Adversarial-attack-on-Person-ReID-With-Deep-Mis-Ranking.
Hongjun Wang 0005, Guangrun Wang, Dongyu Zhang 0002, Liang Lin 0004
CVPR4
2020 Motion-transformer: self-supervised pre-training for skeleton-based action recognition
abstract
With the development of deep learning, skeleton-based action recognition has achieved great progress in recent years. However, most of the current works focus on extracting more informative spatial representations of the human body, but haven't made full use of the temporal dependencies already contained in the sequence of human action. To this end, we propose a novel transformer-based model called Motion-Transformer to sufficiently capture the temporal dependencies via self-supervised pre-training on the sequence of human action. Besides, we propose to predict the motion flow of human skeletons for better learning the temporal dependencies in sequence. The pre-trained model is then fine-tuned on the task of action recognition. Experimental results on the large scale NTU RGB+D dataset shows our model is effective in modeling temporal relation, and the flow prediction pre-training is beneficial to expose the inherent dependencies in time dimensional. With this pre-training and fine-tuning paradigm, our final model outperforms previous state-of-the-art methods.
Yi-Bin Cheng, Xipeng Chen, Dongyu Zhang 0002, Liang Lin 0004
MMAsia3
2019 Cost-Effective Object Detection: Active Sample Mining With Switchable Selection Criteria
abstract
Though quite challenging, leveraging large-scale unlabeled or partially labeled data in learning systems (e.g., model/classifier training) has attracted increasing attentions due to its fundamental importance. To address this problem, many active learning (AL) methods have been proposed that employ up-to-date detectors to retrieve representative minority samples according to predefined confidence or uncertainty thresholds. However, these AL methods cause the detectors to ignore the remaining majority samples (i.e., those with low uncertainty or high prediction confidence). In this paper, by developing a principled active sample mining (ASM) framework, we demonstrate that cost-effective mining samples from these unlabeled majority data are a key to train more powerful object detectors while minimizing user effort. Specifically, our ASM framework involves a switchable sample selection mechanism for determining whether an unlabeled sample should be manually annotated via AL or automatically pseudolabeled via a novel self-learning process. The proposed process can be compatible with mini-batch-based training (i.e., using a batch of unlabeled or partially labeled data as a one-time input) for object detection. In this process, the detector, such as a deep neural network, is first applied to the unlabeled samples (i.e., object proposals) to estimate their labels and output the corresponding prediction confidences. Then, our ASM framework is used to select a number of samples and assign pseudolabels to them. These labels are specific to each learning batch based on the confidence levels and additional constraints introduced by the AL process and will be discarded afterward. Then, these temporarily labeled samples are employed for network fine-tuning. In addition, a few samples with low-confidence predictions are selected and annotated via AL. Notably, our method is suitable for object categories that are not seen in the unlabeled data during the learning process. Extensive experiments on two public benchmarks (i.e., the PASCAL VOC 2007/2012 data sets) clearly demonstrate that our ASM framework can achieve performance comparable to that of the alternative methods but with significantly fewer annotations.
Keze Wang, Liang Lin 0004, Xiaopeng Yan, Ziliang Chen 0001, Dongyu Zhang 0002, Lei Zhang 0006
IEEE Trans. Neural Networks Learn. Syst.5
2018 Towards Human-Machine Cooperation: Self-Supervised Sample Mining for Object Detection
abstract
Though quite challenging, leveraging large-scale unlabeled or partially labeled images in a cost-effective way has increasingly attracted interests for its great importance to computer vision. To tackle this problem, many Active Learning (AL) methods have been developed. However, these methods mainly define their sample selection criteria within a single image context, leading to the suboptimal robustness and impractical solution for large-scale object detection. In this paper, aiming to remedy the drawbacks of existing AL methods, we present a principled Self-supervised Sample Mining (SSM) process accounting for the real challenges in object detection. Specifically, our SSM process concentrates on automatically discovering and pseudo-labeling reliable region proposals for enhancing the object detector via the introduced cross image validation, i.e., pasting these proposals into different labeled images to comprehensively measure their values under different image contexts. By resorting to the SSM process, we propose a new AL framework for gradually incorporating unlabeled or partially labeled data into the model learning while minimizing the annotating effort of users. Extensive experiments on two public benchmarks clearly demonstrate our proposed framework can achieve the comparable performance to the state-of-the-art methods with significantly fewer annotations.
Keze Wang, Xiaopeng Yan, Dongyu Zhang 0002, Lei Zhang 0006, Liang Lin 0004
CVPR3
2018 Person Re-Identification with Weighted Spatial-Temporal Features
abstract
Person re-identification (re-id) which resolves to recognize a person from the non-overlapped cameras has received increasing research. In this paper, we addressed a new problem of person re-id, i.e., image-to-video (ImtoV) person re-id, in which the probe is an image and the gallery consists of videos from nonoverlapping cameras with different views of probe image as shown in Fig. 1. It is different from the traditional image-based person re-id in which the probe and gallery are all images. Although more information in the video is brought into ImtoV, it remains a challenging problem because of the large variations of light conditions, viewing angles, body pose, and occlusions in different views of videos. One problem is that most of the current models ignore that different frames play different importance in the matching, and assign equal weights to feature vector of each frame of videos. However, frames with serious occlusion and dramatical illumination change have the negative effect in improving the re-id performance. In order to overcome this problem, we proposed a novel framework for this task. We adopted CNNs for the feature extraction of images and videos, and further employed LSTM network for the spatiotemporal feature representation of videos. We added a weight modular to learn the weights for different frames of videos adaptively. We evaluated the proposed framework on three public person re-id datasets, and the experimental results showed that the proposed approach was effective for the ImtoV person re-id.
Dongyu Zhang 0002, Rongcong Chen, Zhilin Qiu, Qing Wang 0018
ICPR1
2018 Image-to-Video Person Re-Identification With Temporally Memorized Similarity Learning
abstract
With the development of video surveillance in public safety field, there is an increasing research on person re-identification (re-id). In this paper, we address the image-to-video person re-id, in which the probe is an image and the gallery is consists of videos captured by nonoverlapping cameras. Compared with image, video sequence contains more temporal information that can be explored to improve the performance of re-identification system. However, it is challenging to model temporal information in the matching process of image-to-video person re-id. In this paper, we proposed a novel temporally memorized similarity learning neural network for this problem. In specific, the proposed network mainly consisted of two parts, including feature representation sub-network and similarity sub-network. In the first part, we adopted a convolutional neural network (CNN) to extract features from the input image. Given a video sequence of a person, features were first extracted from each its frame by using CNN and further forward to a long shot term memory (LSTM) network to encode the temporal information of video sequence. The outputs of LSTM were concatenated together as the feature vector of video sequences. Finally, the feature vectors of probe image and the video sequence were further forward to the similarity sub-network for distance metric learning. In the proposed framework, the feature representation and the similarity metric learning can be learned and optimized simultaneously. We evaluated the proposed framework on three public person re-id data sets, and the experimental results showed that the proposed approach is effective for the image-to-video person re-id.
Dongyu Zhang 0002, Wenxi Wu, Hui Cheng 0002, Ruimao Zhang, Zhenjiang Dong, Zhaoquan Cai 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Look into Person: Self-Supervised Structure-Sensitive Learning and a New Benchmark for Human Parsing
abstract
Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances and the coverage of challenging cases in unconstrained environment. In this paper, we introduce a new benchmark Look into Person (LIP) that makes a significant advance in terms of scalability, diversity and difficulty, a contribution that we feel is crucial for future developments in human-centric analysis. This comprehensive dataset contains over 50,000 elaborately annotated images with 19 semantic part labels, which are captured from a wider range of viewpoints, occlusions and background complexity. Given these rich annotations we perform detailed analysis of the leading human parsing approaches, gaining insights into the success and failures of these methods. Furthermore, in contrast to the existing efforts on improving the feature discriminative capability, we solve human parsing by exploring a novel self-supervised structure-sensitive learning approach, which imposes human pose structures into parsing results without resorting to extra supervision (i.e., no need for specifically labeling human joints in model training). Our self-supervised learning framework can be injected into any advanced neural networks to help incorporate rich high-level knowledge regarding human joints from a global perspective and improve the parsing results. Extensive evaluations on our LIP and the public PASCAL-Person-Part dataset demonstrate the superiority of our method.
Xiaodan Liang, Dongyu Zhang 0002, Xiaohui Shen, Liang Lin 0004
CVPR3
2017 Multiple human tracking based on distributed collaborative cameras
Zhaoquan Cai 0001, Shiyi Hu, Yukai Shi, Qing Wang 0018, Dongyu Zhang 0002
Multim. Tools Appl.5
2017 Cost-Effective Active Learning for Deep Image Classification
abstract
Recent successes in learning-based image classification, however, heavily rely on the large number of annotated training samples, which may require considerable human effort. In this paper, we propose a novel active learning (AL) framework, which is capable of building a competitive classifier with optimal feature representation via a limited amount of labeled training instances in an incremental learning manner. Our approach advances the existing AL methods in two aspects. First, we incorporate deep convolutional neural networks into AL. Through the properly designed framework, the feature representation and the classifier can be simultaneously updated with progressively annotated informative samples. Second, we present a cost-effective sample selection strategy to improve the classification performance with less manual annotations. Unlike traditional methods focusing on only the uncertain samples of low prediction confidence, we especially discover the large amount of high-confidence samples from the unlabeled set for feature learning. Specifically, these high-confidence samples are automatically selected and iteratively assigned pseudolabels. We thus call our framework cost-effective AL (CEAL) standing for the two advantages. Extensive experiments demonstrate that the proposed CEAL framework can achieve promising results on two challenging image classification data sets, i.e., face recognition on the cross-age celebrity face recognition data set database and object categorization on Caltech-256.
Keze Wang, Dongyu Zhang 0002, Ruimao Zhang, Liang Lin 0004
IEEE Trans. Circuits Syst. Video Technol.2
2017 Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning
abstract
Sketch portrait generation benefits a wide range of applications such as digital entertainment and law enforcement. Although plenty of efforts have been dedicated to this task, several issues still remain unsolved for generating vivid and detail-preserving personal sketch portraits. For example, quite a few artifacts may exist in synthesizing hairpins and glasses, and textural details may be lost in the regions of hair or mustache. Moreover, the generalization ability of current systems is somewhat limited since they usually require elaborately collecting a dictionary of examples or carefully tuning features/components. In this paper, we present a novel representation learning framework that generates an end-to-end photo-sketch mapping through structure and texture decomposition. In the training stage, we first decompose the input face photo into different components according to their representational contents (i.e., structural and textural parts) by using a pre-trained convolutional neural network (CNN). Then, we utilize a branched fully CNN for learning structural and textural representations, respectively. In addition, we design a sorted matching mean square error metric to measure texture patterns in the loss function. In the stage of sketch rendering, our approach automatically generates structural and textural representations for the input photo and produces the final result via a probabilistic fusion scheme. Extensive experiments on several challenging benchmarks suggest that our approach outperforms example-based synthesis algorithms in terms of both perceptual and objective metrics. In addition, the proposed method also has better generalization ability across data set without additional training.
Dongyu Zhang 0002, Liang Lin 0004, Tianshui Chen, Xian Wu 0007, Wenwei Tan, Ebroul Izquierdo
IEEE Trans. Image Process.1
2016 Class relatedness oriented-discriminative dictionary learning for multiclass image classification
Dongyu Zhang 0002, Qing Wang 0018, Xiaoyuan Jing
Pattern Recognit.1
2010 Distinguishing Patients with Gastritis and Cholecystitis from the Healthy by Analyzing Wrist Radial Arterial Doppler Blood Flow Signals
abstract
This paper tries to fill the gap between Traditional Chinese Pulse Diagnosis (TCPD) and Doppler diagnosis by applying digital signal analysis and pattern classification techniques to wrist radial arterial Doppler blood flow signals. Doppler blood flows signals (DBFS) of patients with cholecystitis, gastritis and healthy people are classified by L2-soft margin SVM and 5 linear classifiers using the proposed feature - piecewise axially integrated bispectra (PAIB). A 5-fold cross validation is used for performance evaluation. The classification accuracies between either two groups of subjects are greater than 93%. Gastritis can be recognized with higher accuracy than cholecystitis. Cholecystitis can be recognized with higher accuracy on left hand data than right. The findings in this paper partly conform to the theory of TCPD. Though the sample size is relatively small, we could still argue that the methods proposed here are effective and could serve as an assistive tool for TCPD.
Xiaorui Jiang, Dongyu Zhang 0002, Kuanquan Wang, Wangmeng Zuo
ICPR2
2010 Gaussian ERP Kernel Classifier for Pulse Waveforms Classification
abstract
While advances in sensor and signal processing techniques have provided effective tools for quantitative research on traditional Chinese pulse diagnosis (TCPD), the automatic classification of pulse waveforms is remained a difficult problem. To address this issue, this paper proposed a novel edit distance with real penalty (ERP)-based k-nearest neighbors (KNN) classifier by referring to recent progresses in time series matching and KNN classifier. Taking advantage of the metric property of ERP, we first develop a Gaussian ERP kernel, and then embed it into kernel difference-weighted KNN classifier. The proposed Gaussian ERP kernel classifier is evaluated on a dataset which includes 2470 pulse waveforms. Experimental results show that the proposed classifier is much more accurate than several other pulse waveform classification approaches.
Dongyu Zhang 0002, Wangmeng Zuo, David Zhang 0001, Yanlai Li, Naimin Li
ICPR1
2010 Time Series Classification Using Support Vector Machine with Gaussian Elastic Metric Kernel
abstract
Motivated by the great success of dynamic time warping (DTW) in time series matching, Gaussian DTW kernel had been developed for support vector machine (SVM)-based time series classification. Counter-examples, however, had been subsequently reported that Gaussian DTW kernel usually cannot outperform Gaussian RBF kernel in the SVM framework. In this paper, by extending the Gaussian RBF kernel, we propose one novel class of Gaussian elastic metric kernel (GEMK), and present two examples of GEMK: Gaussian time warp edit distance (GTWED) kernel and Gaussian edit distance with real penalty (GERP) kernel. Experimental results on UCR time series data sets show that, in terms of classification accuracy, SVM with GEMK is much superior to SVM with Gaussian RBF kernel and Gaussian DTW kernel, and the state-of-the-art similarity measure methods.
Dongyu Zhang 0002, Wangmeng Zuo, David Zhang 0001
ICPR1