Dongyan Guo

dblp:124/9294 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 ProtoIENet: An interpretable approach for medical image classification
Jianhao Yu, Yating Zhu, Dongyan Guo
Pattern Recognit.8
2025 DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
abstract
Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data. Recently, diffusion models trained on expansive datasets have been confirmed to maintain the capability to generate diverse, high-quality images. This success suggests a strong potential of the models to effectively understand varied visual information. In this work, we leverage the comprehensive visual knowledge embedded in pre-trained diffusion models to enable more robust and accurate monocular camera intrinsic estimation. Specifically, we reformulate the problem of estimating the four degrees of freedom (4-DoF) of camera intrinsic parameters as a dense incident map generation task. The map details the angle of incidence for each pixel in the RGB image, and its format aligns well with the paradigm of diffusion models. The camera intrinsic then can be derived from the incident map with a simple non-learning RANSAC algorithm during inference. Moreover, to further enhance the performance, we jointly estimate a depth map to provide extra geometric information for the incident map estimation. Extensive experiments on multiple testing datasets demonstrates that our model achieves state-of-the-art performance, gaining up to a 40% reduction in prediction errors. Besides, the experiments also show that the precise camera intrinsic and depth maps estimated by our pipeline can greatly benefit practical applications such as 3D reconstruction from a single in-the-wild image.
Xiankang He, Guangkai Xu, Dongyan Guo
AAAI6
2025 Graph Anomaly Detection via Diffusion Enhanced Multi-View Contrastive Learning
Xiangjie Kong 0001, Jin Liu 0034, Dongyan Guo, Guojiang Shen
Knowl. Based Syst.6
2025 Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation
Xiaoyan Wang 0007, Yating Zhu, Dongyan Guo, Pan Mu, Ming Xia 0005, Cong Bai, Zhongzhao Teng, Shengyong Chen
Medical Image Anal.5
2025 Graph Attention Network for Context-Aware Visual Tracking
abstract
Siamese-network-based trackers convert the general object tracking as a similarity matching task between a template and a search region. Using convolutional feature cross correlation (Xcorr) for similarity matching, a large number of Siamese trackers are proposed and achieved great success. However, due to the predefined size of the target feature, these trackers suffer from either retaining much background information or losing important foreground information. Moreover, the global matching between the target and search region also largely neglects the part-level structural information and the contextual information of the target. To tackle the aforementioned obstacles, in this article, we propose a simple context-aware Siamese graph attention network, which establishes part-to-part correspondence between the Siamese branches with a complete bipartite graph. The object information from the template is propagated to the search region via a graph attention mechanism. With such a design, a target-aware template input is enabled to replace the prefixed template region, which can adaptively fit the size and aspect ratio variations in different objects. Based on it, we further construct a context-aware feature matching mechanism to embed both the target and the contextual information in the search region. Experiments on challenging benchmarks including GOT-10k, TrackingNet, LaSOT, VOT2020, and OTB-100 demonstrate that the proposed SiamGAT* outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT.
Yanyan Shao, Dongyan Guo, Zhenhua Wang 0003, Liyan Zhang 0001, Jianhua Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2024 PointAttN: You Only Need Attention for Point Cloud Completion
abstract
Point cloud completion referring to completing 3D shapes from partial 3D point clouds is a fundamental problem for 3D point cloud analysis tasks. Benefiting from the development of deep neural networks, researches on point cloud completion have made great progress in recent years. However, the explicit local region partition like kNNs involved in existing methods makes them sensitive to the density distribution of point clouds. Moreover, it serves limited receptive fields that prevent capturing features from long-range context information. To solve the problems, we leverage the cross-attention and self-attention mechanisms to design novel neural network for point cloud completion with implicit local region partition. Two basic units Geometric Details Perception (GDP) and Self-Feature Augment (SFA) are proposed to establish the structural relationships directly among points in a simple yet effective way via attention mechanism. Then based on GDP and SFA, we construct a new framework with popular encoder-decoder architecture for point cloud completion. The proposed framework, namely PointAttN, is simple, neat and effective, which can precisely capture the structural information of 3D shapes and predict complete point clouds with detailed geometry. Experimental results demonstrate that our PointAttN outperforms state-of-the-art methods on multiple challenging benchmarks. Code is available at: https://github.com/ohhhyeahhh/PointAttN
Dongyan Guo, Junxia Li, Qingshan Liu 0001, Chunhua Shen
AAAI3
2024 TDCL: Dense Semantic Contrastive Learning for Vision-Language Tracking
abstract
Traditional single-object tracking tasks are undergoing a new wave of transformation, especially with the emergence of the lack of semantics, which has led to the rise of the vision-language tracking task. However, previous approaches that combine the visual tracker with natural language descriptions tend to rely on a global representation of the text description, considering less about the fine-grained connections between the text description and the visual appearance. This paper proposes to utilize a bi-directional cross-attention module to capture the connections between language and visual features, which are further projected as dense semantic representations for alignment. In order to keep the semantic consistency between the search region and the coupled natural language and align the fused feature, this paper proposes a novel dense semantic contrastive learning loss to bridge the semantic gap between text and visual modalities and align them in a dense form. The proposed framework achieves promising results in tracking datasets that contain natural language descriptions, such as TNL2K, and OTB99-LANG. Our approach provides a novel solution for representing and aligning cross-modal information for the single object tracking task and may inspire further research in this field.
Xiankang He, Kaiyang Lan, Dongyan Guo
ECAI5
2024 Semantic Bridging and Feature Anchoring for Class Incremental Learning
abstract
The purpose of class-incremental learning is to continually assimilate new classes and preserving the knowledge of learned classes. An important issue is that the learned knowledge would be catastrophically forgotten while the model updated to adapting to new classes. In this paper, we introduce two strategies, Knowledge Bridging and Category Anchoring, to balance the old and new classes. Knowledge Bridging aims to build the semantic correlation of old and new classes, which uses feature-level distillation to apply learned knowledge to new information. Category Anchoring focuses on learning class-specific feature centers that are crucial for distinctively categorizing all classes. Finally, we incorporate the proposed strategies with six prominent class-incremental learning approaches and conduct comprehensive experiments on the CIFAR100 and ImageNetSubset datasets. The results demonstrate that the proposed strategies are helpful to enhancing performance in Class-Incremental Learning (CIL) tasks.
Kanghui Wu, Dongyan Guo
ICME2
2024 Object discriminability re-extraction for distractor-aware visual object tracking
Dongyan Guo, Xiangjie Kong 0001, Zhenhua Wang 0003, Jianhua Zhang 0002
Comput. Vis. Image Underst.3
2024 Adaptive Activation Network for Weakly Supervised Semantic Segmentation
abstract
Class activation maps generated by image classifiers are widely used as priors for image-level weakly supervised semantic segmentation. However, these activation maps mainly focus on the sparse discriminative regions, which has been a bottleneck for the segmentation task. Based on our observations, the activation maps actually capture almost the entire target regions, and some regions with lower activation values are easily to be neglected. Thus, to solve the issue, we propose an adaptive activation network with two branches to recalibrate the low-confidence regions in the activation maps. Specifically, an activation enhancement branch is designed to redistribute the activation values by leveraging attention mechanism. Since multi-scale images can provide complementary information, a scale adaptation branch is paralleled to supervise the activation enhancement branch. The mutual supervision and fusion of the two branches can promote the less-discriminative parts, and deactivate the background regions. Based on them, a simple yet effective denoising module is proposed to further improve the quality of pseudo masks, which makes use of the large scale predictions of the trained segmentation network. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 benchmarks show that our method achieves state-of-the-art performance, demonstrating the effectiveness of our algorithm. Code will be made publicly available.
Junxia Li, Deshuo Shi, Dongyan Guo, Qingshan Liu 0001
IEEE Trans. Multim.4
2024 Self-Supervised Enhancement for Named Entity Disambiguation via Multimodal Graph Convolution
abstract
Named entity disambiguation (NED) finds the specific meaning of an entity mention in a particular context and links it to a target entity. With the emergence of multimedia, the modalities of content on the Internet have become more diverse, which poses difficulties for traditional NED, and the vast amounts of information make it impossible to manually label every kind of ambiguous data to train a practical NED model. In response to this situation, we present MMGraph, which uses multimodal graph convolution to aggregate visual and contextual language information for accurate entity disambiguation for short texts, and a self-supervised simple triplet network (SimTri) that can learn useful representations in multimodal unlabeled data to enhance the effectiveness of NED models. We evaluated these approaches on a new dataset, MMFi, which contains multimodal supervised data and large amounts of unlabeled data. Our experiments confirm the state-of-the-art performance of MMGraph on two widely used benchmarks and MMFi. SimTri further improves the performance of NED methods. The dataset and code are available at https://github.com/LanceZPF/NNED_MMGraph.
Kaining Ying, Zhenhua Wang 0003, Dongyan Guo, Cong Bai
IEEE Trans. Neural Networks Learn. Syst.4
2023 Multi-Stage Aggregation Transformer for Medical Image Segmentation
abstract
Capturing rich multi-scale features is essential for resolving complex variations in medical image segmentation. In this paper, we explore how to fully utilize the advantages of Convolutional neural networks (CNN) and Transformer, and propose a novel multi-stage aggregation architecture named MA-Transformer for accurate segmentation of medical images with large variations and blurs. Specifically, an encoder module is introduced in each stage, which is a dual-branch structure parallelly combining Transformers and convolutions. By such design, the self-attention can provide a global context for CNN to extract multi-resolution complementary features stage by stage, thus the feature representations are gradually enhanced with local details and contextual information. Multi-scale semantic features are then combined with skip connections in the decoder to produce the final result. Extensive experiments on public medical imaging datasets demonstrate our superior segmentation performance, compared to the state-of-the-art CNN-based, Transformer-based approaches and CNN-Transformer combined approaches. Code will be made publicly available.
Xiaoyan Wang 0007, Minghan Shao, Dongyan Guo, Ming Xia 0005, Cong Bai
ICASSP3
2023 Latent-space Unfolding for MRI Reconstruction
abstract
To circumvent the problems caused by prolonged acquisition periods, compressed sensing MRI enjoys a high usage profile to accelerate the recovery of high-quality images from under-sampled k-space data. Most current solutions dedicate to solving this issue with the pursuit of certain prior properties, yet the treatments are all enforced in the original space, resulting in limited feature information. To achieve a performance promotion yet with the guarantee of running efficiency, in this work, we propose a latent-space unfolding network (LsUNet). Specifically, by an elaborately designed reversible network, the inputs are first mapped to a channel-lifted latent space, which taps the potential of capturing spatial-invariant features sufficiently. Within the latent space, we then unfold an accelerated optimization algorithm to iterate an efficient and feasible solution, in which a parallelly dual-domain update is equipped for better feature fusion. Finally, an inverse embedding transformation of the recovered high-dimensional representation is applied to achieve the expected estimation. LsUNet enjoys high interpretability due to the physically induced modules, which not only facilitates an intuitive understanding of the internal operating mechanism but also endows it with high generalization ability. Comprehensive experiments on different datasets and various sampling rates/patterns demonstrate the advantages of our proposal over the latest methods both visually and numerically.
Jiawei Jiang 0002, Yuchao Feng, Dongyan Guo, Jianwei Zheng 0001
ACM Multimedia4
2023 Human Interaction Understanding With Consistency-Aware Learning
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions.
Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 DSP-Based Traffic Target Detection for Intelligent Transportation
abstract
Internet of Things (IoT)-based intelligent transportation is attracting more and more attention. As a key component of intelligent transportation, traffic video monitoring is very important, in which vehicle and pedestrian detection on the road is a crucial task. Although vehicle and pedestrian detection through deep learning (DL) may achieve high accuracy, it tends to require high computing resources, which hinders its use on IoT devices. As an important class of IoT devices, digital signal processor (DSP) has the characteristics of low energy consumption, small size, and strong performance, which has been widely used in intelligent transportation. In order to use DL on DSP for accurate vehicle and pedestrian detection, we first propose a series of general tactics to optimize the object detection convolutional neural network (CNN) model, including convolution layer optimization, cache optimization, compiler optimization, intrinsics optimization and direct memory access (DMA) acceleration, and then a parallel scheme to extend the model to run on multicore, and further quantize the implementation of the model. We evaluate it on UA-DETRAC and KITTI datasets. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with only 0.06% accuracy loss.
Jianhua Zhang 0002, Rucen Wang, Ruyu Liu, Dongyan Guo, Bo Li 0090, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.4
2023 Looking at Boundary: Siamese Densely Cooperative Fusion for Salient Object Detection
abstract
Though deep learning-based saliency detection methods have achieved gratifying performance recently, the predicted saliency maps still suffer from the boundary challenge. From the perspective of foreground-background separation, this article attempts to extract the edge information of objects by exploiting the difference between different color channels in the RGB color space and establishes a novel multicolor contrast extraction (MCE) mechanism to improve the learning ability of exquisite boundary information of the network. To make full use of the MCE outputs and RGB colors, and well depict and capture the complementary information between them, we devise a novel Siamese densely cooperative fusion (DCF) network (SDFNet) for saliency detection, which consists of two effective components: boundary-directed feature learning (BDFL) and DCF. The BDFL provides joint learning for both MCE and RGB modalities through a Siamese network, while the DCF module is devised for complementary feature discovery, in order to effectively combine the features learned from two modalities. Experiments on five well-known benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches in terms of different evaluation metrics. We provide a detailed analysis of these results and indicate that our joint modeling of MCE and RGB colors helps to better capture the object details, especially in the object boundaries.
Junxia Li, Qingshan Liu 0001, Dongyan Guo
IEEE Trans. Neural Networks Learn. Syst.5
2022 Joint Classification and Regression for Visual Tracking with Fully Convolutional Siamese Networks
abstract
Abstract Visual tracking of generic objects is one of the fundamental but challenging problems in computer vision. Here, we propose a novel fully convolutional Siamese network to solve visual tracking by directly predicting the target bounding box in an end-to-end manner. We first reformulate the visual tracking task as two subproblems: a classification problem for pixel category prediction and a regression task for object status estimation at this pixel. With this decomposition, we design a simple yet effective Siamese architecture based classification and regression framework, termed SiamCAR, which consists of two subnetworks: a Siamese subnetwork for feature extraction and a classification-regression subnetwork for direct bounding box prediction. Since the proposed framework is both proposal- and anchor-free, SiamCAR can avoid the tedious hyper-parameter tuning of anchors, considerably simplifying the training. To demonstrate that a much simpler tracking framework can achieve superior tracking results, we conduct extensive experiments and comparisons with state-of-the-art trackers on a few challenging benchmarks. Without bells and whistles, SiamCAR achieves leading performance with a real-time speed. Furthermore, the ablation study validates that the proposed framework is effective with various backbone networks, and can benefit from deeper networks. Code is available at https://github.com/ohhhyeahhh/SiamCAR .
Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Chunhua Shen, Liyan Zhang 0001, Shengyong Chen
Int. J. Comput. Vis.2
2022 SSA-Net: Spatial self-attention network for COVID-19 pneumonia infection segmentation with semi-supervised few-shot learning
Xiaoyan Wang 0007, Yiwen Yuan, Dongyan Guo, Ming Xia 0005, Zhenhua Wang 0003, Cong Bai, Shengyong Chen
Medical Image Anal.3
2021 Graph Attention Tracking
abstract
Siamese network based trackers formulate the visual tracking task as a similarity matching problem. Almost all popular Siamese trackers realize the similarity learning via convolutional feature cross-correlation between a target branch and a search branch. However, since the size of target feature region needs to be pre-fixed, these cross-correlation base methods suffer from either reserving much adverse background information or missing a great deal of foreground information. Moreover, the global matching be-tween the target and search region also largely neglects the target structure and part-level information.In this paper, to solve the above issues, we propose a simple target-aware Siamese graph attention network for general object tracking. We propose to establish part-to-part correspondence between the target and the search region with a complete bipartite graph, and apply the graph attention mechanism to propagate target information from the template feature to the search feature. Further, instead of using the pre-fixed region cropping for template-feature-area selection, we investigate a target-aware area selection mechanism to fit the size and aspect ratio variations of different objects. Experiments on challenging benchmarks including GOT-10k, UAV123, OTB-100 and LaSOT demonstrate that the proposed SiamGAT outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT
Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Liyan Zhang 0001, Chunhua Shen
CVPR1
2021 Consistency-Aware Graph Network for Human Interaction Understanding
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models, which is inadequate to model complicated human interactions. In this paper, we propose a consistency-aware graph network, which combines the representative ability of graph network and the consistency-aware reasoning to facilitate HIU. Our network consists of three components, a backbone CNN to extract image features, a factor graph network to learn third-order interactive relations among participants, and a consistency-aware reasoning module to enforce labeling and grouping consistencies. Our key observation is that the consistency-aware-reasoning bias for HIU can be embedded into an energy, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained jointly in an end-to-end manner. Experimental results show that our approach achieves leading performance on three benchmarks. Code is available at https://git.io/CAGNet.
Zhenhua Wang 0003, Jiajun Meng, Dongyan Guo, Jianhua Zhang 0002, Qinfeng Shi, Shengyong Chen
ICCV3
2021 Collaborative Visual Inertial SLAM for Multiple Smart Phones
abstract
The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can complete tasks that a single agent cannot do. However, it depends on robust communication, efficient location detection, robust mapping, and efficient information sharing among agents. We propose a multi-intelligence collaborative monocular visual-inertial SLAM deployed on multiple ios mobile devices with a centralized architecture. Each agent can independently explore the environment, run a visual-inertial odometry module online, and then send all the measurement information to a central server with higher computing resources. The server manages all the information received, detects overlapping areas, merges and optimizes the map, and shares information with the agents when needed. We have verified the performance of the system in public datasets and real environments. The accuracy of mapping and fusion of the proposed system is comparable to VINS-Mono which requires higher computing resources.
Ruyu Liu, Kaiqi Chen 0001, Jianhua Zhang 0002, Dongyan Guo
ICRA5
2021 End-to-end feature fusion Siamese network for adaptive visual tracking
abstract
Abstract According to observations, different visual objects have different salient features in different scenarios. Even for the same object, its salient shape and appearance features may change greatly from time to time in a long‐term tracking task. Motivated by them, an end‐to‐end feature fusion framework was proposed based on the Siamese network, named FF‐Siam, which can effectively fuse different features for adaptive visual tracking. The framework consists of four layers. A feature extraction layer is designed to extract the different features of the target region and search region. The extracted features are then put into a weight generation layer to obtain the channel weights, which indicate the importance of different feature channels. Both features and the channel weights are utilised in a template generation layer to generate a discriminative template. Finally, the corresponding response maps created by the convolution of the search region features and the template are applied with a fusion layer to obtain the final response map for locating the target. Experimental results demonstrate that the proposed framework achieves state‐of‐the‐art performance on the popular Temple‐Colour, OTB50 and UAV123 benchmarks.
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen
IET Image Process.1
2021 Human Interaction Understanding With Joint Graph Decomposition and Node Labeling
abstract
The task of human interaction understanding involves both recognizing the action of each individual in the scene and decoding the interaction relationship among people, which is useful to a series of vision applications such as camera surveillance, video-based sports analysis and event retrieval. This paper divides the task into two problems including grouping people into clusters and assigning labels to each of them, and presents an approach to solving these problems in a joint manner. Our method does not assume the number of groups is known beforehand as this will substantially restrict its application. With the observation that the two challenges are highly correlated, the key idea is to model the pairwise interacting relations among people via a complete graph and its associated energy function such that the labeling and grouping problems are translated into the minimization of the energy function. We implement this joint framework by fusing both deep features and rich contextual cues, and learn the fusion parameters from data. An alternating search algorithm is developed in order to efficiently solve the associated inference problem. By combining the grouping and labeling results obtained with our method, we are able to achieve the semantic-level understanding of human interactions. Extensive experiments are performed to qualitatively and quantitatively evaluate the effectiveness of our approach, which outperforms state-of-the-art methods on several important benchmarks. An ablation study is also performed to verify the effectiveness of different modules within our approach.
Zhenhua Wang 0003, Jinchao Ge, Dongyan Guo, Jianhua Zhang 0002, Yanjing Lei, Shengyong Chen
IEEE Trans. Image Process.3
2020 SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
abstract
By decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR consists of two simple subnetworks: one Siamese subnetwork for feature extraction and one classification-regression subnetwork for bounding box prediction. Different from state-of-the-art trackers like Siamese-RPN, SiamRPN++ and SPM, which are based on region proposal, the proposed framework is both proposal and anchor free. Consequently, we are able to avoid the tricky hyper-parameter tuning of anchors and reduce human intervention. The proposed framework is simple, neat and effective. Extensive experiments and comparisons with state-of-the-art trackers are conducted on challenging benchmarks including GOT-10K, LaSOT, UAV123 and OTB-50. Without bells and whistles, our SiamCAR achieves the leading performance with a considerable real-time speed. The code is available at https://github.com/ohhhyeahhh/SiamCAR.
Dongyan Guo, Zhenhua Wang 0003, Shengyong Chen
CVPR1
2020 Improving auto-encoder novelty detection using channel attention and entropy minimization
abstract
Novelty detection is a important research area which mainly solves the classification problem of inliers which usually consists of normal samples and outliers composed of abnormal samples. Auto-encoder is often used for novelty detection. However, the generalization ability of the auto-encoder may cause the undesirable reconstruction of abnormal elements and reduce the identification ability of the model. To solve the problem, we focus on the perspective of better reconstructing the normal samples as well as retaining the unique information of normal samples to improve the performance of auto-encoder for novelty detection. Firstly, we introduce attention mechanism into the task. Under the action of attention mechanism, auto-encoder can pay more attention to the representation of inlier samples through adversarial training. Secondly, we apply the information entropy into the latent layer to make it sparse and constrain the expression of diversity. Experimental results on three public datasets show that the proposed method achieves comparable performance compared with previous popular approaches.
Dongyan Guo, Shengyong Chen
MMAsia2
2018 Deep CRF-Graph Learning for Semantic Image Segmentation
Fuguang Ding, Zhenhua Wang 0003, Dongyan Guo, Shengyong Chen, Jianhua Zhang 0002, Zhanpeng Shao
PRICAI3
2018 Siamese Network Based Features Fusion for Adaptive Visual Tracking
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen, Jian Zhang 0002
PRICAI (1)1
2018 Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao
Neurocomputing8
2017 User relationship strength modeling for friend recommendation on Instagram
Dongyan Guo, Jingsong Xu, Jian Zhang 0002, Min Xu 0001, Xiangjian He
Neurocomputing1
2015 Saliency-based content-aware lifestyle image mosaics
Dongyan Guo, Jinhui Tang 0001, Jundi Ding, Chunxia Zhao
J. Vis. Commun. Image Represent.1
2015 NIF-based seam carving for image resizing
Dongyan Guo, Jundi Ding, Jinhui Tang 0001, Min Xu 0001, Chunxia Zhao
Multim. Syst.1
2015 Robust facial landmark localization using classified random ferns and pose-based initialization
Jian Zhang 0002, Dongyan Guo, Zhong Jin
Signal Process.3
2014 Multiple Kernel Learning Based Multi-view Spectral Clustering
abstract
For a given data set, exploring their multi-view instances under a clustering framework is a practical way to boost the clustering performance. This is because that each view might reflect partial information for the existing data. Furthermore, due to the noise and other impact factors, exploring these instances from different views will enhance the mining of the real structure and feature information within the data set. In this paper, we propose a multiple kernel spectral clustering algorithm through the multi-view instances on the given data set. By combining the kernel matrix learning and the spectral clustering optimization into one process framework, the algorithm can determine the kernel weights and cluster the multi-view data simultaneously. We compare the proposed algorithm with some recent published methods on real-world datasets to show the efficiency of the proposed algorithm.
Dongyan Guo, Jian Zhang 0002, Xinwang Liu 0002, Chunxia Zhao
ICPR1
2013 Saliency-Based Content-Aware Image Mosaics
Dongyan Guo, Jinhui Tang 0001, Jundi Ding, Chunxia Zhao
MMM (1)1