VLDB 2026 Research / reviewers in the wild / expert
Xiaoyun Yang
dblp:54/230
· DBLP profile ↗
35ranked-venue papers
2as first author
24since 2021 · last 2025
0000-0002-7326-4387ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 20 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multi-Degradation Dataset and a Universal Method for Image Deraining in All-Time Driving ScenesabstractImproving the visual quality of rainy driving scenes poses a significant challenge, as the rain streaks in the distance and the raindrops attached to nearby surfaces exhibit different characteristics under varying lighting conditions, both during the daytime and nighttime. We note that existing image deraining approaches are trained independently for specific types of rain degradation, which limits the model’s ability to adapt to dynamic driving scenes. In this paper, we introduce a new task: all-time rainy driving scene reconstruction, which aims to simultaneously address both daytime and nighttime rain degradation using a universal mix-trained model. Firstly, we construct a high-quality benchmark dataset termed RainDrive-10K, which contains four patterns: daytime rain streak, daytime raindrop, nighttime rain streak and nighttime raindrop. Furthermore, we also develop an effective Mamba-based baseline de-raining model, which employs a multi-patch progressive learning strategy to better help image restoration. Unlike existing Mamba-based methods that use fixed-scale scanning for feature extraction, we design a new multi-patch hierarchical scanning block that improves the model’s robustness to diverse rain appearances. Extensive experiments demonstrate the effectiveness of our proposed model, and show that it achieves favorable performance against state-of-the-art ones. The dataset is available at https://github.com/ZXXaaaa/MP-RainMamba. Yonghong Song, Xinyue Su, Xiaoyun Yang |
ECAI | 4 |
| 2025 | Y-Net-ECG: A Multi-Lead informed and interpretable architecture for ECG segmentation across diverse rhythms
Peng Zhang 0106, Xiaoli Feng, Kaibiao Huang, Yinuo Zhao, Zuoming Fu, Zhigang Ye, Tao Wang 0138, Xiaoyun Yang, Fan Lin, Qiang Li 0018 |
Expert Syst. Appl. | 13 |
| 2025 | Deep Reinforcement Learning-Based Cooperative Frequency Controller for Hydropower Dominated SystemsabstractHydropower dominated systems often experience negative damping and severe power flow oscillations due to improperly tuned governor parameters, even with power system stabilizers. Conventional methods optimize these parameters in a single-machine infinite bus model, neglecting interactions between generations during primary frequency control. This article proposes a deep reinforcement learning based cooperative primary frequency controller using a goal representation heuristic dynamic programming model to address this challenge. The controller minimizes power flow oscillations and frequency deviation through cooperative action among hydrogenerations. Case studies on an IEEE 9 bus system demonstrate that the proposed method achieves superior damping of power oscillations and achieves a much shorter regulation time in the primary frequency control compared to the traditional PID controllers. Xiaoyun Yang, Xing Nie, Chi Yao, Zhi-Wei Liu 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractGraph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasible to represent the semantic information. In this paper, we proposed a dynamic semantic-based graph convolution network (DS-GCN) for skeleton-based human action recognition, where the joints and edge types were encoded in the skeleton topology in an implicit way. Specifically, two semantic modules, the joints type-aware adaptive topology and the edge type-aware adaptive topology, were proposed. Combining proposed semantics modules with temporal convolution, a powerful framework named DS-GCN was developed for skeleton-based action recognition. Extensive experiments in two datasets, NTU-RGB+D and Kinetics-400 show that the proposed semantic modules were generalized enough to be utilized in various backbones for boosting recognition accuracy. Meanwhile, the proposed DS-GCN notably outperformed state-of-the-art methods. The code is released here https://github.com/davelailai/DS-GCN Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
AAAI | 5 |
| 2024 | Multi-dimensional Spatio-temporal Prediction Network for Pre-hospital Emergency CareabstractPredictive analysis of pre-hospital care can optimize resource allocation, address challenges in assessing ambulance demand, and significantly improve emergency care efficiency and patient survival rates. However, the lack of high-quality publicly available datasets, combined with the influence of complex factors such as urban spatial patterns, weather, and time, makes predictive analysis of pre-hospital emergency care a challenge. To address this, we introduce the Pre-hospital Emergency Information Mart, a large-scale, publicly accessible repository of authentic emergency records. Given the complexity of such analysis, we propose a Multivariate Spatiotemporal Prediction Network. Specifically, we design a multifactor fine-grained sequence learning branch to model historical sequences, time factors, and external factors. The spatial distribution prediction branch, based on spatiotemporal attention, effectively captures spatial features with hyperconnections. The global knowledge transfer module transfers global spatial knowledge to the sequence prediction module, increasing spatial dependencies for reliable predictions under multidimensional constraints. Experiments on the Pre-hospital Emergency Information Mart dataset show that our proposed Multivariate Spatiotemporal Prediction Network outperforms six well-known methods. Xifeng Hu, Jingchuan Wang, Wenmiao Wang, Xiaoyun Yang |
BIBM | 5 |
| 2024 | Dynamic Semantic-Based Spatial-Temporal Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractHuman action recognition is an essential topic in computer vision and image processing. Graph convolutional networks (GCNs) have attracted significant attention and achieved noteworthy performance in skeleton-based human action recognition tasks. However, most of the previous graph-based works are designed to refine skeleton topology without considering the types of different joints and edges and the occurrence order of the frames. Such a limitation makes them insufficient to represent intrinsic semantic information. Differently, we proposed a dynamic semantic-based spatial-temporal graph convolution network (DS-STGCN) to address the challenge. DS-STGCN has two dynamic semantic modules for spatial and temporal contexts respectively. Specifically, the joints and edge types were encoded in the spatial module implicitly, and the occurrence order of frames was encoded in the temporal module implicitly. Extensive experiments on four datasets including NTU-RGB+D 60(120), Kinetics-400, and FineGYM show that our proposed two semantic modules can bring consistent recognition performance improvement with various backbones. Meanwhile, the proposed DS-STGCN notably surpassed state-of-the-art methods on these datasets. Notably, in the more challenging dataset, such as Kinetics-400, our model significantly outperformed other state-of-the-art GCN-based methods by a large margin. The code has been released at https://github.com/davelailai/DS-STGCN. Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
IEEE Trans. Image Process. | 5 |
| 2024 | A Video Is Worth Three Views: Trigeminal Transformers for Video-Based Person Re-IdentificationabstractVideo-based person Re-Identification (Re-ID) is a hot research topic in intelligent transportation systems, which aims to retrieve video sequences of the same person under non-overlapping surveillance cameras. Compared with static images, video sequences contain more visual information from multiple views, such as spatial and temporal views. However, previous Re-ID methods usually focus on single limited views, lacking diverse observations from different views. To capture richer perceptions and extract more comprehensive representations, we propose a novel learning framework namedTrigeminal Transformers (TMT)to tackle video-based person Re-ID. More specifically, we first design aView-wise Projector (VP)to jointly transform raw videos from spatial, temporal and spatial-temporal views. In addition, inspired by the great success of Vision Transformers (ViT), we introduce the Transformer structure for information enhancement and aggregation. In our work, threeSelf-view Transformers (ST)are proposed to exploit the relationships of local features for information enhancement in spatial, temporal and spatial-temporal. Moreover, aCross-view Transformer (CT)is proposed to aggregate the multi-view features for comprehensive representations. Experimental results indicate that our approach can obtain better performance than some other state-of-the-art approaches on four public Re-ID benchmarks. Xuehu Liu, Chenyang Yu, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Label correlation embedding guided network for multi-label ECG arrhythmia diagnosis
Shaolin Ran, Beizhen Zhao, Yinuo Jiang, Xiaoyun Yang, Cheng Cheng 0010 |
Knowl. Based Syst. | 5 |
| 2023 | Effective Local and Global Search for Fast Long-Term TrackingabstractCompared with short-term tracking, long-term tracking remains a challenging task that usually requires the tracking algorithm to track targets within a local region and re-detect targets over the entire image. However, few works have been done and their performances have also been limited. In this paper, we present a novel robust and real-time long-term tracking framework based on the proposed local search module and re-detection module. The local search module consists of an effective bounding box regressor to generate a series of candidate proposals and a target verifier to infer the optimal candidate with its confidence score. For local search, we design a long short-term updated scheme to improve the target verifier. The verification capability of the tracker can be improved by using several templates updated at different times. Based on the verification scores, our tracker determines whether the tracked object is present or absent and then chooses the tracking strategies of local or global search, respectively, in the next frame. For global re-detection, we develop a novel re-detection module that can estimate the target position and target size for a given base tracker. We conduct a series of experiments to demonstrate that this module can be flexibly integrated into many other tracking algorithms for long-term tracking and that it can improve long-term tracking performance effectively. Numerous experiments and discussions are conducted on several popular tracking datasets, including VOT, OxUvA, TLP, and LaSOT. The experimental results demonstrate that the proposed tracker achieves satisfactory performance with a real-time speed. Code is available at https://github.com/difhnp/ELGLT. Haojie Zhao, Bin Yan 0004, Dong Wang 0004, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Transportation Object Counting With Graph-Based Adaptive Auxiliary LearningabstractThis paper proposes an adaptive auxiliary task learning-based approach for transport object counting problems such as humans and vehicles. These problems are essential in many real-world tasks such as video surveillance, traffic monitoring, public security, and urban planning, to aid intelligent transportation systems. Unlike existing auxiliary task learning-based methods, we develop an attention-enhanced adaptively shared backbone network to enable both task-shared and task-tailored features that are learned in an end-to-end manner. The network seamlessly combines a standard Convolution Neural Network (CNN) and a Graph Convolution Network (GCN) for feature extraction and feature reasoning among different domains of tasks. Our approach gains enriched contextual information by iteratively and hierarchically fusing features across different task branches of the adaptive CNN backbone. The whole framework pays special attention to objects’ spatial locations and varied density levels, informed by object (or crowd) segmentation and density level segmentation auxiliary tasks. In particular, thanks to the proposed dilated contrastive density loss function, our network benefits from individual and regional context supervision, along with strengthened robustness. Experiments on six challenging multi-domain datasets demonstrate that our method achieves superior performance compared with state-of-the-art auxiliary task learning-based counting methods. Our code is publicly available. Yanda Meng, Joshua Bridge, Yitian Zhao, Martha Joddrell, Yihong Qiao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | 3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature AggregationabstractEstimating 3D human body shapes and poses from videos is a challenging computer vision task. The intrinsic temporal information embedded in adjacent frames is helpful in making accurate estimations. Existing approaches learn temporal features of the target frames simply by aggregating features of their adjacent frames, using off-the-shelf deep neural networks. Consequently these approaches cannot explicitly and effectively use the correlations between adjacent frames to help infer the parameters of the target frames. In this paper, we propose a novel framework that can measure the correlations amongst adjacent frames in the form of an estimated confidence metric. The confidence value will indicate to what extent the adjacent frames can help predict the target frames’ 3D shapes and poses. Based on the estimated confidence values, temporally aggregated features are then obtained by adaptively allocating different weights to the temporal predicted features from the adjacent frames. The final 3D shapes and poses are estimated by regressing from the temporally aggregated features. Experimental results on three benchmark datasets show that the proposed method outperforms state-of-the-art approaches (even without the motion priors involved in training). In particular, the proposed method is more robust against corrupted frames. Hongrun Zhang, Yanda Meng, Yitian Zhao, Xuesheng Qian, Yihong Qiao, Xiaoyun Yang, Yalin Zheng |
IEEE Trans. Multim. | 6 |
| 2022 | Using deep learning to incorporate longitudinal health factors into 10-year atherosclerotic cardiovascular disease (ASCVD) risk prediction in the Lifetime Risk Pooled Project (LRPP)
Jingzhi Yu, Xiaoyun Yang, Amy E. Krefman, Hongyan Ning, Lucia C. Petito, Lihui Zhao, Norrina B. Allen |
AMIA | 2 |
| 2022 | DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationabstractMultiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those related to small sample cohorts. In these, there are limited number of WSI slides (bags), while the resolution of a single WSI is huge, which leads to a large number of patches (instances) cropped from this slide. To address this issue, we propose to virtually enlarge the number of bags by introducing the concept of pseudo-bags, on which a double-tier MIL framework is built to effectively use the intrinsic features. Besides, we also contribute to deriving the instance probability under the framework of attentionbased MIL, and utilize the derivation to help construct and analyze the proposed framework. The proposed method outperforms other latest methods on the CAMELYON-16 by substantially large margins, and is also better in performance on the TCGA lung cancer dataset. The proposed framework is ready to be extended for wider MIL applications. The code is available at: https://github. com/hrzhang1123/DTFD-MIL. Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E. Coupland, Yalin Zheng |
CVPR | 5 |
| 2022 | Semi-Supervised Learning for Automatic Atrial Fibrillation Detection in 24-Hour Holter MonitoringabstractParoxysmal atrial fibrillation (AF) is generally diagnosed by long-term dynamic electrocardiogram (ECG) monitoring. Identifying AF episodes from long-term ECG data can place a heavy burden on clinicians. Many machine-learning-based automatic AF detection methods have been proposed to solve this issue. However, these methods require numerous annotated data to train the model, and the annotation of AF in long-term ECG is extremely time-consuming. Reducing the demand for labeled data can effectively improve the clinical practicability of automatic AF detection methods. In this study, we developed a novel semi-supervised learning method that generated modified low-entropy labels of unlabeled samples for training a deep learning model to automatically detect paroxysmal AF in 24 h Holter monitoring data. Our method employed a 1D CNN-LSTM neural network with RR intervals as input and used few labeled training data with numerous unlabeled data for training the neural network. This method was evaluated using a 24 h Holter monitoring dataset collected from 1000 paroxysmal AF patients. Using labeled samples from only 10 patients for model training, our method achieved a sensitivity of 97.8%, specificity of 97.9%, and accuracy of 97.9% in five-fold cross-validation. Compared to the supervised learning method with complete labeled samples, the detection accuracy of our method was only 0.5% lower, while the workload of data annotation was significantly reduced by more than 98%. In general, this is the first study to apply semi-supervised learning techniques for automatic AF detection using ECG. Our method can effectively reduce the demand for AF data annotations and can improve the clinical practicability of automatic AF detection. Peng Zhang 0106, Fan Lin, Xiaoyun Yang, Qiang Li 0018 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Graph-Based Region and Boundary Aggregation for Biomedical Image SegmentationabstractSegmentation is a fundamental task in biomedical image analysis. Unlike the existing region-based dense pixel classification methods or boundary-based polygon regression methods, we build a novel graph neural network (GNN) based deep learning framework with multiple graph reasoning modules to explicitly leverage both region and boundary features in an end-to-end manner. The mechanism extracts discriminative region and boundary features, referred to as initialized region and boundary node embeddings, using a proposed Attention Enhancement Module (AEM). The weighted links between cross-domain nodes (region and boundary feature domains) in each graph are defined in a data-dependent way, which retains both global and local cross-node relationships. The iterative message aggregation and node update mechanism can enhance the interaction between each graph reasoning module's global semantic information and local spatial characteristics. Our model, in particular, is capable of concurrently addressing region and boundary feature reasoning and aggregation at several different feature levels due to the proposed multi-level feature node embeddings in different parallel graph reasoning modules. Experiments on two types of challenging datasets demonstrate that our method outperforms state-of-the-art approaches for segmentation of polyps in colonoscopy images and of the optic disc and optic cup in colour fundus images. The trained models will be made available at: https://github.com/smallmax00/Graph_Region_Boudnary. Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Yihong Qiao, Ian J. C. MacCormick, Xiaowei Huang 0001, Yalin Zheng |
IEEE Trans. Medical Imaging | 4 |
| 2021 | BI-GCN: Boundary-Aware Input-Dependent Graph Convolution Network for Biomedical Image Segmentation
Yanda Meng, Hongrun Zhang, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng |
BMVC | 5 |
| 2021 | Transformer TrackingabstractCorrelation acts as a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion manner to consider the similarity between the template and the search region. However, the correlation operation itself is a local linear matching process, leading to lose semantic information and fall into local optimum easily, which may be the bottleneck of designing high-accuracy tracking algorithms. Is there any better feature fusion method than correlation? To address this issue, inspired by Transformer, this work presents a novel attention-based feature fusion network, which effectively combines the template and search region features solely using attention. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. Finally, we present a Transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression head. Experiments show that our TransT achieves very promising results on six challenging datasets, especially on large-scale LaSOT, TrackingNet, and GOT-10k benchmarks. Our tracker runs at approximatively 50 fps on GPU. Code and models are available at https://github.com/chenxin-dlut/TransT. Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu |
CVPR | 5 |
| 2021 | Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos. Existing methods usually focus on the most conspicuous image regions, thus they may easily miss out fine-grained clues due to the person varieties in image sequences. To address above issues, in this paper, we propose a novel Global-guided Reciprocal Learning (GRL) framework for video-based person Re-ID. Specifically, we first propose a Global-guided Correlation Estimation (GCE) to generate feature correlation maps of local features and global features, which help to localize the high- and low-correlation regions for identifying the same person. After that, the discriminative features are disentangled into high-correlation features and low-correlation features under the guidance of the global representations. Moreover, a novel Temporal Reciprocal Learning (TRL) mechanism is designed to sequentially enhance the high-correlation semantic information and accumulate the low-correlation sub-critical clues. Extensive experiments are conducted on three public benchmarks. The experimental results indicate that our approach can achieve better performance than other state-of-the-art approaches. The code is released at https://github.com/flysnowtiger/GRL. Xuehu Liu, Chenyang Yu, Huchuan Lu, Xiaoyun Yang |
CVPR | 5 |
| 2021 | Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box EstimationabstractVisual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage strategy to improve bounding box estimation. These methods first coarsely locate the target and then refine the initial prediction in the following stages. However, existing approaches still suffer from limited precision, and the coupling of different stages severely restricts the method’s transferability. This work proposes a novel, flexible, and accurate refinement module called Alpha-Refine (AR), which can significantly improve the base trackers’ box estimation quality. By exploring a series of design options, we conclude that the key to successful refinement is extracting and maintaining detailed spatial information as much as possible. Following this principle, Alpha-Refine adopts a pixel-wise correlation, a corner prediction head, and an auxiliary mask head as the core components. Comprehensive experiments on TrackingNet, LaSOT, GOT-10K, and VOT2020 benchmarks with multiple base trackers show that our approach significantly improves the base tracker’s performance with little extra latency. The proposed Alpha-Refine method leads to a series of strengthened trackers, among which the ARSiamRPN (AR strengthened SiamRPNpp) and the ARDiMP50 (AR strengthened DiMP50) achieve good efficiency-precision trade-off, while the ARDiMPsuper (AR strengthened DiMPsuper) achieves very competitive performance at a realtime speed. Code and pretrained models are available at https://github.com/MasterBin-IIAU/AlphaRefine. Bin Yan 0004, Xinyu Zhang 0017, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 5 |
| 2021 | Video Annotation for Visual Tracking via Selection and RefinementabstractDeep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data. Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang |
ICCV | 8 |
| 2021 | Spatial Uncertainty-Aware Semi-Supervised Crowd CountingabstractSemi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via regularized surrogate task (binary segmentation) for crowd counting problems. Different from existing semi-supervised learning-based crowd counting methods, to exploit the unlabeled data, our proposed spatial uncertainty-aware teacher-student framework focuses on high confident regions’ information while addressing the noisy supervision from the unlabeled data in an end-to-end manner. Specifically, we estimate the spatial uncertainty maps from the teacher model’s surrogate task to guide the feature learning of the main task (density regression) and the surrogate task of the student model at the same time. Besides, we introduce a simple yet effective differential transformation layer to enforce the inherent spatial consistency regularization between the main task and the surrogate task in the student model, which helps the surrogate task to yield more reliable predictions and generates high-quality uncertainty maps. Thus, our model can also address the task-level perturbation problems that occur spatial inconsistency between the primary and surrogate tasks in the student model. Experimental results on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised methods. Code is available at : https://github.com/smallmax00/SUA_crowd_counting Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng |
ICCV | 4 |
| 2021 | DeepNetBim: deep learning model for predicting HLA-epitope interactions based on network analysis by harnessing binding and immunogenicity informationabstractBACKGROUND: Epitope prediction is a useful approach in cancer immunology and immunotherapy. Many computational methods, including machine learning and network analysis, have been developed quickly for such purposes. However, regarding clinical applications, the existing tools are insufficient because few of the predicted binding molecules are immunogenic. Hence, to develop more potent and effective vaccines, it is important to understand binding and immunogenic potential. Here, we observed that the interactive association constituted by human leukocyte antigen (HLA)-peptide pairs can be regarded as a network in which each HLA and peptide is taken as a node. We speculated whether this network could detect the essential interactive propensities embedded in HLA-peptide pairs. Thus, we developed a network-based deep learning method called DeepNetBim by harnessing binding and immunogenic information to predict HLA-peptide interactions. RESULTS: Quantitative class I HLA-peptide binding data and qualitative immunogenic data (including data generated from T cell activation assays, major histocompatibility complex (MHC) binding assays and MHC ligand elution assays) were retrieved from the Immune Epitope Database database. The weighted HLA-peptide binding network and immunogenic network were integrated into a network-based deep learning algorithm constituted by a convolutional neural network and an attention mechanism. The results showed that the integration of network centrality metrics increased the power of both binding and immunogenicity predictions, while the new model significantly outperformed those that did not include network features and those with shuffled networks. Applied on benchmark and independent datasets, DeepNetBim achieved an AUC score of 93.74% in HLA-peptide binding prediction, outperforming 11 state-of-the-art relevant models. Furthermore, the performance enhancement of the combined model, which filtered out negative immunogenic predictions, was confirmed on neoantigen identification by an increase in both positive predictive value (PPV) and the proportion of neoantigen recognition. CONCLUSIONS: We developed a network-based deep learning method called DeepNetBim as a pan-specific epitope prediction tool. It extracted the attributes of the network as new features from HLA-peptide binding and immunogenic models. We observed that not only did DeepNetBim binding model outperform other updated methods but the combination of our two models showed better performance. This indicates further applications in clinical practice. Xiaoyun Yang, Liyuan Zhao, Jing Li 0107 |
BMC Bioinform. | 1 |
| 2021 | Learning Adaptive Attribute-Driven Representation for Real-Time RGB-T Tracking
Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
Int. J. Comput. Vis. | 4 |
| 2021 | Jointly Modeling Motion and Appearance Cues for Robust RGB-T TrackingabstractIn this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB and thermal (T) modalities. The fusion weights are determined by using offline-trained global and local multimodal fusion networks, and then adopted to linearly combine the response maps of RGB and T modalities. Second, when the appearance cue is unreliable, we comprehensively take motion cues, i.e., target and camera motions, into account to make the tracker robust. We further propose a tracker switcher to switch the appearance and motion trackers flexibly. Numerous results on three recent RGB-T tracking datasets show that the proposed tracker performs significantly better than other state-of-the-art algorithms. Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
IEEE Trans. Image Process. | 6 |
| 2020 | High-Performance Long-Term Tracking With Meta-UpdaterabstractLong-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online update. However, it is quite risky to straightforwardly introduce online-update-based trackers to solve the long-term problem, due to long-term uncertain and noisy observations. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed meta-updater can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Our meta-updater learns a binary output to guide the tracker’s update and can be easily embedded into different trackers. This work also introduces a long-term tracking framework consisting of an online local tracker, an online verifier, a SiamRPN-based re-detector, and our meta-updater. Numerous experimental results on the VOT2018LT,VOT2019LT, OxUvALT, TLP, and LaSOT benchmarks show that our tracker performs remarkably better than other competing algorithms. Our project is available on the website: https://github.com/Daikenan/LTMU. Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 6 |
| 2020 | Cooling-Shrinking Attack: Blinding the Tracker With Imperceptible NoisesabstractAdversarial attack of CNN aims at deceiving models to misbehave by adding imperceptible perturbations to images. This feature facilitates to understand neural networks deeply and to improve the robustness of deep learning models. Although several works have focused on attacking image classifiers and object detectors, an effective and efficient method for attacking single object trackers of any target in a model-free way remains lacking. In this paper, a cooling-shrinking attack method is proposed to deceive state-of-the-art SiameseRPN-based trackers. An effective and efficient perturbation generator is trained with a carefully designed adversarial loss, which can simultaneously cool hot regions where the target exists on the heatmaps and force the predicted bounding box to shrink, making the tracked target invisible to trackers. Numerous experiments on OTB100, VOT2018, and LaSOT datasets show that our method can effectively fool the state-of-the-art SiameseRPN++ tracker by adding small perturbations to the template or the search regions. Besides, our method has good transferability and is able to deceive other top-performance trackers such as DaSiamRPN, DaSiamRPN-UpdateNet, and DiMP. The source codes are available at https://github.com/MasterBin-IIAU/CSA. Bin Yan 0004, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 4 |
| 2020 | Regression of Instance Boundary by Aggregated CNN and GCN
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
ECCV (8) | 5 |
| 2020 | CNN-GCN Aggregation Enabled Boundary Regression for Biomedical Image Segmentation
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
MICCAI (4) | 5 |
| 2020 | Online Filtering Training Samples for Robust Visual TrackingabstractIn recent years, discriminative trackers show its great tracking performance, that is mainly due to the online updating using samples collected during tracking. The model could adapt appearance changes of objects and the background well after updating. But these trackers have a serious disadvantage that wrong samples may cause severe model degradation. Most of the training samples in the tracking phase are obtained according to the tracking result of the current frame. Wrong training samples will be collected when the tracking result is inaccurate, seriously affecting the discrimination ability of the model. Besides, partial occlusion also leads to the same problem. In this paper, we propose an optimization module named MetricNet for online filtering training samples. It applies a matching network containing the classification and distance branches, and uses multiple metric methods for different type samples. MetricNet optimizes the training sample set by recognizing wrong and redundant samples, thereby improving the tracking performance. The proposed MetricNet can be regarded as an independent optimization module and integrated into all discriminative trackers updated online. Extensive experiments on three tracking datasets show its effectiveness and generalization ability. After applying MetricNet to MDNet, the tracking result is increased by 5.3% in terms of the success plot on the LaSOT dataset. Our project is available at https://github.com/zj5559/MetricNet. Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
ACM Multimedia | 5 |
| 2019 | GradNet: Gradient-Guided Network for Visual Object TrackingabstractThe fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cannot capture the temporal variations of targets or background clutter. In this work, we propose a novel gradient-guided network to exploit the discriminative information in gradients and update the template in the siamese network through feed-forward and backward operations. To be specific, the algorithm can utilize the information from the gradient to update the template in the current frame. In addition, a template generalization training method is proposed to better use gradient information and avoid overfitting. To our knowledge, this work is the first attempt to exploit the information in the gradient for template update in siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than other state-of-the-art trackers. Peixia Li, Wanli Ouyang, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu |
ICCV | 5 |
| 2019 | 'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term TrackingabstractCompared with traditional short-term tracking, long-term tracking poses more challenges and is much closer to realistic applications. However, few works have been done and their performance have also been limited. In this work, we present a novel robust and real-time long-term tracking framework based on the proposed skimming and perusal modules. The perusal module consists of an effective bounding box regressor to generate a series of candidate proposals and a robust target verifier to infer the optimal candidate with its confidence score. Based on this score, our tracker determines whether the tracked object being present or absent, and then chooses the tracking strategies of local search or global search respectively in the next frame. To speed up the image-wide global search, a novel skimming module is designed to efficiently choose the most possible regions from a large number of sliding windows. Numerous experimental results on the VOT-2018 long-term and OxUvA long-term benchmarks demonstrate that the proposed method achieves the best performance and runs in real-time. The source codes are available at https://github.com/iiau-tracker/SPLT. Bin Yan 0004, Haojie Zhao, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
ICCV | 5 |
| 2019 | Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. It helps intelligent devices to understand and interact with the surrounding scenes. Due to the high-memory requirement, current methods only produce low-resolution completion predictions, and generally lose the object details. Furthermore, they also ignore the multi-scale spatial contexts, which play a vital role for the 3D inference. To address these issues, in this work we propose a novel deep learning framework, named Cascaded Context Pyramid Network (CCPNet), to jointly infer the occupancy and semantic labels of a volumetric 3D scene from a single depth image. The proposed CCPNet improves the labeling coherence with a cascaded context pyramid. Meanwhile, based on the low-level features, it progressively restores the fine-structures of objects with Guided Residual Refinement (GRR) modules. Our proposed framework has three outstanding advantages: (1) it explicitly models the 3D spatial context for performance improvement; (2) full-resolution 3D volumes are produced with structure-preserving details; (3) light-weight models with low-memory requirements are captured with a good extensibility. Extensive experiments demonstrate that in spite of taking a single-view depth map, our proposed framework can generate high-quality SSC results, and outperforms state-of-the-art approaches on both the synthetic SUNCG and real NYU datasets. Wei Liu 0044, Yinjie Lei, Huchuan Lu, Xiaoyun Yang |
ICCV | 5 |
| 2016 | Provably Secure Threshold Paillier Encryption Based on Hyperplane Geometry
Zhe Xia, Xiaoyun Yang, Debiao He |
ACISP (2) | 2 |
| 2013 | Erosion band signatures for spatial extraction of features
Eduard Vazquez, Xiaoyun Yang, Gregory Slabaugh |
Mach. Vis. Appl. | 2 |
| 2006 | Optic Nerve Head Segmentation in HRT ImagesabstractAccurate segmentation of the optic nerve head (or optic disk) is very useful during the analysis and assessment of glaucoma in the eye. The Heidelberg retinal tomograph (HRT) can acquire high quality images of the optic disk and also allows three dimensional topographic measurements to be made. However, there are significant problems in optic disk segmentation due to having to deal with issues such as distractors along blood vessel edges; the extent of pallor in the optic disk or the very variable appearance of the optic nerve head itself. We propose a multi-scale region and boundary hybrid snake method to extract the optic disk. This model takes account of the vessel edge gradient direction and tries to avoid its force influence when evolving the snake. Experimental results are assessed using an overlap ratio. Xiaoyun Yang, Philip J. Morrow, Bryan W. Scotney |
ICIP | 1 |