EDBT 2026 Demo / reviewers in the wild / expert
Ke Zhang 0014
dblp:20/4152-14
· DBLP profile ↗
37ranked-venue papers
9as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-LLIE: Semi-supervised contrastive learning with Mamba-based low-light enhancement
Ke Zhang 0014, Bin Zhao 0001, Xuelong Li 0001 |
Neural Networks | 2 |
| 2024 | An adaptive self-correction joint training framework for person re-identification with noisy labels
Ke Zhang 0014, Jingyu Wang 0002 |
Expert Syst. Appl. | 2 |
| 2024 | WTVI: A Wavelet-Based Transformer Network for Video InpaintingabstractVideo inpainting aims to complete missing frames visually convincingly by balancing high-frequency detailed textures and low-frequency semantic structures. Conventional approaches utilize generative adversarial and reconstruction losses for optimizing output frames, each favoring different frequency aspects, to achieve this equilibrium. However, employing both loss types concurrently often results in a conflict between perceptual and distortion qualities, mainly due to their distinct frequency preferences. In response, this letter introduces the Waveletbased Transformer network for Video Inpainting (WTVI). WTVI employs a 2D discrete wavelet transform (DWT) to decompose frames into various frequency bands, ensuring the preservation of spatial information. It then independently completes missing regions in each band using Transformer network. To mitigate inter-frequency conflicts, we apply reconstruction loss to the low-frequency bands and adversarial loss to the high-frequency bzands. Additionally, we innovate High-frequency Cross-Attention (HCA) and Low-frequency Cross-Attention (LCA) modules to enhance frequency dependency learning beyond the spatialtemporal scope and to align features across bands. Our experiments confirm that WTVI surpasses previous methods, significantly improving both quantitative and qualitative performance. Ke Zhang 0014, Guanxiao Li, Jingyu Wang 0002 |
IEEE Signal Process. Lett. | 1 |
| 2024 | A Multi-Information Fusion Algorithm to Fault Diagnosis of Power Converter in Wind Power Generation SystemsabstractPower electronics-based converters are the major and most vulnerable components in wind power generation systems. Converter faults will affect power quality,damage expensive equipment such asgenerators, or even pose a massive threat to the entire power grid. Fault diagnosis is considered as a powerful means to improve system reliability and reduce maintenance costs. Existing data-driven fault diagnosis methods related to improving feature extraction approaches to obtain high diagnostic accuracy are mostly based on single-scale feature of signals; this neglects the potentially valuable information of other scales. This paper proposes an algorithm of Dempster-Shafer and Deng entropy fusion multi-scale approximate entropy (DSDEMAE) for wind power converter fault diagnosis. Firstly, it calculates the multi-scale approximate entropy of fault signals, mining more potentially valuable information. Secondly, Dempster-Shafer is used to fuse features of various scales and effectively handles the conflicts and uncertainties between different features. Finally,Deng entropyis adopted to measure the uncertainty to adjust weight distribution, reducing the influence of highly conflicting features and mutually benefitting features of different scales. Extensive experimental results on simulated and experimental data demonstrate the effectiveness of the proposed algorithm. Compared with advanced methods, this method can diagnose the faults with higher accuracy and stronger robustness. Jinping Liang, Ke Zhang 0014, Ahmed Al-Durra, Daming Zhou |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Learning Cross-Attention Discriminators via Alternating Time-Space Transformers for Visual TrackingabstractIn the past few years, visual tracking methods with convolution neural networks (CNNs) have gained great popularity and success. However, the convolution operation of CNNs struggles to relate spatially distant information, which limits the discriminative power of trackers. Very recently, several Transformer-assisted tracking approaches have emerged to alleviate the above issue by combining CNNs with Transformers to enhance the feature representation. In contrast to the methods mentioned above, this article explores a pure Transformer-based model with a novel semi-Siamese architecture. Both the time-space self-attention module used to construct the feature extraction backbone and the cross-attention discriminator used to estimate the response map solely leverage attention without convolution. Inspired by the recent vision transformers (ViTs), we propose the multistage alternating time-space Transformers (ATSTs) to learn robust feature representation. Specifically, temporal and spatial tokens at each stage are alternately extracted and encoded by separate Transformers. Subsequently, a cross-attention discriminator is proposed to directly generate response maps of the search region without additional prediction heads or correlation filters. Experimental results show that our ATST-based model attains favorable results against state-of-the-art convolutional trackers. Moreover, it shows comparable performance with recent "CNN + Transformer" trackers on various benchmarks while our ATST requires significantly less training data. Wuwei Wang, Ke Zhang 0014, Jingyu Wang 0002, Qi Wang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Feature pre-inpainting enhanced transformer for video inpainting
Guanxiao Li, Ke Zhang 0014, Jingyu Wang 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | SOKS: Automatic Searching of the Optimal Kernel Shapes for Stripe-Wise Network PruningabstractIn spite of the remarkable performance, deep convolutional neural networks (CNNs) are typically over-parameterized and computationally expensive. Network pruning has become a popular approach to reducing the storage and calculations of CNN models, which commonly prunes filters in a structured way or discards single weights without structural constraints. However, the redundancy in convolution kernels and the influence of kernel shapes on the performance of CNN models have attracted little attention. In this article, we develop a framework, termed searching of the optimal kernel shape (SOKS), to automatically search for the optimal kernel shapes and perform stripe-wise pruning (SWP). To be specific, we introduce coefficient matrices regularized by a variety of regularization terms to locate important kernel positions. The optimal kernel shapes not only provide appropriate receptive fields for each convolution layer, but also remove redundant parameters in convolution kernels. SWP is also achieved by utilizing these irregular kernels and actual inference speedups on the graphics processing unit (GPU) are obtained. Comprehensive experimental results demonstrate that SOKS searches high-efficiency kernel shapes and achieves superior performance in terms of both compression ratio and inference latency. Embedding the searched kernels into VGG-16 increases the accuracy from 93.53% to 94.26% on CIFAR-10, while pruning 59.27% model parameters and reducing 27.07% inference latency. Guangzhe Liu, Ke Zhang 0014, Meibo Lv |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | ProbNet: Bayesian deep neural network for point cloud analysis
Ke Zhang 0014, Hua Luo, Jingyu Wang 0002 |
Comput. Graph. | 2 |
| 2022 | MRRNet: Learning multiple region representation for video person re-identification
Ke Zhang 0014, Jingyu Wang 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | RUFP: Reinitializing unimportant filters for soft pruning
Ke Zhang 0014, Guangzhe Liu, Meibo Lv |
Neurocomputing | 1 |
| 2022 | A novel system identification algorithm for nonlinear Markov jump system
Ke Zhang 0014, Minghu Tan |
Inf. Sci. | 2 |
| 2022 | Spatial temporal and channel aware network for video-based person re-identification
Ke Zhang 0014, Jingyu Wang 0002, Zhen Wang 0004 |
Image Vis. Comput. | 2 |
| 2022 | Hyperspectral Anomaly Detection via S1/2 and Total Variation Low Rank Matrix DecompositionabstractAnomaly detection (AD) on hyperspectral images has been widely researched in recent decades due to its high practicability and wide range of application scenarios. Such AD methods derived from low-rank matrix decomposition (LRMD) have appeared rapidly and been applied effectively. However, most of them focused on the use of spectral information and neglected the abundant spatial characteristics. In this letter, a spectral-spatial total variation (TV) (SSTV) regularized low-rank matrix decomposition method with a Schatten 1/2 quasi-norm ($S_{1/2}$) and denoising is proposed. First, to exploit the hyperspectral imagery (HSI) characteristics from the spectral perspective, we propose the low-rank matrix decomposition method with$S_{1/2}$norm and image denoising modules. Second, we incorporate the SSTV regularization by employing a 2-D TV (TV) spatially and 1-D TV along the spectral dimension to realize the maximized utilization of spatial characteristics of HSI. Finally, the alternating direction multiplier method (ADMM) is brought in the calculating process to attain the consequent detection results. The superiority of the proposed method has been demonstrated by the excellent performance on three real datasets. Jingyu Wang 0002, Ke Zhang 0014, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Hyperspectral Anomaly Detection via S1/2 Regularized Low Rank RepresentationabstractAnomaly detection has been drawing a great deal of attention by virtue of its practicability among the hyperspectral research area. Low-rank representation (LRR) has been widely employed to detect anomalies from hyperspectral imagery (HSI) effectively while a great number of methods derived from LRR replace rank function with a nuclear norm, which gives rise to a certain amount of error. In this letter, we propose a Schatten 1/2 quasi-norm ($S_{1/2}$) regularized LRR (SRLRR) method with an improved algorithm of establishing the dictionary for hyperspectral anomaly detection. First,$S_{1/2}$regularization is proposed to substitute the initial nuclear norm to approximate the rank function. Second, an improved dictionary construction algorithm based on K-Means++ clustering is presented to integrate the model and improve the performance. Finally, the optimization algorithm through alternating direction multiplier method (ADMM) incorporating a half threshold operator is introduced to attain the eventual results. Our method has been testified on three typical data sets and demonstrates the eminent performance. Jingyu Wang 0002, Ke Zhang 0014, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | FRPNet: A Feature-Reflowing Pyramid Network for Object Detection of Remote Sensing ImagesabstractAs a significant and fundamental task in the remote sensing field, object detection has received increasing attention and research studies. However, geospatial object detection is still a challenge owing to the dramatic variation in object scales, intraclass differences, and interclass similarity from multiscale and multiclass objects. To deal with these problems, an end-to-end feature-reflowing pyramid network (FRPNet) is proposed in this letter. FRPNet has two advantages that contribute to improve object detection accuracy. First, we embed a nonlocal block into the backbone in order to get the relevancy between different regions of the geospatial image for obtaining discriminative features. Furthermore, a feature-reflowing pyramid structure is proposed to generate high-quality feature presentation for each scale through fusing fine-grained features from the adjacent lower level, which improves the detection capability for multiscale and multiclass objects. Experiments on a public remote sensing data set DIOR illustrate that FRPNet can significantly improve the performance when compared to several state-of-the-art detection approaches in terms of mean average precision (mAP). Jingyu Wang 0002, Yezi Wang, Ke Zhang 0014, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | CDD-Net: A Context-Driven Detection Network for Multiclass Object DetectionabstractUnlike object detection in natural images that usually achieved great success, remote sensing imagery has its own challenges to detect and localize multiclass objects, such as large-scale change, uncertain direction, and high density. The context information of the objects is very worthwhile for solving these challenges in remote sensing images. In this letter, we propose a context-driven detection network (CDD-Net) to improve the accuracy of multiclass object detection in remote sensing images. For capturing the local neighboring objects and features, a local context feature network (LCFN) is proposed to learn the local context of the region of interest. Meanwhile, a hybrid attention pyramid network (HAPN) is designed, which can steer the focus to more valuable features. The HAPN inserts a squeeze and excitation block (SEB) and three asymmetric convolution blocks (ACBs) in the feature pyramid network (FPN). The experimental results over the DOTA-v1.5 data set demonstrate that the proposed CDD-Net yields promising results. Ke Zhang 0014, Jingyu Wang 0002, Yezi Wang, Qi Wang 0009, Qiang Li 0042 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Robust Correlation Tracking in Unmanned Aerial Vehicle Videos via Deep Target-Specific Rectification NetworksabstractUnmanned Aerial Vehicle (UAV) Videos have received active research attention in the remote sensing field by taking full advantage of the bird’s eye view offered by UAV. At the same time, visual tracking approaches based on discriminative correlation filters (DCF) have recently achieved increasing popularity and success. Despite their success, the tracking robustness and accuracy of existing DCF-based trackers are hard to promote in challenging tracking scenarios of aerial videos due to excessive reliance on the response map and fixed linear model update strategy. To resolve these issues, we propose a robust DCF-based tracking framework via an effective pretrained rectification network for UAV-based remote sensing. Specifically, the target-specific rectification network is offline trained to discriminatively classify the target and background. During the online tracking stage, the DCF module performs fast inference to obtain the potential locations of the target. After that, the deep rectification network evaluates the correlation-specific proposals offered by the DCF module and provides precise tracking results. Besides, to achieve a robust and adaptive model update strategy, we propose to finetune both the DCF module and rectification network according to the classification confidence of the estimated result. Extensive experimental results on recent UAV benchmarks demonstrate that our method achieves better performance than other competing algorithms. Ke Zhang 0014, Wuwei Wang, Jingyu Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Hierarchical Context Embedding Network for Object Detection in Remote Sensing ImagesabstractCompared with general optical images, remote sensing images (RSIs) capture large areas from high altitudes with a bird’s eye view, which is responsible for the many categories and scale variations of objects in the images, as well as the abundant scene information. Although the complexity of the RSIs presents a significant challenge to the object detection task, the complexity presents opportunities as well. The RSIs contain plenty of object-related context information, which is valuable for boosting the object detection performance. To address the existing issue of poor context utilization in RSIs, we propose a hierarchical context embedding network (HCENet) in this letter. First, we construct a semantic feature pyramid, in which the semantic context aggregation module (SFAM) integrates the semantic contexts included in the adjacent layers of features with a novel feature fusion mechanism. Furthermore, the scene-level context embedding module (SLCEM) extracts the scene context of the overall image by a simple design and is utilized to guide feature classification. Finally, we outperform the popular object detectors on the publicly available DOTA-v1.5 dataset, achieving superior performance. Ke Zhang 0014, Jingyu Wang 0002, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Semantic Context-Aware Network for Multiscale Object Detection in Remote Sensing ImagesabstractAccurate object detection in remote sensing images is an essential part of automatic extraction, analysis, and understanding of image information, which potentially plays a significant role in a number of practical applications. However, the scale diversity in remote sensing images presents a substantial challenge for object detection, regarded as one of the crucial problems to be solved. To extract multiscale feature representations and sufficiently exploit semantic context information, this letter proposes a semantic context-aware network (SCANet) model for multiscale object detection. We propose two novel modules, called receptive field-enhancement module (RFEM) and semantic context fusion module (SCFM), to enhance the performance of SCANet. The RFEM dedicates to more robust multiscale feature extraction by paying attention to distinct receptive fields through multibranch different convolutions. For the purpose of utilizing the semantic context information contained in the scene to guide the network to better detection accuracy, the SCFM integrates the semantic context features from the upper level with the lower level features and delivers them hierarchically. Experiments demonstrate that, compared with the state-of-the-art approaches, the SCANet yields superior detection results on the DOTA-v1.5 data set. Ke Zhang 0014, Jingyu Wang 0002, Yezi Wang, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Layer Pruning for Obtaining Shallower ResNetsabstractNetwork pruning has become an effective scheme to cut down the network complexity and speed up the inference. Current work mainly focuses on filter pruning, which deletes filters in a structured way. Although the calculation amount is reduced after pruning, the acceleration effect is not obvious. In this letter, we present a layer pruning method to obtain shallower ResNets (LPSR) and significantly reduce the inference latency. Taking advantage of the characteristics of the residual module, we disconnect the unimportant residual mapping and reserve only the identity mapping, thereby equivalently removing the corresponding residual module and achieving the purpose of layer pruning. We adopt Taylor expansion to estimate the impact of disconnecting BN layers in the residual mapping. These impacts are then normalized and used to identify unimportant residual blocks to prune. Comprehensive experimental results demonstrate that LPSR not only reduces the number of parameters and calculations, but also significantly speed up the inference. LPSR improves the accuracy of ResNet-56 on CIFAR-10 from 93.21% to 93.40%, while reducing 44.71% parameters, 52.68% calculations and 44.89% inference latency. Ke Zhang 0014, Guangzhe Liu |
IEEE Signal Process. Lett. | 1 |
| 2022 | Spatio-Temporal Online Matrix Factorization for Multi-Scale Moving Objects DetectionabstractDetecting moving objects from the video sequences has been treated as a challenging computer vision task, since the problems of dynamic background, multi-scale moving objects and various noise interference impact the corresponding feasibility and efficiency. In this paper, a novel spatio-temporal online matrix factorization (STOMF) method is proposed to detect multi-scale moving objects under dynamic background. To accommodate a wide range of the real noise distractions, we apply a specific mixture of exponential power (MoEP) distributions to the framework of low-rank matrix factorization (LRMF). For the optimization of solution algorithm, a temporal difference motion prior (TDMP) model is proposed, which estimates the motion matrix and calculates the weight matrix. Moreover, a partial spatial motion information (PSMI) post-processing method is further designed to implement multi-scale objects extraction in varieties of complex dynamic scenes, which utilizes partial background and motion information. The superiority of the STOMF method is validated by massive experiments on practical datasets, as compared with state-of-the-art moving objects detection approaches. Jingyu Wang 0002, Yue Zhao 0038, Ke Zhang 0014, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Real-Time Video Emotion Recognition Based on Reinforcement Learning and Domain KnowledgeabstractMultimodal emotion recognition in conversational videos (ERC) develops rapidly in recent years. To fully extract the relative context from video clips, most studies build their models on the entire dialogues which make them lack of real-time ERC ability. Different from related researches, a novel multimodal emotion recognition model for conversational videos based on reinforcement learning and domain knowledge (ERLDK) is proposed in this paper. In ERLDK, the reinforcement learning algorithm is introduced to conduct real-time ERC with the occurrence of conversations. The collection of history utterances is composed as an emotion-pair which represents the multimodal context of the following utterance to be recognized. Dueling deep-Q-network (DDQN) based on gated recurrent unit (GRU) layers is designed to learn the correct action from the alternative emotion categories. Domain knowledge is extracted from public dataset based on the former information of emotion-pairs. The extracted domain knowledge is used to revise the results from the RL module and is transformed into other dataset to examine the rationality. The experimental results on datasets show that ERLDK achieves the state-of-the-art results on weighted average and most of the specific emotion categories. Ke Zhang 0014, Yuanqing Li 0003, Jingyu Wang 0002, Erik Cambria, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Learning Adaptive Target-and-Surrounding Soft Mask for Correlation Filter Based Visual TrackingabstractVisual tracking is a very critical issue in computer vision and video processing. For Discriminative Correlation Filter (DCF)-based tracking methods, it is very essential and meaningful to adaptively incorporate reliable target and surrounding information from video frames. However, most existing DCF-based trackers solely rely on pre-defined and fixed constraints such as a binary mask or quadratic function-based regularization to improve the discrimination. Unfortunately, such attempts fail to adjust the constraints according to the change of tracking circumstance in the video sequence, and thus lead to the lack of reliability of learned filters. To mitigate these problems, we present a novel DCF-based tracking method that introduces an adaptive target-and-surrounding soft mask (ATSM) into the learning formula. The adaptive soft mask that is represented by float numbers contains the detail information for both target region and its surrounding information: first, for the background area, it introduces meaningful background information and suppressing uninformative one; second, for the target area inside the bounding box, it helps to focus on the reliable area and repress the rapidly changing area; third, the target-and-surrounding soft mask is adaptively adjusted based on the variations of the target and its surrounding during the tracking process. By jointly modeling the filter and the adaptive soft mask, our ATSM tracker achieves an efficient integration of meaningful information of both foreground and background and performs favorably against state-of-the-art algorithms on seven well-known benchmarks. Ke Zhang 0014, Wuwei Wang, Jingyu Wang 0002, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Detection of Small Aerial Object Using Random Projection Feature With Region ClusteringabstractSmall aerial object detection plays an important role in numerous computer vision tasks, including remote sensing, early warning systems, and visual tracking. Despite existing moving object detection techniques that can achieve reasonable results in normal size objects, they fail to distinguish the small objects from the dynamic background. To cope with this issue, a novel method is proposed for accurate small aerial object detection under different situations. Initially, the block segmentation is introduced for reducing frame information redundancy. Meanwhile, a random projection feature (RPF) is proposed for characterizing blocks into feature vectors. Subsequently, a moving direction estimation based on feature vectors is presented to measure the motions of blocks and filter out the major directions. Finally, variable search region clustering (VSRC), together with the color feature difference, is designed for extracting pixelwise targets from the remaining moving direction blocks. The comprehensive experiments demonstrate that our approach outperforms the level of state-of-the-art methods upon the integrity of small aerial objects, especially on the dynamic background and scale variation targets. Jingyu Wang 0002, Ke Zhang 0014, Yue Zhao 0038, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | GCWNet: A Global Context-Weaving Network for Object Detection in Remote Sensing ImagesabstractWith practical applications such as environment surveillance, agricultural production, and disaster assessment, accurate object detection in remote sensing images is in high demand. Precise detection of object instances in remote sensing images remains considerably challenging due to dense instance stacking, large-scale variations, and complex backgrounds. To solve the mentioned issues, a novel global context-weaving network (GCWNet) is developed for object detection in remote sensing images. We propose two novel modules for feature extraction and refinement, which include the global context aggregation module (GCAM) and the feature refinement module (FRM). GCAM assembles a global context with high-level and low-level features through feature weaving, which facilitates dense object detection. Meanwhile, FRM convolves multiple receptive fields by combining different branches, thereby further refining the features and improving the feature distinction at different scales. Furthermore, we design to alleviate the sample imbalanced problem during training using focal loss and balanced L1 loss to improve object classification and regression, respectively. The experimental results indicate that GCWNet achieves superior performance in object classification and localization on the DOTA-v1.5 dataset, which illustrates the superiority of GCWNet. Ke Zhang 0014, Jingyu Wang 0002, Yezi Wang, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | A cognitive brain model for multimodal sentiment analysis based on attention neural networks
Yuanqing Li 0003, Ke Zhang 0014, Jingyu Wang 0002, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2021 | ASKs: Convolution with any-shape kernels for efficient neural networks
Guangzhe Liu, Ke Zhang 0014, Meibo Lv |
Neurocomputing | 2 |
| 2021 | Discriminative visual tracking via spatially smooth and steep correlation filters
Wuwei Wang, Ke Zhang 0014, Meibo Lv, Jingyu Wang 0002 |
Inf. Sci. | 2 |
| 2021 | Feature Fusion for Multimodal Emotion Recognition Based on Deep Canonical Correlation AnalysisabstractFusion of multimodal features is a momentous problem for video emotion recognition. As the development of deep learning, directly fusing feature matrixes of each mode through neural networks at feature level becomes mainstream method. However, unlike unimodal issues, for multimodal analysis, finding the correlations between different modal is as important as discovering effective unimodal features. To make up the deficiency in unearthing the intrinsic relationships between multimodal, a novel modularized multimodal emotion recognition model based on deep canonical correlation analysis (MERDCCA) is proposed in this letter. In MERDCCA, four utterances are gathered as a new group and each utterance contains text, audio and visual information as multimodal input. Gated recurrent unit layers are used to extract the unimodal features. Deep canonical correlation analysis based on encoder-decoder network is designed to extract cross-modal correlations by maximizing the relevance between multimodal. The experiments on two public datasets show that MERDCCA achieves the better results. Ke Zhang 0014, Yuanqing Li 0003, Jingyu Wang 0002, Zhen Wang 0004, Xuelong Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Hierarchical Spatiotemporal Context-Aware Correlation Filters for Visual TrackingabstractDiscriminative correlation filters (DCF)-based trackers have been increasingly applied to visual tracking due to their high precision while running at high frame rates. However, most recent DCF-based methods solely concentrate on learning the correlation filter with spatial information and thus do not have sufficient descriptive power to discriminate the target from the background in the complex circumstances, such as full occlusion (OCC) and rapid target variation. In this article, we introduce a novel tracking framework that exploits the relationship between the target and its spatiotemporal context to improve tracking accuracy and robustness. Especially, we present our spatiotemporal context model in a hierarchical way, where each layer of the context pyramid is a spatial correlation filter learned from different temporal instances. For gaining an accurate spatiotemporal model, we propose an optimization fusion approach that can adaptively and efficiently learn the effect of each hierarchical layer and exploit these multiple temporal levels of correlation filters for visual tracking. Moreover, an adaptive model update strategy for correlation filters is introduced into the framework to dynamically select proper hierarchical layers, which boosts the temporal diversity of the target appearance, while radically reduces the number of model parameters and guarantees the real-time performance of the tracking method. The experimental results show that, with conventional handcrafted features, our tracker achieves the best success rates among available state-of-the-art trackers with handcrafted features, and provides state-of-the-art performance comparable to those of deep-learning-based trackers on OTB-2013, OTB-2015, VOT-2016, and UAV-20L benchmarks but runs significantly faster than deep trackers. Wuwei Wang, Ke Zhang 0014, Meibo Lv, Jingyu Wang 0002 |
IEEE Trans. Cybern. | 2 |
| 2018 | Computational Modelling Auditory AwarenessabstractInternational audience Jingyu Wang 0002, Ke Zhang 0014, Kurosh Madani |
IJCCI | 3 |
| 2018 | Morphological Band Selection for Hyperspectral ImageryabstractIn this letter, a novel morphological band selection method is proposed to obtain the most representative bands from hyperspectral image (HSI) in an unsupervised manner. In order to sufficiently process the HSI, we propose to use only a small set of data instead of using the original full data. For the obtained clusters, the differences of spectral response curves are applied to measure the local discrimination capability of bands, of which the local maximum value point is yielded based on the morphological processing. To verify the performance of the proposed method, the robustness of the parameters has been evaluated, while the effectiveness and superiority have been tested on three popular hyperspectral data sets. The experiment results have shown that the proposed method outperforms other methods. Jingyu Wang 0002, Ke Zhang 0014, Kurosh Madani, Christophe Sabourin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Research on the Mission Critical Parameters Identification by using Kinematic Boundaries
Jingyu Wang 0002, Usman Fareed, Ke Zhang 0014, Pei Wang 0011 |
ICINCO (1) | 3 |
| 2017 | Unsupervised Band Selection Using Block-Diagonal Sparsity for Hyperspectral Image ClassificationabstractIn order to alleviate the negative effect of curse of dimensionality, band selection is a crucial step for hyperspectral image (HSI) processing. In this letter, we propose a novel unsupervised band selection approach to reduce the dimensionality for hyperspectral imagery. In order to obtain the most representative bands, the correlation matrix computed from the original HSI is used to describe the correlation characteristics among bands, while the block-diagonal structure is measured to segment all bands into a series of subspace. After applying the spectral clustering algorithm, the optimal combination of band is finally selected. To verify the effectiveness and superiority of the proposed band selection method, experiments have been conducted on three widely used real-world hyperspectral data. The results have shown that the proposed method outperforms other methods in HSI classification application. Jingyu Wang 0002, Ke Zhang 0014, Pei Wang 0011, Kurosh Madani, Christophe Sabourin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | The Design of Finite-time Convergence Guidance Law for Head Pursuit based on Adaptive Sliding Mode ControlabstractThe high-speed target interception plays a vital role in the modern industry with many application scenarios. Due to the difficulties of direct interception in high speed, the head pursuit intercept is frequently considered for the target of re-entering flight vehicle. In this paper, a novel terminal sliding mode control method is proposed for the interception of high-speed maneuvering target, in which the finite convergence guidance law is initially designed under the constraint of intercept angle. By introducing the mathematical model of correlative motion between interceptor and target, the sliding surface is researched and designed to meet the critical conditions of head pursuit intercepting. Meanwhile, considering the dynamic characteristics of both interceptor and target, an adaptive guidance law is therefore proposed to compensate the modelling errors, which goal is to improve the accuracy of interception. The stability analysis is theoretically proved in terms of the Lyapunov method. Numerical simulations are presented to validate the robustness and effectiveness of the proposed guidance law, by which good intercepting performance can be supported. Ke Zhang 0014, Jingyu Wang 0002 |
ICINCO (2) | 2 |
| 2015 | Salient Foreground Object Detection based on Sparse Reconstruction for Artificial AwarenessabstractArtificial awareness is an interesting way of realizing artificial intelligent perception for machines. Since the foreground object can provide more useful information for perception and informative description of the environment than background regions, the informative saliency characteristics of the foreground object can be treated as a important cue of the objectness property. Thus, a sparse reconstruction error based detection approach is proposed in this paper. To be specific, the overcomplete dictionary is trained by using the image features derived from randomly selected background images, while the reconstruction error is computed in several scales to obtain better detection performance. Experiments on popular image dataset are conducted by applying the proposed approach, while comparison tests by using a state of the art visual saliency detection method are demonstrated as well. The experimental results have shown that the proposed approach is able to detect the foreground object which is distinct for awareness, and has better performance in detecting the information salient foreground object for artificial awareness than the state of the art visual saliency method. Jingyu Wang 0002, Ke Zhang 0014, Kurosh Madani, Christophe Sabourin |
ICINCO (2) | 2 |
| 2015 | Salient environmental sound detection framework for machine awareness
Jingyu Wang 0002, Ke Zhang 0014, Kurosh Madani, Christophe Sabourin |
Neurocomputing | 2 |