EDBT 2026 Demo / reviewers in the wild / expert
Hu Zhu
dblp:195/2123
· DBLP profile ↗
36ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0002-5528-8721ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object DetectionabstractCamera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost. Hu Zhu, Weihao Gu, Yang Yang 0062, Yanyan Liang 0001 |
AAAI | 2 |
| 2025 | UniScene: Unified Occupancy-centric Driving Scene GenerationabstractGenerating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/. Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014 |
CVPR | 7 |
| 2025 | PVSSNet: Progressive Feature Interaction Visual State-Space Network for Multispectral PansharpeningabstractPansharpening involves extracting spectral information from multispectral images and structural details from panchromatic images, then fusing them to produce high-resolution multispectral remote sensing images. However, high-resolution multispectral images often suffer from spectral or structural information loss. In this paper, we introduce a pansharpening algorithm based on a Progressive Feature Interaction Visual State Space Network. It enables interaction between local and global features of multispectral and panchromatic images and facilitates the injection of spectral and spatial details through distinct attention modules. This approach effectively preserves both spectral characteristics and spatial structure through inter-branch information interaction and complementation. Additionally, by integrating a visual state space network, the proposed model achieves deep reconstruction of multi-scale global information, enhancing robustness and generalization. Extensive experimental results demonstrate that the proposed network achieves highly competitive performance in both visual assessments and objective metric evaluations. Guoxia Xu, Zhenwei Xu, Lizhen Deng, Hu Zhu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | SSP-Net: A Siamese-Based Structure-Preserving Generative Adversarial Network for Unpaired Medical Image EnhancementabstractRecently, unpaired medical image enhancement is one of the important topics in medical research. Although deep learning-based methods have achieved remarkable success in medical image enhancement, such methods face the challenge of low-quality training sets and the lack of a large amount of data for paired training data. In this article, a dual input mechanism image enhancement method based on Siamese structure (SSP-Net) is proposed, which takes into account the structure of target highlight (texture enhancement) and background balance (consistent background contrast) from unpaired low-quality and high-quality medical images. Furthermore, the proposed method introduces the mechanism of the generative adversarial network to achieve structure-preserving enhancement by jointly iterating adversarial learning. Experiments comprehensively illustrate the performance in unpaired image enhancement of the proposed SSP-Net compared with other state-of-the-art techniques. Guoxia Xu, Hao Wang 0003, Marius Pedersen, Meng Zhao 0001, Hu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Deep Tensor Evidence Fusion Network for Sentiment ClassificationabstractRecently, a multimodal sentiment analysis of social media has attracted increasing attention, and its core idea is to discovery heuristic fusion strategy to analyze the sentiment orientations over heterogeneous multimodal source from a learned compact multimodal representation. The existing multimodal fusion techniques not only struggle to achieve full heterogeneous data interaction, but also they are unable to dynamically assess the quality of various modal data to determine predictability. In this article, we present a novel deep tensor evidence fusion (DTEF) network for multimodal sentiment classification. First, we propose a common view evaluation network that uses a long short-term memory (LSTM) network and a tensor-based neural network to extract rich intermodal and intramodal information. Then, we propose a unique time cue evaluation network that takes advantage of the temporal granularity associated with numerous pattern sequences. To make reliable decisions, we finally incorporate uncertainty through the trusted fusion layer, which improves the accuracy and robustness of sentimental classification. Our model is validated using the CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) and CMU Multimodal Corpus of Sentiment Intensity (CMU-MOSI) datasets, and the experimental findings demonstrate the superior performance of the proposed network in terms of accuracy compared with the state-of-the-art methods. Guoxia Xu, Xiaokang Zhou, Jung Yoon Kim, Hu Zhu, Lizhen Deng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Trusted Multimodal Socio-Cyber Sentiment Analysis Based on Disentangled Hierarchical Representation LearningabstractThe rapid development of the digital age has led to a qualitative leap in social media. To meet the cognitive needs of users, social media platforms have been mining users’ private information and disseminating information through various means. However, these platforms lack effective management of information release and various forms of emotional expressions make public propaganda increasingly diverse and complex. Therefore, accurately identifying the relationships between multimodal data poses a challenge. An effective modal representation must consider both the consistency of multimodal data and the complementarity of single-modal data. However, existing methods focus on fusing different modal features into a unified feature representation, while neglecting to evaluate the reliability of prediction results. In this article, we disentangle the consistency and complementarity in the fused representation problem of multimodal data. We construct the modal private task (unique) by using the Dirichlet distribution and evidence theory to solve the uncertainty of each modal prediction. The model can output the uncertainty of prediction and learn complementary information through the fusion of decision layers. At the same time, we construct the modal common task using a low-rank tensor fusion model to learn consistent features. Finally, we compare the model with the current mainstream methods on three public datasets, and the experimental results show that the performance of our method reaches the level of current advanced algorithms. Guoxia Xu, Lizhen Deng, Yansheng Li 0001, Yantao Wei, Xiaokang Zhou, Hu Zhu |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | DM-Fusion: Deep Model-Driven Network for Heterogeneous Image FusionabstractHeterogeneous image fusion (HIF) is an enhancement technique for highlighting the discriminative information and textural detail from heterogeneous source images. Although various deep neural network-based HIF methods have been proposed, the most widely used single data-driven manner of the convolutional neural network always fails to give a guaranteed theoretical architecture and optimal convergence for the HIF problem. In this article, a deep model-driven neural network is designed for this HIF problem, which adaptively integrates the merits of model-based techniques for interpretability and deep learning-based methods for generalizability. Unlike the general network architecture as a black box, the proposed objective function is tailored to several domain knowledge network modules to model the compact and explainable deep model-driven HIF network termed DM-fusion. The proposed deep model-driven neural network shows the feasibility and effectiveness of three parts, the specific HIF model, an iterative parameter learning scheme, and data-driven network architecture. Furthermore, the task-driven loss function strategy is proposed to achieve feature enhancement and preservation. Numerous experiments on four fusion tasks and downstream applications illustrate the advancement of DM-fusion compared with the state-of-the-art (SOTA) methods both in fusion quality and efficiency. The source code will be available soon. Guoxia Xu, Chunming He, Hao Wang 0003, Hu Zhu, Weiping Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Mobile User Pairing Scheme in NOMA-Enabled Backscatter Communication NetworksabstractIn this paper, we study the problem of mobile user pairing in non-orthogonal multi-access (NOMA) backscatter communication networks, in which re-pairing users after a change in the communication link can bring significant computational overhead. Firstly, We propose a user pairing scheme that employs the Kuhn-Munkres (KM) algorithm for the uplink of the monostatic backscatter communication network, which can effectively avoid the problem of frequent sorting mobile users. Secondly, to address the NOMA principle violations problem (NPVP) caused by mobile users, we propose a modified role switching (RS) technique based on the difference of user channel gain in which paired users switch their roles based on the magnitude of their channel gain after moving. Finally, we propose a dynamic power allocation scheme for mobile users to maximize the system throughput after pairing. In backscatter communication, the difference in power levels between paired users is achieved by setting different reflection coefficients. At the same time, it is necessary to consider the balance between data transmission and energy harvesting, and the reflection coefficient must not exceed the predetermined upper limit. Simulation results indicate that our proposed user pairing scheme outperforms conventional schemes in terms of the number of decoded users and system throughput. Tingpei Huang, Hu Zhu, Jianhang Liu, Shibao Li |
MSN | 3 |
| 2023 | Semantic-Aware Attack and Defense on Deep Hashing Networks for Remote-Sensing Image RetrievalabstractDeep hashing networks have been successful in retrieving interesting images from massive remote sensing images. There is no doubt that security and reliability are critical in remote sensing image retrieval. Recent studies about natural image retrieval have shown the vulnerability of deep hashing networks to adversarial examples, but there do not exist any researches about the attack and defense on deep hashing networks in remote sensing image retrieval. Due to the large intra-class difference and high inter-class similarity of remote sensing images, the attack and defense methods on deep hashing networks for natural images cannot be directly applied to the remote sensing images. Different from the widely adopted instance-aware hash codes which often present the suboptimum performance of the attack and defense on deep hashing networks, this paper recommends the usage of semantic-aware hash codes, which take into account multiple samples in the given semantic categories, in both attack and defense. To pursue the strongest attack on remote sensing image retrieval, a novel semantic-aware attack with weights via multiple random initialization (RWC) is proposed. To alleviate the retrieval degradation caused by adversarial attacks, a new adversarial training defense method on deep hashing networks with the adversarial semantic-aware consistency constraint (ACN) is proposed. Extensive experiments on three typical open remote sensing image datasets (i.e., UCM, AID, NWPU-RESISC45) show the proposed attack and defense methods on various deep hashing networks achieve better performance compared with the state-of-the-art methods. The source code will be made publicly available along with this paper. Yansheng Li 0001, Mengze Hao, Hu Zhu, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Learning the Distribution-Based Temporal Knowledge With Low Rank Response Reasoning for UAV Visual TrackingabstractIn recent years, the constraint based correlation filter has shown good performance in unmanned aerial vehicle (UAV) tracking, which gains a lot popularity in many intelligence transportation applications. In this work, a distribution-based temporal knowledge driven method is proposed to leverage the temporal translation property in UAV tracking. Instead of focusing on the traditional issues in the correlation filter, we provide a new method of learning parametric distribution on temporal knowledge by Wasserstein distance which is successfully embedded to solve the problem of temporal degeneration in learning process of tracking. Furthermore, we approximate optimal response reasoning with low-rank constraint over response consistency. Furthermore, the proposed method is solved by a simple iterative scheme with alternating direction multiplication ADMM algorithm. We demonstrate the superior tracking performance in several public standard UAV tracking benchmarks compared with state-of-the-art algorithms. Guoxia Xu, Hao Wang 0003, Meng Zhao 0001, Marius Pedersen, Hu Zhu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Unpaired Self-supervised Learning for Industrial Cyber-Manufacturing Spectrum Blind DeconvolutionabstractCyber-Manufacturing combines industrial big data with intelligent analysis to find and understand the intangible problems in decision-making, which requires a systematic method to deal with rich signal data. With the development of spectral detection and photoelectric imaging technology, spectral blind deconvolution has achieved remarkable results. However, spectral processing is limited by one-dimensional signal, and there is no available structural information with few training samples. Moreover, in the majority of practical applications, it is entirely feasible to gather unpaired spectrum dataset for training. This training method of unpaired learning is practical and valuable. Therefore, a two-stage deconvolution scheme combining self supervised learning and feature extraction is proposed in this paper, which generates two complementary paired sets through self supervised learning to extract the final deconvolution network. In addition, a new deconvolution network is designed for feature extraction. The spectrum is pre-trained through spectral feature extraction and noise estimation network to improve the training efficiency and meet the assumed noise characteristics. Experimental results show that this method is effective in dealing with different types of synthetic noise. Lizhen Deng, Guoxia Xu, Jiaqi Pi, Hu Zhu, Xiaokang Zhou |
ACM Trans. Internet Techn. | 4 |
| 2022 | A Self-paced Learning based Transfer Model for Hypergraph Matching
Hu Zhu, Guoxia Xu, Lizhen Deng |
Inf. Sci. | 1 |
| 2022 | Kernel embedding transformation learning for graph matching
Yu-Feng Yu 0001, Long Chen 0001, Ke-Kun Huang, Hu Zhu, Guoxia Xu |
Pattern Recognit. Lett. | 4 |
| 2022 | A Dual Stream Spectrum Deconvolution Neural NetworkabstractWith the development of spectral detection and photoelectric imaging, multiband spectrum is always degraded by the random noise and band overlap during the acquisition of spectrum devices. Owing to the fixed spectrum degradation model, the existing spectrum deconvolution technologies are sensitive to the handcrafted model designed and manually selected parameters. The fundamental cause of these limitations during spectral analysis is that spectral processing is limited by 1-D signal without structural information available and insufficient training samples. In this article, a dual stream neural network is proposed to reconstruct the original infrared spectroscopy, which effectively strengthens the capability to represent the feature of infrared spectrum. A novel activation function is proposed to realize the function of the dual stream network. Furthermore, a heuristic learning strategy from the perspective of balanced self-paced learning is exploited to help network train from simple to difficult, resolving the problem of high sample repeatability. Compared with other traditional methods, the experimental results show that our network can achieve state-of-the-art reconstruction result and fairly excellent performance in terms of the corresponding index within synectics and real spectrum experiments. Lizhen Deng, Guoxia Xu, Yanyu Dai, Hu Zhu |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | PcGAN: A Noise Robust Conditional Generative Adversarial Network for One Shot LearningabstractTraffic sign classification plays a vital role in autonomous vehicles for its powerful capability in information representation. However, the low-quality data of traffic signs captured by in-vehicle cameras often inevitably bring inherent challenges to the one-shot classification task. Apart from the problem of data degradation, learning-based classification techniques of real traffic signs also come across the challenges of intra-class and inter-class data imbalance from the training data. To overcome the aforementioned problems, we propose an end-to-end degradation robust deep model, termed PcGAN, to classify traffic signs in a manner of few-shot learning. The proposed PcGAN models the joint distribution between the degraded traffic signal data and the corresponding prototypes from both degradation removal and generation perspectives by two alternating optimized modules, which ensures the generalization of the learned embedding of latent space for novel tasks. A multi-task loss function is designed to improve the robustness of PcGAN. Numerous experiments comprehensively demonstrate that the accuracy of our proposed PcGAN is improved by 5% compared with other state-of-the-art (SOTA) approaches in few-shot classification. Lizhen Deng, Chunming He, Guoxia Xu, Hu Zhu, Hao Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Digital Twins in Unmanned Aerial Vehicles for Rapid Medical Resource Delivery in EpidemicsabstractThe purposes are to explore the effect of Digital Twins (DTs) in Unmanned Aerial Vehicles (UAVs) on providing medical resources quickly and accurately during COVID-19 prevention and control. The feasibility of UAV DTs during COVID-19 prevention and control is analyzed. Deep Learning (DL) algorithms are introduced. A UAV DTs information forecasting model is constructed based on improved AlexNet, whose performance is analyzed through simulation experiments. As end-users and task proportion increase, the proposed model can provide smaller transmission delays, lesser energy consumption in throughput demand, shorter task completion time, and higher resource utilization rate under reduced transmission power than other state-of-art models. Regarding forecasting accuracy, the proposed model can provide smaller errors and better accuracy in Signal-to-Noise Ratio (SNR), bit quantizer, number of pilots, pilot pollution coefficient, and number of different antennas. Specifically, its forecasting accuracy reaches 95.58% and forecasting velocity stabilizes at about 35 Frames-Per-Second (FPS). Hence, the proposed model has stronger robustness, making more accurate forecasts while minimizing the data transmission errors. The research results can reference the precise input of medical resources for COVID-19 prevention and control. Zhihan Lyu, Hailin Feng, Hu Zhu, Haibin Lv |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Bilateral Weighted Regression Ranking Model With Spatial-Temporal Correlation Filter for Visual TrackingabstractMany discriminative correlation filter (DCF)-based methods have successfully leveraged the guidance for solving two problems (i.e., the boundary effect and temporal filtering degradation) as a model prior to visual tracking. The intuitive motivation of these methods is to control the degeneration of the updating loss of the objective function with a structural framework. While these methods rely mostly on various regularization items, they always ignore the loss from data fidelity term. Therefore, we propose a bilateral weighted regression ranking model termed as BWRR. Here, we resort to two procedures for solving the above problems. First, BWRR introduces a bilateral constraint into the data fidelity term to control the loss of rows and columns of the filter learning data term. The weighted matrices could impose an adaptive penalty for large data loss during the learning process to avoid the model degradation problem. Second, the data of the updated weighted matrices is not directly applied to the calculation of the filter during each iteration. Instead, a new weighted product matrix is obtained by ranking and numerical transformation for updating the filter. We show that the proposed model converts the original correlation filter regression problem into a regression-with-ranking problem, thus avoiding the problem of positive and negative sample imbalance. Overall, the BWRR model is iteratively solved by the alternating direction method of multipliers(ADMM). Qualitative and quantitative evaluations demonstrate the effectiveness and superiority of our proposed method by extensive and quantitative experiments on the OTB, VOT, and UAV datasets. Hu Zhu, Guoxia Xu, Lizhen Deng, Yueying Cheng, Aiguo Song |
IEEE Trans. Multim. | 1 |
| 2021 | M-UPS: A multi-user Pairing Scheme for NOMA-enabled Backscatter Communication NetworksabstractIn this paper, we investigate the user pairing schemes for NOMA-enabled backscatter communication network. Since the successful decoding of high channel gain users can significantly improve the system performance, to make more high channel gain users be decoded successfully, we propose a pairing scheme, M-UPS, which is based on the user channel gain difference. Firstly, in the case of two-user pairing, we divide users into the high channel gain group and the low channel gain group based on channel gain, then we provide a formula of the pairing distance threshold for the high channel gain user to select low channel gain users for pairing. Secondly, we extend the proposed paring scheme to the multi-user case, a dynamic user pairing scheme, which can adaptively form the different pairs with different numbers of users based on the channel gain differences. Simulation results show that M-UPS can outperform the Conventional-NOMA (C-NOMA), Uniform Channel Gain Difference-NOMA (UCGD-NOMA), and OMA in terms of the number of paired users and total throughput. Tingpei Huang, Hu Zhu, Shibao Li, Jianhang Liu |
ICPADS | 2 |
| 2021 | Video smoke removal based on low-rank tensor completion via spatial-temporal continuity constraintabstractAbstract Smoke has a very bad effect on the outdoor vision system. Not only are the videos with poor visual effects obtained, but also the quality and structure of the videos are reduced. In this paper, we propose a video smoke removal method based on low‐rank tensor completion via spatial‐temporal continuity constraint. The proposed method is based on the smoke mixing model and consider the sparseness of smoke and the global and local consistency of clean video. Then, the optimal solution of the smoke removal algorithm model is quickly realized by the Alternating Direction Method of Multiplier. Finally, we evaluate the experiment results of real‐world data and simulated data from the visual effects and objective indicators. And the experiment results show that our proposed algorithm can achieve better smoke removal results. Hu Zhu, Guoxia Xu, Lizhen Deng |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Vector co-occurrence morphological edge detection for colour imageabstractAbstract Morphological edge detection is a principal component in pattern recognition and machine vision. Traditional edge detection operators only take pixel mutual into consideration. However, the edges are influenced not only by pixel mutual but also by the boundary characteristics. Here, the vector co‐occurrence morphological edge detection operator is proposed, which takes the pixel and boundary information both into consideration. The vector co‐occurrence algorithm is exploited to resist the influence of the noise points and detect the edges from the colour image rather than the grey image. And, we lead to define a precise definition of the manner of sorting high‐dimensional data for the colour image. The experiment results always illustrate the advancement and practicability of our methods against the baseline method. In terms of experiments, the BSDS500 dataset is introduced to compare and analyse with other algorithms. Based on the standard benchmark index evaluation in the BSDS500 dataset, the ODS and AP of various algorithms are compared and analysed. Chunming He, Yu-Feng Yu 0001, Guoxia Xu, Hu Zhu, Lizhen Deng |
IET Image Process. | 5 |
| 2021 | A parallel multi-block alternating direction method of multipliers for tensor completionabstractAbstract This paper proposes an algorithm for the tensor completion problem of estimating multi‐linear data under the limitation of observation rate. Many tensor completion methods are based on nuclear norm minimization, they may fail to achieve the global solution for solving nuclear norm minimization in tensor completion problem with high missing ratio. To tackle this issue, an adaptive tensor completion method based on parallel multi‐block alternating direction method of multipliers (ADMM) algorithm is proposed, it can derive the model from the initial estimate and compute the next estimate from the current solution. The parallel multi‐block ADMM with global convergence is adopted to solve the dual problem, which greatly improves the processing power and reliability of the algorithm. Hu Zhu, Taiyu Yan, Yu-Feng Yu 0001, Lizhen Deng, Bing-Kun Bao |
IET Image Process. | 1 |
| 2021 | RoDeRain: Rotational Video Derain via Nonconvex and Nonsmooth Optimization
Lizhen Deng, Guoxia Xu, Hu Zhu, Bing-Kun Bao |
Mob. Networks Appl. | 3 |
| 2021 | Infrared small target detection via adaptive M-estimator ring top-hat transformation
Lizhen Deng, Jieke Zhang, Guoxia Xu, Hu Zhu |
Pattern Recognit. | 4 |
| 2021 | Dual Calibration Mechanism Based L2, p-Norm for Graph MatchingabstractUnbalanced geometric structure caused by variations with deformations, rotations and outliers is a critical issue that hinders correspondence establishment between image pairs in existing graph matching methods. To deal with this problem, in this work, we propose a dual calibration mechanism (DCM) for establishing feature points correspondence in graph matching. In specific, we embed two types of calibration modules in the graph matching, which model the correspondence relationship in point and edge respectively. The point calibration module performs unary alignment over points and the edge calibration module performs local structure alignment over edges. By performing the dual calibration, the feature points correspondence between two images with deformations and rotations variations can be obtained. To enhance the robustness of correspondence establishment, the L2,p-norm is employed as the similarity metric in the proposed model, which is a flexible metric due to setting the different p values. Finally, we incorporate the dual calibration and L2,p-norm based similarity metric into the graph matching model which can be optimized by an effective algorithm, and theoretically prove the convergence of the presented algorithm. Experimental results in the variety of graph matching tasks such as deformations, rotations and outliers evidence the competitive performance of the presented DCM model over the state-of-the-art approaches. Yu-Feng Yu 0001, Guoxia Xu, Ke-Kun Huang, Hu Zhu, Long Chen 0001, Hao Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Tensor Field Graph-Cut for Image Segmentation: A Non-Convex PerspectiveabstractImage segmentation is a key component of image analysis, which refers to the process of partitioning the image into multiple segments. Graph cut is widely used in image segmentation by constructing a graph that the minimal cut of this graph would lead to partition the corresponding pixels of the different objects. In this paper, we reconstruct the graph cut problem as a special non-convex optimization problem instead of the traditional maximum flow problem. We extend this non-convex problem to the hypergraph method and combine it with a tensor field based on a directional bilateral filter bank to achieve segmentation in grayscale images. Accordingly, an efficient minimization algorithm is proposed to solve this non-convex problem with global convergence. Furthermore, we have selected the data of BSDS300 and BSDS500 as tests. Experimental results and evaluation index tests further demonstrate the superiority of the proposed method. Hu Zhu, Jieke Zhang, Guoxia Xu, Lizhen Deng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Joint Transformation Learning via the L2, 1-Norm Metric for Robust Graph MatchingabstractEstablishing correspondence between two given geometrical graph structures is an important problem in computer vision and pattern recognition. In this paper, we propose a robust graph matching (RGM) model to improve the effectiveness and robustness on the matching graphs with deformations, rotations, outliers, and noise. First, we embed the joint geometric transformation into the graph matching model, which performs unary matching over graph nodes and local structure matching over graph edges simultaneously. Then, the L2,1-norm is used as the similarity metric in the presented RGM to enhance the robustness. Finally, we derive an objective function which can be solved by an effective optimization algorithm, and theoretically prove the convergence of the proposed algorithm. Extensive experiments on various graph matching tasks, such as outliers, rotations, and deformations show that the proposed RGM model achieves competitive performance compared to the existing methods. Yu-Feng Yu 0001, Guoxia Xu, Min Jiang 0003, Hu Zhu, Dao-Qing Dai, Hong Yan 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | Elastic Net Constraint-Based Tensor Model for High-Order Graph MatchingabstractThe procedure of establishing the correspondence between two sets of feature points is important in computer vision applications. In this article, an elastic net constraint-based tensor model is proposed for high-order graph matching. To control the tradeoff between the sparsity and the accuracy of the matching results, an elastic net constraint is introduced into the tensor-based graph matching model. Then, a nonmonotone spectral projected gradient (NSPG) method is derived to solve the proposed matching model. During the optimization of using NSPG, we propose an algorithm to calculate the projection on the feasible convex sets of elastic net constraint. Further, the global convergence of solving the proposed model using the NSPG method was proved. The superiority of the proposed method is verified through experiments on the synthetic data and natural images. Hu Zhu, Chunfeng Cui, Lizhen Deng, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Infrared Small Target Detection via Low-Rank Tensor Completion With Top-Hat RegularizationabstractInfrared small target detection technology is one of the key technologies in the field of computer vision. In recent years, several methods have been proposed for detecting small infrared targets. However, the existing methods are highly sensitive to challenging heterogeneous backgrounds, which are mainly due to: 1) infrared images containing mostly heavy clouds and chaotic sea backgrounds and 2) the inefficiency of utilizing the structural prior knowledge of the target. In this article, we propose a novel approach for infrared small target detection in order to take both the structural prior knowledge of the target and the self-correlation of the background into account. First, we construct a tensor model for the high-dimensional structural characteristics of multiframe infrared images. Second, inspired by the low-rank background and morphological operator, a novel method based on low-rank tensor completion with top-hat regularization is proposed, which integrates low-rank tensor completion and a ring top-hat regularization into our model. Third, a closed solution to the optimization algorithm is given to solve the proposed tensor model. Furthermore, the experimental results from seven real infrared sequences demonstrate the superiority of the proposed small target detection method. Compared with traditional baseline methods, the proposed method can not only achieve an improvement in the signal-to-clutter ratio gain and background suppression factor but also provide a more robust detection model in situations with low false-positive rates. Hu Zhu, Shiming Liu, Lizhen Deng, Yansheng Li 0001, Fu Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | DSPNet: A Lightweight Dilated Convolution Neural Networks for Spectral Deconvolution With Self-Paced LearningabstractIn the fields of industry research, infrared spectrometers are widely used in diverse applications. However, the spectrum often suffers from band overlap and random noise due to the distortion caused by the point spread function, especially for aging instruments. The problem of reconstructing the clear spectrum from the degraded spectrum is called spectrum deconvolution. Traditional partial differential equation (PDE) methods rely on distribution assumptions in the reconstructed process. This restriction makes PDE methods sensitive to tackle complex instrumental broadening effect in the dispersive IR spectrometers. Also, we need to spend much time setting the parameters of PDE models manually. These problems intuitively degrade the performances of PDE methods. In this article, we propose an end-to-end neural network framework for spectral deconvolution problem. The novelty of this article lies in its strong robustness from dilated deconvolution and self-paced learning procedure to challenge the complicated degraded spectra. Actually, the deconvolution problem is tailored to a dense prediction problem in this article. Inspired by the extensive use and excellent effects of dilated convolutions in dense prediction, a lightweight dilated convolution module is given to detect the overlaps of degraded spectra. Experimental results demonstrate that the proposed solution has an outstanding performance against many other approaches. Such improvements have the potential to facilitate industrial applications and further exploration of an unknown chemical mixture. Our framework has a good performance on feature extracting and spectrum reconstruction, even in the case of low signal-to-noise ratio. Hu Zhu, Yiming Qiao, Guoxia Xu, Lizhen Deng, Yu-Feng Yu 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | TNLRS: Target-Aware Non-Local Low-Rank Modeling With Saliency Filtering Regularization for Infrared Small Target DetectionabstractRecently, infrared small target detection problem has attracted substantial attention. Many works based on local low-rank model have been proven to be very successful for enhancing the discriminability during detection. However, these methods construct patches by traversing local images and ignore the correlations among different patches. Although the calculation is simplified, some texture information of the target is ignored, and targets of arbitrary forms cannot be accurately identified. In this paper, a novel target-aware method based on a non-local low-rank model and saliency filter regularization is proposed, with which the newly proposed detection framework can be tailored as a non-convex optimization problem, therein enabling joint target saliency learning in a lower dimensional discriminative manifold. More specifically, non-local patch construction is applied for the proposed target-aware low-rank model. By combining similar patches, we reconstruct them together to achieve a better generalization of non-local spatial sparsity constraints. Furthermore, to encourage target saliency learning, our proposed saliency filtering regularization term based on entropy is restricted to lie between the background and foreground. The regularization of the saliency filtering locally preserves the contexts from the target and surrounding areas and avoids the deviated approximation of the low-rank matrix. Finally, a unified optimization framework is proposed and solved with the alternative direction multiplier method (ADMM). Experimental evaluations of real infrared images demonstrate that the proposed method is more robust under different complex scenes compared with some state-of-the-art methods. Hu Zhu, Haopeng Ni, Shiming Liu, Guoxia Xu, Lizhen Deng |
IEEE Trans. Image Process. | 1 |
| 2020 | Multimodal Fusion Method Based on Self-Attention MechanismabstractMultimodal fusion is one of the popular research directions of multimodal research, and it is also an emerging research field of artificial intelligence. Multimodal fusion is aimed at taking advantage of the complementarity of heterogeneous data and providing reliable classification for the model. Multimodal data fusion is to transform data from multiple single-mode representations to a compact multimodal representation. In previous multimodal data fusion studies, most of the research in this field used multimodal representations of tensors. As the input is converted into a tensor, the dimensions and computational complexity increase exponentially. In this paper, we propose a low-rank tensor multimodal fusion method with an attention mechanism, which improves efficiency and reduces computational complexity. We evaluate our model through three multimodal fusion tasks, which are based on a public data set: CMU-MOSI, IEMOCAP, and POM. Our model achieves a good performance while flexibly capturing the global and local connections. Compared with other multimodal fusions represented by tensors, experiments show that our model can achieve better results steadily under a series of attention mechanisms. Hu Zhu, Yingying Hua, Guoxia Xu, Lizhen Deng |
Wirel. Commun. Mob. Comput. | 1 |
| 2019 | Dilated-aware discriminative correlation filter for visual tracking
Guoxia Xu, Hu Zhu, Lizhen Deng, Lixin Han, Yujie Li 0001, Huimin Lu 0001 |
World Wide Web | 2 |
| 2018 | Discriminative tracking via supervised tensor learning
Guoxia Xu, Sheheryar Khan, Hu Zhu, Lixin Han, Michael Kwok-Po Ng, Hong Yan 0001 |
Neurocomputing | 3 |
| 2018 | Adaptive top-hat filter based on quantum genetic algorithm for infrared small target detection
Lizhen Deng, Hu Zhu, Quan Zhou 0004, Yansheng Li 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Face recognition via fast dense correspondence
Quan Zhou 0004, Wenbin Yu 0002, Yawen Fan, Hu Zhu, Xiaofu Wu, Weihua Ou, Wei-Ping Zhu 0001, Longin Jan Latecki |
Multim. Tools Appl. | 5 |
| 2018 | Large-Scale Remote Sensing Image Retrieval by Deep Hashing Neural NetworksabstractAs one of the most challenging tasks of remote sensing big data mining, large-scale remote sensing image retrieval has attracted increasing attention from researchers. Existing large-scale remote sensing image retrieval approaches are generally implemented by using hashing learning methods, which take handcrafted features as inputs and map the high-dimensional feature vector to the low-dimensional binary feature vector to reduce feature-searching complexity levels. As a means of applying the merits of deep learning, this paper proposes a novel large-scale remote sensing image retrieval approach based on deep hashing neural networks (DHNNs). More specifically, DHNNs are composed of deep feature learning neural networks and hashing learning neural networks and can be optimized in an end-to-end manner. Rather than requiring to dedicate expertise and effort to the design of feature descriptors, we can automatically learn good feature extraction operations and feature hashing mapping under the supervision of labeled samples. To broaden the application field, DHNNs are evaluated under two representative remote sensing cases: scarce and sufficient labeled samples. To make up for a lack of labeled samples, DHNNs can be trained via transfer learning for the former case. For the latter case, DHNNs can be trained via supervised learning from scratch with the aid of a vast number of labeled samples. Extensive experiments on one public remote sensing image data set with a limited number of labeled samples and on another public data set with plenty of labeled samples show that the proposed remote sensing image retrieval approach based on DHNNs can remarkably outperform state-of-the-art methods under both of the examined conditions. Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Hu Zhu, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |