EDBT 2026 Demo / reviewers in the wild / expert
Zhiquan He
dblp:21/1073
· DBLP profile ↗
36ranked-venue papers
9as first author
31since 2021 · last 2026
0000-0003-2255-4293ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hear to reveal: Stealing keystroke content from keyboard acoustic side-channel
Zhiquan He, Zhihai Yang, Zicheng Cui, Pinghui Wang, Zhiquan Liu |
Knowl. Based Syst. | 1 |
| 2026 | Attribute affinity coordination debiasing for generalized zero-shot learning
Zhiquan He |
Neural Networks | 2 |
| 2026 | Align-then-generate: An effective cross-modal generation paradigm for multi-label zero-shot learning
Peirong Ma, Wu Ran, Yanhui Gu, Huaqiu Chen, Zhiquan He, Hong Lu 0001 |
Pattern Recognit. | 5 |
| 2025 | PLNet: Entropy-Guided Pseudo Label Refinement with Sensitivity-Specificity Enhancement for Medical Image Segmentation
Huibin Weng, Zhiquan He, Wuzhen Shi |
CGI (1) | 5 |
| 2025 | Implicit Retinex Decomposition with Chromaticity Disentanglement for Low-Light Image Enhancement
Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma |
ACM Multimedia | 3 |
| 2025 | Enhancing 3D medical image registration with cross attention, residual skips, and cascade attentionabstractAt the core of Deep Learning-based Deformable Medical Image Registration (DMIR) lies a strong foundation. Essentially, this network compares features in two images to identify their mutual correspondence, which is necessary for precise image registration. In this paper, we use three novel techniques to increase the registration process and enhance the alignment accuracy between medical images. First, we propose cross attention over multi-layers of pairs of images, allowing us to take out the correspondences between them at different levels and improve registration accuracy. Second, we introduce a skip connection with residual blocks between the encoder and decoder, helping information flow and enhancing overall performance. Third, we propose the utilization of cascade attention with residual block skip connections, which enhances information flow and empowers feature representation. Experimental results on the OASIS data set and the LPBA40 data set show the effectiveness and superiority of our proposed mechanism. These novelties contribute to the enhancement of 3D DMIR-based on unsupervised learning with potential implications in clinical practice and research. Zhiquan He, Wenming Cao 0001 |
Intell. Data Anal. | 2 |
| 2025 | Defect image generation through feature disentanglement using StyleGAN2-ADA
Zhiquan He, Kangxing Wu |
Neurocomputing | 1 |
| 2025 | Low-Light Image Enhancement via Multi-Exposure Progressive Contrastive RegularizationabstractLow-light image enhancement (LLIE) aims to restore low-light images to their normal-light counterparts with optimal global illumination distribution and clear local details. With the advancement of deep learning, deep learning-based methods have become the mainstream in the LLIE community. However, most deep learning-based method cannot yet fully exploit the global and local contextual information in the low-light image. In this paper, we introduce a dual-branch module to simultaneously restore global and local features from spatial and frequency domain. To fuse these multi-level features, we propose a perception module to perform feature interaction between global and local features via cross attention and self-gating. By integrating the two developed modules into a U-Net backbone, we present a global-local interaction network for LLIE. Furthermore, recent studies have shown that contrastive learning can be an effective paradigm for the LLIE task. However, previous works typically use semantically-inconsistent under-/over-exposed images as negative samples. These images are very dissimilar to the ground-truth and cannot provide sufficient regularization in contrastive learning. To address this limitation, we explore a practical multi-exposure progressive contrastive regularization framework for LLIE. With a customized sample generation, sample selection, and progressive learning strategy, our proposed framework progressively narrows down the solution space around the optimum, and helps to improve the performance of LLIE methods without additional inference overhead. Combining the proposed network and contrastive regularization, our proposed method achieves favorable results compared to state-of-the-art LLIE methods on benchmark datasets. Extensive experiments further demonstrate the generalization ability of our proposed method. Zuojie Xie, Hao Ren 0002, Junjian Huang, Zhiquan He, Hong Lu 0001, Lvfan Yuan, Changyong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Unleashing the Potential of Hierarchical Region Clues for Open-Vocabulary Multi-Label ClassificationabstractOpen-vocabulary multi-label classification (OV-MLC) aims to leverage the rich multi-modal knowledge from Vision-language pre-training (VLP) models to further improve the recognition ability for unseen (novel) classes beyond the training set in multi-label scenarios. Existing OV-MLC methods only perform predictions on single hierarchical regions, and aggregate the prediction scores of these regions through simpletop-kmean pooling. This fails to unleash the potential of rich hierarchical region clues in multi-label images and does not fully exploit the discriminative information from all regions in the image, resulting in sub-optimal performance. In this work, we propose a novel OV-MLC framework to fully harness the power of multiple hierarchical region clues. Specifically, we first design a hierarchical clue gathering (HCG) module to gather different hierarchical clues, enabling more precise recognition of multiple object categories with different sizes in a multi-label image. Then, by viewing multi-label classification as single-label classification of each region within the image, we present a novel hierarchical score aggregation (HSA) approach, thereby better utilizing the predictions of each image region for each class. We also utilize a well-designed region selection strategy (RSS) to eliminate noise or background regions in an image that are irrelevant to classification, achieving higher multi-label classification accuracy. In addition, we propose a hybrid prompt learning (HPL) strategy to enhance visual-semantic consistency while preserving the generalization capability of label embeddings for unseen classes. Extensive experiments on public benchmark datasets demonstrate that our method significantly outperforms the current state-of-the-art. Peirong Ma, Wu Ran, Zhiquan He, Jian Pu, Hong Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multi-prior guided depth map super-resolution based on a diffusion model
Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Zhiquan He |
Vis. Comput. | 6 |
| 2024 | Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate RainsabstractRecent advances in image deraining have focused on training powerful models on mixed multiple datasets comprising diverse rain types and backgrounds. However, this approach tends to overlook the inherent differences among rainy images, leading to suboptimal results. To overcome this limitation, we focus on addressing various rainy images by delving into meaningful representations that encapsulate both the rain and background components. Leveraging these representations as instructive guidance, we put forth a Context-based Instance-level Modulation (CoI-M) mechanism adept at efficiently modulating CNN- or Transformer-based models. Furthermore, we devise a rain-/detail-aware contrastive learning strategy to help extract joint rain-/detail-aware representations. By integrating CoI-M with the rain-/detail-aware Contrastive learning, we develop [CoIC](https://github.com/Schizophreni/CoIC), an innovative and potent algorithm tailored for training models on mixed datasets. Moreover, CoIC offers insight into modeling relationships of datasets, quantitatively assessing the impact of rain and details on restoration, and unveiling distinct behaviors of models given diverse inputs. Extensive experiments validate the efficacy of CoIC in boosting the deraining ability of CNN and Transformer models. CoIC also enhances the deraining prowess remarkably when real-world dataset is included. Wu Ran, Peirong Ma, Zhiquan He, Hao Ren 0002, Hong Lu 0001 |
ICLR | 3 |
| 2024 | Rainmer: Learning Multi-view Representations for Comprehensive Image Deraining and BeyondabstractWe address image deraining under complex backgrounds, diverse rain scenarios, and varying illumination conditions, representing a highly practical and challenging problem. Our approach utilizes synthetic, real-world, and nighttime datasets, wherein rich backgrounds, multiple degradation types, and diverse illumination conditions coexist. The primary challenge in training models on these datasets arises from the discrepancies among them, potentially leading to conflicts or competition during the training period. To address this issue, we first align the distribution of synthetic, real-world and nighttime datasets. Then we propose a novel contrastive learning strategy to extract multi-view (multiple) representations that effectively capture image details, degradations, and illuminations, thereby facilitating training across all datasets. Regarding multiple representations as profitable prompts for deraining, we devise a prompting strategy to integrate them into the decoding process. This contributes to a potent deraining model, dubbed Rainmer. Additionally, a spatial-channel interaction module is introduced to fully exploit cues when extracting multi-view representations. Extensive experiments on synthetic, real-world, and nighttime datasets demonstrate that Rainmer outperforms current representative methods. Moreover, Rainmer achieves superior performance on the All-in-One image restoration dataset, underscoring its effectiveness. Furthermore, quantitative results reveal that Rainmer significantly improves object detection performance on both daytime and nighttime rainy datasets. These observations substantiate the potential of Rainmer for practical applications. Wu Ran, Peirong Ma, Zhiquan He, Hong Lu 0001 |
ACM Multimedia | 3 |
| 2024 | Representation modeling learning with multi-domain decoupling for unsupervised skeleton-based action recognition
Zhiquan He, Jiantu Lv, Shizhang Fang |
Neurocomputing | 1 |
| 2024 | ragBERT: Relationship-aligned and grammar-wise BERT model for image captioning
Hengyou Wang, Kani Song, Xiang Jiang 0008, Zhiquan He |
Image Vis. Comput. | 4 |
| 2024 | Low-rank matrix recovery with total generalized variation for defending adversarial examplesabstractLow-rank matrix decomposition with first-order total variation (TV) regularization exhibits excellent performance in exploration of image structure. Taking advantage of its excellent performance in image denoising, we apply it to improve the robustness of deep neural networks. However, although TV regularization can improve the robustness of the model, it reduces the accuracy of normal samples due to its over-smoothing. In our work, we develop a new low-rank matrix recovery model, called LRTGV, which incorporates total generalized variation (TGV) regularization into the reweighted low-rank matrix recovery model. In the proposed model, TGV is used to better reconstruct texture information without over-smoothing. The reweighted nuclear norm and L 1 -norm can enhance the global structure information. Thus, the proposed LRTGV can destroy the structure of adversarial noise while re-enhancing the global structure and local texture of the image. To solve the challenging optimal model issue, we propose an algorithm based on the alternating direction method of multipliers. Experimental results show that the proposed algorithm has a certain defense capability against black-box attacks, and outperforms state-of-the-art low-rank matrix recovery methods in image restoration. Wen Li 0015, Hengyou Wang, Qiang He 0003, Zhiquan He, Wing W. Y. Ng |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2024 | Gradient multi-foci networks for 3D skeleton-based human motion prediction
Junyu Shi, Jianqi Zhong, Zhiquan He, Wenming Cao 0001 |
Neural Comput. Appl. | 3 |
| 2024 | PANet: Pluralistic Attention Network for Few-Shot Image ClassificationabstractAbstract Traditional deep learning methods require a large amount of labeled data for model training, which is laborious and costly in real word. Few-shot learning (FSL) aims to recognize novel classes with only a small number of labeled samples to address these challenges. We focus on metric-based few-shot learning with improvements in both feature extraction and metric method. In our work, we propose the Pluralistic Attention Network (PANet), a novel attention-oriented framework, involving both a local encoded intra-attention(LEIA) module and a global encoded reciprocal attention(GERA) module. The LEIA is designed to capture comprehensive local feature dependencies within every single sample. The GERA concentrates on the correlation between two samples and learns the discriminability of representations obtained from the LEIA. The two modules are complementary to each other and ensure the feature information within and between images can be fully utilized. Furthermore, we also design a dual-centralization (DC) cosine similarity to eliminate the disparity of data distribution in different dimensions and enhance the metric accuracy between support and query samples. Our method is thoroughly evaluated with extensive experiments, and the results demonstrate that with the contribution of each component, our model can achieve high-performance on four widely used few-shot classification benchmarks of miniImageNet, tieredImageNet, CUB-200-2011 and CIFAR-FS. Wenming Cao 0001, Tianyuan Li, Qifan Liu, Zhiquan He |
Neural Process. Lett. | 4 |
| 2024 | Low-Light Image Enhancement With Multi-Scale Attention and Frequency-Domain OptimizationabstractLow-light image enhancement aims to improve the perceptual quality of images captured in conditions of insufficient illumination. However, such images are often characterized by low visibility and noise, making the task challenging. Recently, significant progress has been made using deep learning-based approaches. Nonetheless, existing methods encounter difficulties in balancing global and local illumination enhancement and may fail to suppress noise in complex lighting conditions. To address these issues, we first propose a multi-scale illumination adjustment network to balance both global illumination and local contrast. Furthermore, to effectively suppress noise potentially amplified by the illumination adjustment, we introduce a wavelet-based attention network that efficiently perceives and removes noise in the frequency domain. We additionally incorporate a discrete wavelet transform loss to supervise the training process. Particularly, the proposed wavelet-based attention network has been shown to enhance the performance of existing low-light image enhancement methods. This observation indicates that the proposed wavelet-based attention network can be flexibly adapted to current approaches to yield superior enhancement results. Furthermore, extensive experiments conducted on benchmark datasets and downstream object detection task demonstrate that our proposed method achieves state-of-the-art performance and generalization ability. Zhiquan He, Wu Ran, Kehua Li, Chang-Yong Xie, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | A Transferable Generative Framework for Multi-Label Zero-Shot LearningabstractMulti-label zero-shot learning (MLZSL) is a more realistic and challenging task than single-label zero-shot learning (SLZSL), which aims to recognize multiple unseen classes in a single image. To adapt generative models to the MLZSL task and better recognize multiple unseen object categories in an image, this paper proposes a Transferable Generative Framework (TGF), which consists of a Multi-Label Semantic Embedding Autoencoders (SEAs), a Semantic-Related Multi-Label Feature Transformation Network (FTN) and a Multi-Label Feature Generation Networks (FGNs). First, SEAs adaptively encodes the class-level word vectors corresponding to each sample containing different number of classes into sample-level semantic embeddings with the same dimension. Then, FTN transforms global features extracted by a CNN pre-trained on single-label images into features that are semantic-related and more suitable for multi-label classification. Finally, FGNs generates both global and local features to better recognize the dominant and minor object categories in a multi-label image, respectively. Extensive experiments on three benchmark datasets show that TGF significantly outperforms state-of-the-arts. Specifically, compared with the previous best generative MLZSL method (i.e., Gen-MLZSL), TGF improves the mAP of the ZSL (GZSL) task by 5.4% (6.9%), 20.5% (27.9%), and 2.4% (3.9%) on NUS-WIDE, Open Images, and MS-COCO datasets, respectively. Peirong Ma, Zhiquan He, Wu Ran, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MSCAP: DNA Methylation Age Predictor based on Multiscale Convolutional Neural NetworkabstractDNA methylation can reflect age-related issues in individuals, and age prediction tools developed based on DNA methylation are called Epigenetic clocks. Epigenetic clocks also serve as potent tools for enhancing researchers’ understanding of the aging process and aim to contribute to the advancement of aging research. The current prevailing approach involves creating clocks using Elastic Net regression modeling, wherein high-age-correlation CpG sites were selected as features in advance. However, pre-selecting CpG sites in advance may inadvertently overlook interactions between distant CpG sites. Deep learning has significant advantages in dealing with vast feature sets due to the tremendous opportunities that the development of deep learning presents for research. Deep learning architectures have demonstrated greater efficiency in acquiring and representing large numbers of intricate features compared to traditional machine learning models. In this study, we proposed a deep learning model based on a multiscale convolutional neural network (MSCNN) for pan-tissue age prediction, called MSCAP. First, for the model to efficiently predict datasets from different sequencing platforms, we chose a common intersection of sites from the Illumina 27K and Illumina 450K platforms (25,789 CpG sites), then trained and tested the model on 93 datasets. Different tissues of the human body are used in the datasets for training and testing. And independently tested on 7 independent datasets against an Epigenetic clock developed by Horvath in 2013 with high prediction accuracy. The results showed that MSCAP could effectively extract key features from a vast number of potential features without advanced feature selection. In addition, MSCAP outperformed the Horvath clock on all seven independent test datasets. This research helped boost the development of epigenetic clock methods and reinforced the value of deep learning in computational biology. Han Wang 0028, Ruirui Cai, Xizeng Zong, Zhiquan He |
BIBM | 4 |
| 2023 | Predicting Protein-Ligand Binding Affinity with Multi-Scale Structural FeaturesabstractPredicting protein-ligand binding affinity is important in areas such as drug discovery, gene regulation and signal transduction. The DTA(Drug-Target Affinity) method based on protein structure can not only effectively compensates for the lack of binding information, but also more in line with real biological processes. Although the structure-based DTA methods have achieved good performance, the existing methods still have the problem of only considering single-scale structural features and ignoring multi-scale structural features. In order to solve this problem, we propose the MSSDTA (Multi-Scale Structural Representation Drug-Target Affinity Prediction), which extracts multi-scale protein features by integrating the surface node features and structural node features of proteins. At the same time, the drug representation network is used to fuse the 2D molecular structure characteristics and chemical characteristics of the drug to effectively distinguish the drug molecules with similar planar structures. Finally, the affinity prediction network is used to generate protein-ligand binding affinity scores. We verify the performance of this model on the PDBbind v.2019 dataset. The experimental results show that the proposed method achieves excellent performance. Han Wang 0028, Jingtong Zhao, Shengkun Wang, Zhiquan He, Xike Ouyang |
BIBM | 4 |
| 2023 | Deformable image registration with attention-guided fusion of multi-scale deformation fieldsabstractAbstract Deformable medical image registration plays a crucial role in theoretical research and clinical application. Traditional methods suffer from low registration accuracy and efficiency. Recent deep learning-based methods have made significant progresses, especially those weakly supervised by anatomical segmentations. However, the performance still needs further improvement, especially for images with large deformations. This work proposes a novel deformable image registration method based on an attention-guided fusion of multi-scale deformation fields. Specifically, we adopt a separately trained segmentation network to segment the regions of interest to remove the interference from the uninterested areas. Then, we construct a novel dense registration network to predict the deformation fields of multiple scales and combine them for final registration through an attention-weighted field fusion process. The proposed contour loss and image structural similarity index (SSIM) based loss further enhance the model training through regularization. Compared to the state-of-the-art methods on three benchmark datasets, our method has achieved significant performance improvement in terms of the average Dice similarity score (DSC), Hausdorff distance (HD), Average symmetric surface distance (ASSD), and Jacobian coefficient (JAC). For example, the improvements on the SHEN dataset are 0.014, 5.134, 0.559, and 359.936, respectively. Zhiquan He, Yupeng He, Wenming Cao 0001 |
Appl. Intell. | 1 |
| 2023 | Multi-layer noise reshaping and perceptual optimization for effective adversarial attack of imagesabstractAbstract Adversarial attack aims to fail the deep neural network by adding a small amount of perturbation to the input image, in which the attack success rate and resulting image quality are maximized under the lp norm perturbation constraint. However, the lp norm is not accurately correlated to human perception of image quality. Attack methods based on l0 norm constraint usually suffer from the high computational cost due to the iterative search for candidate pixels to modify. In this work, we explore how perceptual quality optimization can be incorporated into the adversarial attack design and propose a two-stage attack method to reshape the adversarial noise by an initial attack and optimize the visual quality of the attacked images without sacrificing the attack success rate. Specifically, we construct a visual attention network to generate a perceptual attention map to modulate the adversarial noise generated by a base attack method. The network is trained to maximize the visual quality in Structural Similarity Index Metric (SSIM) while achieving the same attack success rate. To improve the image perceptual quality further, we propose a fast search algorithm to perform an iterative block-wise pruning of the adversarial noise. We evaluate our method on the mini-ImageNet dataset against three different defense schemes. The results have demonstrated that our method can achieve better attack performance in image quality, attack success rate, and efficiency than the state-of-the-art attack methods. Zhiquan He, Xujia Lan, Jianhe Yuan, Wenming Cao 0001 |
Appl. Intell. | 1 |
| 2023 | A novel sample and feature dependent ensemble approach for Parkinson's disease detectionabstractAbstract Parkinson’s disease (PD) is a neurological disease that has been reported to have affected most people worldwide. Recent research pointed out that about 90% of PD patients possess voice disorders. Motivated by this fact, many researchers proposed methods based on multiple types of speech data for PD prediction. However, these methods either face the problem of low rate of accuracy or lack generalization. To develop an approach that will be free of these issues, in this paper we propose a novel ensemble approach. These paper contributions are two folds. First, investigating feature selection integration with deep neural network (DNN) and validating its effectiveness by comparing its performance with conventional DNN and other similar integrated systems. Second, development of a novel ensemble model namely EOFSC (Ensemble model with Optimal Features and Sample Dependant Base Classifiers) that exploits the findings of recently published studies. Recent research pointed out that for different types of voice data, different optimal models are obtained which are sensitive to different types of samples and subsets of features. In this paper, we further consolidate the findings by utilizing the proposed integrated system and propose the development of EOFSC. For multiple types of vowel phonations, multiple base classifiers are obtained which are sensitive to different subsets of features. These features and sample-dependent base classifiers are integrated, and the proposed EOFSC model is constructed. To evaluate the final prediction of the EOFSC model, the majority voting methodology is adopted. Experimental results point out that feature selection integration with neural networks improves the performance of conventional neural networks. Additionally, feature selection integration with DNN outperforms feature selection integration with conventional machine learning models. Finally, the newly developed ensemble model is observed to improve PD detection accuracy by 6.5%. Chinmay Chakraborty, Zhiquan He, Wenming Cao 0001, Yakubu Imrana, Joel J. P. C. Rodrigues |
Neural Comput. Appl. | 3 |
| 2023 | Correction to: Multi-level context-driven interaction modeling for human future trajectory prediction
Zhiquan He, Hao Sun 0024, Wenming Cao 0001, Henry Z. He |
Neural Comput. Appl. | 1 |
| 2023 | Contrastive Bayesian Analysis for Deep Metric LearningabstractRecent methods for deep metric learning have been focusing on designing different contrastive loss functions between positive and negative pairs of samples so that the learned feature embedding is able to pull positive samples of the same class closer and push negative samples from different classes away from each other. In this work, we recognize that there is a significant semantic gap between features at the intermediate feature layer and class labels at the final output layer. To bridge this gap, we develop a contrastive Bayesian analysis to characterize and model the posterior probabilities of image labels conditioned by their features similarity in a contrastive learning setting. This contrastive Bayesian analysis leads to a new loss function for deep metric learning. To improve the generalization capability of the proposed method onto new classes, we further extend the contrastive Bayesian loss with a metric variance constraint. Our experimental results and ablation studies demonstrate that the proposed contrastive Bayesian metric learning method significantly improves the performance of deep metric learning in both supervised and pseudo-supervised scenarios, outperforming existing methods by a large margin. Shichao Kan, Zhiquan He, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Implicit user relationships across sessions enhanced graph for session-based recommendationabstractSession-based recommendation aims to predict users’ next preference based on the sequence of their own history preferences in a short period. Most state-of-the-art methods model the session as a graph using graph neural networks (GNN) to capture the dynamic transitions between items within sessions. However, the complex and hidden correlations between different sessions are not adequately addressed, especially during the testing stage. We argue that session-based recommendation tasks can be improved by exploiting the correlations between different sessions in both training and testing. To this end, we propose a novel three-GNN-based recommendation framework to exploit the intra- and inter-session item correlations and the session-session correlations. The first one is a graph-based multi-layer perceptron to learn the inter-session item representations in a contrastive learning scheme guided by a contrastive loss function. The second one is a multi-relation graph attention network for intra-session item representations. The two item embeddings are combined through a position attention scheme to form the session representation, which is modulated and enhanced by the third extra session GNN by capturing the session-session correlations. The three levels of correlations are used in the training and testing stages in a joint prediction manner. To alleviate the data sparsity issue faced by the session GNN, we expand the session items by incorporating the neighboring items in the global item graph built from the entire training sessions. We have evaluated our method on multiple benchmark datasets. The results have shown that our joint recommendation method based on session correlation has significantly improved the recommendation accuracy over the state-of-the-art by more than 10%. Wenming Cao 0001, Yishan Liu, Guitao Cao, Zhiquan He |
Inf. Sci. | 4 |
| 2022 | Multi-level context-driven interaction modeling for human future trajectory predictionabstractAbstract Human trajectory prediction is a challenging task with important applications such as intelligent surveillance and autonomous driving. We recognize that pedestrians in close and distant neighborhoods have different impacts on the person’s decision of future movements. Local scene context and global scene layout also affect the movement decision differently. Existing methods have not adequately addressed these interactions between humans and the multi-level contexts occurring at different spatial and temporal scales. To this end, we propose a multi-level context-driven interaction modeling (MCDIM) method for human future trajectory learning and prediction. Specifically, we construct a multilayer graph attention network (GAT) to model the hierarchical human–human interactions. An extra set of long short-term memory networks is designed to capture the correlations of these human–human interactions at different temporal scales. To model the human–scene interactions, we explicitly extract and encode the global scene layout features and local context features in the neighborhood of the person at each time step and capture the spatial–temporal information of the interactions between human and the local scene contexts. The human–human and human–scene interactions are incorporated into the multi-level GAT-based network for accurate prediction of future trajectories. We have evaluated the method on benchmark datasets: the walking pedestrians dataset provided by ETH Zurich (ETH) and the crowd data provided by the University of Cyprus. The results demonstrate that our MCDIM method outperforms existing methods, being able to generate more accurate and plausible trajectories for pedestrians. The average performance gain is 2 and 3 percentage points in terms of the average displacement error and final displacement error, respectively. Zhiquan He, Hao Sun 0024, Wenming Cao 0001, Henry Z. He |
Neural Comput. Appl. | 1 |
| 2021 | Progressive anatomically constrained deep neural network for 3D deformable medical image registrationabstractThe 3D deformable image registration is one of the most challenging tasks in medical image analysis. Due to the large and complex deformation in 3D medical images, many deep neural network based methods have been proposed to improve the image similarity after registration, among which recursive cascading network structure is one of the state-of-the-art. However, most existing works rely on the pixel-level image similarities to achieve anatomical rationality and overlook the global-level resemblance between the two structures. Therefore, the resulting registration is not quite clinically valuable. To this end, in this work, we propose a Progressive Anatomically Constrained deep neural Network (PACN) to incorporate the anatomical priors into a progressive cascading registration network to improve the anatomical plausibility as well as the pixel-level similarity of the registration results. Specifically, an Anatomical Constraint Encoder (ACE) network is proposed to encode the global context of the anatomical segmentations and attached to the dense registration network to form a registration unit. Repeated such units forming a cascading framework progressively warps the moving image toward the fixed one, with the output warped image of one unit as the input of the next unit. In this design, the global anatomical priors along with the pixel-level local information are used to guide the model learning process to produce high quality deformation field. Based on this, we explore two frameworks to investigate their registration effectiveness, one attaches the anatomical constraint encoder (ACE) to every dense registration sub-network and the other one attaches ACE only to the last dense registration unit. We test the two frameworks on benchmarks of three liver image datasets SLIVER, LiTS and LSPIG, and one brain dataset LPBA. Our two frameworks have achieved significantly better results in terms of average Dice score than the state-of-the-art baseline method on three liver datasets and comparable on LPBA when both tested with up to three cascades. Wenming Cao 0001, Zhiquan He |
Neurocomputing | 3 |
| 2021 | A cascaded registration network RCINet with segmentation mask
Wenlan Zou, Wenming Cao 0001, Zhiquan He, Zhihai He |
Neural Comput. Appl. | 4 |
| 2021 | Mask Cross-Modal Hashing NetworksabstractDue to the rapid development of deep learning, cross-modal retrieval has achieved significant progress in recent years. Moreover, cross-modal hashing has recently attracted considerable attention to multi-modal retrieval applications due to its advantages of low storage costs and fast retrieval speed. However, it is still a challenging problem due to an existing semantic heterogeneity gap between different modalities. In order to further narrow the gap and obtain more effective hash codes, we put forward a novel mask deep cross-modal hashing (MDCH) approach to explore the similarity between inter-modal instances. The main contributions of this paper are that: (1) we attempt to introduce semantic mask information into cross-modal hashing retrieval, (2) we alternately train intra-modal and inter-modal networks to fully mine the semantic relationship between different modalities. The semantic mask can improve the semantic information of the image feature. While inter-modal similarity, explored by inter-modal networks, focuses on enforcing images and their corresponding text tags to have similar hash codes, intra-modal similarity, explored by intra-modal networks, can retain local structural information embedded in each modality to achieve internal similarity. A large number of experiments conducted on three datasets demonstrate that our proposed MDCH approach is superior to several state-of-the-art cross-modal hashing approaches. Qiubin Lin, Wenming Cao 0001, Zhiquan He, Zhihai He |
IEEE Trans. Multim. | 3 |
| 2020 | Semantic deep cross-modal hashing
Qiubin Lin, Wenming Cao 0001, Zhihai He, Zhiquan He |
Neurocomputing | 4 |
| 2019 | Hybrid representation learning for cross-modal retrieval
Wenming Cao 0001, Qiubin Lin, Zhihai He, Zhiquan He |
Neurocomputing | 4 |
| 2018 | Reweighted Low-Rank Matrix Analysis With Structural Smoothness for Image DenoisingabstractIn this paper, we develop a new low-rank matrix recovery algorithm for image denoising. We incorporate the total variation (TV) norm and the pixel range constraint into the existing reweighted low-rank matrix analysis to achieve structural smoothness and to significantly improve quality in the recovered image. Our proposed mathematical formulation of the low-rank matrix recovery problem combines the nuclear norm, TV norm, and norm, thereby allowing us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Using the iterative alternating direction and fast gradient projection methods, we develop an algorithm to solve the proposed challenging non-convex optimization problem. We conduct extensive performance evaluations on single-image denoising, hyper-spectral image denoising, and video background modeling from corrupted images. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for large random noise. For example, when the density of random sparse noise is 30%, for single-image denoising, our proposed method is able to improve the quality of the restored image by up to 4.21 dB over existing methods. Hengyou Wang, Yi-Gang Cen, Zhiquan He, Zhihai He, Ruizhen Zhao, Fengzhen Zhang |
IEEE Trans. Image Process. | 3 |
| 2018 | Knowledge-Guided Deep Fractal Neural Networks for Human Pose EstimationabstractHuman pose estimation using deep neural networks aims to map input images with large variations into multiple body keypoints, which must satisfy a set of geometric constraints and interdependence imposed by the human body model. This is a very challenging nonlinear manifold learning process in a very high dimensional feature space. We believe that the deep neural network, which is inherently an algebraic computation system, is not the most efficient way to capture highly sophisticated human knowledge, for example those highly coupled geometric characteristics and interdependence between keypoints in human poses. In this work, we propose to explore how external knowledge can be effectively represented and injected into the deep neural networks to guide its training process using learned projections that impose proper prior. Specifically, we use the stacked hourglass design and inception-resnet module to construct a fractal network to regress human pose images into heatmaps with no explicit graphical modeling. We encode external knowledge with visual features, which are able to characterize the constraints of human body models and evaluate the fitness of intermediate network output. We then inject these external features into the neural network using a projection matrix learned using an auxiliary cost function. The effectiveness of the proposed inception-resnet module and the benefit in guided learning with knowledge projection is evaluated on two widely used human pose estimation benchmarks. Our approach achieves state-of-the-art performance on both datasets. Guanghan Ning, Zhi Zhang 0005, Zhiquan He |
IEEE Trans. Multim. | 3 |
| 2006 | Rate-Distortion Optimized Transmission Power Adaptation for Video Streaming over Wireless ChannelsabstractAn important characteristic of video transmission over wireless channels is that the channel is error-prone and the compressed video data is highly sensitive to errors. The transmission errors will cause decoding failure which will distort the reconstructed pictures. One of the challenging issues in wireless video communication system design is to minimize the overall energy consumption and maximize the operational lifetime of the devices. In this work, we focus on the energy consumption in wireless video transmission. Based on a transmission distortion model, we develop a rate-distortion optimized transmission power adaptation scheme for video streaming over wireless channels. Our experimental results demonstrate that this scheme is able to significantly improve the video quality under the transmission power constraints. Zhiquan He, Wenjun Zeng 0001 |
ICIP | 1 |