EDBT 2026 Demo / reviewers in the wild / expert
Zhiheng Zhou 0001
dblp:45/3558-1 · also Zhi-Heng Zhou 0001
· DBLP profile ↗
56ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0003-4040-0175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 14 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffinformer: Diffusion informer model for long sequence time-series forecasting
Wei Chen 0165, Yican Liu, Junmei Yang, Zhiheng Zhou 0001, Delu Zeng |
Expert Syst. Appl. | 5 |
| 2026 | SDCos_ACLS: Active contours with local blocks similarity guided by co-saliency maps based on sparse decomposition
Zhiheng Zhou 0001, Guoqi Liu |
Signal Process. | 2 |
| 2026 | Revisiting Semantic Gap in U-Shaped Networks From a Performance Oriented Perspective: Quantification, Visualization, and OptimizationabstractAutomatic medical image segmentation is a fundamental component of computer-aided diagnosis. U-shaped networks (U-Nets) remain the most widely adopted architecture due to their suitability for the unique challenges of medical imaging. However, recent studies have shown that U-Nets fuse low-level visual features from the encoder with high-level semantic features from the decoder using direct skip connections (DSC), which are likely to degrade segmentation performance due to the semantic gap, thereby limiting their ability to meet high-precision clinical requirements. This paper revisits the semantic gap from a performance-oriented perspective and conceptualizes it as a learnable task. A key characteristic of this semantic gap is revealed through comprehensive quantification and visualization, demonstrating its significant negative impact on segmentation performance. Further analysis indicates that a contributing factor is the substantial channel noise present in low-level pixel features, which is transmitted to the decoder via DSC, thereby disrupting the modeling of high-level semantic representations. In response, this paper proposes a self-disambiguating skip connection (SDSC), which incorporates a self-guided filter leveraging spatial features for channel filtering, a multi-layer fusion Transformer to capture long-range contextual dependencies, and Jensen-Shannon divergence as a constraint to enhance learning. The proposed method, referred to as SDSC-UNet, is evaluated through extensive experiments on four challenging benchmarks. The results demonstrate that replacing DSC with our SDSC yields an improvement of 5.91% in mean Intersection over Union (mIoU), 4.58% in Dice coefficient, and 1.32 in Hausdorff Distance, achieving state-of-the-art performance and highlighting the effectiveness of SDSC. Xiaoshan Xie, Zhiheng Zhou 0001, Chang Niu, Zhelin Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Cross-View and Multi-Step Interaction for Change CaptioningabstractChange captioning is a task that describes changes in image pairs using natural language. This task is more complex than single-image captioning as it requires a comprehensive understanding of each image and the ability to recognize and describe the semantic changes in image pairs. The key challenge lies in making the network generate an accurate and stable change representation under the interference of viewpoint shift. In this paper, we propose a cross-view and multi-step interaction network to generate robust change representation to resist pseudo-change. Specifically, in the intra-image representation learning stage, a cross-view interaction encoder is designed to enhance internal relationships by cross-referencing in image pairs. In the change feature learning stage, a multi-step change perceptron is employed to capture the change semantics from coarse to fine progressively. Then, a fusion module dynamically combines them as a fine-grained change representation. Besides, we propose a backward representation reconstruction module that facilitates the capture of semantic changes, thus improving the quality of captions in a self-supervised manner. Extensive experiments have shown that the method effectively captures real semantic changes under the interference of viewpoint shift and achieves state-of-the-art performance on five public datasets. The code is available at https://github.com/TTXiann/CVMSI Tiantao Xian, Zhiheng Zhou 0001, Wenlve Zhou, Delu Zeng, Bo Li 0111 |
IEEE Trans. Multim. | 2 |
| 2025 | Spatial-angular features based no-reference light field quality assessment
Zerui Yu, Zhiheng Zhou 0001, Xiyuan Tao |
Expert Syst. Appl. | 3 |
| 2025 | BiPC: Bidirectional Probability Calibration for Unsupervised Domain Adaption
Wenlve Zhou, Zhiheng Zhou 0001, Junyuan Shang, Chang Niu, Xiyuan Tao, Tianlei Wang |
Expert Syst. Appl. | 2 |
| 2025 | Noisy image segmentation utilizing entropy-adaptive fractional differential-driven active contours
Shang Zhuge, Zhiheng Zhou 0001, Wenlue Zhou, Jiangfeng Wu |
Multim. Tools Appl. | 2 |
| 2025 | Refining visual token sequence for efficient image captioning
Tiantao Xian, Zhiheng Zhou 0001, Wenlve Zhou |
Neural Networks | 2 |
| 2025 | Dynamic and Asymmetric Enhancement for Remote Sensing Image Change CaptioningabstractRemote Sensing Image Change Captioning (RSICC) plays a critical role in automated environmental monitoring by generating natural language descriptions that provide intuitive interpretations of changes between bi-temporal remote sensing images. Despite recent advancements, existing methods suffer from two fundamental limitations: (1) static encoding architectures fail to account for scene complexity and semantic diversity during feature extraction, leading to inflexible representation learning; and (2) symmetric computational structures are inherently unsuitable for modeling the asymmetric temporal dependencies inherent in “before-to-after” image pairs. To address these challenges, we propose a Dynamic Asymmetric Encoder (DAE), which introduces two key innovations. First, we design a difference-guided dynamic convolution module that adaptively adjusts convolutional parameters using input-driven scaling factors and offsets, thereby enabling scene-aware intra-image feature enhancement. Second, we develop a Multi-expert Temporal Interaction (METI) module that establishes an asymmetric computational topology: the “after” image branch actively perceives change information relative to the “before” image through three heterogeneous interaction experts, followed by feature fusion. This design allocates greater computational capacity to the “after” image branch while preserving temporal coherence. Furthermore, we introduce a Multi-level Feature Aggregator (MFA) that enhances the representation of salient changed regions across multiple scales via an iterative reinforcement mechanism. Experimental results validate the effectiveness of the proposed method, demonstrating state-of-the-art performance on two benchmark RSICC datasets. We publicly release our code repository at https://github.com/TTXiann/Dynamic-Asymmetric to facilitate future research. Tiantao Xian, Zhiheng Zhou 0001, Delu Zeng, Bo Li 0111 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Multimodal Feature Fusion Network With Text Difference Enhancement for Remote Sensing Change DetectionabstractAlthough deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization—especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision-language model (VLM) to generate semantic descriptions of bi-temporal images. A Textual Difference Enhancement (TDE) module then captures fine-grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image-Text Feature Fusion (ITFF) module that enables deep cross-modal integration. Extensive experiments on LEVIR-CD, WHU-CD, and SYSU-CD demonstrate that MMChange consistently surpasses state-of-the-art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange. Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou 0001, Xudong Jia 0001, Hongsheng Zhang 0001, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Balanced Active Sampling for Person Re-identificationabstractActive learning is attracting more and more attention in person re-identification (Re-ID), as it is promising in the scalability of Re-ID models to satisfy performance with reduced labeling cost. Active sampling of pair-wise images in Re-ID is a highly imbalanced problem, where negative pairs are the vast majority. To avoid sampled pairs being dominated by the negative relationship, previous works tend to sample pairs with confident positive relationships in various ways. However, it is a waste of the labeling budget as most sampled pairs will be positive and already have a very close distance. Thus, there is no significant improvement in the model performance. In this paper, we first argue that balanced sampling is the key to active learning for Re-ID. Along this line, we propose a naïve balanced sampling method based on the global estimation of the most confusing distance. It is further improved by the label-wise estimation and diversity measurement. We also formulate the training of Re-ID models as a constrained clustering problem, where labeled positive and negative pairs are as must-link and cannot-link. Then the model training is based on the pseudo labels. Extensive experiments on benchmarks evaluate the effectiveness and superiority of the proposed methods. Specifically, it achieves comparable performance with supervised counterparts with less than 0.1% pair-wise annotation, which significantly surpasses the state-of-the-art. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 4 |
| 2024 | Camera Bias Regularization for Person Re-identificationabstractPerson re-identification (Re-ID) is to match persons captured by non-overlapping cameras. Due to the discrepancies between cameras caused by illumination, background, or viewpoint, the underlying difficulty for Re-ID is the camera bias problem, which leads to the large gap of within-identity features from different cameras. With limited cross-camera annotation, Re-ID models tend to learn camera-related features, instead of identity-related features. Consequently, Re-ID models suffer from poor transfer ability from seen to unseen domains. In this paper, we investigate the camera bias problem in both supervised and unsupervised learning. In particular, we propose a novel Camera Bias Regularization (CBR) term to reduce the feature distribution gap between cameras. The CBR works by simultaneously enlarging the distance of intra-camera distributions between positive and negative pairs, and reducing the distance of positive pairs’ distributions between intra-camera and cross-camera. In addition, a Cross-Camera (CC) clustering method is also designed for unsupervised learning, which puts more emphasis on cross-camera pairs than intra-camera ones during the clustering process. Extensive experiments are conducted to validate the effectiveness of the proposed CBR and CC. Specifically, with only a plain ResNet-50, it achieves 56.7% mAP and 40.7% mAP on the challenging MSMT17 dataset in supervised and unsupervised settings respectively, which surpasses most state-of-the-arts. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 4 |
| 2024 | Superclass-aware visual feature disentangling for generalized zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang |
Expert Syst. Appl. | 3 |
| 2024 | EARNet: Error-Aware Reconstruction Network for no-reference image quality assessment
Zhiheng Zhou 0001, Zenan Zhou, Xiyuan Tao, Zerui Yu, Yinglie Cao |
Expert Syst. Appl. | 1 |
| 2024 | Consistent representation joint adaptive adjustment for incremental zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang |
Neurocomputing | 3 |
| 2024 | Neural Ordinary Differential Equation Networks for Fintech Applications Using Internet of ThingsabstractThe Internet-of-Things (IoT) technology is becoming increasingly pivotal in the financial services sector, with a growing number of algorithms being employed in high-frequency trading. High-frequency prediction in financial time series prediction presents a promising avenue of research. From convolutional neural networks to recurrent neural networks, deep learning have demonstrated exceptional capabilities in capturing the nonlinear characteristics of stock markets, thereby achieving high performance in stock index prediction. In this paper, we employ ODE-LSTM model for high-frequency price forecasting, predicting stock price data across various time scales, including 1-minute, 5-minutes, and 30-minutes frequencies. This approach introduces a novel concept, wherein the LSTM (Long Short-Term Memory) model is integrated with Neural ODE (Ordinary Differential Equations) to manage the hidden state and augment model interpretability. Over the course of 7 months, we achieved a 41.79% excess return on a simulated trading platform, with a daily average excess return of 0.30%, showcasing the commendable performance of our model and strategy. Wei Chen 0165, Yican Liu, Junmei Yang, Delu Zeng, Zhiheng Zhou 0001 |
IEEE Internet Things J. | 6 |
| 2024 | Attributes-Assisted Joint Contrastive Learning for Person Re-IdentificationabstractPerson re-identification (Re-ID) is a crucial technology for intelligent security in Internet of Things (IoT) systems. Recently, unsupervised learning has been widely used for person Re-ID due to its generalization property. However, the effectiveness of commonly used unsupervised clustering methods heavily relies on the quality of the clustered pseudo-labels. Moreover, pedestrian shots in real scenes are prone to factors such as occlusion. In this paper, we propose a novel Global and Local Joint Contrastive Learning (GLCL) framework based on the memory bank. Specifically, we establish separate memory banks for global and local features, which are updated using global simple samples and local hard samples. The GLCL module helps excavate information from simple and hard samples, aiming to overcome the effects of poor retrieval scenarios such as background clutter and occlusion. Additionally, we design an Attributes-Assisted Clustering (AAC)) module that utilizes pedestrian attributes to refine the clustering results. The AAC module can effectively reduce the impact of pseudo-label noise owing to the supplementary information offered by attributes. Our approach shows improved performance in person Re-ID tasks in complex scenarios, providing a promising solution for intelligent security systems in the IoT. Experimental results demonstrate the superiority of our proposed method. Qingru Wu, Zhiheng Zhou 0001, Chang Niu, Xiaosheng Liu, Bo Li 0111 |
IEEE Internet Things J. | 2 |
| 2024 | Generalized zero-shot action recognition through reservation-based gate and semantic-enhanced contrastive learning
Junyuan Shang, Chang Niu, Xiyuan Tao, Zhiheng Zhou 0001, Junmei Yang |
Knowl. Based Syst. | 4 |
| 2024 | Cross-modal domain generalization semantic segmentation based on fusion featuresabstractThe primary techniques for domain generalization in semantic segmentation revolve around domain randomization and feature whitening. Although less commonly employed, methods based on cross-modality have demonstrated effective outcomes. This paper introduces enhancements to cross-modal feature alignment by redesigning the feature alignment module. This redesign facilitates alignment across different modalities by leveraging fusion features derived from both visual and textual inputs. These fusion features provide a more effective anchor point for alignment, enhancing the transfer of semantic information from textual to visual domains. Furthermore, the decoder plays a crucial role in the model as its ability to categorize features directly impacts the segmentation performance of the entire model. To enhance the decoder’s capability, this study employs the fusion features as the input for the decoder, with image labels providing the supervision. Experimental results indicate that our approach significantly enhances the model’s generalization capabilities. Wanlin Yue, Zhiheng Zhou 0001, Yinglie Cao, Liuman |
Knowl. Based Syst. | 2 |
| 2024 | Fast CU patition based on image similarity using neural network
Yinglie Cao, Wenjin Wu, Zhiheng Zhou 0001, Haoqi Xu, Wanlin Yue, Shang Zhuge |
Multim. Tools Appl. | 3 |
| 2024 | Super-resolution reconstructed video coding scheme based on inter-frame information
Yinglie Cao, Haoqi Xu, Zhiheng Zhou 0001, Wanlin Yue, Shang Zhuge |
Multim. Tools Appl. | 3 |
| 2024 | An adaptive multi-level-sets active contour model based on block search
Zhiheng Zhou 0001, Guoqi Liu, Tianlei Wang |
Multim. Tools Appl. | 1 |
| 2024 | DeepAR-Attention probabilistic prediction for stock price series
Wei Chen 0165, Zhiheng Zhou 0001, Junmei Yang, Delu Zeng |
Neural Comput. Appl. | 3 |
| 2024 | Unsupervised Domain Adaption Harnessing Vision-Language Pre-TrainingabstractThis paper addresses two vital challenges in Unsupervised Domain Adaptation (UDA) with a focus on harnessing the power of Vision-Language Pre-training (VLP) models. Firstly, UDA has primarily relied on ImageNet pre-trained models. However, the potential of VLP models in UDA remains largely unexplored. The rich representation of VLP models holds significant promise for enhancing UDA tasks. To address this, we propose a novel method called Cross-Modal Knowledge Distillation (CMKD), leveraging VLP models as teacher models to guide the learning process in the target domain, resulting in state-of-the-art performance. Secondly, current UDA paradigms involve training separate models for each task, leading to significant storage overhead and impractical model deployment as the number of transfer tasks grows. To overcome this challenge, we introduce Residual Sparse Training (RST) exploiting the benefits conferred by VLP’s extensive pre-training, a technique that requires minimal adjustment (approximately 0.1%~0.5%) of VLP model parameters to achieve performance comparable to fine-tuning. Combining CMKD and RST, we present a comprehensive solution that effectively leverages VLP models for UDA tasks while reducing storage overhead for model deployment. Furthermore, CMKD can serve as a baseline in conjunction with other methods like FixMatch, enhancing the performance of UDA. Our proposed method outperforms existing techniques on standard benchmarks. Our code will be available at: https://github.com/Wenlve-Zhou/VLP-UDA. Wenlve Zhou, Zhiheng Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Visual representations with texts domain generalization for semantic segmentation
Wanlin Yue, Zhiheng Zhou 0001, Yinglie Cao, Weikang Wu |
Appl. Intell. | 2 |
| 2023 | Visual-Textual Alignment for Generalizable Person Reidentification in Internet of ThingsabstractPerson reidentification (re-id) has gained increased attention for its important application in surveillance. Yet one of the major obstacles for the re-id approaches deploying in practical Internet of Things systems is their weak generalization ability. In spite that the current methods have a high performance under supervised setting, their discrimination ability in unseen domains meets decline. Due to the immutability of attributes among different domains, we attempt to exploit the alignment between the pedestrians’ attributes and visual features to enhance our model’s generalization ability. Furthermore, for the existing methods cannot fully extract the attribute information, we formulate a more effective NLP-based method for attribute feature extraction. Thus, the generated features are termed as textual features and our proposed method are called visual–textual alignment (VTA). As for alignment, two strategies are adopted: 1) metric learning-based alignment and 2) adversarial learning-based alignment. The former is designed to adjust the metric relationship of different persons in feature space. And the latter is aimed to guide our model’s domain-invariant feature learning. The experimental results demonstrate the effectiveness and superiority of our proposed method compared to the state-of-the-art methods. Xiaosheng Liu, Zhiheng Zhou 0001, Chang Niu, Qingru Wu |
IEEE Internet Things J. | 2 |
| 2022 | Asymmetric Adversarial-based Feature Disentanglement Learning for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition (MER) has gained tremendous progress. However, most methods are based on individual-database micro-expression recognition and are difficult to generalize into complicated scenarios. Therefore, cross-database micro-expression recognition (CDMER) has drawn growing attention due to its robustness and generalizability. In this paper, we propose a novel CDMER algorithm with asymmetric adversarial-based feature disentanglement learning, which implements the disentanglement of domain features and emotion features aiming to learn domain-invariant and discriminative representation. Furthermore, to facilitate the feature disentanglement learning, a Domain Information Filtering (DF) module is designed to filter out the domain component of the micro-expression features (emotion feature). Extensive experiments on the SMIC and CASME II databases have shown that our proposed method outperforms the state-of-the-art method and has superior performance against excessive domain discrepancies. Zhiheng Zhou 0001, Junyuan Shang |
ACM Multimedia | 2 |
| 2022 | Few-shot domain adaptation through compensation-guided progressive alignment and bias reduction
Junyuan Shang, Chang Niu, Junchu Huang, Zhiheng Zhou 0001, Junmei Yang |
Appl. Intell. | 4 |
| 2022 | Discriminative distribution alignment for domain adaptive object detection
Junchu Huang, Shifu Shen, Zhiheng Zhou 0001, Kefeng Fan |
Neurocomputing | 3 |
| 2022 | Unbiased feature generating for generalized zero-shot learning
Chang Niu, Junyuan Shang, Junchu Huang, Junmei Yang, Yuting Song, Zhiheng Zhou 0001, Guoxu Zhou |
J. Vis. Commun. Image Represent. | 6 |
| 2022 | Cascaded hierarchical CNN for 2D hand pose estimation from a single color image
Zhiheng Zhou 0001 |
Multim. Tools Appl. | 2 |
| 2022 | Superpixel attention guided network for accurate and real-time salient object detection
Zhiheng Zhou 0001, Yongfan Guo, Junchu Huang, Qingjun Yu |
Multim. Tools Appl. | 1 |
| 2022 | Label-guided heterogeneous domain adaptation
Zhiheng Zhou 0001, Chang Niu, Junyuan Shang |
Multim. Tools Appl. | 1 |
| 2021 | STA3DCNN: Spatial-Temporal Attention 3D Convolutional Neural Network for Citywide Crowd Flow Prediction
Gaozhong Tang, Zhiheng Zhou 0001, Bo Li 0111 |
ICONIP (4) | 2 |
| 2021 | Weakly supervised salient object detection via double object proposals guidanceabstractAbstract The weakly supervised methods for salient object detection are attractive, since they greatly release the burden of annotating time‐consuming pixel‐wise masks. However, the image‐level annotations utilized by current weakly supervised salient object detection models are too weak to provide sufficient supervision for this dense prediction task. To this end, a weakly supervised salient object detection method is proposed via double object proposals guidance, which is generated under the supervision of double bounding boxes annotations. With the double object proposals, the authors' method is capable of capturing both accurate but incomplete salient foreground and background information, which contributes to generating saliency maps with uniformly highlighted saliency regions and effectively suppressed background. In addition, an unsupervised salient object segmentation method is proposed, taking advantage of the non‐parametric statistical active contour model (NSACM), for segmenting salient objects with complete and compact boundaries. Experiments on five benchmark datasets show that the authors' weakly supervised salient object detection approach consistently outperforms other weakly supervised and unsupervised methods by a considerable margin, and even has comparable performance to the fully supervised ones. Zhiheng Zhou 0001, Yongfan Guo, Junchu Huang, Xiangwei Li |
IET Image Process. | 1 |
| 2021 | Domain compensatory adversarial networks for partial domain adaptation
Junchu Huang, Zhiheng Zhou 0001, Kefeng Fan |
Multim. Tools Appl. | 3 |
| 2021 | Asymmetric alignment joint consistent regularization for multi-source domain adaptation
Junyuan Shang, Chang Niu, Zhiheng Zhou 0001, Junchu Huang, Zhiwei Yang 0010, Xiangwei Li |
Multim. Tools Appl. | 3 |
| 2020 | Frame-Guided Region-Aligned Representation for Video Person Re-IdentificationabstractPedestrians in videos are usually in a moving state, resulting in serious spatial misalignment like scale variations and pose changes, which makes the video-based person re-identification problem more challenging. To address the above issue, in this paper, we propose a Frame-Guided Region-Aligned model (FGRA) for discriminative representation learning in two steps in an end-to-end manner. Firstly, based on a frame-guided feature learning strategy and a non-parametric alignment module, a novel alignment mechanism is proposed to extract well-aligned region features. Secondly, in order to form a sequence representation, an effective feature aggregation strategy that utilizes temporal alignment score and spatial attention is adopted to fuse region features in the temporal and spatial dimensions, respectively. Experiments are conducted on benchmark datasets to demonstrate the effectiveness of the proposed method to solve the misalignment problem and the superiority of the proposed method to the existing video-based person re-identification methods. Zengqun Chen, Zhiheng Zhou 0001, Junchu Huang, Bo Li 0111 |
AAAI | 2 |
| 2020 | DAGNet: Exploring the Structure of Objects for Saliency DetectionabstractFully Convolutional Neural Networks (FCNs) greatly promote the development of saliency detection. However, most of the FCN-based models have suffered from the structure of salient objects challenges. The extracted multi-scale features by previous models could help locate the objects with various scales, but they cannot contribute to effectively locating the objects with complex shapes, especially the salient regions that might intertwine with non-salient regions. Moreover, the style of decoder in previous models cannot adequately filter out the disturbance in low-level features, which is sub-optimal to sharpen the boundary of salient objects. In this paper, we propose DAGNet that explores the structure of salient objects from multi-level features to precisely detect salient objects. Firstly, the new dense multi-scale context extraction modules (DMCEMs) are implemented to transmit the rich structural information flow of salient objects from shallower layers to deeper layers, by which our model can locate the objects with complex shapes. Secondly, attention-based deeply refining modules (ADRMs) are designed in an effective attention-based style to effectively restore the boundary of objects stage-by-stage. In the style, the semantic information of high-level features is utilized to guide the shallow layer to filter out disturbance and refine the high-level features. Considering the salient objects surrounded by a cluttered scene, we propose a global context extraction module (GCEM) that can sufficiently understand the cluttered scene of an image from a global view. Comprehensive experiments indicate that our model is superior to 13 state-of-the-art models on 5 benchmark datasets under different evaluation metrics. Haobo Rao, Zhiheng Zhou 0001, Bo Li 0111 |
IJCNN | 2 |
| 2020 | Common-specific feature learning for multi-source domain adaptationabstractMulti‐source domain adaptation (MDA) aims to leverage knowledge from multiple source domains to improve the classification performance on target domains. Different degrees of distribution discrepancies between every two domains pose a huge challenge to MDA tasks. Most works focus on extracting features shared by all domains, which is critical but not enough to reduce distribution discrepancies. In this paper, we propose a method named as common‐specific feature learning (CSFL). Constituting a framework of feature learning, CSFL explores a subspace where the combination of common and specific features makes learned representations comprehensive. Based on this framework, we conduct a metric learning method for learning a discriminative feature representation. Considering redundant information caused by source domains is likely to hurt the performance, we impose an effective low‐rank constraint to remove the redundant information. Further, we adopt structure consistent constraint to preserve the local structure in each domain. CSFL has obtained about 1–5% improvement of mean accuracy, compared to the state‐of‐the‐art shallow methods. Further, compared with 90.2% and 89.4% of the best baseline deep method, CSFL achieves mean accuracy of 90.8% and 89.7% on the Office‐31 and ImageCLEF‐DA datasets respectively. The encouraging results validate the effectiveness of our method. Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junchu Huang, Tianlei Wang, Xiangwei Li |
IET Image Process. | 3 |
| 2020 | Heterogeneous domain adaptation with label and structural consistency
Junchu Huang, Zhiheng Zhou 0001, Junyuan Shang, Chang Niu |
Multim. Tools Appl. | 2 |
| 2019 | Delving into the Impact of Saliency Detector: A GeminiNet for Accurate Saliency Detection
Bo Li 0111, Delu Zeng, Zhiheng Zhou 0001 |
ICANN (3) | 4 |
| 2019 | Transfer metric learning for unsupervised domain adaptationabstractDomain adaptation is still a challenging task due to the fact that the distribution discrepancy between source domain and target domain weakens the transfer ability. Intuitively, it is crucial to discover a more discriminative feature representation across domains. However, previous methods do not take the target discriminative information into account since (most) target data are unlabelled. Here, the authors propose a transfer metric learning method which decreases intra‐class distance and increases inter‐class distance simultaneously even in the case of target data are unlabelled. The shared features are more discriminative, hence the model could be more robust for target data. Specially, the global optimal solution can be obtained by solving a generalised eigen‐decomposition problem. Extensive experiments on image datasets demonstrate that compared to several state‐of‐the‐art methods, authors’ method achieves significant improvement of 9.0% in average classification accuracy. Junchu Huang, Zhiheng Zhou 0001 |
IET Image Process. | 2 |
| 2019 | Prior distribution-based statistical active contour model
Zhiheng Zhou 0001, Tianlei Wang, Ruzheng Zhao |
Multim. Tools Appl. | 1 |
| 2018 | Practical Incremental Gradient Method for Large-Scale ProblemsabstractStochastic algorithms have become more and more popular in the minimization of finite sums due to their efficiency and effectiveness. Recent advances include the stochastic average gradient algorithm, the stochastic variance reduced gradient algorithm, and the SAGA algorithm, a set of incremental gradient algorithm. However, both the stochastic average gradient algorithm and the SAGA algorithm require to store gradients for each sample, which is expensive and impractical especially in the case of large scale problems. To the best of our knowledge, existing memory-free algorithm like the stochastic variance reduced gradient algorithm might not be efficient (fast) enough in this case. Taking these into account, we propose a new optimisation algorithm in this class with low memory requirement but still achieves faster convergence rate than the state-of-the-art, called Practical SAGA. Remarkly, as a variant of the SAGA algorithm, the Practical SAGA algorithm enjoys the advantages of the SAGA algorithm, for example, supports non-strongly convex problems directly. Extensive experiments on four benchmarks show the efficiency and effectiveness of the Practical SAGA. Junchu Huang, Zhiheng Zhou 0001, Zhiwei Yang 0010 |
TENCON | 2 |
| 2017 | Incremental Extreme Learning Machine via Fast Random Search Method
Zhihui Lao, Zhiheng Zhou 0001, Junchu Huang |
ICONIP (1) | 2 |
| 2017 | Accelerating Stochastic Variance Reduced Gradient Using Mini-Batch Samples on Estimation of Average Gradient
Junchu Huang, Zhiheng Zhou 0001, Bingyuan Xu |
ISNN (1) | 2 |
| 2017 | Static Hand Gesture Recognition Based on RGB-D Image and Arm Removal
Bingyuan Xu, Zhiheng Zhou 0001, Junchu Huang |
ISNN (1) | 2 |
| 2017 | Video error concealment scheme based on tensor model
Zhiheng Zhou 0001, Ruzheng Zhao, Bo Li 0111, Huiqiang Zhong, Yiming Wen |
Multim. Tools Appl. | 1 |
| 2014 | Gradient descent with adaptive momentum for active contour modelsabstractIn active contour models (snakes), various vector force fields replacing the gradient of the original external energy in the equations of motion are a popular way to extract the object boundary. Gradient descent method is usually used to obtain the equations of motion by minimising the energy functional. However, it always suffers from local minimum in extracting complex geometries because of non‐convex functional. Gradient descent method with adaptive momentum term is proposed in this study. First, an acceleration function of evolution is defined. Then, the adaptive momentum term is obtained by calculating the product between the edge stopping function and the defined acceleration function. Finally, adaptive momentum is compatible with the snakes. The edge stopping function is used to decide the influence region of the momentum, whereas the defined acceleration function determines the magnitude of the momentum. It is used to extract the complex geometries (such as deep concavity) when adding the adaptive momentum into some snakes, such as gradient vector field or vector field convolution snakes. On the other hand, the proposed method also accelerates the rate of convergence. It can be applied to extract a single object in real images. The experimental results show that the proposed method is effective and efficient. Guoqi Liu, Zhiheng Zhou 0001, Huiqiang Zhong, Shengli Xie 0001 |
IET Comput. Vis. | 2 |
| 2012 | Image Segmentation Based on the Poincaré Map MethodabstractActive contour models (ACMs) integrated with various kinds of external force fields to pull the contours to the exact boundaries have shown their powerful abilities in object segmentation. However, local minimum problems still exist within these models, particularly the vector field's "equilibrium issues." Different from traditional ACMs, within this paper, the task of object segmentation is achieved in a novel manner by the Poincaré map method in a defined vector field in view of dynamical systems. An interpolated swirling and attracting flow (ISAF) vector field is first generated for the observed image. Then, the states on the limit cycles of the ISAF are located by the convergence of Newton-Raphson sequences on the given Poincaré sections. Meanwhile, the periods of limit cycles are determined. Consequently, the objects' boundaries are represented by integral equations with the corresponding converged states and periods. Experiments and comparisons with some traditional external force field methods are done to exhibit the superiority of the proposed method in cases of complex concave boundary segmentation, multiple-object segmentation, and initialization flexibility. In addition, it is more computationally efficient than traditional ACMs by solving the problem in some lower dimensional subspace without using level-set methods. Delu Zeng, Zhiheng Zhou 0001, Shengli Xie 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | An efficient spatio-temporal boundary matching algorithm for video error concealment
Youjun Xiang, Liangmou Feng, Shengli Xie 0001, Zhiheng Zhou 0001 |
Multim. Tools Appl. | 4 |
| 2010 | Coarse-to-fine boundary location with a SOM-like methodabstractA coarse-to-fine boundary location with a self-organizing map (SOM)-like method is proposed in this paper. Inspired from the conventional SOM and universal gravitation, given a small quantity of supervision seeds from the desired boundaries, neurons are used to evolve to the desired boundaries in a coarse-to-fine framework. The major components of this framework are the designs of union action and evolving rate. In the course of neuron evolution, the union actions acting on these neurons will offer them the evolving directions. Also controlled by the corresponding referenced gradients, the neurons' evolving rates are adaptively adjusted at different positions. With the union actions and evolving rates, the neurons will evolve with appropriate manners to expand the set of feature points on the desired boundaries. The newly expanded feature points will cause the generation updates for feature points and neurons, and offer new information to guide the new generation of neurons to the boundaries. What is more, the proposed multiround evolution is as well a coarse-to-fine way for boundary location. Experiments and comparisons show that the proposed method performs well in complex long concavities, inhomogeneous and weak boundary location with good initialization flexibility. Delu Zeng, Zhiheng Zhou 0001, Shengli Xie 0001 |
IEEE Trans. Neural Networks | 2 |
| 2009 | Arranging and Interpolating Sparse Unorganized Feature Points With Geodesic Circular ArcabstractA novel method to reconstruct object boundaries with geodesic circular arc is proposed in this paper. Within this framework, an energy of circular arc spline is utilized to simultaneously arrange and interpolate each member in the set of sparse unorganized feature points from the desired boundaries. A general form for a family of parametric circular arc spline is firstly derived and followed by a novel method of arranging these feature points by minimizing an energy term depending on the circular arc spline configuration defined on these feature points. With regard to the fact that the energy function is usually nonconvex and nondifferentiable at its critical points, an improved scheme of particle swarm optimizer is given to find the minimum for the energy in this paper. With this improved scheme, each pair of neighboring feature points along the boundaries of the desired objects are picked out from the set of sparse unorganized feature points, and the corresponding directional chord tangent angles are computed simultaneously to finish interpolation. We show experimentally and comparatively that the proposed method can perform effectively to restrict leakage on weak boundaries and premature convergence on long concave boundaries. Besides, it has good noise robustness and can as well extract multiple and open boundaries. Shengli Xie 0001, Delu Zeng, Zhiheng Zhou 0001, Jun Zhang 0003 |
IEEE Trans. Image Process. | 3 |
| 2006 | Improved Clustering and Anisotropic Gradient Descent Algorithm for Compact RBF Network
Delu Zeng, Shengli Xie 0001, Zhiheng Zhou 0001 |
ICONIP (2) | 3 |
| 2004 | Video sequences error concealment based on texture detectionabstractA new error concealment algorithm based on average motion vector (AVMV) method is proposed as a post-processing tool at the decoder side for recovering the lost blocks and their motion vectors as well incurred during the video transmission. In our proposed scheme, LOG operator is used to detect the texture of the blocks around the lost block and find the more precise estimation of the motion vectors, so as to improve the AVMV method. Simulation results show that the proposed method can recover the higher quality image in different rates of lost block, comparing to the existing traditional concealment algorithms. Zhiheng Zhou 0001, Shengli Xie 0001 |
ICARCV | 1 |