VLDB 2026 Research / reviewers in the wild / expert
Shihua Zhou
dblp:145/4854
· DBLP profile ↗
25ranked-venue papers
1as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Structurally Stabilized Representations for Lossless DNA StorageabstractThis paper presents Reed-Solomon coded single-stranded representation learning (RSRL), a novel end-to-end model for learning representations for lossless DNA data storage. In contrast to existing learning-based methods, RSRL is inspired by both error-correction codec and structural biology. Specifically, RSRL first learns the representations for the subsequent storage from the binary data transformed by the Reed-Solomon codec (RS code). Then, the representations are masked by an RS-code-informed mask to focus on correcting the burst errors occurring in the learning process. The synergy of RS masks and graph attention enables active error localization, breaking through the limitations of traditional passive error correction. With the decoded representations with error corrections, a novel biologically stabilized loss is formulated to regularize the data representations to possess stable single-stranded structures. By incorporating these novel strategies, RSRL can learn highly durable, dense, and lossless representations for subsequent storage tasks in DNA sequences. The proposed RSRL has been compared with a number of baselines in real-world tasks of multi-type data storage. The experimental results obtained demonstrate that RSRL can store diverse types of data with much higher information density and durability, but much lower error rates. Ben Cao, Xue Li 0019, Tiantian He 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008 |
AAAI | 5 |
| 2026 | Predicting protein-protein interaction sites based on dynamic perception mechanism within a hierarchical E(n)-equivariant graphabstractAccurate prediction of protein-protein interaction sites is crucial to understanding biological processes, elucidating disease mechanisms, and accelerating drug discovery. Although graph neural network methods have shown potential in this field, but existing methods are limited by the static integrate multi-group features and insufficient perception of hierarchical 3D spatial geometric information, leading to insufficient predictive ability of orphan sites. To address these issues, this paper proposes a Dperception mechanism within a Hierarchical E(n)-equivariant Graph architecture (DHEG). DHEG introduces a dynamic feature importance perception mechanism that adaptively perceives the contextual inter-dependencies of features and assigns weights to feature groups based on their relevance to the interaction relationship. And a hierarchical gated architecture based on E(n)-equivariant graph neural networks that effectively captures protein 3D spatial structures while mitigating over-smoothing problems. The results show that DHEG achieves improvements in 11 of 13 key metrics, with an enhancement 8% in Matthews correlation coefficient, indicating that DHEG not only predicts more interaction sites but also does so with greater reliability. Furthermore, case studies and visualization analyzes show that DHEG aligns better with the biological mechanism and has excellent predictive capabilities for both orphan sites and continuous regions, demonstrating interpretability, and application potential. Xue Li 0019, Suheng Qiao, Shihua Zhou, Jianmin Wang 0016, Bin Wang 0005, Tao Song 0001, Ben Cao |
Briefings Bioinform. | 4 |
| 2026 | A gear surface defect detection approach based on class-linked mechanism
Shihua Zhou, Zichun Zhou, Kaibo Ji, Tingshuo Zhang, Zhaohui Ren |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | An end-to-end DNA storage coding method based on a low-complexity multiple biological constraints loss and RL-inspired differentiable solver
Yanfen Zheng, Xue Li 0019, Bin Wang 0005, Shihua Zhou, Ben Cao, Pan Zheng 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction
Liuming Shi, Xue Li 0019, Bin Wang 0005, Shihua Zhou, Ben Cao, Pan Zheng 0001 |
J. Biomed. Informatics | 6 |
| 2026 | Multimodal prompt-guided vision transformer for precise image manipulation localization
Yafang Xiao, Wei Jiang 0016, Shihua Zhou, Bin Wang 0005, Pengfei Wang 0013, Pan Zheng 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | DNA Sequence Clustering in High Error Rates via Hash Sketches Fuzzy Clustering for Efficient Stored Data Reconstruction
Yanfen Zheng, Ben Cao, Zhenlu Liu, Bin Wang 0005, Shihua Zhou, Pan Zheng 0001 |
PAKDD (3) | 6 |
| 2025 | Frequency-Assisted Local Attention in Lower Layers of Visual TransformersabstractSince vision transformers excel at establishing global relationships between features, they play an important role in current vision tasks. However, the global attention mechanism restricts the capture of local features, making convolutional assistance necessary. This paper indicates that transformer-based models can attend to local information without using convolutional blocks, similar to convolutional kernels, by employing a special initialization method. Therefore, this paper proposes a novel hybrid multi-scale model called Frequency-Assisted Local Attention Transformer (FALAT). FALAT introduces a Frequency-Assisted Window-based Positional Self-Attention (FWPSA) module that limits the attention distance of query tokens, enabling the capture of local contents in the early stage. The information from value tokens in the frequency domain enhances information diversity during self-attention computation. Additionally, the traditional convolutional method is replaced with a depth-wise separable convolution to downsample in the spatial reduction attention module for long-distance contents in the later stages. Experimental results demonstrate that FALAT-S achieves 83.0% accuracy on IN-1k with an input size of [Formula: see text] using 29.9[Formula: see text]M parameters and 5.6[Formula: see text]G FLOPs. This model outperforms the Next-ViT-S by 0.9[Formula: see text]APb/0.8[Formula: see text]APm with Mask-R-CNN [Formula: see text] on COCO and surpasses the recent FastViT-SA36 by 3.1% mIoU with FPN on ADE20k. Shihua Zhou, Zhaohui Ren, Yongchao Zhang 0004, Tianzhuang Yu |
Int. J. Neural Syst. | 3 |
| 2025 | RQVR: A multi-exposure image fusion network that optimizes rendering quality and visual realism
Enlong Wang, Huizi Man, Shihua Zhou, Yueping Wang |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | A dual-aligned knowledge self-distillation framework for visible-infrared cross-modal person re-identificationabstract• Dual alignment knowledge self-distillation to better capture modality-invariant/specific features for VI-ReID • Temperature-modulated alignment and confidence-based selective masking to enhance model reliability. • CutSwap augmentation to improve model robustness against intra-class variations and modality discrepancies. • State-of-the-art performance on SYSU-MM01 and RegDB benchmarks. Visible-infrared person re-identification (VI-ReID) significantly enhances identity retrieval across different illumination conditions by matching visible and infrared modalities. However, existing contrastive-learning-based approaches predominantly focus on cross-modal feature alignment, thus undermining model reliability in complex scenarios. To address this challenge, we introduce a Dual Alignment Knowledge Distillation (DAKD) framework that leverages comprehensive self-distillation at both instance and class levels. Our framework incorporates a temperature-modulated alignment strategy, capturing rich modality-invariant generalities as well as modality-specific discriminative details. Additionally, we propose a confidence-based selective masking mechanism that guides the distillation towards confident and informative teacher predictions. To further enhance robustness against modality discrepancies and intra-class variations, we develop a dedicated augmentation technique, CutSwap, which exchanges image channels to simulate realistic cross-modality variations. Extensive experiments on the benchmark SYSU-MM01 and RegDB datasets demonstrate superior performance compared to other state-of-the-art methods, achieving rank-1 accuracies of 76.31% and 94.83%, respectively and validating the efficacy of DAKD in maintaining robust cross-modal alignment while preserving essential identity-specific discriminative information. Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Shihua Zhou, George Vogiatzis, Yifan Wang 0008, Hui Fang 0003 |
Knowl. Based Syst. | 4 |
| 2025 | ELAFormer: Early Local Attention in multi-scale vision transFormers
Zhaohui Ren, Yongchao Zhang 0004, Tianzhuang Yu, Hengfa Luo, Shihua Zhou |
Knowl. Based Syst. | 7 |
| 2025 | LarTap: A Luminance-Aware Framework With Text-Correlation Priors for Multi-Exposure Image FusionabstractConventional imaging devices often struggle to produce high-dynamic-range (HDR) images that accurately represent natural scenes. To overcome this limitation, multi-exposure image fusion (MEF) techniques have been introduced as a viable solution. Existing MEF approaches aim to enhance performance by optimizing or searching architectures. However, they face challenges in precise feature extraction and scene reconstruction, leading to distortion in the fused images. Additionally, most methods do not adequately address luminance variations across different image regions, which may result in the loss of essential details. To address these challenges, we present a novel luminance-aware MEF framework that integrates text-correlation priors (LarTap). By embedding textual information into fusion process, the proposed framework enhances content extraction and comprehension. Specifically, it consist of two key components: the text-image correlation network (N1) and the multi-exposure fusion network (N2). First, N1 performs correlation training to achieve a holistic alignment between text and image pairs. Its iterative vision encoders (VEs) generate text-correlated prior knowledge to facilitate the fusion process in N2. Second, N2 leverages these priors for scene reconstruction and dynamically adjusts luminance based on comparative perception. Extensive experiments on multiple datasets demonstrate that LarTap outperforms state-of-the-art methods. The source code is available at https://github.com/EnLong-wang/LarTap. Enlong Wang, Jiawei Li 0016, Tiantian Yan, Jia Lei 0001, Shihua Zhou, Bin Wang 0005, Jinyuan Liu 0001, Nikola K. Kasabov |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MLFuse: Multi-Scenario Feature Joint Learning for Multi-Modality Image FusionabstractMulti-modality image fusion (MMIF) entails synthesizing images with detailed textures and prominent objects. Existing methods tend to use general feature extraction to handle different fusion tasks. However, these methods have difficulty breaking fusion barriers across various modalities owing to the lack of targeted learning routes. In this work, we propose a multi-scenario feature joint learning architecture, MLFuse, that employs the commonalities of multi-modality images to deconstruct the fusion progress. Specifically, we construct a cross-modal knowledge reinforcing network that adopts a multipath calibration strategy to promote information communication between different images. In addition, two professional networks are developed to maintain the salient and textural information of fusion results. The spatial-spectral domain optimizing network can learn the vital relationship of the source image context with the help of spatial attention and spectral attention. The edge-guided learning network utilizes the convolution operations of various receptive fields to capture image texture information. The desired fusion results are obtained by aggregating the outputs from the three networks. Extensive experiments demonstrate the superiority of MLFuse for infrared-visible image fusion and medical image fusion. The excellent results of downstream tasks (i.e., object detection and semantic segmentation) further verify the high-quality fusion performance of our method. Jia Lei 0001, Jiawei Li 0016, Jinyuan Liu 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008, Xiaopeng Wei, Nikola K. Kasabov |
IEEE Trans. Multim. | 5 |
| 2025 | TFFD-Net: an effective two-stage mixed feature fusion and detail recovery dehazing network
Wei Qi Yan 0001, Shihua Zhou, Yueping Wang |
Vis. Comput. | 4 |
| 2024 | DiffDD: A surface defect detection framework with diffusion probabilistic model
Yongchao Zhang 0004, Zhaohui Ren, Tianchuan Mi, Ke Feng 0004, Shihua Zhou |
Adv. Eng. Informatics | 6 |
| 2024 | A Unet-inspired spatial-attention transformer model for segmenting gear tooth surface defects
Yongchao Zhang 0004, Zhaohui Ren, Tianchuan Mi, Tianzhuang Yu, Shihua Zhou |
Adv. Eng. Informatics | 7 |
| 2024 | SDFuse: Semantic-injected dual-flow learning for infrared and visible image fusion
Enlong Wang, Jiawei Li 0016, Jia Lei 0001, Jinyuan Liu 0001, Shihua Zhou, Bin Wang 0005, Nikola K. Kasabov |
Expert Syst. Appl. | 5 |
| 2024 | Human Factors Design and Evaluation of China's Space Manipulator Teleoperation SystemabstractSpace manipulators are important for the assembly, construction, operation, and maintenance of space stations. To ensure that the manipulator system on the China Space Station (CSS) is user-friendly, the human factors engineering team has carried out extensive studies on the human factors design and evaluation of the space manipulator teleoperation system from three aspects: the technical process, research on ergonomic requirements, and ergonomic evaluation. This article systematically combs the technical process and index system of the human factors design and requirements research on CSS manipulator teleoperation system, which has guided the optimization of space manipulator engineering. This article expounds the research work on ergonomic requirements and suggestions related to manipulator teleoperation through experimental research on cognitive characteristics and human factors in display and control interfaces. We focus on human factors experimental research for camera type and handle polarity, and a recommended scheme for camera display and handle polarity has been screened out accordingly. Finally, the human-in-the-loop ergonomic evaluation method is demonstrated with a cooperative manual control task. The research results have facilitated the human factors design and optimization improvement of the space manipulator manipulator system, and have contributed significantly to the safe and smooth completion of astronaut extravehicular activities on the CSS. Weicai Tang, Lifen Tan, Shuqi Xue, Shihua Zhou |
Int. J. Hum. Comput. Interact. | 6 |
| 2024 | Rethinking Position Embedding Methods in the Transformer ArchitectureabstractAbstract In the transformer architecture, as self-attention reads entire image patches at once, the context of the sequence between patches is omitted. Therefore, the position embedding method is employed to assist the self-attention layers in computing the ordering information of tokens. While many papers simply add the position vector to the corresponding token vector rather than concatenating them, few papers offer a thorough explanation and comparison beyond dimension reduction. However, the addition method is not meaningful because token vectors and position vectors are different physical quantities that cannot be directly combined through addition. Hence, we investigate the disparity in learnable absolute position information between the two embedding methods (concatenation and addition) and compare their performance on models. Experiments demonstrate that the concatenation method can learn more spatial information (such as horizontal, vertical, and angle) than the addition method. Furthermore, it reduces the attention distance in the final few layers. Moreover, the concatenation method exhibits greater robustness and leads to a performance gain of 0.1–0.5% for existing models without additional computation overhead. Zhaohui Ren, Shihua Zhou, Tianzhuang Yu, Hengfa Luo |
Neural Process. Lett. | 3 |
| 2024 | GeSeNet: A General Semantic-Guided Network With Couple Mask Ensemble for Medical Image FusionabstractAt present, multimodal medical image fusion technology has become an essential means for researchers and doctors to predict diseases and study pathology. Nevertheless, how to reserve more unique features from different modal source images on the premise of ensuring time efficiency is a tricky problem. To handle this issue, we propose a flexible semantic-guided architecture with a mask-optimized framework in an end-to-end manner, termed as GeSeNet. Specifically, a region mask module is devised to deepen the learning of important information while pruning redundant computation for reducing the runtime. An edge enhancement module and a global refinement module are presented to modify the extracted features for boosting the edge textures and adjusting overall visual performance. In addition, we introduce a semantic module that is cascaded with the proposed fusion network to deliver semantic information into our generated results. Sufficient qualitative and quantitative comparative experiments (i.e., MRI-CT, MRI-PET, and MRI-SPECT) are deployed between our proposed method and ten state-of-the-art methods, which shows our generated images lead the way. Moreover, we also conduct operational efficiency comparisons and ablation experiments to prove that our proposed method can perform excellently in the field of multimodal medical image fusion. The code is available at https://github.com/lok-18/GeSeNet. Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Unsupervised Saliency Detection via kNN Mechanism and Object-Biased Prior
Zhaohui Ren, Shihua Zhou, Tianzhuang Yu |
Neural Process. Lett. | 3 |
| 2023 | Learning a Coordinated Network for Detail-Refinement Multiexposure Image FusionabstractNowadays, deep learning has made rapid progress in the field of multi-exposure image fusion. However, it is still challenging to extract available features while retaining texture details and color. To address this difficult issue, in this paper, we propose a coordinated learning network for detail-refinement in an end-to-end manner. Firstly, we obtain shallow feature maps from extreme over/under-exposed source images by a collaborative extraction module. Secondly, smooth attention weight maps are generated under the guidance of a self-attention module, which can draw a global connection to correlate patches in different locations. With the cooperation of the two aforementioned used modules, our proposed network can obtain a coarse fused image. Moreover, by assisting with an edge revision module, edge details of fused results are refined and noise is suppressed effectively. We conduct subjective qualitative and objective quantitative comparisons between the proposed method and twelve state-of-the-art methods on two available public datasets, respectively. The results show that our fused images significantly outperform others in visual effects and evaluation metrics. In addition, we also perform ablation experiments to verify the function and effectiveness of each module in our proposed method. The source code can be achieved athttps://github.com/lok-18/LCNDR. Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A meta-inspired termite queen algorithm for global optimization and engineering design problems
Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
Eng. Appl. Artif. Intell. | 2 |
| 2018 | Constructing DNA Barcode Sets Based on Particle Swarm OptimizationabstractFollowing the completion of the human genome project, a large amount of high-throughput bio-data was generated. To analyze these data, massively parallel sequencing, namely next-generation sequencing, was rapidly developed. DNA barcodes are used to identify the ownership between sequences and samples when they are attached at the beginning or end of sequencing reads. Constructing DNA barcode sets provides the candidate DNA barcodes for this application. To increase the accuracy of DNA barcode sets, a particle swarm optimization (PSO) algorithm has been modified and used to construct the DNA barcode sets in this paper. Compared with the extant results, some lower bounds of DNA barcode sets are improved. The results show that the proposed algorithm is effective in constructing DNA barcode sets. Bin Wang 0005, Xuedong Zheng, Shihua Zhou, Changjun Zhou, Xiaopeng Wei, Qiang Zhang 0008, Ziqi Wei 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | A Method of Discriminative Features Extraction for Restricted Boltzmann Machines
Song Guo 0002, Changjun Zhou, Bin Wang 0005, Shihua Zhou |
IDEAL | 4 |