VLDB 2026 Research / reviewers in the wild / expert
Shenglong Zhou 0002
dblp:133/8633-2
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-2575-8104ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training-Free Test-Time Adaptation via Shape and Style Guidance for Vision-Language ModelsabstractTest-time adaptation with pre-trained vision-language models shows impressive zero-shot classification abilities, and training-free methods further improve the performance without any optimization burden. However, existing training-free test-time adaptation methods typically rely on entropy criteria to select the visual features and update the visual caches, while ignoring the generalizable factors, such as shape-sensitive and style-insensitive factors. In this paper, we propose a novel shape and style guidance method (SSG) for training-free test-time adaptation in vision-language models, aiming to highlight the shape-sensitive (SHS) and style-insensitive (STI) factors in addition to entropy criteria.
Specifically, SSG perturbs the raw test image with shape and style corruption operations, and measures the prediction difference between the raw and corrupted one as perturbed prediction difference (PPD). Based on the PPD measurement, SSG reweights the high-confidence visual features and corresponding predictions, aiming to highlight the effect of SHS and STI factors during the test-time procedure. Furthermore, SSG takes both PPD and entropy into consideration to update the visual cache, aiming to maintain the stored sample with high entropy and generalizable factors. Extensive experimental results on out-of-distribution and cross-domain benchmark datasets demonstrate that our proposed SSG consistently outperforms previous state-of-the-art methods while also exhibiting promising computational efficiency. Shenglong Zhou 0002, Manjiang Yin, Leiyu Sun, Shicai Yang, Di Xie |
NeurIPS | 1 |
| 2024 | Test-Time Adaptation via Style and Structure Guidance for Histological Image RegistrationabstractImage registration plays a crucial role in histological image analysis, encompassing tasks like multi-modality fusion and disease grading. Traditional registration methods optimize objective functions for each image pair, yielding reliable accuracy but demanding heavy inference burdens. Recently, learning-based registration methods utilize networks to learn the optimization process during training and apply a one-step forward process during testing. While these methods offer promising registration performance with reduced inference time, they remain sensitive to appearance variances and local structure changes commonly encountered in histological image registration scenarios. In this paper, for the first time, we propose a novel test-time adaptation method for histological image registration, aiming to improve the generalization ability of learning-based methods. Specifically, we design two operations, style guidance and shape guidance, for the test-time adaptation process. The former leverages style representations encoded by feature statistics to address the issue of appearance variances, while the latter incorporates shape representations encoded by HOG features to improve registration accuracy in regions with structural changes. Furthermore, we consider the continuity of the model during the test-time adaptation process. Different from the previous methods initialized by a given trained model, we introduce a smoothing strategy to leverage historical models for better generalization. We conduct experiments with several representative learning-based backbones on the public histological dataset, demonstrating the superior registration performance of our test-time adaptation method. Shenglong Zhou 0002, Zhiwei Xiong, Feng Wu 0001 |
AAAI | 1 |
| 2023 | Learning Fine-Grained Features for Pixel-wise Video CorrespondencesabstractVideo analysis tasks rely heavily on identifying the pixels from different frames that correspond to the same visual target. To tackle this problem, recent studies have advocated feature learning methods that aim to learn distinctive representations to match the pixels, especially in a self-supervised fashion. Unfortunately, these methods have difficulties for tiny or even single-pixel visual targets. Pixel-wise video correspondences were traditionally related to optical flows, which however lead to deterministic correspondences and lack robustness on real-world videos. We address the problem of learning features for establishing pixel-wise correspondences. Motivated by optical flows as well as the self-supervised feature learning, we propose to use not only labeled synthetic videos but also unlabeled real-world videos for learning fine-grained representations in a holistic framework. We adopt an adversarial learning scheme to enhance the generalization ability of the learned features. Moreover, we design a coarse-to-fine framework to pursue high computational efficiency. Our experimental results on a series of correspondence-based tasks demonstrate that the proposed method outperforms state-of-the-art rivals in both accuracy and efficiency. Shenglong Zhou 0002, Dong Liu 0002 |
ICCV | 2 |
| 2023 | Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance SegmentationabstractSparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision. In this paper, we propose a sparsely supervised biomedical instance segmentation framework via cross-representation affinity consistency regularization. Specifically, we adopt two individual networks to enforce the perturbation consistency between an explicit affinity map and an implicit affinity map to capture both feature-level instance discrimination and pixel-level instance boundary structure. We then select the highly confident region of each affinity map as the pseudo label to supervise the other one for affinity consistency learning. To obtain the highly confident region, we propose a pseudo-label noise filtering scheme by integrating two entropy-based decision strategies. Extensive experiments on four biomedical datasets with sparse instance annotations show the state-of-the-art performance of our proposed framework. For the first time, we demonstrate the superiority of sparse instance-level supervision on 3D volumetric datasets, compared to common semi-supervision under the same annotation cost. Code is available at https://github.com/liuxy1103/CRAC. Xiaoyu Liu 0006, Wei Huang 0036, Zhiwei Xiong, Shenglong Zhou 0002, Yueyi Zhang 0001, Xuejin Chen, Zhengjun Zha, Feng Wu 0001 |
ICCV | 4 |
| 2023 | Generalized Lightness Adaptation with Channel Selective NormalizationabstractLightness adaptation is vital to the success of image processing to avoid unexpected visual deterioration, which covers multiple aspects, e.g., low-light image enhancement, image retouching, and inverse tone mapping. Existing methods typically work well on their trained lightness conditions but perform poorly in unknown ones due to their limited generalization ability. To address this limitation, we propose a novel generalized lightness adaptation algorithm that extends conventional normalization techniques through a channel filtering design, dubbed Channel Selective Normalization (CSNorm). The proposed CSNorm purposely normalizes the statistics of lightness-relevant channels and keeps other channels unchanged, so as to improve feature generalization and discrimination. To optimize CSNorm, we propose an alternating training strategy that effectively identifies lightness-relevant channels. The model equipped with our CSNorm only needs to be trained on one lightness condition and can be well generalized to unknown lightness conditions. Experimental results on multiple benchmark datasets demonstrate the effectiveness of CSNorm in enhancing the generalization ability for the existing lightness adaptation methods. Code is available at https://github.com/mdyao/CSNorm. Mingde Yao, Jie Huang 0017, Ruikang Xu, Shenglong Zhou 0002, Man Zhou 0003, Zhiwei Xiong |
ICCV | 5 |
| 2023 | Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement LearningabstractThe performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performance of downstream tasks, among which the mask image model (MIM) has been widely used due to its simplicity and effectiveness in recovering original information from masked images. However, due to the high degree of structural locality in EM images, as well as the existence of considerable noise, many voxels contain little discriminative information, making MIM pretraining inefficient on the neuron segmentation task. To overcome this challenge, we propose a decision-based MIM that utilizes reinforcement learning (RL) to automatically search for optimal image masking ratio and masking strategy. Due to the vast exploration space, using single-agent RL for voxel prediction is impractical. Therefore, we treat each input patch as an agent with a shared behavior policy, allowing for multi-agent collaboration. Furthermore, this multi-agent model can capture dependencies between voxels, which is beneficial for the downstream segmentation task. Experiments conducted on representative EM datasets demonstrate that our approach has a significant advantage over alternative self-supervised methods on the task of neuron segmentation. Code is available at https://github.com/ydchen0806/dbMiM. Yinda Chen, Wei Huang 0036, Shenglong Zhou 0002, Qi Chen 0014, Zhiwei Xiong |
IJCAI | 3 |
| 2023 | Self-Distilled Hierarchical Network for Unsupervised Deformable Image RegistrationabstractUnsupervised deformable image registration benefits from progressive network structures such as Pyramid and Cascade. However, existing progressive networks only consider the single-scale deformation field in each level or stage and ignore the long-term connection across non-adjacent levels or stages. In this paper, we present a novel unsupervised learning approach named Self-Distilled Hierarchical Network (SDHNet). By decomposing the registration procedure into several iterations, SDHNet generates hierarchical deformation fields (HDFs) simultaneously in each iteration and connects different iterations utilizing the learned hidden state. Specifically, hierarchical features are extracted to generate HDFs through several parallel gated recurrent units, and HDFs are then fused adaptively conditioned on themselves as well as contextual features from the input image. Furthermore, different from common unsupervised methods that only apply similarity loss and regularization loss, SDHNet introduces a novel self-deformation distillation scheme. This scheme distills the final deformation field as the teacher guidance, which adds constraints for intermediate deformation fields on deformation-value and deformation-gradient spaces respectively. Experiments on five benchmark datasets, including brain MRI and liver CT, demonstrate the superior performance of SDHNet over state-of-the-art methods with a faster inference speed and a smaller GPU memory. Code is available at https://github.com/Blcony/SDHNet. Shenglong Zhou 0002, Bo Hu 0014, Zhiwei Xiong, Feng Wu 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Cross-Resolution Distillation for Efficient 3D Medical Image RegistrationabstractImages captured in clinic such as MRI scans are usually in 3D formats with high spatial resolutions. Existing learning-based models for medical image registration consume large GPU memories and long inference time, which is difficult to be deployed in resource-limited diagnosis scenarios. To address this problem, instead of shrinking the model size as in previous works, we turn to reducing the input resolution of existing registration models and boosting their performance through knowledge distillation. Specifically, we propose a cross-resolution distillation (CRD) scheme, which is designed to train low-resolution models under the guidance of corresponding high-resolution models. Nevertheless, due to the resolution gap between features in high/low-resolution models, straightforward distillation is difficult to apply. To overcome this challenge, we first introduce a feature-shifted teacher (FST) to shift and fuse features of high/low-resolution models. Then, we exploit this teacher model to guide the learning of the low-resolution student model with distillation losses on both features and deformation fields. Finally, we only need to use the distilled student model during inference. Experimental results on four 3D medical image datasets demonstrate that the low-resolution models trained through our CRD scheme use fewer than 20% GPU memories and less than 20% inference time while achieving competitive performance compared with corresponding high-resolution models. Bo Hu 0014, Shenglong Zhou 0002, Zhiwei Xiong, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Recursive Decomposition Network for Deformable Image RegistrationabstractDeformation decomposition serves as a good solution for deformable image registration when the deformation is large. Current deformation decomposition methods can be categorized into cascade-based methods and pyramid-based methods. However, cascade-based methods suffer from heavy computational burdens and long inference time due to their structures of repeated subnetworks, while the effectiveness of pyramid-based methods is constrained by their limited numbers of resolution levels. In this paper, to address both the insufficient and inefficient decomposition problems in current deformation decomposition methods, we propose a recursive decomposition network (RDN) to offer a novel solution for deformable image registration. Stage-wise recursion can efficiently decompose a large deformation into different pyramid estimation stages without using repeated subnetworks like in cascade-based methods. Level-wise recursion can sufficiently decompose the deformation inside each resolution level instead of only one-time estimation like in pyramid-based methods. Extensive experiments and ablation studies on two representative datasets validate the effectiveness and efficiency of our proposed RDN. Bo Hu 0014, Shenglong Zhou 0002, Zhiwei Xiong, Feng Wu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Fast and Accurate Electron Microscopy Image Registration with 3D Convolution
Shenglong Zhou 0002, Zhiwei Xiong, Chang Chen 0004, Xuejin Chen, Dong Liu 0002, Yueyi Zhang 0001, Zhengjun Zha, Feng Wu 0001 |
MICCAI (1) | 1 |