VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Fang
dblp:90/10409
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-0120-4585ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cas-OVD: Cascaded Open-Vocabulary Detection of Small Objects Using Multi-Refined Region Proposal Network in Autonomous DrivingabstractAlthough text information has aided existing models to achieve promising results in open vocabulary object detection (OVD), the lack of semantic information has led to the difficulty in small objects detection (SOD). Moreover, such semantic gap also causes failure when matching texts and image features, resulting in false negative instances being detected. To address these issues, we propose a Cascade Open Vocabulary Detector (Cas-OVD), which builds upon existing multi-stage detection pipelines but specializes in text-vision alignment for small objects. In particular, we adapt a multi-refined region proposal network, guided by a non-sampled anchor strategy, to reduce the missing and false detections of small objects. Meanwhile, a deformable convolution network based feature conversion module is proposed to enhance the semantic information of small objects even the potential ones with low confidence. Unlike existing methods that rely on coarse-grained image-based features for image-text matching, Cas-OVD refines these features through a cascade alignment process, allowing each stage to build on the results of the previous one. This can progressively enhance the feature correlation between the image regions and the textual descriptions through successive error correction. On the joint BDD100K-SODA-D dataset, Cas-OVD achieved 17.95% AP$_{\mathrm{all}}$and 14.6% AP$_{\mathrm{s}}$, outperforming RegionCLIP by 3.5% AP$_{\mathrm{all}}$and 3.0% AP$_{\mathrm{s}}$, respectively. On the OV_COCO dataset, Cas-OVD has the 32.71% AP$_{\mathrm{all}}$and 17.26% AP$_{\mathrm{s}}$, surpassing the RegionCLIP by 6.6% AP$_{\mathrm{all}}$and 6.1% AP$_{\mathrm{s}}$, respectively. Zhenyu Fang, Jinchang Ren, Jiangbin Zheng 0001, Yijun Yan, Lixiang Zhang |
IEEE Trans. Multim. | 1 |
| 2025 | Aligning local features from multi-view (ALFM): A hybrid self-Supervised framework for object detection via contextual distillation and global representation learning
Zhenyu Fang, Zhuowei Wang 0006, Jinchang Ren, Jiangbin Zheng 0001, Rongjun Chen 0001, Huimin Zhao 0001 |
Knowl. Based Syst. | 1 |
| 2025 | Dual Teacher: Improving the Reliability of Pseudo Labels for Semi-Supervised Oriented Object DetectionabstractOriented object detection in remote sensing is a critical task for accurately location and measurement of the interested targets. Despite of its success in object detection, deep learning-based detectors rely heavily on extensive data annotation. However, variations in object appearance significantly increase the difficulty and the cost of creating large-scale annotated datasets. Semi-supervised learning (SSL) aims to utilize unlabeled data to enhance object detectors. Among these, pseudo-label-based methods have shown promising results recently. Nonetheless, as training progresses, the accumulation of errors in pseudo labels leads to prediction bias without corrections. To tackle this particular challenge, we present a SSL pipeline, named “dual teacher,” for improving the reliability of pseudo labels in the semi-supervised oriented object detection. First, to mitigate the bias caused by limited annotated data, a global burn-in (GBI) strategy is introduced at the beginning of training, which guides the student detector to learn the feature extraction on a global scale. In addition, an online bounding box (bbox) correction module is proposed to decrease the occurrence of mislabeled instances and enhance the reliability of detection. These improvements are facilitated by an additional detector, instead of a single teacher model in the teacher-student architecture. Dual teacher reduces the dependency on the quality of pseudo labels related to the model complexity and combines the strengths of both the two-stage and one-stage detectors. With only 20% labeled data, dual teacher outperforms fully supervised rotated fully convolutional one-stage object detection (R-FCOS), you only look once X-small (YOLOX-s), and rotated region-based convolutional neural network (R-RCNN) by up to 2% on both a large-scale dataset for object detection in aerial images (DOTA) and SODA-A datasets. This reveals its potential in reducing labor-intensive tasks and enhancing robustness against environmental interference and noisy labels. The code is available at:https://github.com/ZYFFF-CV/DualTeacher-semisup.git. Zhenyu Fang, Jinchang Ren, Jiangbin Zheng 0001, Rongjun Chen 0001, Huimin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Two-Stream Vision Swin Transformer for Video-based Eye Movement Detection
Xiaowei Bai, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
CogSci | 4 |
| 2024 | Complex-Valued Gabor-Attention Residual Fusion Network for Iris RecognitionabstractIris recognition has gained significant attention in identity verification due to the unique, stable texture patterns in iris. Successfully extracting these patterns is essential for quick and precise identification. Although deep learning methods have automated the iris recognition, they predominantly rely on real-valued networks that overlook the complex-valued representation of iris texture. This means they cannot effectively process phase and amplitude information, and fail to integrate domain-specific knowledge of iris, thereby not fully capturing the intricate details of the iris texture. Inspired by classical manual methods that efficiently harness the complex-valued representation of the iris to extract both amplitude and phase information. We integrate Gabor filters with complex-valued neural networks, propose a Complex-Valued Gabor-Attention Residual Fusion Network (GRFN) tailored for iris recognition, aiming to comprehensively capture the iris texture’s multi-scale and multi-orientation phase and amplitude features. The GRFN incorporates adaptive Gabor Complex-Valued Convolution Kernels (GCVK) to introduce a Gabor attention mechanism focused on iris biometric characteristics. Furthermore, we propose a novel residual feature fusion approach that selects and merges local and global features across multiple directions and scales, mitigating model degradation and enhancing the network’s ability to extract iris texture features effectively. Extensive experiments show that the proposed network outperforms the state-of-the-art performance on two benchmark datasets. Zhuoru Li, Xiaowei Bai, Yingxi Li, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ECAI | 6 |
| 2024 | Two-Click-Based Fast Small Object Annotation in Remote Sensing ImagesabstractIn the remote sensing field, detecting small objects is a pivotal task, yet achieving high performance in deep learning-based detectors heavily relies on extensive data annotation. The challenge intensifies as small objects in remote sensing imagery are typically densely distributed and numerous, leading to a substantial increase in the cost of creating large-scale annotated datasets. This elevated cost poses significant limitations on the application and advancement of small object detection. To address this issue, a point-based annotation (PBA) method is proposed, which generates bounding boxes (BBOXs) through graph-based segmentation. In this framework, user annotations categorize nodes into three distinct classes—positive, negative, and to-cut—facilitating a more intuitive and efficient annotation process. Utilizing the max-flow algorithm, our method seamlessly generates oriented BBOXs (OBBOXs) from these classified nodes. The efficacy of PBA is underscored by our empirical findings. Notably, annotation efficiency is enhanced by at least 40%, a significant leap forward. Moreover, the intersection over union (IoU) metric of our OBBOX outperforms existing methods like “segment anything model (SAM)” by 10%. Finally, when applied in training, models annotated with PBA exhibit a 3% increase in the mean average precision (mAP) compared with those using traditional annotation methods. These results not only affirm the technical superiority of PBA but also its practical impact on advancing small object detection in remote sensing. Lu Lei, Zhenyu Fang, Jinchang Ren, Paolo Gamba, Jiangbin Zheng 0001, Huimin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Semi-supervised Histological Image Segmentation via Hierarchical Consistency Enforcement
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leyi Wei, Zhenyu Fang, Zhaopeng Meng, Ran Su |
MICCAI (2) | 6 |
| 2022 | Adaptive Distance-Based Band Hierarchy (ADBH) for Effective Hyperspectral Band SelectionabstractBand selection has become a significant issue for the efficiency of the hyperspectral image (HSI) processing. Although many unsupervised band selection (UBS) approaches have been developed in the last decades, a flexible and robust method is still lacking. The lack of proper understanding of the HSI data structure has resulted in the inconsistency in the outcome of UBS. Besides, most of the UBS methods are either relying on complicated measurements or rather noise sensitive, which hinder the efficiency of the determined band subset. In this article, an adaptive distance-based band hierarchy (ADBH) clustering framework is proposed for UBS in HSI, which can help to avoid the noisy bands while reflecting the hierarchical data structure of HSI. With a tree hierarchy-based framework, we can acquire any number of band subset. By introducing a novel adaptive distance into the hierarchy, the similarity between bands and band groups can be computed straightforward while reducing the effect of noisy bands. Experiments on four datasets acquired from two HSI systems have fully validated the superiority of the proposed framework. He Sun 0009, Jinchang Ren, Huimin Zhao 0001, Genyun Sun, Wenzi Liao, Zhenyu Fang, Jaime Zabalza |
IEEE Trans. Cybern. | 6 |
| 2022 | SC2Net: A Novel Segmentation-Based Classification Network for Detection of COVID-19 in Chest X-Ray ImagesabstractThe pandemic of COVID-19 has become a global crisis in public health, which has led to a massive number of deaths and severe economic degradation. To suppress the spread of COVID-19, accurate diagnosis at an early stage is crucial. As the popularly used real-time reverse transcriptase polymerase chain reaction (RT-PCR) swab test can be lengthy and inaccurate, chest screening with radiography imaging is still preferred. However, due to limited image data and the difficulty of the early-stage diagnosis, existing models suffer from ineffective feature extraction and poor network convergence and optimisation. To tackle these issues, a segmentation-based COVID-19 classification network, namely SC2Net, is proposed for effective detection of the COVID-19 from chest x-ray (CXR) images. The SC2Net consists of two subnets: a COVID-19 lung segmentation network (CLSeg), and a spatial attention network (SANet). In order to supress the interference from the background, the CLSeg is first applied to segment the lung region from the CXR. The segmented lung region is then fed to the SANet for classification and diagnosis of the COVID-19. As a shallow yet effective classifier, SANet takes the ResNet-18 as the feature extractor and enhances high-level feature via the proposed spatial attention module. For performance evaluation, the COVIDGR 1.0 dataset is used, which is a high-quality dataset with various severity levels of the COVID-19. Experimental results have shown that, our SC2Net has an average accuracy of 84.23% and an average F1 score of 81.31% in detection of COVID-19, outperforming several state-of-the-art approaches. Huimin Zhao 0001, Zhenyu Fang, Jinchang Ren, Calum MacLellan, Yong Xia 0001, Shuo Li 0001, Meijun Sun, Kevin Ren |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Topological optimization of the DenseNet with pretrained-weights inheritance and genetic channel selection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Song Wang 0002, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2020 | Triple loss for hard face detection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Zheng Wang 0008, Kaizhu Huang, Bing Xiao 0005 |
Neurocomputing | 1 |