EDBT 2026 Demo / reviewers in the wild / expert
Shibai Yin
dblp:140/0713
· DBLP profile ↗
20ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | V-Sparse: From temporal-spatial visual semantic compression to coarse-to-fine interaction for text-video retrieval
Shibai Yin, Jun Wang 0089, Xingyang Wang, Yubing Shen, Yee-Hong Yang |
Neural Networks | 2 |
| 2025 | Enhancing the Adversarial Robustness via Manifold ProjectionabstractDeep learning has been widely applied to various aspects of computer vision, but the emergence of adversarial attacks raises concerns about its reliability. Adversarial training (AT) is one of the most effective defense methods, which incorporates adversarial examples into the training data. However, AT is typically employed in a discriminative learning manner, i.e., learning the mapping (conditional probability) from samples to labels, it essentially reinforces this mapping without considering the underlying data distribution. It is notable that adversarial examples often deviate from the distribution of normal (clean) samples. Therefore, building upon existing adversarial defense schemes, we propose to further exploit the distribution of normal samples, partly from the generative learning perspective, resulting in a novel robustness enhancement paradigm. We train a simple autoencoder (AE) autoregressively on normal samples to learn their prior distribution, effectively serving as an image manifold. This AE is then used as a manifold projection operator to incorporate the distribution information of normal samples. Specifically, we organically integrate the pretrained AE into the training process of both AT and adversarial distillation (AD), a method aiming at improving the robustness of small models with low capacity. Since the AE captures the distribution of normal samples, it can adaptively pull adversarial examples closer to the normal sample manifold, weakening the attack strength of adversarial samples and easing the learning of mappings from adversarial samples to correct labels. From the Pearson correlation coefficient (PCC) between the statistics on normal and adversarial examples, it’s validated that the AE indeed pulls adversarial samples closer to normal samples. Extensive experiments illustrate that our proposed adversarial defense paradigm significantly improves the robustness compared with previous state-of-the-art AT and AD methods. Zhiting Li, Shibai Yin, Tai-Xiang Jiang, Yexun Hu, Jia-Mian Wu, Guowei Yang 0001, Guisong Liu |
AAAI | 2 |
| 2025 | DUQ: Dual Uncertainty Quantification for Text-Video RetrievalabstractText-video retrieval establishes accurate similarity relationships between text and video through feature enhancement and granularity alignment. However, relying solely on similarity to associate intra-pair features and distinguish inter-pair features is insufficient, \textit{e.g.}, when querying a multi-scene video with sparse text or selecting the most relevant video from many similar candidates. In this paper, we propose a novel Dual Uncertainty Quantification (DUQ) model that separately handles uncertainties in intra-pair interaction and inter-pair exclusion. Specifically, to enhance intra-pair interaction, we propose an intra-pair similarity uncertainty module to provide similarity-based trustworthy predictions and explicitly model this uncertainty. To increase inter-pair exclusion, we propose an inter-pair distance uncertainty module to construct a distance-based diversity probability embeding, thereby widening the gap between similar features. The two components work synergistically, jointly improving the calculation of similarity between features. We evaluate our model on six benchmark datasets: MSRVTT (51.2%), DiDeMo, MSVD, LSMDC, Charades, and VATEX, achieving state-of-the-art retrieval performance. Shibai Yin, Xingyang Wang, Yee-Hong Yang |
IJCAI | 2 |
| 2025 | Dual-Branch Wavelet Diffusion models with Dual-Prior Refinement for Underwater Image Enhancement
Yiwei Shi, Shibai Yin, Yanfang Fu, Yee-Hong Yang |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | FishDetectLLM: Multimodal instruction tuning with large language models for fish detectionabstractAquatic species play crucial roles in global ecosystems but are increasingly threatened by factors such as overfishing, coastal development and climate change . Existing deep learning methods address these challenges by employing powerful networks and large-scale, diverse datasets, separately tackling species recognition and trait identification during ongoing monitoring. However, they often exhibit limited generalization ability. Inspired by the human ability to quickly identify fish species and their locations with just a glance at an underwater image or scene, we introduce FishDetectLLM—a framework built on the lightweight TinyLLaVA architecture. FishDetectLLM utilizes the powerful reasoning capabilities and vast world knowledge of large language models (LLMs) to address the fish detection problem, providing both fish classification results and predicted bounding boxes for fish. Specifically, we create instruction dialogues for fish detection that connect fish taxonomy with classification descriptions and map location descriptions to the corresponding coordinates of bounding box in the input images from the recently released large-scale FishNet dataset. Then, we pretrain and fine-tune FishDetectLLM to achieve fish detection using the created dataset, leveraging the principle of augmenting human knowledge. Our results show that FishDetectLLM significantly outperforms existing multimodal LLMs and task-specific methods. Unlike conventional detection architectures that struggle to generalize beyond the training data, FishDetectLLM exhibits strong generalization capabilities, achieving robust performance on unseen data. This innovation paves the way for future applications of MLLMs in full research and offers valuable tools for the conservation of fish biodiversity. Shibai Yin, Xingyang Wang, Yee-Hong Yang |
Knowl. Based Syst. | 2 |
| 2025 | When Aware Haze Density Meets Diffusion Model for Synthetic-to-Real DehazingabstractImage dehazing is an important preliminary step for downstream vision tasks. Existing deep learning-based methods have limited generalization capabilities for real hazy images because they are trained on synthetic data and exhibit high domain-specific properties. This work proposes a new Diffusion Model for Synthetic-to-Real dehazing (DMSR) based on the haze-aware density. DMSR mainly comprises of a physics-based dehazing model and a Conditional Denoising Diffusion Model (CDDM)-based model. The coarse transmission map and coarse dehazing result estimated by the physics-based dehazing model serve as conditions for the subsequent CDDM-based model. In this process, the CDDM-based dehazing model progressively refines the coarse transmission map while generating the dehazing result, enabling the model to remove haze with accurate haze density information. Next, we propose a haze density-aware resampling strategy that incorporates the coarse dehazed result into the resampling process using the transmission map, thereby fully leveraging the diffusion model for heavy haze removal. Moreover, a new synthetic-to-real training strategy with the prior-based loss function and the memory loss function is applied to DMSR for improving generalization capabilities and narrowing the gap between the synthetic and real domains with low computational cost. Extensive experiments on various real datasets demonstrate the effectiveness and superiority of the proposed DMSR over state-of-the-art methods. Shibai Yin, Yiwei Shi, Yibin Wang 0001, Yee-Hong Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Visual Attention and ODE-inspired Fusion Network for image dehazing
Shibai Yin, Ruyuan Lu, Zhen Deng, Yee-Hong Yang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | A multi-level wavelet-based underwater image enhancement network with color compensation prior
Yibin Wang 0001, Shuhao Hu, Shibai Yin, Zhen Deng, Yee-Hong Yang |
Expert Syst. Appl. | 3 |
| 2024 | Convolution-transformer blend pyramid network for underwater image enhancement
Lunpeng Ma, Dongyang Hong, Shibai Yin, Wanqiu Deng, Yang Yang 0229, Yee-Hong Yang |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Multilevel Semantic Interaction Alignment for Video-Text Cross-Modal RetrievalabstractVideo–text cross-modal retrieval (VTR) is more natural and challenging than image–text retrieval, which has attracted increasing interest from researchers in recent years. To align VTR more closely with real-world scenarios, i.e., weak semantic text description as a query, we propose a multilevel semantic interaction alignment (MSIA) model. We develop a two-stream network, which decomposes video and text alignment into multiple dimensions. Specifically, in the video stream, to better align heterogeneity data, redundant video information is suppressed via the designed frame adaptation attention mechanism, and richer semantic interaction is achieved through a text-guided attention mechanism. Then, for text alignment in the video local region, we design a distinctive anchor frame strategy and a word selection method. Finally, a cross-granularity alignment approach is designed to learn more and finer semantic features. With the above schema, the alignment between video and weak semantic text descriptions is reinforced, further alleviating the issues of difficult alignment caused by weak semantic text descriptions. The experimental results on VTR benchmark datasets show the competitive performance of our approach in comparison to that of state-of-the-art methods. The code is available at: https://github.com/jiaranjintianchism/MSIA. Zhen Deng, Shibai Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Adams-based hierarchical features fusion network for image dehazing
Shibai Yin, Shuhao Hu, Yibin Wang 0001, Weixing Wang 0001, Yee-Hong Yang |
Neural Networks | 1 |
| 2022 | Multi-step implicit Adams predictor-corrector network for fire detectionabstractAbstract Fire detection methods based on the Convolutional Neural Networks (CNN) have advantages of high accuracy, wide coverage and robustness, receiving significant attention from researchers. Among CNN‐based methods, ResNet has achieved better performance than other CNN frameworks in fire detection system, since it uses stacked residual blocks to enlarge the receptive field to overcome the vanishing gradient problem with residual learning. The merits of ResNet can be attributed to the similarity between ResNet and the single‐step explicit solver for Ordinary Differential Equations (ODEs), for example, the Euler method. Motivated by the theory of numerical ODE that a multi‐step implicit solver has higher accuracy than a single‐step explicit solver, the Multi‐step Implicit Adams predictor‐corrector (MIAPC) network for fire detection is proposed. The MIAPC method is first mapped to a corresponding predictor‐corrector Adams block which achieves higher accuracy than a single‐step explicit solver. Then, Adaptive Feature Fusion (AFF) and the Spatial Attention Layer (SAL) are utilized to extract hierarchical features from stacked predictor‐corrector Adams blocks, forming the corresponding Adams module. Finally, the 4 Adams modules which are made of 4, 6, 8, 10 predictor‐corrector Adams blocks and followed by AFF and SAL form the crucial ODE‐based approximation part in the proposed network. By adding a simple feature extraction and detection in front of and after the ODE‐based approximation part, the MIAPC network is built. Experiments demonstrate that the method achieves 87% accuracy in the challenging test dataset, outperforming existing methods by at least 6%. Besides, the 5.3M model size with inference speed of 4.7 frames/second in CPU and 65.7 frames/second in GPU enables the proposed method to be used in practical applications. Zhen Deng, Shuhao Hu, Shibai Yin, Yibin Wang 0001, Anup Basu, Irene Cheng 0001 |
IET Image Process. | 3 |
| 2022 | Degradation-aware and color-corrected network for underwater image enhancement
Shibai Yin, Shuhao Hu, Yibin Wang 0001, Weixing Wang 0001, Yee-Hong Yang |
Knowl. Based Syst. | 1 |
| 2021 | Attentive U-recurrent encoder-decoder network for image dehazing
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang |
Neurocomputing | 1 |
| 2021 | A multi-scale attentive recurrent network for image dehazing
Yibin Wang 0001, Shibai Yin, Anup Basu |
Multim. Tools Appl. | 2 |
| 2021 | Visual Attention Dehazing Network with Multi-level Features Refinement and Fusion
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2020 | Image dehazing with uneven illumination prior by dense residual channel attention networkabstractExisting dehazing methods based on convolutional neural networks estimate the transmission map by treating channel‐wise features equally, which lacks flexibility in handling different types of haze information, leading to the poor representational ability of the network. Besides, the scene lights are predicted by an even illumination prior which does not work for a real situation. To solve these problems, the authors propose a dense residual channel attention network (DRCAN) for estimating the transmission map and use an image segmentation strategy to predict scene lights. Specifically, DRCAN is built based on the proposed dense residual block (DRB) and dense residual channel attention block (DRCAB). DRB extracts the hierarchical features with increasing receptive fields. DRCAB makes the network focus on the features containing heavy haze information. After the transmission map is estimated, fuzzy partition entropy combined with graph cuts is used to segment the transmission map into scene regions covered with varying scene lights. This strategy not only considers the fuzzy intensities of the low‐contrast transmission map but also takes spatial correlation into account. Finally, a clear image is obtained by the transmission map and varying scene lights. Extensive experiments demonstrate that our method is comparable to most of existing methods. Shibai Yin, Jin Xin, Yibin Wang 0001, Anup Basu |
IET Image Process. | 1 |
| 2020 | A novel image-dehazing network with a parallel attention block
Shibai Yin, Yibin Wang 0001, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2017 | Unsupervised hierarchical image segmentation through fuzzy entropy maximization
Shibai Yin, Yiming Qian, Minglun Gong |
Pattern Recognit. | 1 |
| 2014 | Efficient multilevel image segmentation through fuzzy entropy maximization and graph cut optimization
Shibai Yin, Xiangmo Zhao, Weixing Wang 0001, Minglun Gong |
Pattern Recognit. | 1 |