Zhibin Yu 0002

dblp:70/8170-2 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-4372-1767ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unlocking the power of vision foundation models via multi-expert collaboration for cross-domain few-shot segmentation
Yuxin Li 0004, Zihao Zhu 0002, Zhibin Yu 0002
Knowl. Based Syst.4
2025 Boost the Inference with Co-training: A Depth-guided Mutual Learning Framework for Semi-supervised Medical Polyp Segmentation
abstract
Semi-Supervised polyp segmentation has made significant progress in recent years as a potential solution for computer-assisted treatment. Since depth images can provide extra information other than RGB images to help segment these problematic areas, depth-assisted polyp segmentation has gained much attention. However, the utilization of depth information is still worth studying. The existing RGB-D segmentation methods rely on depth data in the inference stage, limiting their clinical applications. To tackle this problem, we propose a semi-supervised polyp segmentation framework based on the mean teacher architecture. We establish an auxiliary student network with depth images as input in the training stage, and we propose a depth-guided cross-modal mutual learning strategy to promote the learning of complementary information between different student networks. Meanwhile, we use the high-confidence pseudo-labels generated by the auxiliary student network to guide the learning progress of the main student network from different perspectives. Our model does not need depth data in the inference phase. In addition, we introduce a depth-guided patch augmentation method to improve the model’s learning performance in difficult regions of unlabeled polyp images. Experimental results show that our method achieves state-of-the-art performance under different label conditions on five polyp datasets. The code is available at https://github.com/pingchuan/RD-Net.
Yuxin Li 0004, Zihao Zhu 0002, Zhibin Yu 0002
CVPR5
2025 Multi-object detection and tracking algorithm for fry counting based on DV-YOLO and FryMOT
abstract
The fry counting is essential in aquaculture. Precisely counting the fry can be used to estimate the quantity of aquaculture products. Deep-learning-based multi-object tracking and counting technologies have received increasing attention in recent years. However, detecting and tracking fry in complex scenarios remains challenging due to fry’s high density and irregular movements. To address these issues, we introduce a deformable convolutional block (DCB) to capture more fine-grained features and vertex distance intersection over union (VDIoU) loss for better bounding box regression localization. Thus, our model can efficiently deal with the high-density scenarios caused by neighboring or overlapping fry. Furthermore, we introduce the enlarged intersection over union (IoU) to address the issue of unmatched tracking caused by the irregular movements of the fry. The complementary experimental results show that our methods achieve state-of-the-art performance.
Zhibin Yu 0002, Qiusheng Li, Tianning Fu
ICASSP2
2025 DeepMatch: Navigating the Complexities of Underwater Textures for Enhanced Keypoint Matching
abstract
Driven by demands for oceanic exploration, advancements in 3D visual tasks based on video frames are essential. Keypoint matching, essential for camera pose and motion, is hindered by the unique challenges of underwater imagery, such as sparse and repetitive textures. To tackle these issues, we introduce the DeepMatch framework, tailored for aquatic environments. It comprises two main components: a feature extraction network with a multi-level feature fusion module for enhanced keypoint detection in sparse textures, and an attention-based graph neural network with relative position encoding to correct matches in repetitive texture areas. Our experiments demonstrate that DeepMatch surpasses traditional and learning-based methods in underwater image matching, providing a robust solution for accurate camera pose estimation.
Zhibin Yu 0002
ICASSP4
2025 Selective Enhanced Attention and Cross-Component Interaction for Long-Term Time Series Forecasting
abstract
Transformer-based methods have demonstrated promising performance in long-term time series forecasting. However, existing methods often consider interactions among all variables when modeling cross-variable correlations, which may introduce noise and redundant features from unrelated variables. These features distract the model's attention from relevant ones and increase the risk of overfitting. Moreover, the intertwined nature of long-term trends and short-term fluctuations in time series makes it challenging to model temporal dependencies effectively. To address these challenges, we propose the Selective Enhanced Time Series Transformer (SETST), which incorporates two key modules: the Selective Enhanced Attention (SEA) module and the Decomposition Component Fusion (DCF) module. SEA selectively captures correlations between highly related variables while filtering out meaningless information, facilitating cross-variable interaction learning. DCF integrates time series decomposition with cross-component interaction, enabling effective modeling of temporal dependencies. SETST demonstrates state-of-the-art performance on five real-world datasets.
Hongchi Hao, Shouyu Ren, Hanru Li, Zhibin Yu 0002
IEEE Signal Process. Lett.4
2024 UVEB: A Large-scale Benchmark and Baseline Towards Real-World Underwater Video Enhancement
abstract
Learning-based underwater image enhancement (UIE) methods have made great progress. However, the lack of large-scale and high-quality paired training samples has become the main bottleneck hindering the development of UIE. The inter-frame information in underwater videos can accelerate or optimize the UIE process. Thus, we constructed the first large-scale high-resolution underwater video enhancement benchmark (UVEB) to promote the development of underwater vision. It contains 1,308 pairs of video sequences and more than 453,000 high-resolution with 38% Ultra-High-Definition (UHD) 4K frame pairs. UVEB comes from multiple countries, containing various scenes and video degradation types to adapt to diverse and complex underwater environments. We also propose the first supervised underwater video enhancement method, UVE-Net. UVE-Net converts the current frame information into convolutional kernels and passes them to adjacent frames for efficient inter-frame information exchange. By fully utilizing the redundant degraded information of underwater videos, UVE-Net completes video enhancement better. experiments show the effective network design and good performance of UVE-Net.
Yaofeng Xie, Lingwei Kong, Ziqiang Zheng, Zhibin Yu 0002
CVPR6
2024 UDUIE: Unpaired Domain-Irrelevant Underwater Image Enhancement
Zhibin Yu 0002
PRICAI (4)3
2024 Multispectral Semantic Segmentation for UAVs: A Benchmark Dataset and Baseline
abstract
Solidago canadensis L. is a typical invasive plant that has become a significant threat worldwide and profoundly impacts local ecosystems. An unmanned aerial vehicle (UAV)-based semantic segmentation (SS) system can help in monitoring the spread and location of Solidago canadensis L. To identify the growth range of this species with greater efficiency, we employ a high-speed multispectral camera, which provides richer color information and features with limited resolution, in conjunction with a high-quality RGB camera to construct a segmentation dataset. We construct a validated UAV multispectral (UAVM) dataset comprising 3260 pairs of calibrated RGB and multispectral images. All the images in the dataset underwent semantic annotation at a fine-grained pixel level, with 12 categories being covered. In addition, other plant categories can be employed in precision agriculture and ecological conservation. Moreover, we propose a benchmark model, UAVM semantic segmentation network (UAVMNet). With the aid of the feature alignment module and the UAVMFuse module, UAVMNet efficiently integrates multispectral and high-quality RGB image information, enhancing its ability to perform semantic segmentation tasks effectively. To the best of our knowledge, this is the first model to colearn semantic representations via high-quality RGB and paired multispectral information on a UAV platform. We conduct comprehensive experiments on the proposed UAVM dataset.
Qiusheng Li, Tianning Fu, Zhibin Yu 0002, Shuguo Chen
IEEE Trans. Geosci. Remote. Sens.4
2023 Robust Perception Under Adverse Conditions for Autonomous Driving Based on Data Augmentation
abstract
Many existing advanced deep learning-based autonomous systems have recently been used for autonomous vehicles. In general, a deep learning-based visual perception system heavily relies on visual perception to recognize and localize dynamic interest objects (e.g., pedestrians and cars) and indicative traffic signs and lights to assist autonomous vehicles in maneuvering safely. However, the performance of existing object recognition algorithms could degrade significantly under some adverse and challenging scenarios including rainy, foggy, and rainy night conditions. The raindrops, light reflection, and low illumination pose a great challenge to robust object recognition. Thus, A robust and accurate autonomous driving system has attracted growing attention from the computer vision community. To achieve robust and accurate visual perception, we target to build effective and efficient augmentation and fusion techniques based on visual perception under various adverse conditions. The unpaired image-to-image (I2I) synthesis is integrated for visual perception enhancement and effective synthesis-based augmentation. Besides, we design a two-branch architecture to utilize the information from both the original image and the enhanced image synthesized by I2I. We comprehensively and hierarchically investigate the performance improvement and limitation of the proposed system based on visual recognition tasks and network backbones. An extensive experimental analysis of various adverse weather conditions is also included. The experimental results have demonstrated the proposed system could promote the ability of autonomous vehicles for robust and accurate perception under adverse weather conditions.
Ziqiang Zheng, Yujie Cheng, Zhichao Xin, Zhibin Yu 0002
IEEE Trans. Intell. Transp. Syst.4
2022 Not every sample is efficient: Analogical generative adversarial network for unpaired image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yubo Wang 0001, Zhijian Sun
Neural Networks3
2022 One-Shot Image-to-Image Translation via Part-Global Learning With a Multi-Adversarial Framework
abstract
It is well known that humans can learn and recognize objects effectively from several limited image samples. However, learning from just a few images is still a tremendous challenge for existing main-stream deep neural networks. Inspired by analogical reasoning in the human mind, a feasible strategy is to “translate” the abundant images of a rich source domain to enrich the relevant yet different target domain with insufficient image data. To achieve this goal, we propose a novel, effective multi-adversarial framework (MA) based on part-global learning, which accomplishes the one-shot cross-domain image-to-image translation. In specific, we first devise a part-global adversarial training scheme to provide an efficient way for feature extraction and prevent discriminators from being overfitted. Then, a multi-adversarial mechanism is employed to enhance the image-to-image translation ability to unearth the high-level semantic representation. Moreover, a balanced adversarial loss function is presented, which aims to balance the training data and stabilize the training process. Extensive experiments demonstrate that the proposed approach can obtain impressive results on various datasets between two extremely imbalanced image domains and outperform state-of-the-art methods on one-shot image-to-image translation. Our code will be released with this paper athttps://github.com/zhengziqiang/OST.
Ziqiang Zheng, Zhibin Yu 0002, Haiyong Zheng, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Multim.2
2021 Learning spectral normalized adversarial systems with stacked structure for high-quality 3D object generation
abstract
Summary This paper proposes a new method for generating 3D objects based on generative adversarial networks (GANs). Recently, GANs have been used in 3D object generation, but it is still very challenging to generate high‐quality 3D objects because of the complex data distribution over 3D objects. In this paper, we propose a system based on GAN that makes the generated objects more realistic. We use multiple generators and discriminators to enhance the ability of the model for learning complex distributions. Such a stacked structure can be considered as a coarse‐to‐fine or low‐to‐high–resolution mechanism. We employ the spectral normalization technology to control the Lipschitz constant of the discriminators by literally constraining the spectral norm of each layer to get a more stable training process. In this way, the proposed model can generate realistic and high‐quality 3D objects. Moreover, our system can also recover incomplete 3D objects into complete 3D objects. Experiments demonstrate that our model performs better in the quality of the generated objects than the baselines.
Haoxu Zhang, Chenchen Qiu, Chao Wang 0022, Zhibin Yu 0002, Haiyong Zheng
Concurr. Comput. Pract. Exp.5
2021 Generative Adversarial Network with Multi-branch Discriminator for imbalanced cross-species image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yang Wu 0001, Haiyong Zheng, Minho Lee 0001
Neural Networks2
2020 Discriminative Region Proposal Adversarial Network for High-Quality Image-to-Image Translation
Chao Wang 0022, Wenjie Niu, Haiyong Zheng, Zhibin Yu 0002, Zhaorui Gu
Int. J. Comput. Vis.5
2020 KA-Ensemble: towards imbalanced image classification ensembling under-sampling and over-sampling
Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng
Multim. Tools Appl.4
2020 Depth map prediction from a single image with generative adversarial nets
Shaoyong Zhang, Chenchen Qiu, Zhibin Yu 0002, Haiyong Zheng
Multim. Tools Appl.4
2020 Fine-grained facial image-to-image translation with an attention based pipeline generative adversarial framework
Ziqiang Zheng, Chao Wang 0022, Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013
Multim. Tools Appl.6
2019 Unpaired photo-to-caricature translation on faces in the wild
Ziqiang Zheng, Chao Wang 0022, Zhibin Yu 0002, Nan Wang 0013, Haiyong Zheng
Neurocomputing3
2018 Discriminative Region Proposal Adversarial Networks for High-Quality Image-to-Image Translation
Chao Wang 0022, Haiyong Zheng, Zhibin Yu 0002, Ziqiang Zheng, Zhaorui Gu
ECCV (1)3
2018 Unsupervised pixel-wise classification for Chaetoceros image segmentation
Fei Zhou 0007, Zhaorui Gu, Haiyong Zheng, Zhibin Yu 0002
Neurocomputing5
2017 CGAN-plankton: Towards large-scale imbalanced class generation and fine-grained classification
abstract
Plankton classification is becoming critically important as people concentrate more on oceans and global environment changing. Data of plankton species naturally exhibit imbalance in their class distribution. Meanwhile, it arouses fine-grained classification challenge. Although Convolutional Neural Networks (CNNs) have human-level performance on image classification task, they tend to be biased to large classes without considering the imbalance issue. In this paper, we introduce Generative Adversarial Network (GAN) based generative model to overcome these challenges. Our proposed model consists of fully convolutional layers, and includes three parts: a generative model G, a discriminative model D and a classification model C. We train generative and discriminative models on small classes data to learn a discriminative features through D model and reduce mode missing problem to some extent. We implement classification task using shared CNN layers of D model on whole data. Experimental results show that our model significantly improved the F1 score on an imbalanced plankton dataset with well-generated plankton images.
Chao Wang 0022, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013
ICIP2
2017 Automatic plankton image classification combining multiple view features via multiple kernel learning
abstract
BACKGROUND: Plankton, including phytoplankton and zooplankton, are the main source of food for organisms in the ocean and form the base of marine food chain. As the fundamental components of marine ecosystems, plankton is very sensitive to environment changes, and the study of plankton abundance and distribution is crucial, in order to understand environment changes and protect marine ecosystems. This study was carried out to develop an extensive applicable plankton classification system with high accuracy for the increasing number of various imaging devices. Literature shows that most plankton image classification systems were limited to only one specific imaging device and a relatively narrow taxonomic scope. The real practical system for automatic plankton classification is even non-existent and this study is partly to fill this gap. RESULTS: Inspired by the analysis of literature and development of technology, we focused on the requirements of practical application and proposed an automatic system for plankton image classification combining multiple view features via multiple kernel learning (MKL). For one thing, in order to describe the biomorphic characteristics of plankton more completely and comprehensively, we combined general features with robust features, especially by adding features like Inner-Distance Shape Context for morphological representation. For another, we divided all the features into different types from multiple views and feed them to multiple classifiers instead of only one by combining different kernel matrices computed from different types of features optimally via multiple kernel learning. Moreover, we also applied feature selection method to choose the optimal feature subsets from redundant features for satisfying different datasets from different imaging devices. We implemented our proposed classification system on three different datasets across more than 20 categories from phytoplankton to zooplankton. The experimental results validated that our system outperforms state-of-the-art plankton image classification systems in terms of accuracy and robustness. CONCLUSIONS: This study demonstrated automatic plankton image classification system combining multiple view features using multiple kernel learning. The results indicated that multiple view features combined by NLMKL using three kernel functions (linear, polynomial and Gaussian kernel functions) can describe and use information of features better so that achieve a higher classification accuracy.
Haiyong Zheng, Ruchen Wang, Zhibin Yu 0002, Nan Wang 0013, Zhaorui Gu
BMC Bioinform.3
2017 Robust and automatic cell detection and segmentation from microscopic images of non-setae phytoplankton species
abstract
Saliency‐based marker‐controlled watershed method was proposed to detect and segment phytoplankton cells from microscopic images of non‐setae species. This method first improved IG saliency detection method by combining saturation feature with colour and luminance feature to detect cells from microscopic images uniformly and then produced effective internal and external markers by removing various specific noises in microscopic images for efficient performance of watershed segmentation automatically. The authors built the first benchmark dataset for cell detection and segmentation, including 240 microscopic images across multiple phytoplankton species with pixel‐wise cell regions labelled by a taxonomist, to evaluate their method. They compared their cell detection method with seven popular saliency detection methods and their cell segmentation method with six commonly used segmentation methods. The quantitative comparison validates that their method performs better on cell detection in terms of robustness and uniformity and cell segmentation in terms of accuracy and completeness. The qualitative results show that their improved saliency detection method can detect and highlight all cells, and the following marker selection scheme can remove the corner noise caused by illumination, the small noise caused by specks, and debris, as well as deal with blurred edges.
Haiyong Zheng, Nan Wang 0013, Zhibin Yu 0002, Zhaorui Gu
IET Image Process.3
2017 Understanding human intention by connecting perception and action learning in artificial agents
Zhibin Yu 0002, Minho Lee 0001
Neural Networks2
2015 Human-Robot Interaction using Intention Recognition
abstract
Recognition of human intention is an important issue in human-robot interaction research and allows a robot to respond adequately according to human's wish. In this paper, we discuss how robots can infer human intention by learning affordance, a concept used to represent the relation between an agent and its environment. Learning of the robot, to understand human and its interaction with environment, is achieved within the framework of action-perception cycle. The action-perception cycle explains how an intelligent agent learns and enhances its ability continuously by interacting with its surrounding. The proposed intention recognition and recommendation system includes several key functions such as joint attention, object recognition, affordance model, motion understanding module and so on. The experimental results show high successful recognition performance and the plausibility of the proposed system.
Zhibin Yu 0002, Jonghong Kim, Amitash Ojha, Minho Lee 0001
HAI2
2015 A Fast Training Algorithm of Multiple-Timescale Recurrent Neural Network for Agent Motion Generation
abstract
Motion understanding and regeneration are two basic aspects of human-agent interaction. One important function of agents is to represent human's activities. For better interaction with human, robot agents should not only do something following human's order, but also be able to understand or even play some actions. Multiple Timescale Recurrent Neural Networks (MTRNN) is believed to be an efficient tool for robots action generation. In our previous work, we extended the concept of MTRNN and developed Supervised MTRNN for motion recognition. In this paper, we use Conditional Restricted Boltzmann Machine (CRBM) to initialize Supervised MTRNN and accelerate the training speed of Supervised MTRNN. Experiment results show that our method can greatly increase the training speed without losing much performance.
Zhibin Yu 0002, Rammohan Mallipeddi, Minho Lee 0001
HAI1
2015 Human intention understanding based on object affordance and action classification
abstract
Intention understanding is a basic requirement for human-machine interaction. Action classification and object affordance recognition are two possible ways to understand human intention. In this study, Multiple Timescale Recurrent Neural Network (MTRNN) is adapted to analyze human action. Supervised MTRNN, which is an extension of Continuous Timescale Recurrent Neural Network (CTRNN), is used for action and intention classification. On the other hand, deep learning algorithms proved to be efficient in understanding complex concepts in complex real world environment. Stacked denoising auto-encoder (SDA) is used to extract human implicit intention related information from the observed objects. A feature based object detection method namely Speeded Up Robust Features (SURF) is also used to find the object information. Object affordance describes the interactions between agent and the environment. In this paper, we propose an intention recognition system using `action classification' and `object affordance information'. Experimental result shows that supervised MTRNN is able to use different information in different time period and improve the intention recognition rate by cooperating with the SDA.
Zhibin Yu 0002, Rammohan Mallipeddi, Minho Lee 0001
IJCNN1
2015 Deep learning of support vector machines with class probability output networks
Zhibin Yu 0002, Rhee Man Kil, Minho Lee 0001
Neural Networks2
2015 Real-time human action classification using a dynamic neural model
Zhibin Yu 0002, Minho Lee 0001
Neural Networks1
2013 Multiple Timescale Recurrent Neural Network with Slow Feature Analysis for Efficient Motion Recognition
Jihun Kim 0003, Sungmoon Jeong, Zhibin Yu 0002, Minho Lee 0001
ICONIP (2)3
2013 Supervised Multiple Timescale Recurrent Neuron Network Model for Human Action Classification
Zhibin Yu 0002, Rammohan Mallipeddi, Minho Lee 0001
ICONIP (2)1
2013 Continuous Motion Recognition Using Multiple Time Constant Recurrent Neural Network with a Deep Network Model
Zhibin Yu 0002, Minho Lee 0001
IDEAL1