EDBT 2026 Demo / reviewers in the wild / expert
Yongbin Gao
dblp:166/2757 · also Yong-Bin Gao
· DBLP profile ↗
58ranked-venue papers
5as first author
47since 2021 · last 2027
0000-0001-9930-0502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 2 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 13 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | ClauseRoute: Risk-aware clause-guided routing for retrieval-augmented legal reasoning
Zexuan Du, Chenmou Wu, Junxin Lu, Yongbin Gao, Mingxuan Chen |
Expert Syst. Appl. | 4 |
| 2026 | Monocular Unsupervised Depth Estimation of Residual Stratification Based on Ordinal Relation NetworksabstractABSTRACT Depth estimation has been widely applied in the field of computer vision, primarily using unsupervised deep neural networks, which often rely on deeper neural networks. However, the addition of layers can result in slower convergence and suboptimal performance. To overcome these issues, we introduce a novel architecture employing model distillation, wherein a teacher network enhances the learning process of a preceding student network. To improve network speed, we integrate an ordinal module in the decoder of the teacher network for weight normalization. This module can classify weights and filter out those with the lowest information content. After the weight classification is completed, as the category value increases, the necessity of useful information decreases accordingly. Furthermore, we incorporate a residual stratification module, which adapts 2D image feature extraction methods to 3D depth, facilitating finer, multi‐scale feature representation, to expand the receptive field size at each layer of the network, thereby enhancing the accuracy and robustness of depth estimation. Experimental results using the publicly available KITTI dataset demonstrate that the proposed method accelerates network training compared to the benchmark algorithm, reducing the relative squared error by 2.3% and the root‐mean‐square error by 3.3%, thus validating the effectiveness of our approach. Ye Kuang, Yongbin Gao |
IET Image Process. | 3 |
| 2026 | MIF-gaus: Monocular implicit feature-driven generalizable Gaussian splatting reconstruction
Ying Li 0020, Chenmou Wu, Jiuqing Dong, Xihe Qiu, Yongbin Gao |
Neurocomputing | 7 |
| 2026 | Continual few-shot relation extraction via multi-task balanced dual-branch network
Chenyang Shan, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014 |
Neurocomputing | 4 |
| 2026 | GeoDiffuser: A geometry-aware extension of pretrained diffusion models for consistent multi-view synthesis
Jiahao Tang, Mingxuan Chen, Ying Li 0020, Zuolei Sun, Yongbin Gao |
Neurocomputing | 6 |
| 2026 | MAFIFusion: a multi-attention and feature interaction network for infrared and visible image fusion
Haochen Yu, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014, Yadong Zhu |
Multim. Syst. | 4 |
| 2025 | GPSP-CLIP: learning generic pseudo-state prompts for flexible zero-shot anomaly detection
Weiyu Hu, Shubo Zhou, Yongbin Gao, Xueqin Jiang 0001 |
Appl. Intell. | 3 |
| 2025 | Human-object interaction detection based on adaptive contrastive learning and class-specific feature enhancement
Huanchun Peng, Kejun Xue, Xincheng Wang 0001, Yongbin Gao, Zhijun Fang 0001, Chenmou Wu |
Appl. Intell. | 4 |
| 2025 | Knowledge guided relation enhancement for human-object interaction detection
Yongbin Gao, Chenmou Wu, Shubo Zhou |
Appl. Intell. | 2 |
| 2025 | LLM-GAODE: Large-language-model augmented neural ordinary differential equation network for video nystagmography classificationabstractBenign paroxysmal positional vertigo (BPPV), a common type of vertigo with complex etiologies, is traditionally diagnosed using video nystagmography (VNG). Current automated methods lack diagnostic precision owing to subjective interpretation of eye movement characteristics. To address these challenges, we introduce a l arge l anguage m odel-augmented G ram-based a ttentive neural o rdinary d ifferential e quation ( LLM-GAODE ), an innovative and data-driven framework integrating eye-tracking technology with a Gram-based attention mechanism and a neural ordinary differential equation network to improve BPPV classification. Furthermore, when the neural network exhibits low confidence in its predictions, an LLM can supplement the process with advanced reasoning in natural language. LLM-GAODE was evaluated using an extensive VNG dataset provided by a collaborative university hospital. Results suggest that LLM-GAODE significantly outperforms existing benchmarks in trajectory classification for BPPV diagnosis. The framework enhances BPPV diagnostic accuracy and achieves state-of-the-art performance in open-source trajectory classification benchmarks. The code is available at https://github.com/XiheQiu/LLM-GAODE . Xihe Qiu, Shaojie Shi, Bin Li 0091, Xiaoyu Tan, Yongbin Gao, Shuo Li 0001 |
Knowl. Based Syst. | 5 |
| 2025 | ECKT: enhancing cross-task knowledge transfer in continual few-shot relation extraction
Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao |
J. Supercomput. | 4 |
| 2025 | Depth-guided color correction and multi-scale Retinex network for underwater image enhancement
Zhan Hu, Juan Zhang 0001, Yongbin Gao, Bo Huang 0014, Zhijun Fang 0001 |
Vis. Comput. | 3 |
| 2024 | Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation
Anjie Wang, Xujun Wei, Mingxuan Chen, Yongbin Gao, Zhijun Fang 0001, Siwei Ma 0001 |
ICONIP (9) | 5 |
| 2024 | Dynamic Convolution Based Intelligent Algorithm for YOLOv5 Underwater Target DetectionabstractWith the growing importance of marine resources and the increasing demand for exploration of underwater environments, underwater target detection technology has become one of the key technologies in the fields of ocean engineering, underwater archaeology, and intelligent agriculture. However, due to the complexity and uncertainty of underwater environments, such as light attenuation, water turbidity, and dynamic changes of water currents, current target detection methods often perform poorly in underwater scenes. To solve the corresponding problems, this paper proposes the YOLOv5_OD_Conv model, which aims to improve the accuracy and generalisation ability of the model by introducing the OD_Conv full-dimensional dynamic convolution in the YOLOv5 Neck part. Simulation and experimental results show that the proposed method increases the detection accuracy P by 1.05%, the precision mAP0.5 by 1.5%, and the recall R by 0.43% compared to YOLOv5s. The improvement of detection effect is obvious, which proves the effectiveness of the method. Jialing Jiang, Bo Huang 0014, Zhijun Fang 0001, Yongbin Gao |
SoMeT | 4 |
| 2024 | QLDT: adaptive Query Learning for HOI Detection via vision-language knowledge Transfer
Xincheng Wang 0001, Yongbin Gao, Chenmou Wu, Mingxuan Chen, Honglei Ma |
Appl. Intell. | 2 |
| 2024 | Adaptive multimodal prompt for human-object interaction with local feature enhanced transformer
Kejun Xue, Yongbin Gao, Zhijun Fang 0001, Mingxuan Chen, Chenmou Wu |
Appl. Intell. | 2 |
| 2024 | Transformer-based end-to-end attack on text CAPTCHAs with triplet deep attention
Yujie Xiong, Chunming Xia, Yongbin Gao |
Comput. Secur. | 4 |
| 2024 | A lightweight RGB superposition effect adjustment network for low-light image enhancement and denoising
Pei-Dong Chen, Juan Zhang 0001, Yongbin Gao, Zhijun Fang 0001, Jenq-Neng Hwang |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Entity alignment based on informative neighbor sampling and multi-embedding graph matching
Yongbin Gao, Zhijun Fang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Self-Enhanced Attention for Image CaptioningabstractAbstract Image captioning, which involves automatically generating textual descriptions based on the content of images, has garnered increasing attention from researchers. Recently, Transformers have emerged as the preferred choice for the language model in image captioning models. Transformers leverage self-attention mechanisms to address gradient accumulation issues and eliminate the risk of gradient explosion commonly associated with RNN networks. However, a challenge arises when the input features of the self-attention mechanism belong to different categories, as it may result in ineffective highlighting of important features. To address this issue, our paper proposes a novel attention mechanism called Self-Enhanced Attention (SEA), which replaces the self-attention mechanism in the decoder part of the Transformer model. In our proposed SEA, after generating the attention weight matrix, it further adjusts the matrix based on its own distribution to effectively highlight important features. To evaluate the effectiveness of SEA, we conducted experiments on the COCO dataset, comparing the results with different visual models and training strategies. The experimental results demonstrate that when using SEA, the CIDEr score is significantly higher compared to the scores obtained without using SEA. This indicates the successful addressing of the challenge of effectively highlighting important features with our proposed mechanism. Qingyu Sun, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao |
Neural Process. Lett. | 4 |
| 2024 | F-SCP: An automatic prompt generation method for specific classes based on visual language pre-training models
Baihong Han, Zhijun Fang 0001, Hamido Fujita, Yongbin Gao |
Pattern Recognit. | 5 |
| 2024 | Enhancing Few-Shot Out-of-Distribution Detection With Pre-Trained Model FeaturesabstractEnsuring the reliability of open-world intelligent systems heavily relies on effective out-of-distribution (OOD) detection. Despite notable successes in existing OOD detection methods, their performance in scenarios with limited training samples is still suboptimal. Therefore, we first construct a comprehensive few-shot OOD detection benchmark in this paper. Remarkably, our investigation reveals that Parameter-Efficient Fine-Tuning (PEFT) techniques, such as visual prompt tuning and visual adapter tuning, outperform traditional methods like fully fine-tuning and linear probing tuning in few-shot OOD detection. Considering that some valuable information from the pre-trained model, which is conducive to OOD detection, may be lost during the fine-tuning process, we reutilize features from the pre-trained models to mitigate this issue. Specifically, we first propose a training-free approach, termed uncertainty score ensemble (USE). This method integrates feature-matching scores to enhance existing OOD detection methods, significantly narrowing the gap between traditional fine-tuning and PEFT techniques. However, due to its training-free property, this method is unable to improve in-distribution accuracy. To this end, we further propose a method called Domain-Specific and General Knowledge Fusion (DSGF) to improve few-shot OOD detection performance and ID accuracy under different fine-tuning paradigms. Experiment results demonstrate that DSGF enhances few-shot OOD detection across different fine-tuning strategies, shot settings, and OOD detection methods. We believe our work can provide the research community with a novel path to leveraging large-scale visual pre-trained models for addressing FS-OOD detection. The code will be released. Jiuqing Dong, Yongbin Gao, Zhijun Fang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | TV-Net: A Structure-Level Feature Fusion Network Based on Tensor Voting for Road Crack SegmentationabstractPavement cracks are a common and significant problem for intelligent pavement maintainment. However, the features extracted in pavement images are often texture-less, and noise interference can be high. Segmentation using traditional convolutional neural network training can lose feature information when the network depth goes larger, which makes accurate prediction a challenging topic. To address these issues, we propose a new approach that features an enhanced tensor voting module and a customized pixel-level pavement crack segmentation network structure, called TV-Net. We optimize the tensor voting framework and find the relationship between tensor scale factors and crack distributions. A tensor voting fusion module is introduced to enhance feature maps by incorporating significant domain maps generated by tensor voting. Additionally, we propose a structural consistency loss function to improve segmentation accuracy and ensure consistency with the structural characteristics of the cracks obtained through tensor voting. The sufficient experimental analysis demonstrates that our method outperforms existing mainstream pixel-level segmentation networks on the same road crack dataset. Our proposed TV-Net has an excellent performance in avoiding noise interference and strengthening the structure of the fracture site of pavement cracks. Code is available at https://github.com/sues-vision/ TV-Net.git. Wenwen Zheng, Zhijun Fang 0001, Yongbin Gao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Monocular Depth and Ego-motion Estimation with Scale Based on Superpixel and Normal ConstraintsabstractThree-dimensional perception in intelligent virtual and augmented reality (VR/AR) and autonomous vehicles (AV) applications is critical and attracting significant attention. The self-supervised monocular depth and ego-motion estimation serves as a more intelligent learning approach that provides the required scene depth and location for 3D perception. However, the existing self-supervised learning methods suffer from scale ambiguity, boundary blur, and imbalanced depth distribution, limiting the practical applications of VR/AR and AV. In this article, we propose a new self-supervised learning framework based on superpixel and normal constraints to address these problems. Specifically, we formulate a novel 3D edge structure consistency loss to alleviate the boundary blur of depth estimation. To address the scale ambiguity of estimated depth and ego-motion, we propose a novel surface normal network for efficient camera height estimation. The surface normal network is composed of a deep fusion module and a full-scale hierarchical feature aggregation module. Meanwhile, to realize the global smoothing and boundary discriminability of the predicted normal map, we introduce a novel fusion loss which is based on the consistency constraints of the normal in edge domains and superpixel regions. Experiments are conducted on several benchmarks, and the results illustrate that the proposed approach outperforms the state-of-the-art methods in depth, ego-motion, and surface normal estimation. Junxin Lu, Yongbin Gao, Jieyu Chen, Jenq-Neng Hwang, Hamido Fujita, Zhijun Fang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Self-Supervised Learning of Depth and Ego-Motion for 3D Perception in Human Computer Interactionabstract3D perception of depth and ego-motion is of vital importance in intelligent agent and Human Computer Interaction (HCI) tasks, such as robotics and autonomous driving. There are different kinds of sensors that can directly obtain 3D depth information. However, the commonly used Lidar sensor is expensive, and the effective range of RGB-D cameras is limited. In the field of computer vision, researchers have done a lot of work on 3D perception. While traditional geometric algorithms require a lot of manual features for depth estimation, Deep Learning methods have achieved great success in this field. In this work, we proposed a novel self-supervised method based on Vision Transformer (ViT) with Convolutional Neural Network (CNN) architecture, which is referred to as ViT-Depth . The image reconstruction losses computed by the estimated depth and motion between adjacent frames are treated as supervision signal to establish a self-supervised learning pipeline. This is an effective solution for tasks that need accurate and low-cost 3D perception, such as autonomous driving, robotic navigation, 3D reconstruction, and so on. Our method could leverage both the ability of CNN and Transformer to extract deep features and capture global contextual information. In addition, we propose a cross-frame loss that could constrain photometric error and scale consistency among multi-frames, which lead the training process to be more stable and improve the performance. Extensive experimental results on autonomous driving dataset demonstrate the proposed approach is competitive with the state-of-the-art depth and motion estimation methods. Shanbao Qiao, Naixue Xiong, Yongbin Gao, Zhijun Fang 0001, Juan Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Fast-Fading Channel and Power Optimization of the Magnetic Inductive Cellular NetworkabstractThe cellular network of magnetic Induction (MI) communication holds promise in long-distance underground environments. In the traditional MI communication, there is no fast-fading channel since the MI channel is treated as a quasi-static channel. However, for the vehicle (mobile) MI (VMI) communication, the unpredictable antenna vibration brings the remarkable fast-fading. As such fast-fading cannot be modeled by the central limit theorem, it differs radically from other wireless fast-fading channels. Unfortunately, few studies focus on this phenomenon. In this paper, using a novel space modeling based on the electromagnetic field theorem, we propose a 3-dimension model of the VMI antenna vibration. By proposing “conjugate pseudo-piecewise functions” and boundary$p(x)$distribution, we derive the cumulative distribution function (CDF), probability density function (PDF) and the expectation of the VMI fast-fading channel. We also theoretically analyze the effects of the VMI fast-fading on the network throughput, including the VMI outage probability which can be ignored in the traditional MI channel study. We draw several intriguing conclusions different from those in wireless fast-fading studies. For instance, the fast-fading brings more uniformly distributed channel coefficients. Finally, we propose the power control algorithm using the non-cooperative game and multiagent Q-learning methods to optimize the throughput of the cellular VMI network. Simulations validate the derivation and the proposed algorithm. Honglei Ma, Erwu Liu, Zhijun Fang 0001, Rui Wang 0001, Yongbin Gao, Dongming Zhang 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Gram-based Attentive Neural Ordinary Differential Equations Network for Video Nystagmography ClassificationabstractVideo nystagmography (VNG) is the diagnostic gold standard of benign paroxysmal positional vertigo (BPPV), which requires medical professionals to examine the direction, frequency, intensity, duration, and variation in the strength of nystagmus on a VNG video. This is a tedious process heavily influenced by the doctor’s experience, which is error-prone. Recent automatic VNG classification methods approach this problem from the perspective of video analysis without considering medical prior knowledge, resulting in unsatisfactory accuracy and limited diagnostic capability for nystagmographic types, thereby preventing their clinical application. In this paper, we propose an end-to-end data-driven novel BPPV diagnosis framework (TC-BPPV) by considering this problem as an eye trajectory classification problem due to the disease’s symptoms and experts’ prior knowledge. In this framework, we utilize an eye movement tracking system to capture the eye trajectory and propose the Gram-based attentive neural ordinary differential equations network (Gram-AODE) to perform classification. We validate our framework using the VNG dataset provided by the collaborative university hospital and achieve state-of-the-art performance. We also evaluate Gram-AODE on multiple open-source benchmarks to demonstrate its effectiveness in trajectory classification. Code is available at https://github.com/XiheQiu/Gram-AODE. Xihe Qiu, Shaojie Shi, Xiaoyu Tan, Chao Qu, Zhijun Fang 0001, Yongbin Gao, Peixia Wu |
ICCV | 7 |
| 2023 | Depth Estimation of Multi-Modal Scene Based on Multi-Scale ModulationabstractAs multimodal information is complementary, effectively utilizing scene multimodal information has become an increasingly important research topic for many scholars. This paper proposes a novel multi-scale global learning strategy that utilizes both echo and visual modal data as inputs to estimate scene depth. The framework involves constructing a multi-scale feature extraction method using pyramid pooling modules to aggregate contextual information from different regions and improve global information acquisition ability. Furthermore, a recurrent multi-scale feature modulation module is introduced to generate more semantic and accurate spatial representations in each iteration update process. Additionally, a multi-scale fusion method is constructed for the fusion of echo and visual modalities. The proposed method's superior performance is demonstrated through sufficient experiments conducted on the Replica dataset. Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Gaofeng Cao, Siwei Ma 0001 |
ICIP | 4 |
| 2023 | Spatial graph attention network-based object tracking with adaptive cosine window
Liuyi Fan, Bo Huang 0014, Juan Zhang 0001, Yongbin Gao |
Appl. Intell. | 5 |
| 2023 | Online object-level SLAM with dual bundle adjustment
Yongbin Gao, Zhijun Fang 0001 |
Appl. Intell. | 2 |
| 2023 | A fine-grained causality extraction model incorporating relative location coding
Weibing Wan, Yongbin Gao, Chen Shao |
Appl. Intell. | 3 |
| 2023 | Transformer networks with adaptive inference for scene graph generation
Yini Wang, Yongbin Gao, Ruyan Guo, Weibing Wan, Shuqun Yang, Bo Huang 0014 |
Appl. Intell. | 2 |
| 2023 | MACFNet: multi-attention complementary fusion network for image denoising
Jiaolong Yu, Juan Zhang 0001, Yongbin Gao |
Appl. Intell. | 3 |
| 2023 | Monocular 3-D Object Detection Based on Depth-Guided Local Convolution for Smart Payment in D2D Systemsabstract3-D object detection from mobile phones in Device-to-Device (D2D) system provides a new smart payment tool for the next generation of fintech, which is more flexible and efficient than the traditional barcode. In this article, we propose a monocular 3-D object detection method based on depth-guided local convolution. The method combines the information of RGB image mode and depth mode by using a convolution kernel through depth image and works on a single RGB image locally. According to the multiscale input information, the convolution kernel is adaptively adjusted to capture the target objects of different scales, so as to improve the performance of 3-D object detection. In addition, we use the soft-non-maximum suppression algorithm instead of traditional non-maximum suppression to select the best prediction box. In order to further improve the accuracy of 3-D object detection, the depth estimation network and 3-D object detection network are jointly trained in this method to make the two networks constrain each other and achieve the best performance. Jun Li 0036, Yongbin Gao, Huixing Wang, Yier Yan, Bo Huang 0014, Jun Zhang 0004, Wei Wang 0030 |
IEEE Internet Things J. | 3 |
| 2023 | An angular shrinkage BERT model for few-shot relation extraction with none-of-the-above detection
Junwen Wang, Yongbin Gao, Zhijun Fang 0001 |
Pattern Recognit. Lett. | 2 |
| 2023 | LFT-Net: Local Feature Transformer Network for Point Clouds Analysisabstract6G network enables the rapid connection of autonomous vehicles, the generated internet of vehicles establishes a large-scale point cloud, which requires automatic point cloud analysis to build an intelligent transportation system in terms of the 3D object detection and segmentation. Recently, a great variety of deep convolution networks have been proposed for 3D data analysis, making significant progress in the application of deep learning in 3D computer vision. Inspired by the application of transformer network in 2D computer visual tasks, and in order to increase the expression ability of local fine-grained features, we propose an effective local feature transformer network to learn local feature information and correlations between point clouds. Our network is adaptive to the arrangement of set elements through transformer module, so it is suitable for the feature extraction of local point clouds. In addition, experimental results demonstrate that our LFT-network outperforms the state-of-the-art in 3D model classification tasks on ModelNet40 dataset and segmentation tasks on S3DIS dataset. Yongbin Gao, Xuebing Liu, Jun Li 0036, Zhijun Fang 0001, Kazi Mohammed Saidul Huq |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Joint Optimization of Depth and Ego-Motion for Intelligent Autonomous VehiclesabstractThe three-dimensional (3D) perception of autonomous vehicles is crucial for localization and analysis of the driving environment, while it involves massive computing resources for deep learning, which can’t be provided by vehicle-mounted devices. This requires the use of seamless, reliable, and efficient massive connections provided by the 6G network for computing in the cloud. In this paper, we propose a novel deep learning framework with 6G enabled transport system for joint optimization of depth and ego-motion estimation, which is an important task in 3D perception for autonomous driving. A novel loss based on feature map and quadtree is proposed, which uses feature value loss with quadtree coding instead of photometric loss to merge the feature information at the texture-less region. Besides, we also propose a novel multi-level V-shaped residual network to estimate the depths of the image, which combines the advantages of V-shaped network and residual network, and solves the problem of poor feature extraction results that may be caused by the simple fusion of low-level and high-level features. Lastly, to alleviate the influence of image noise on pose estimation, we propose a number of parallel sub-networks that use RGB image and its feature map as the input of the network. Experimental results show that our method significantly improves the quality of the depth map and the localization accuracy and achieves the state-of-the-art performance. Yongbin Gao, Jun Li 0036, Zhijun Fang 0001, Saba Al-Rubaye, Yier Yan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Sparse-attentive meta temporal point process for clinical decision support
Yajun Ru, Xihe Qiu, Xiaoyu Tan, Yongbin Gao, Yaochu Jin |
Neurocomputing | 5 |
| 2022 | Depth Estimation Using a Self-Supervised Network Based on Cross-Layer Feature Fusion and the Quadtree ConstraintabstractDepth estimation from a camera is an important task for 3D perception. Recently, without using the labeled ground truth of depth map, a self-supervised deep learning network can use relative pose to synthesize the target image from the reference image, and the photometric error between synthesized reference image and real one is used as self-supervisory signal. In this paper, we propose a novel self-supervised depth estimation network, which takes advantage of the quadtree constraint to optimize the depth estimation network. Based on the quadtree constraint, the photometric loss and depth loss of quadtree are proposed. In order to solve the problem that multiple depth values in repeated structures and uniform texture regions can cause relatively low photometric loss, we use quadtree-based photometric loss, which calculates the averaged photometric loss in quadtree blocks instead of the pixel-wise loss. For the problem of imbalanced depth distribution, we use quadtree depth loss, which constrains the depth inconsistency within quadtree blocks. The depth estimation network is composed of deep fusion module and cross-layer feature fusion module, which can better extract the feature information of RGB image and sparse keypoints depths, and makes full use of the detail information of the shallow feature map and the semantic information of the deep feature map to enrich the feature information extraction. Experimental results demonstrate that our method outperforms the state-of-the-art approaches of depth estimation. Yongbin Gao, Zhijun Fang 0001, Yuming Fang 0001, Hamido Fujita, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Point AE-DCGAN: A deep learning model for 3D point cloud lossy geometry compressionabstract3D point cloud has been widely applied in virtual reality and augmented reality. A complex 3D scene always needs a large number of the point cloud to represent and demands a lot of space to store. Thus, point cloud compression becomes a crucial issue to research. In this paper, we propose a novel lossy geometric compression method of autoencoder based on DCGAN optimization. This method can reconstruct a high-quality point cloud and solves a large area of missing points in the process of compression and decompression. To improve the point cloud codec performance, we propose a multi-scale 3D deconvolution hopping connection structure to obtain a better-quality reconstructed point cloud under low bit rates. Our approach is the first GAN-based point cloud compression algorithm to our knowledge. Compared with state-of-the-art methods on the MVUB dataset, our approach achieves a better rate-distortion performance and visual quality. Zhijun Fang 0001, Yongbin Gao, Siwei Ma 0001, Yaochu Jin, Anjie Wang |
DCC | 3 |
| 2021 | The analysis of isolation measures for epidemic control of COVID-19
Bo Huang 0014, Yongbin Gao, Guohui Zeng, Juan Zhang 0001, Jin Liu 0016 |
Appl. Intell. | 3 |
| 2021 | Celiac trunk segmentation incorporating with additional contour constraint
Xianhua Tang, Bo Huang 0014, Qingping Cai, Ziran Wei, Yongbin Gao, Huilin Tong, Pan Liang, Cengsi Zhong |
Appl. Intell. | 5 |
| 2021 | Automatic coronary artery segmentation algorithm based on deep learning and digital image processing
Yongbin Gao, Zhijun Fang 0001 |
Appl. Intell. | 2 |
| 2021 | SAT-Net: a side attention network for retinal image segmentation
Huilin Tong, Zhijun Fang 0001, Ziran Wei, Qingping Cai, Yongbin Gao |
Appl. Intell. | 5 |
| 2021 | Line-based visual odometry using local gradient fitting
Junxin Lu, Zhijun Fang 0001, Yongbin Gao, Jieyu Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | 3D reconstruction with auto-selected keyframes based on depth completion correction and pose fusion
Yongbin Gao, Zhijun Fang 0001, Shuqun Yang |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Photometric transfer for direct visual odometry
Kaiying Zhu, Zhijun Fang 0001, Yongbin Gao, Hamido Fujita, Jenq-Neng Hwang |
Knowl. Based Syst. | 4 |
| 2020 | Feature fusion network based on attention mechanism for 3D semantic segmentation of point clouds
Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014, Cengsi Zhong, Ruoxi Shang |
Pattern Recognit. Lett. | 3 |
| 2020 | Adversarial Learning for Joint Optimization of Depth and Ego-MotionabstractIn recent years, supervised deep learning methods have shown a great promise in dense depth estimation. However, massive high-quality training data are expensive and impractical to acquire. Alternatively, self-supervised learning-based depth estimators can learn the latent transformation from monocular or binocular video sequences by minimizing the photometric warp error between consecutive frames, but they suffer from the scale ambiguity problem or have difficulty in estimating precise pose changes between frames. In this paper, we propose a joint self-supervised deep learning pipeline for depth and ego-motion estimation by employing the advantages of adversarial learning and joint optimization with spatial-temporal geometrical constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with an absolute scale and a pre-trained pose network serves as a good starting point for direct visual odometry (DVO). DVO optimization based on spatial geometric constraints can result in a fine-grained ego-motion estimation with the additional backpropagation signals provided to the depth estimation network. Finally, the spatial and temporal domain-based reconstructed views are concatenated, and the iterative coupling optimization process is implemented in combination with the adversarial learning for accurate depth and precise ego-motion estimation. The experimental results show superior performance compared with state-of-the-art methods for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach. Anjie Wang, Zhijun Fang 0001, Yongbin Gao, Songchao Tan, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 3 |
| 2019 | Unsupervised Learning of Depth and Ego-Motion with Spatial-Temporal Geometric ConstraintsabstractIn this paper, we propose an unsupervised joint deep learning pipeline for depth and ego-motion estimation that explicitly incorporated with traditional spatial-temporal geometric constraints. The stereo reconstruction error provides the spatial geometric constraint to estimate the absolute scale depth. Meanwhile, the depth map with absolute scale and a pre-trained pose network serve as a good starting point for direct visual odometry (DVO), resulting in a fine-grained ego-motion estimation with the additional back-propagation signals provided to the depth estimation network. The proposed joint training pipeline enables an iterative coupling optimization process for accurate depth and precise ego-motion estimation. The experimental results show the state-of-the-art performance for monocular depth and ego-motion estimation on the KITTI dataset and a great generalization ability of the proposed approach. Anjie Wang, Yongbin Gao, Zhijun Fang 0001, Shanshe Wang, Siwei Ma 0001, Jenq-Neng Hwang |
ICME | 2 |
| 2019 | DD-CycleGAN: Unpaired image dehazing via Double-Discriminator Cycle-Consistent Generative Adversarial Network
Jingming Zhao, Juan Zhang 0001, Zhi Li 0049, Jenq-Neng Hwang, Yongbin Gao, Zhijun Fang 0001, Bo Huang 0014 |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Pose-invariant features and personalized correspondence learning for face recognition
Yongbin Gao, Hyo Jong Lee |
Neural Comput. Appl. | 1 |
| 2019 | Transfer deep feature learning for face sketch recognition
Weiguo Wan, Yongbin Gao, Hyo Jong Lee |
Neural Comput. Appl. | 2 |
| 2019 | Unsupervised learning of depth estimation based on attention model and global pose optimization
Renyue Dai, Yongbin Gao, Zhijun Fang 0001, Anjie Wang, Juan Zhang 0001, Cengsi Zhong |
Signal Process. Image Commun. | 2 |
| 2018 | Learning warps based similarity for pose-unconstrained face recognition
Yongbin Gao, Hyo Jong Lee |
Multim. Tools Appl. | 1 |
| 2017 | A New Code Generation Method for Software Engineering: From Requirements Model to Source Codeabstractthe existing software engineering techniques for software synthesis from requirements model to source code have many limitations. The synthesis approach shows that these limitations focused on the refinement relationship between the requirements specification and the desired system. We have to propose a new approach for code generation to overcome such limitations, i.e. refining the software behaviors in requirements model to code, distinguishing function information and architecture information from requirements model, among others. Hence, in this thesis we aim at the problems that how to modeling based on software behaviors, how to delimitate the system architecture and so on. And we also will show a sample, ultimately, to demonstrate our approach. Meanwhile, some additional techniques for synthesis to ensure the correctness of the source code will be recommended. Bo Huang 0014, Zhijun Fang 0001, Yongbin Gao |
SoMeT | 5 |
| 2016 | A general effective rate control system based on matching measurement and inter-quantizer
Zhijun Fang 0001, Yongbin Gao, Naixue Xiong, Athanasios V. Vasilakos, Yuming Fang 0001 |
Inf. Sci. | 2 |
| 2015 | Cross-pose face recognition based on multiple virtual views and alignment errorabstractAlthough studied for decades, effective face recognition remains difficult to accomplish on account of occlusions and pose and illumination variations. Pose variance is a particular challenge in face recognition. Effective local descriptors have been proposed for frontal face recognition. When these descriptors are directly applied to cross-pose face recognition, the performance significantly decreases. To improve the descriptor performance for cross-pose face recognition, we propose a face recognition algorithm based on multiple virtual views and alignment error. First, warps between poses are learned using the Lucas–Kanade algorithm. Based on these warps, multiple virtual profile views are generated from a single frontal face, which enables non-frontal faces to be matched using the scale-invariant feature transform (SIFT) algorithm. Furthermore, warps indicate the correspondence between patches of two faces. A two-phase alignment error is proposed to obtain accurate warps, which contain pose alignment and individual alignment. Correlations between patches are considered to calculate the alignment error of two faces. Finally, a hybrid similarity between two faces is calculated; it combines the number of matched keypoints from SIFT and the alignment error. Experimental results show that our proposed method achieves better recognition accuracy than existing algorithms, even when the pose difference angle was greater than 30°. Yongbin Gao, Hyo Jong Lee |
Pattern Recognit. Lett. | 1 |