EDBT 2026 Demo / reviewers in the wild / expert
Yehua Ling
dblp:278/9068
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-1196-6615ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain adaptive object detection via CLIP-space guidance and LoRA fine-tuning
Enze Qi, Kan Chang, Qingzhi Zhang, Xueyu Zhang, Yehua Ling, Yujian Yuan, Zan Gao 0001 |
Expert Syst. Appl. | 6 |
| 2026 | UECNet: A unified framework for exposure correction utilizing region-level prompts
Shucheng Xia, Kan Chang, Xuxin Tai, Yehua Ling, Yujian Yuan, Zan Gao 0001 |
Knowl. Based Syst. | 6 |
| 2026 | A Self-Attention-Based LiDAR Point Cloud Compression Framework in Autonomous Driving EnvironmentsabstractLight detection and ranging (LiDAR) sensors are crucial for autonomous vehicles to accurately perceive the surrounding environment. However, the sparsity and irregularity of large-scale LiDAR point clouds (LPCs) bring challenges for storage and transmission. Meanwhile, existing works usually adopt insufficient context and bring intolerable computation complexity, especially for high-precision LPC reconstruction. To address these problems, we propose a novel self-attention-based framework for LPC compression and reconstruction in autonomous driving environments. Specifically, our approach employs a robust backbone for octree-based feature extraction, which can be pretrained and easily extended to various tasks, thereby reducing the need for extensive task-specific architectural modifications. The backbone constructs node sequences of octree by nonoverlapping context windows and shares the result of a multihead self-attention (MSA) operation among them. Considering the similarity in features among sibling nodes, we design a locally enhanced module for exploiting sibling features and a positional encoding generator for enhancing the translation invariance of the octree node sequence. During postprocessing, we further propose an offset prediction model to reduce coordinate distortions caused by voxelization. Experimental results indicate that compared to the benchmark geometry-based point cloud compression (GPCC), our approach achieves gains of up to 54.4% for geometry and 6.8% for intensity, while compared to the attention-based baseline, we achieve up to 99% reduction in coding time. We believe that our approach effectively mines the spatial geometric features in LPCs and has low coupling for specific tasks, which will boost the related applications from algorithm optimization to industrial products. Mingyue Cui, Junhua Long, Mingjian Feng, Juncheng Tao, Yuyang Zhong, Yehua Ling, Daosong Hu, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | GAEM: Graph-Driven Attention-Based Entropy Model for LiDAR Point Cloud CompressionabstractHigh-quality LiDAR point cloud (LPC) coding is essential for efficiently transmitting and storing the vast amounts of data required for accurate 3D environmental representation. The Octree-based entropy coding framework has emerged as the predominant method, however, previous study usually overly relies on large-scale attention-based context prediction to encode Octree nodes, overlooking the inherent correlational properties of this structure. In this paper, we propose a novel Graph-driven Attention-based Entropy Model (GAEM), which adopts partitioned graph attention mechanisms to uncover contextual dependencies among neighboring nodes. Different from the Cartesian coordinate-based coding mode with higher redundancy, GAEM uses the multi-level spherical Octree to organize point clouds, improving the quality of LPC reconstruction. GAEM combines graph convolution for node feature embedding and grouped-graph attention for exploiting dependency among contexts, which preserves performance in low-computation using localized nodes. Besides, to further increase the receptive field, we design a high-resolution cross-attention module introducing sibling nodes. Experimental results show that our method achieves state-of-the-art performance on the LiDAR benchmark SemanticKITTI and MPEG-specified dataset Ford, compared to all baselines. Compared to the benchmark GPCC, our method achieves gains of up to 53.9% and 53.6% on SemanticKITTI and Ford while compared to the sibling-introduced methods, we achieve up to 42.3% and 44.7% savings in encoding/decoding time. In particular, our GAEM allows for extension to downstream tasks (i.e.,vehicle detection and semantic segmentation), further demonstrating the practicality of the method. Mingyue Cui, Yuyang Zhong, Mingjian Feng, Junhua Long, Yehua Ling, Kai Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | DAPCC: Diverse Attention-Based Entropy Model for Dynamic LiDAR Point Cloud CompressionabstractLiDAR point cloud (LPC) compression is an indispensable component for 3D vision tasks, especially for dynamic point clouds. However, the existing methods based on traditional spatial-temporal attention are immature, causing little improvement in inter-frame feature extraction. In this paper, we propose Diverse Attention-based Point Cloud Compression (DAPCC), an LPC compression entropy model combining aggregation embedding modules for temporal point matching and spatial-temporal attention blocks for dynamic Octree node encoding, which can effectively utilize the change information of dynamic point clouds. Specifically, we first introduce aggregation embedding to match the Octree sequences from two sweeps to establish temporal correlation. To effectively capture the feature details, we further design local and global combined attention for the spatial-temporal information of point clouds which can focus on the whole context. Finally, we organize a symmetric MLP module capable of strengthening vital features. We conduct experiments of static and dynamic compression on both indoor/outdoor point cloud benchmark datasets (i.e., ScanNet, SemanticKITTI, and MPEG Common Test Conditions (CTC) Category 3 datasets) and downstream applications (i.e., vehicle detection and semantic segmentation). Compared with the previous state-of-the-art methods, our method achieves up to 14.7% bpp and 45% decoding time savings and adapts to the downstream tasks with almost no impact on performance. Mingyue Cui, Yuyang Zhong, Mingjian Feng, Yehua Ling, Junhua Long, Jinhong Xia, Kai Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Accelerated Optimization for Simulation of Brain Spiking Neural Network on GPGPUs
Fangzhou Zhang, Mingyue Cui, Jiakang Zhang, Yehua Ling, Kai Huang 0001 |
ICA3PP (6) | 4 |
| 2023 | Real-Time Instance Segmentation and Tip Detection for Neuroendoscopic Surgical Instruments
Rihui Song, Silu Guo, Yehua Ling, Jin Gong, Kai Huang 0001 |
ICONIP (10) | 4 |
| 2023 | Dense Depth Estimation for Monocular Endoscope Robot with an Adaptive BaselineabstractDepth information is useful to surgeons and surgical assistance systems. However, it is a challenging task to estimate the depth of various surgical scenes based on a monocular endoscope. We propose a depth estimation approach for a monocular endoscope with a stereo matching algorithm. The monocular endoscope is moved horizontally by a robotic endoscope holder to simulate a stereo vision system. The main challenge is how to obtain a proper baseline for better depth information generation as the depth range of a surgical scene is unknown beforehand. We design a baseline evaluation and selection algorithm to search for suitable baselines for surgical scenes with different depth ranges. Experimental results show that our approach improves the average accuracy of different depth scenarios by 10.8% when the error range is 2mm. Rihui Song, Zhidong Tan, Hongli Liang, Yehua Ling, Gang Chen 0023, Kai Huang 0001, Jin Gong |
SMC | 4 |
| 2023 | An efficient real-time accelerator for high-accuracy DNN-based optical flow estimation in FPGA
Yuanxing Yan, Yehua Ling, Kai Huang 0001, Gang Chen 0023 |
J. Syst. Archit. | 2 |
| 2023 | A robust and real-time DNN-based multi-baseline stereo accelerator in FPGAs
Yehua Ling, Haitao Meng, Gang Chen 0023 |
J. Syst. Archit. | 3 |
| 2022 | FlowAcc: Real-Time High-Accuracy DNN-based Optical Flow Accelerator in FPGAabstractRecently, accelerator architectures have been designed to use deep neural networks (DNNs) to accelerate computer vision tasks, possessing the advantages of both accuracy and speed. Optical flow accelerator is however not among these architectures that DNNs have been successfully deployed. Existing hardware accelerators for optical flow estimation are all designed for classic methods and generally perform poorly in estimated accuracy. In this paper, we present FlowAcc, a dedicated hardware accelerator for DNN-based optical flow estimation, adopting a pipelined hardware design for real-time processing of image streams. We design an efficient multiplexing binary neural network (BNN) architecture for pyramidal feature extraction to significantly reduce the hardware cost and make it independent of the pyramid level number. Furthermore, efficient hamming distance calculation and competent flow regularization are utilized for hierarchical optical flow estimation to greatly improve the system efficiency. Comprehensive experimental results demonstrate that FlowAcc achieves state-of-the-art estimation accuracy and real-time performance on the Middlebury dataset when compared with the existing optical flow accelerators. Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023 |
DATE | 1 |
| 2022 | Ultra-Flow: An Ultra-fast and High-quality Optical Flow Accelerator with Deep Feature Matching on FPGAabstractDense and accurate optical flow estimation is an important requirement for dynamic scene perception in autonomous systems. However, most of the existing FPGA accelerators are based on classic methods, which cannot deal with large displacements of moving objects in ultra-fast scenes. In this paper, we present Ultra-Flow, an ultra-fast pipelined architecture for efficient optical flow estimation and refinement. Ultra-Flow utilizes binary neural networks to generate the robust feature map, on which hierarchical matching is directly performed. Therefore, multiple usages of neural networks at hierarchical levels can be avoided to achieve hardware efficiency in Ultra-Flow. Optimizations, including local flow regularization and enhanced matching, are further used to improve the throughput and refine the optical flow to obtain higher accuracy. Evaluation results show that, compared to state-of-the-art FPGA accelerators, Ultra-Flow achieves leading accuracy in the Middlebury sequences at ultra-fast processing speed up to 687.92 frames/s for 640 × 480 pixel images. Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023 |
FPL | 1 |
| 2022 | Lite-Stereo: A Resource-Efficient Hardware Accelerator for Real-Time High-Quality Stereo Estimation Using Binary Neural NetworkabstractStereo estimation plays a key role in many autonomous systems, such as robotics and self-driving cars. Recent work on StereoEngine, an FPGA-based accelerator for deep neural network (DNN)-based stereo estimation, has been demonstrated as a promising solution to achieve both real-time and high accuracy performance for depth sensing. However, this solution still suffers from over-utilizing the hardware resource of FPGAs. In this article, we present Lite-Stereo, a resource-efficient DNN-based stereo vision accelerator to improve the hardware efficiency for StereoEngine running on a resource-constrained FPGA. To achieve this, we design a set of optimized hardware architectures for resource-demanding bottleneck modules. In order to balance the gap between the processing speed and resource efficiency, the process elements in binary neural network modules are shared within and across modules. In addition, we provide reusing strategies on path aggregation and neighbor calculation to improve the resource efficiency of the semi-global matching module. Evaluation results demonstrate that Lite-Stereo reduces the hardware cost of ALUTs and RAM bits by 60% and 29%, respectively, without compromising the accuracy and energy efficiency compared with StereoEngine. Yehua Ling, Haitao Meng, Kai Huang 0001, Gang Chen 0023 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Hardware accelerator for an accurate local stereo matching algorithm using binary neural network
Yehua Ling, Haitao Meng, Gang Chen 0023 |
J. Syst. Archit. | 1 |
| 2020 | StereoEngine: An FPGA-Based Accelerator for Real-Time High-Quality Stereo Estimation With Binary Neural NetworkabstractStereo estimation is essential to many applications such as mobile autonomous robots, most of which ask for real-time response, high energy, and storage efficiency. Deep neural networks (DNNs) have shown to yield significant gains in improving accuracy. However, these DNN-based algorithms are challenging to be deployed on energy and resource-constrained devices due to the high computational complexities of DNNs. In this article, we present StereoEngine, a fully pipelined end-to-end stereo vision accelerator that computes accurate dense depth in a real-time and energy-efficient manner. An efficient stereo algorithm is developed and optimized for a high-quality hardware-friendly implementation, that leverages binary neural network (BNN) to learn discriminative binary descriptors to improve the disparity. The design of StereoEngine is a standalone DNN-based stereo vision system where all processing procedures are implemented on a hardware platform. The effectiveness of StereoEngine is evaluated by comprehensive experiments. Compared with software-based implementations on the highend and embedded Nvidia GPUs, StereoEngine achieves up to 3×, 13×, and 50× speedups, as well as up to 211×, 58×, and 73× energy efficiency improvement, respectively. Furthermore, StereoEngine achieves leading accuracy when compared to state-of-the-art hardware implementations on the challenging KITTI dataset. Gang Chen 0023, Yehua Ling, Haitao Meng, Shengyu He, Kai Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |