EDBT 2026 Demo / reviewers in the wild / expert
Liyue Ge
dblp:220/8078
· DBLP profile ↗
12ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-1102-1053ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCG-SSC: Semantic Scene Completion via Self-and-Cross Gated Fusion of Depth Maps and Semantic Priors
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
ICPR (9) | 6 |
| 2026 | SAP-DQR: Joining Spatial-Adaptive Pyramid and Adaptive Query Reorganization for Speed-Accuracy Instance Segmentation
Jiahao Zou, Congxuan Zhang, Liyue Ge, Jiawen Yang, Zhen Chen 0004, Ke Lu 0002 |
MMM (1) | 3 |
| 2026 | SGTP-Net: Semantic Guidance and Texture Priors-Based Dual-Branch Segmentation Network for Surface Defect DetectionabstractDeep learning-based surface defect segmentation approaches have shown promising performance in recent years. However, segmenting defects with complex shapes, large variations in size, and weakly textured defects with indistinct characteristics still poses significant challenges. In this article, a novel semantic guidance and texture priors based dual-branch surface defect segmentation network (SGTP-Net) is proposed for those issues. Firstly, we construct a feature extraction network combines semantic and texture branches. The semantic branch establishes global contextual relationships, while the texture branch captures local features of defects, this dual-branch ensured the network to extract features from various complex defects. Secondly, we design a feature fusion strategy based on semantic guidance and texture priors. The semantic information is used to guides the output of texture branch. After that, the guided texture information provides valuable edge texture priors for each layers output in semantic branch. The two branches mutually guide each other for improving ability of weak textures feature extraction. Finally, we run our method on the NEU-Seg, MT-Defect and MSD datasets to conduct a comprehensive comparison with some state-of-the-art general object segmentation models and specialized surface defect segmentation methods. The experimental results show that our SGTP-Net performs well in surface defect detection, offering excellent semantic segmentation accuracy and exhibiting good stability and robustness in detecting various surface defects. Leqi Jiang, Liyue Ge, Chengzhong Wu, Yaonan Wang 0001, Ke Lu 0002, Congxuan Zhang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Iter3DDet: Depth-Guided Iterative Fusion and Refinement for Monocular 3D Object DetectionabstractMonocular 3D object detection offers significant potential for autonomous systems due to its inherent cost-effectiveness and scalability. While DETR-based architectures excel in 2D vision tasks, critical limitations persist in extending them effectively to monocular 3D detection, as evidenced in existing frameworks like MonoDETR and MonoDGP. These methods typically suffer from inefficient serial fusion of multimodal features and lack iterative refinement mechanisms, limiting their performance, especially for mid-to-long range targets. To overcome these shortcomings, we propose Iter3DDet, a novel depth-guided iterative refinement framework that integrates fine-grained feature fusion to significantly enhance detection performance. The core novelty of our approach lies in two key innovations: (1) A hybrid feature encoder combining MonoDGP’s region segmentation head with MonoDETR’s visual backbone, augmented by a multi-scale context attention module that dynamically aggregates structural and semantic cues across pyramid levels, eliminating heuristic fusion rules; (2) A depth-guided adaptive cross-modal decoder that iteratively fuses depth and context features through prioritized attention mechanisms, coupled with a novel iterative refinement training strategy that progressively refines 3D detection hypotheses, substantially improving accuracy across targets of varying difficulty levels. Extensive experiments on the KITTI, nuScenes, and Waymo benchmarks demonstrate Iter3DDet’s state-of-the-art performance, validating the effectiveness of our iterative refinement paradigm. The code will be open-sourced at https://github.com/PCwenyue. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | MotionFlow: Joint Motion Priors and Appearance Enhancement for High-Accuracy Optical Flow EstimationabstractAlthough optical flow estimation has improved significantly in recent years, large displacements and occlusions remain challenging for current methods due to motion discontinuities that may hinder accurate feature correspondences in these regions, leading to degraded performance. To address this challenge, we propose a novel method named MotionFlow for high-accuracy optical flow estimation. In the encoding stage, we integrate multi-scale features enhance motion and context appearance information via cross- and inter-enhancement module. Subsequently, cross-frame features are utilized to establish motion priors, thereby providing essential prior knowledge for motion estimation. During the decoding stage, we align appearance features of the target frame with the reference frame through warping to retrieve missing context crucial for motion decoding. Experimental results demonstrate the efficacy of our approach, achieving state-of-the-art performance, particularly outperforming online benchmarks on Sintel Final pass and KITTI-2015 datasets. Congxuan Zhang, Zhen Chen 0004, Hongye Chen, Liyue Ge, Ke Lu 0002 |
ICASSP | 5 |
| 2025 | ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow EstimationabstractAlthough the overall performance of existing optical flow estimation methods has improved rapidly, motion discontinuities caused by large displacements and occlusions remain significant challenges for accurate optical flow estimation. To address this issue, we propose a novel Inject Context Information for Motion Generation Network (ICIMG-Net). Firstly, we design a Context Injection Module (CIM), which enhances the generation of correlation volumes by injecting context information. This process supplements the semantic details needed for accurate pixel matching, improving matching precision. Then, we construct a Dual-Injection-GRU (DIG), which facilitates dual interactions between semantic and motion information to address the motion discontinuities problem. Finally, we conduct a comprehensively evaluat of our ICIMG-Net against state-of-the-art methods on the MPI-Sintel and KITTI benchmarks, demonstrating that our method achieves competitive results, particularly in complex synthetic and real-world traffic scenarios. Wenbo Yin, Congxuan Zhang, Zhen Chen 0004, Liyue Ge, Zige Wang |
ICASSP | 5 |
| 2025 | SRSA-Depth: shape and region similarity awareness for outdoor monocular depth estimation
Liyue Ge, Congxuan Zhang, Zhen Chen 0004, Ke Lu 0002 |
Multim. Syst. | 1 |
| 2025 | Self-Supervised Monocular Depth Estimation With Dual-Path Encoders and Offset Field InterpolationabstractAlthough self-supervised learning approaches have demonstrated tremendous potential in multi-frame depth estimation scenarios, existing methods struggle to perform well in cases involving dynamic targets and static ego-camera conditions. To address this issue, we propose a self-supervised monocular depth estimation method featuring dual-path encoders and learnable offset interpolation (LOI). First, we construct a dual-path encoding scheme that utilizes residual and transformer blocks to extract both single- and multi-frame features from the input frames. We design a contrastive learning strategy to effectively decouple single- and multi-frame features, enabling weighted fusion guided by a confidence map. Next, we explore two distinct decoding heads for simultaneously generating low-resolution predictions and offset fields. We then design an LOI module to directly upsample a low-resolution depth map to a full-resolution map. This one-step decoding framework enables accurate and efficient depth prediction. Finally, we evaluate our proposed method on the KITTI and Cityscapes benchmarks, conducting a comprehensive comparison with state-of-the-art approaches. The experimental results demonstrate that our DualDepth method achieves competitive performance in terms of both estimation accuracy and efficiency. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Image Process. | 6 |
| 2024 | Real-Time Monocular Depth Estimation on Embedded SystemsabstractDepth sensing is of paramount importance for unmanned aerial and autonomous vehicles. Nonetheless, contemporary monocular depth estimation methods employing complex deep neural networks within Convolutional Neural Networks are inadequately expedient for real-time inference on embedded platforms. This paper endeavors to surmount this challenge by proposing two efficient and lightweight architectures, RT-MonoDepth and RT-MonoDepth-S, thereby mitigating computational complexity and latency. Our methodologies not only attain accuracy comparable to prior depth estimation methods but also yield faster inference speeds. Specifically, RT-MonoDepth and RT-MonoDepth-S achieve frame rates of 18.4&30.5 FPS on NVIDIA Jetson Nano and 253.0&364.1 FPS on Jetson AGX Orin, utilizing a single RGB image of resolution $640 \times 192$. The experimental results underscore the superior accuracy and faster inference speed of our methods in comparison to existing fast monocular depth estimation methodologies on the KITTI dataset. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Liyue Ge |
ICIP | 5 |
| 2024 | ACR-Net: Learning High-Accuracy Optical Flow via Adaptive-Aware Correlation Recurrent NetworkabstractAlthough recurrent network-based optical flow estimation methods have shown great success in recent years, most of these methods have difficulty handling large displacements and occlusions because the existing recurrent networks are usually restricted to coarse-resolution single-scale models while ignoring the multiscale features brought by hierarchical concepts in previous coarse-to-fine approaches. In this paper, we propose an adaptive-aware correlation recurrent network for optical flow estimation, named ACR-Net, which preserves fine motion features with a single-scale resolution recurrent framework and adaptively incorporates multiscale features at different stages to achieve high-accuracy optical flow estimation. First, our proposed self-adaptation scale-aware correlation module can incorporate the adaptive correlation of multiscale inter- and intra-motion features, which makes the features more discriminative for capturing long-range dependencies between pixels. Second, our presented adaptive-aware motion module can effectively extract the required features of different kinds of motion from multilevel correspondence. Third, our introduced cross-guide motion and fusion modules can accurately guide the propagation of reliable pixels towards unreliable pixels and dynamically determine the most suitable expression to address the occlusion challenges. Comprehensive experiments demonstrate that ACR-Net outperforms existing two-view models, striking a good balance between speed and accuracy and achieving the best performance on the MPI-Sintel final pass and KITTI-2015 test datasets. The code will be made publicly available. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge, Zige Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | STDC-Flow: large displacement flow field estimation using similarity transformation-based dense correspondenceabstractIn order to improve the accuracy and robustness of optical flow computation under large displacements and motion occlusions, the authors present in this study a large displacement flow field estimation approach using similarity transformation‐based dense correspondence, named STDC‐Flow approach. First, the authors compute an initial nearest‐neighbour field by using the STDC‐Flow of the consecutive two frames, and then extract the consistent regions as the robust nearest‐neighbour field and label the inconsistent regions as the occlusion areas. Second, they improve a non‐local total variation with the L 1 norm optical flow model by using the occlusion information to modify the weighted median filtering optimisation. Third, they fuse the robust nearest‐neighbour field and the computed flow field of the improved variational optical flow model to construct the final flow field by using the quadratic pseudo‐boolean optimisation fusion algorithm. Finally, the authors compare the proposed STDC‐Flow method with several state‐of‐the‐art approaches including the variational and deep learning‐based optical flow models by using the MPI‐Sintel and KITTI evaluation databases. The comparison results demonstrate that the proposed STDC‐Flow method has a high accuracy for flow field computation, especially the capacity of dealing with large displacements and motion occlusions. Congxuan Zhang, Zhen Chen 0004, Fan Xiong, Wen Liu 0004, Ming Li 0056, Liyue Ge |
IET Comput. Vis. | 6 |
| 2020 | Refined TV-L1 Optical Flow Estimation Using Joint FilteringabstractThough the accuracy and robustness of optical flow has been dramatically enhanced over the past few years, the issue of edge-blurring near the image and motion boundaries has remained a challenge in flow field estimation. In this paper, we propose a refined total variation withL1norm (TV-L1) optical flow estimation approach using joint filtering, named JOF. First, we divide the image into three categorized regions: mutual-structure regions, inconsistent structure regions, and smooth regions. The mutual-structure guided filter for optical flow estimation is constructed by extracting the mutual-structure regions of the flow field. Second, the refined TV-L1optical flow model is proposed by incorporating the non-local term and mutual-structure guided filter objective function into the classical TV-L1energy function. Furthermore, the novel TV-L1optical flow objective function is minimized using a joint filtering program composed of a weighted median filter and a mutual-structure guided filter to optimize the estimated flow field during the coarse-to-fine optical flow computation scheme. Finally, we compare the proposed JOF method with several state-of-the-art approaches including variational and deep learning based optical flow models using the Middlebury, MPI-Sintel, and UCF101 test databases. The evaluation results indicate that the proposed method has high accuracy and good robustness for flow field computation and, especially, the significant benefit of edge-preserving. Congxuan Zhang, Liyue Ge, Zhen Chen 0004, Ming Li 0056, Wen Liu 0004, Hao Chen 0060 |
IEEE Trans. Multim. | 2 |