Yijun Liu 0012

dblp:41/3457-12 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KAN or MLP? Point Cloud Shows the Way Forward
Qingdong He, Yijun Liu 0012, Jingyong Su
ICMR3
2026 Failure Detection in Image Segmentation Under Conditions of Semantic and Covariate Shifts
abstract
Whether deep neural networks can provide reliable confidence is of great significance, especially in risk-sensitive scenarios. This work explores the impact of covariate and semantic shifts on segmentation tasks, an area which has not been extensively studied. Covariate shift refers to changes in the data distribution without alterations in the label space, while semantic shift involves changes in both data distribution and label space. We find that model-unknown distributional shifts in test data can transform an overconfidence problem into a situation of making random predictions with arbitrary confidence. The paper proposes a novel approach for effective failure detection that combines holistic image-level analysis and detailed pixel-level information. This approach involves the use of a Gray Level Co-occurrence Matrix (GLCM) to analyze the prediction randomness between adjacent pixels and a Magnitude-Direction Confidence Score Function (MD-CSF) for determining pixel acceptance or rejection. Furthermore, we introduce a new benchmark dataset, the Robot Inspection dataset for Semantic and Covariate shift in Segmentation (RISKS, the homophone of RISCS), to fill the need for datasets capable of evaluating the simultaneous impact of semantic and covariate shifts. Experimental results demonstrate that our method successfully detects image-level failures in segmentation, with MD-CSF outperforming other pluggable CSFs. The code and RISKS dataset will be available at https://github.com/liuyijungoon/MD-CSF.
Yijun Liu 0012, Zhuotao Tian, Hang Zhao 0019, Zipeng Zhu, Jingyong Su
IEEE Trans. Circuits Syst. Video Technol.1
2025 C2AD: Dual Consistency Learning for Zero-Shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) is dedicated to detecting anomalies without having any seen normal or abnormal samples for the target set. Existing approaches utilize the pre-trained CLIP to assess normality/abnormality by exploiting the similarity between images and text with the frozen visual encoder. However, the frozen CLIP visual encoder impedes performance improvements. Additionally, their representations of anomalies are sensitive to contextual variations, leading to poor localization of unseen abnormalities. Therefore, this paper introduces the Dual Consistency Learning for Zero-Shot Anomaly Detection (C2AD), comprising two components: semantic and contextual consistency. Semantic consistency enhances generalization by maintaining correlational semantic consistency, while contextual consistency encourages representations to be robust to contextual changes. C2AD improves the model training without adding extra computational overhead during inference. Comprehensive experiments demonstrate that C2AD can boost the performance of ZSAD in anomaly detection and localization, achieving state-of-the-art results.
Ruilong Xing, Zhuotao Tian, Yijun Liu 0012, Senqiao Yang, Jingyong Su
ICASSP4
2024 Typicalness-Aware Learning for Failure Detection
abstract
Deep neural networks (DNNs) often suffer from the overconfidence issue, where incorrect predictions are made with high confidence scores, hindering the applications in critical systems. In this paper, we propose a novel approach called Typicalness-Aware Learning (TAL) to address this issue and improve failure detection performance. We observe that, with the cross-entropy loss, model predictions are optimized to align with the corresponding labels via increasing logit magnitude or refining logit direction. However, regarding atypical samples, the image content and their labels may exhibit disparities. This discrepancy can lead to overfitting on atypical samples, ultimately resulting in the overconfidence issue that we aim to address. To address this issue, we have devised a metric that quantifies the typicalness of each sample, enabling the dynamic adjustment of the logit magnitude during the training process. By allowing relatively atypical samples to be adequately fitted while preserving reliable logit direction, the problem of overconfidence can be mitigated. TAL has been extensively evaluated on benchmark datasets, and the results demonstrate its superiority over existing failure detection methods. Specifically, TAL achieves a more than 5\% improvement on CIFAR100 in terms of the Area Under the Risk-Coverage Curve (AURC) compared to the state-of-the-art. Code is available at https://github.com/liuyijungoon/TAL.
Yijun Liu 0012, Jiequan Cui, Zhuotao Tian, Senqiao Yang, Qingdong He, Jingyong Su
NeurIPS1
2023 Stereo RGB and Deeper LIDAR-Based Network for 3D Object Detection in Autonomous Driving
abstract
3D object detection has become an emerging task in autonomous driving scenarios. Most of previous works process 3D point clouds using either projection-based or voxel-based models. However, both approaches contain some drawbacks. The voxel-based methods lack semantic information, while the projection-based methods suffer from numerous spatial information loss when projected to different views. In this paper, we propose the Stereo RGB and Deeper LIDAR (SRDL) framework which can utilize semantic and spatial information simultaneously such that the performance of network for 3D object detection can be improved naturally. Specifically, the network generates candidate boxes from stereo pairs and combines different region-wise features using a deep fusion scheme. The stereo strategy offers more information for prediction compared with prior works. Then, several local and global feature extractors are stacked in the segmentation module to capture richer deep semantic geometric features from point clouds. After aligning the interior points with fused features, the proposed network refines the prediction in a more accurate manner and encodes the whole box in a novel compact method. The decent experimental results on the challenging KITTI detection benchmark demonstrate the effectiveness of utilizing both stereo images and point clouds for 3D object detection.
Qingdong He, Zhengning Wang, Yijun Liu 0012, Shuaicheng Liu, Bing Zeng 0001
IEEE Trans. Intell. Transp. Syst.5
2022 SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds
abstract
Accurate 3D object detection from point clouds has become a crucial component in autonomous driving. However, the volumetric representations and the projection methods in previous works fail to establish the relationships between the local point sets. In this paper, we propose Sparse Voxel-Graph Attention Network (SVGA-Net), a novel end-to-end trainable network which mainly contains voxel-graph module and sparse-to-dense regression module to achieve comparable 3D detection tasks from raw LIDAR data. Specifically, SVGA-Net constructs the local complete graph within each divided 3D spherical voxel and global KNN graph through all voxels. The local and global graphs serve as the attention mechanism to enhance the extracted features. In addition, the novel sparse-to-dense regression module enhances the 3D box estimation accuracy through feature maps aggregation at different levels. Experiments on KITTI detection benchmark and Waymo Open dataset demonstrate the efficiency of extending the graph representation to 3D object detection and the proposed SVGA-Net can achieve decent detection accuracy.
Qingdong He, Zhengning Wang, Yijun Liu 0012
AAAI5
2022 SCIR-Net: Structured Color Image Representation Based 3D Object Detection Network from Point Clouds
abstract
3D object detection from point clouds data has become an indispensable part in autonomous driving. Previous works for processing point clouds lie in either projection or voxelization. However, projection-based methods suffer from information loss while voxelization-based methods bring huge computation. In this paper, we propose to encode point clouds into structured color image representation (SCIR) and utilize 2D CNN to fulfill the 3D detection task. Specifically, we use the structured color image encoding module to convert the irregular 3D point clouds into a squared 2D tensor image, where each point corresponds to a spatial point in the 3D space. Furthermore, in order to fit for the Euclidean structure, we apply feature normalization to parameterize the 2D tensor image onto a regular dense color image. Then, we conduct repeated multi-scale fusion with different levels so as to augment the initial features and learn scale-aware feature representations for box prediction. Extensive experiments on KITTI benchmark, Waymo Open Dataset and more challenging nuScenes dataset show that our proposed method yields decent results and demonstrate the effectiveness of such representations for point clouds.
Qingdong He, Yijun Liu 0012
AAAI4
2021 PD-GAN: Perceptual-Details GAN for Extremely Noisy Low Light Image Enhancement
abstract
Extremely noisy low light enhancement suffers from high-level noise, loss of texture detail, and color degradation. When recovering color or illumination for images taken in a dark environment, the challenge for networks is how to balance the enhancement for noise and texture details for a good visual effect. A single network is not suitable for solving the ill-posed problem of mapping the input image's noise to the clear target in the ground truth. To solve the problems, we pro-pose perceptual-details GAN (PD-GAN) utilizing Zero-DCE to initially recover illumination and combine residual dense-block Encoder-Decoder structure to suppress noise while finely adjusting the illumination. Besides, fractional differential gradient masks are integrated into the discriminator to enhance details. Experiment results demonstrate that PD-GAN outperforms other methods on the extremely low-light image dataset.
Yijun Liu 0012, Zhengning Wang, Deming Zhao
ICASSP1
2021 A Detection and Tracking Combined Network for Long-Term Tracking
Zhengning Wang, Deming Zhao, Yijun Liu 0012
ICIG (3)4
2020 Structure-preserving extremely low light image enhancement with fractional order differential mask guidance
abstract
Low visibility and high-level noise are two challenges for low-light image enhancement. In this paper, by introducing fractional order differential, we propose an end-to-end conditional generative adversarial network(GAN) to solve those two problems. For the problem of low visibility, we set up a global discriminator to improve the overall reconstruction quality and restore brightness information. For the high-level noise problem, we introduce fractional order differentiation into both the generator and the discriminator. Compared with conventional end-to-end methods, fractional order can better distinguish noise and high-frequency details, thereby achieving superior noise reduction effects while maintaining details. Finally, experimental results show that the proposed model obtains superior visual effects in low-light image enhancement. By introducing fractional order differential, we anticipate that our framework will enable high quality and detailed image recovery not only in the field of low-light enhancement but also in other fields that require details.
Yijun Liu 0012, Zhengning Wang, Ruixu Geng
MMAsia1