Da Ai

dblp:172/4872 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-2929-516XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cross-distance near-infrared face recognition
Da Ai, Yunqiao Wang, Zhike Ji
J. Vis. Commun. Image Represent.1
2026 PSO-GhostNet: Lightweight CNN discovery via deep particle search and structured redundancy-aware design
Dianwei Wang, Da Ai
Pattern Recognit.5
2026 Rate-Distortion Optimization in Video Coding for Machines: A License Plate Recognition Perspective
abstract
This paper proposes a recognition-oriented rate-distortion optimization (RO-RDO) method. As a novel Video Coding for Machines (VCM) approach, it addresses the challenge of degraded recognition performance for small targets, such as license plates, caused by coding distortion under stringent bitrate constraints, particularly in Unmanned Aerial Vehicle (UAV) applications. To support this research, we constructed the first Distance Calibrated License Plate Dataset (DCLP) in the field of VCM, establishing a solid foundation for future studies investigating the correlation between shooting distance, bitrate, and recognition performance. Experimental results demonstrate that, compared with the H.266/VVC standard, RO-RDO achieves 43.99% bitrate savings and reduces encoding time by 73.76% while maintaining comparable mean Average Precision (mAP). Under the same bitrate, the license plate recognition rate improves by 4.13%, and the mAP increases by 7.17%, with a 2.43 dB PSNR gain in the license plate region, effectively mitigating the impact of encoding distortion on the recognition rate.
Da Ai, Tianzhen Liu, Chenyang Han, Dianwei Wang
IEEE Signal Process. Lett.1
2026 LPCM: Learning-Based Predictive Coding for LiDAR Point Cloud Compression
abstract
In recent years, LiDAR point clouds have been widely used in many applications. Since the data volume of LiDAR point clouds is very huge, efficient compression is necessary to reduce their storage and transmission costs. However, existing learning-based compression methods do not exploit the inherent angular resolution of LiDAR and ignore the significant differences in the correlation of geometry information at different bitrates. The predictive geometry coding method in the geometry-based point cloud compression (G-PCC) standard uses the inherent angular resolution to predict the azimuth angles. However, it only models a simple linear relationship between the azimuth angles of neighboring points. Moreover, it does not optimize the quantization parameters for residuals on each coordinate axis in the spherical coordinate system. To address these issues, we propose a learning-based predictive coding method (LPCM) with both high-bitrate and low-bitrate coding modes. LPCM converts point clouds into predictive trees using the spherical coordinate system. In high-bitrate coding mode, we use a lightweight Long-Short-Term Memory-based predictive (LSTM-P) module that captures long-term geometry correlations between different coordinates to efficiently predict and compress the elevation angles. In low-bitrate coding mode, where geometry correlation degrades, we introduce a variational radius compression (VRC) module to directly compress the point radii. Then, we analyze why the quantization of spherical coordinates differs from that of Cartesian coordinates and propose a differential evolution (DE)-based quantization parameter selection method, which improves rate-distortion performance without increasing coding time. Experimental results show that LPCM achieved a D1-PSNR BD-rate reduction of 21.2% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 5.6% compared with the PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0.
Hui Yuan 0001, Shiqi Jiang 0006, Da Ai, Wei Zhang 0072, Raouf Hamzaoui
IEEE Trans. Image Process.4
2025 Adaptive local neighborhood search and dual attention convolution network for complex semantic segmentation towards indoor point clouds
Da Ai, Siyu Qin, Zihe Nie, Dianwei Wang, Hui Yuan 0001, Ying Liu 0026
Expert Syst. Appl.1
2025 Temporal and Spatial Perception: A Novel Perceptual Rate-Distortion Optimization Method for H.266/VVC Encoding
abstract
Introducing saliency information to mitigate perceptual redundancy and achieve superior compression represents a novel approach to the development of video compression. Existing saliency-based compression coding methods rely on the determination of saliency regions and focus too much on saliency regions while ignoring the perceptible distortion in non-saliency regions. We propose a spatiotemporal visual perceptual rate-distortion optimization (PRDO) algorithm for Versatile Video Coding (H.266/VVC) that is more in line with the human visual system (HVS). Firstly, we establish a linear weighted distortion model based on spatiotemporal and saliency features. The distortion model makes effective use of saliency features while considering image content in non-saliency regions that is still perceptible to the human eye, thereby achieving an overall visual effect that conforms to human subjective perception. Based on this distortion model, we propose a saliency adaptive quantization parameter (SAQP) selection method with a more flexible quantization parameter selection range, adaptively allocating the optimal coding unit quantization parameter according to the saliency regions of the image, ensuring a balanced bitrate allocation between saliency and non-saliency regions. The proposed method is implemented for the first time on the H.266/VVC coding standard, attaining an average bitrate saving of 19.9% across all test sequences and an average PSNR improvement of 2.65 dB in saliency regions compared to VTM16.0. The BD-EWPSNR of the proposed PRDO and SAQP method improves by 1.34 dB and 1.45 dB in the All-Intra and Lowdelay_P encoding modes, respectively. Additionally, the BD-Rate based on EWPSNR is reduced by 25.86% and 33.73%, respectively, with an overall compression coding time saving of 19.76%. The experimental results demonstrate that the proposed method can significantly reduce the bit rate and coding time while improving the subjective perceptive quality, providing a competitive solution for video compression coding.
Da Ai, Hui Yuan 0001, Ying Liu 0026, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.1
2024 NIR-VIS Image Translation for the Cross-Spectral and Cross-Distance Face Recognition
abstract
Near infrared (NIR) video surveillance is not affected by light conditions and plays important role in the field of in public security and criminal investigation. However, the spectral difference between NIR and visible (VIS) light, as well as the shooting distance, are the two main factors that affect the accuracy of face recognition. To this end, we propose an asymmetric cycle generative adversarial network for such Cross-Spectral and Cross-Distance(CSCD) face recognition. The prosed method is able to translate NIR facial images shot at different distances into their corresponding high-quality VIS images, while maintaining enough identity information to allow existing VIS facial recognition models to perform the recognition. Meanwhile, we have created a new large-scale CSCD face dataset, CSCD-F, which was the first to capture NIR face images at different distances with fixed focus NIR camera. The proposed dataset and method will provide a novel training and evaluating platform for CSCD face recognition.
Da Ai, Yunqiao Wang, Ying Liu 0026
ICME1
2024 MGTN: Multi-scale Graph Transformer Network for 3D Point Cloud Semantic Segmentation
abstract
The structural similarity of point clouds presents challenges in accurately recognizing and segmenting semantic information at the demarcation points of complex scenes or objects. In this study, we propose a multi-scale graph transformer network (MGTN) for 3D point cloud semantic segmentation. First, a multi-scale graph convolution (MSG-Conv) is devised to address the limitations faced by existing methods when extracting local and global features of point cloud data with varying densities simultaneously. Subsequently, we employ a graph-transformer (G-T) module to enhance edge details and spatial position information in the point cloud, thereby improving recognition accuracy for small objects and confusing elements such as columns and beams. Extensive testing on ShapeNet parts and S3DIS datasets was conducted to demonstrate the effectiveness of MGTN. Compared to the baseline network DGCNN, our proposed MGTN achieves substantial performance improvements, as evidenced by notable increases in mIoU of 1.5% and 18.5% on the ShapeNet parts and S3DIS datasets respectively. Additionally, MGTN outperforms the recent CFSA- Net by 2.3% and 3.4% on OA and mIoU respectively.
Da Ai, Siyu Qin, Zihe Nie, Hui Yuan 0001, Ying Liu 0026
VCIP1
2023 STVP: A Spatiotemporal Visual Perception Method for User-generated Content Video Quality Assessment
abstract
With the popularity and development of short video applications, the behavior of using mobile devices to shoot and share user-generated content (UGC) videos has become increasingly common. Video quality assessment (VQA) is critical in guaranteeing end-user viewing experiences. UGC-VQA is a challenging problem due to the complexity and variety of distortion types of UGC videos and the absence of reference videos. To improve the consistency of UGC-VQA results and human subjective ratings, in this paper, we propose a UGC-VQA method based on spatiotemporal visual perception (STVP). Firstly, a hierarchical feature fusion module was added to the feature extraction network to realize the fusion of low-level visual features and high-level semantic features, and obtain the quality perception features with rich visual information. Then, we use the self-attention to weight different frames to distinguish their importance. The long short-term memory (LSTM) network and the time pool are used to model long-term dependencies and temporal memory effects. Experimental results on UGC-VQA datasets show that the proposed method achieves a performance improvement of nearly 2%, and its evaluation results are more consistent with human visual perception.
Da Ai, Mingyue Lu, Ying Liu 0026
VCIP1
2022 A Full-Reference Image Quality Assessment Method with Saliency and Error Feature Fusion
abstract
Image quality assessment (IQA) has obtained certain achievements with the help of convolutional neural network (CNN). To promote the evaluation performance, most existing methods focus on optimizing the structure and parameters of neural networks, while some useful features of image are ignored that can easily be acquired. In this paper, we propose a saliency and error feature fusion IQA (SEFF-IQA) method. Instead of the image itself, two image features, the error between the reference image and the distorted image, and the subjective saliency of distorted image are taken as inputs of the CNN for training. The evaluation score of image quality were obtained by a conventional CNN that trained on frequently used public databases. The proposed method possesses one basic architecture of the CNN only and reduces the volume of training data remarkably compared with state-of-art approaches. Experimental results show that the proposed method is more consistent with human subjective perception than other existing deep learning-based methods.
Da Ai, Yunhong Liu, Yurong Yang, Mingyue Lu, Ying Liu 0026, Nam Ling
ISCAS1
2021 An Adaptive Feature-based Quantization Algorithm for Point Cloud Compression
abstract
To reduce over-rasterization distortion caused by global uniform quantization for static surface point cloud, an adaptive quantization coding method based on feature mining is proposed. Combining spatial position and texture feature of point clouds with level of details, the quantization increment is dynamically set according to feature priority, which can reserve the number of effective points to the maximum extent, and reduce the rasterization distortion. Experimental results show that the proposed method can effectively enhance the subjective reconstruction quality of compressed point cloud, gaining better results of rate-distortion optimization.
Da Ai, Hongying Lu, Yurong Yang, Ying Liu 0026
PCS1