Ying Liu 0026

dblp:91/112-26 · DBLP profile ↗
← Back
49ranked-venue papers
16as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Hierarchical semantic alignment heterogeneous knowledge distillation model for smart agriculture crop leaf disease recognition
Daxiang Li 0002, Ying Liu 0026
Expert Syst. Appl.3
2025 Guided progressive learning for room layout estimation: From pixel-level embeddings to refined depth maps
Weidong Zhang 0005, Ying Liu 0026, Yu Hao 0002
Comput. Vis. Image Underst.3
2025 Adaptive local neighborhood search and dual attention convolution network for complex semantic segmentation towards indoor point clouds
Da Ai, Siyu Qin, Zihe Nie, Dianwei Wang, Hui Yuan 0001, Ying Liu 0026
Expert Syst. Appl.6
2025 STRM-KD: Semantic topological relation matching knowledge distillation model for smart agriculture apple leaf disease recognition
Daxiang Li 0002, Ying Liu 0026
Expert Syst. Appl.3
2025 Deep Learning Based Fine-Grained Image Classification: Recent Advances, Applications and Future Outlook
abstract
ABSTRACT Fine‐grained image classification (FGIC) aims to distinguish visually similar categories by capturing subtle differences, yet the coexistence of large intra‐class variation and small inter‐class differences makes this task highly challenging. This paper provides a systematic review of recent deep learning‐based FGIC methods. According to the type of training data, existing approaches are categorized into four groups: (1) conventional models with large‐scale samples (strongly supervised, weakly supervised, semi‐supervised, and unsupervised); (2) few‐shot learning models for limited‐sample scenarios (e.g. meta‐learning and metric learning); (3) models leveraging external information, including multi‐modal and web‐sourced data; and (4) emerging diffusion‐based models. Representative algorithms in each category are summarized and analysed in terms of their advantages and limitations. The paper also reviews mainstream benchmark datasets and introduces a newly proposed application‐oriented dataset, CIIP‐TPID, to support real‐world tasks. Additionally, practical applications of FGIC in public security, medicine, and commerce are discussed. Finally, future research directions are outlined, including diffusion‐based data augmentation, advanced multi‐modal fusion, transformer architecture optimization, lightweight models for edge deployment, and robustness against noisy labels. This review provides a structured and up‐to‐date reference for researchers and practitioners in the field.
Ying Liu 0026, Weidong Zhang 0005, Guojun Lu
IET Image Process.1
2025 MCWANet: A hyperspectral anomaly detection network with multi-stage collaborative optimization of wavelet convolution and attention mask
Yuquan Gan, Ji Zhang 0001, Ying Liu 0026
Knowl. Based Syst.5
2025 Mamba-Wavelet Cross-Modal Fusion Network With Graph Pooling for Hyperspectral and LiDAR Data Joint Classification
abstract
Recently, with the rapid development of deep learning, the collaborative classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) image has become a research hotspot in remote sensing (RS) technology. However, existing methods either only consider complementary learning of spatial-domain information, or do not take into account the intrinsic dependencies between pixels and overlook the importance difference of pixels. In this letter, we propose a Mamba-Wavelet Cross-Modal Fusion Network with Graph Pooling (MW-CMFNet) for HSI and LiDAR joint classification. First, a Two-Branch Feature Extraction (TBFE) is used to extract spatial and spectral features. Then, in order to dig deeper into the complementary information of different modalities and fully fuse them under the guidance of frequency-domain information, a Mamba-Wavelet Cross-Modal Feature Fusion (MW-CMFF) Module is devised, it aims to utilize Mamba’s outstanding long-range modeling ability to learn complementary information in the spatial and frequency domains, Finally, the Graph Pooling module is designed to sense the intrinsic dependencies of neighbouring pixels and explore the importance difference of pixels, rather than assigning the same weight to different pixels. Experiments on the Houston2013 and Trento datasets show that the MW-CMFNet achieves higher classification accuracy compared to other state-of-the-art methods.
Daxiang Li 0002, Bingying Li, Ying Liu 0026
IEEE Geosci. Remote. Sens. Lett.3
2025 C2P-Net: Comprehensive Depth Map to Planar Depth Conversion for Room Layout Estimation
abstract
Room layout estimation seeks to infer the overall spatial configuration of indoor scenes using perspective or panoramic images. As the layout is determined by the dominant indoor planes, this problem inherently requires the reconstruction of these planes. Some studies reconstruct indoor planes from perspective images by learning pixel-level or instance-level plane parameters. However, directly learning these parameters has the problems of susceptibility to occlusions and position dependency. In this paper, we introduce the Comprehensive depth map to Planar depth (C2P) conversion, which reformulates planar depth reconstruction into the prediction of a comprehensive depth map and planar visibility confidence. Based on the parametric representation of planar depth we propose, the C2P conversion is applicable to both panoramic and perspective images. Accordingly, we present an effective framework for room layout estimation that jointly learns the comprehensive depth map and planar visibility confidence. Due to the differentiability of the C2P conversion, our network autonomously learns planar visibility confidence by constraining the estimated plane parameters and reconstructed planar depth map. We further propose a novel approach for 3D layout generation through sequential planar depth map integration. Experimental results demonstrate the superiority of our method across all evaluated panoramic and perspective datasets.
Weidong Zhang 0005, Mengjie Zhou, Jiyu Cheng, Ying Liu 0026, Wei Zhang 0021
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Temporal and Spatial Perception: A Novel Perceptual Rate-Distortion Optimization Method for H.266/VVC Encoding
abstract
Introducing saliency information to mitigate perceptual redundancy and achieve superior compression represents a novel approach to the development of video compression. Existing saliency-based compression coding methods rely on the determination of saliency regions and focus too much on saliency regions while ignoring the perceptible distortion in non-saliency regions. We propose a spatiotemporal visual perceptual rate-distortion optimization (PRDO) algorithm for Versatile Video Coding (H.266/VVC) that is more in line with the human visual system (HVS). Firstly, we establish a linear weighted distortion model based on spatiotemporal and saliency features. The distortion model makes effective use of saliency features while considering image content in non-saliency regions that is still perceptible to the human eye, thereby achieving an overall visual effect that conforms to human subjective perception. Based on this distortion model, we propose a saliency adaptive quantization parameter (SAQP) selection method with a more flexible quantization parameter selection range, adaptively allocating the optimal coding unit quantization parameter according to the saliency regions of the image, ensuring a balanced bitrate allocation between saliency and non-saliency regions. The proposed method is implemented for the first time on the H.266/VVC coding standard, attaining an average bitrate saving of 19.9% across all test sequences and an average PSNR improvement of 2.65 dB in saliency regions compared to VTM16.0. The BD-EWPSNR of the proposed PRDO and SAQP method improves by 1.34 dB and 1.45 dB in the All-Intra and Lowdelay_P encoding modes, respectively. Additionally, the BD-Rate based on EWPSNR is reduced by 25.86% and 33.73%, respectively, with an overall compression coding time saving of 19.76%. The experimental results demonstrate that the proposed method can significantly reduce the bit rate and coding time while improving the subjective perceptive quality, providing a competitive solution for video compression coding.
Da Ai, Hui Yuan 0001, Ying Liu 0026, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.5
2025 Mamba Cross-Modal Information Fusion Self-Distillation Model for Joint Classification of LiDAR and Hyperspectral Data
abstract
Recent studies have found that compared to single-modal data, the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) multimodal data can utilize their complementary information to further improve the accuracy of land-cover classification. However, due to the significant differences between multimodal data, the complementarity among them is difficult to be fully exploited and utilized, and the features after fusion are not refined and optimized, which limits the further improvement of land-cover classification accuracy. To alleviate these issues, a novel Mamba Cross-Modal Information Fusion Self-Distillation (Mb-CMIFSD) model is designed. Specifically, Mb-CMIFSD first uses conventional convolutional neural networks (CNN) to transform each patch into a token sequence. Second, a Mamba Cross Modal Information Fusion (MCMIF) module is developed to combine cross-modal attention with bidirectional Mamba mechanism, which can better explore the complementarity of multimodal remote sensing (RS) data and obtain more discriminative multimodal fusion features. Finally, a Prototype Constrained Self-Distillation (PCSD) module is designed to utilize the constructed prototype orthogonal regularization knowledge distillation function to further refine cross-modal fusion features, thereby enhancing the robustness and adaptability of feature extraction. The experimental results on three benchmark HSI and LiDAR datasets show that the designed Mb-CMIFSD model has higher classification accuracy compared to other state-of-the-art methods, and the ablation experiments also confirm the positive effect of the designed two key modules.
Daxiang Li 0002, Bingying Li, Ying Liu 0026
IEEE Trans. Geosci. Remote. Sens.3
2025 Fast 3D Room Layout Estimation Based on Compact High-Level Representation
abstract
3D room layout estimation aims to reconstruct the holistic 3D structure from an indoor RGB image. For most of the deep learning-based methods, layout inference is guided by a kind of learned 2D mid-level representation such as pixel-wise surface labels. However, learning such high-resolution 2D representation might suffer from information redundancy and memory consumption, and will increase the runtime of estimation and deployment cost for practical applications. In this paper, we attempt to learn a compact high-level representation with only 29 real numbers for estimating the 3D layout using general regression networks. The learned compact high-level representation contains three components: instance-wise plane parameters, camera intrinsic parameters, and plane location indicators. With the learned representation, the inverse depth map of each plane can be calculated to reconstruct the 3D layout. We further design a set of order-agnostic loss functions to restrict the produced inverse depth maps, with which the model can be trained with either weak 2D layout labels or full 3D layout supervision. Moreover, by jointly learning the plane parameters and locations, the model is benefited from 3D reasoning. Experimental results show that our method is much faster than the existing layout estimation methods and obtains competitive performance on benchmark datasets, showing its potential for real-time applications.
Weidong Zhang 0005, Yu Qiao 0001, Ying Liu 0026, Ran Song 0001, Wei Zhang 0021
IEEE Trans. Image Process.3
2024 NIR-VIS Image Translation for the Cross-Spectral and Cross-Distance Face Recognition
abstract
Near infrared (NIR) video surveillance is not affected by light conditions and plays important role in the field of in public security and criminal investigation. However, the spectral difference between NIR and visible (VIS) light, as well as the shooting distance, are the two main factors that affect the accuracy of face recognition. To this end, we propose an asymmetric cycle generative adversarial network for such Cross-Spectral and Cross-Distance(CSCD) face recognition. The prosed method is able to translate NIR facial images shot at different distances into their corresponding high-quality VIS images, while maintaining enough identity information to allow existing VIS facial recognition models to perform the recognition. Meanwhile, we have created a new large-scale CSCD face dataset, CSCD-F, which was the first to capture NIR face images at different distances with fixed focus NIR camera. The proposed dataset and method will provide a novel training and evaluating platform for CSCD face recognition.
Da Ai, Yunqiao Wang, Ying Liu 0026
ICME4
2024 MGTN: Multi-scale Graph Transformer Network for 3D Point Cloud Semantic Segmentation
abstract
The structural similarity of point clouds presents challenges in accurately recognizing and segmenting semantic information at the demarcation points of complex scenes or objects. In this study, we propose a multi-scale graph transformer network (MGTN) for 3D point cloud semantic segmentation. First, a multi-scale graph convolution (MSG-Conv) is devised to address the limitations faced by existing methods when extracting local and global features of point cloud data with varying densities simultaneously. Subsequently, we employ a graph-transformer (G-T) module to enhance edge details and spatial position information in the point cloud, thereby improving recognition accuracy for small objects and confusing elements such as columns and beams. Extensive testing on ShapeNet parts and S3DIS datasets was conducted to demonstrate the effectiveness of MGTN. Compared to the baseline network DGCNN, our proposed MGTN achieves substantial performance improvements, as evidenced by notable increases in mIoU of 1.5% and 18.5% on the ShapeNet parts and S3DIS datasets respectively. Additionally, MGTN outperforms the recent CFSA- Net by 2.3% and 3.4% on OA and mIoU respectively.
Da Ai, Siyu Qin, Zihe Nie, Hui Yuan 0001, Ying Liu 0026
VCIP5
2024 Micro-expression recognition based on a novel GCN-transformer cooperation model for IoT-eHealth
Daxiang Li 0002, Nannan Qiao, Ying Liu 0026
Expert Syst. Appl.3
2024 Image recognition based on lightweight convolutional neural network: Recent advances
abstract
Image recognition is an important task in computer vision with broad applications. In recent years, with the advent of deep learning, lightweight convolutional neural network (CNN) has brought new opportunities for image recognition, which allows high-performance recognition algorithms to run on resource-constrained devices with strong representation and generalization capabilities. This paper first presents an overview of several classical lightweight CNN models. Then, a comprehensive review is provided on recent image recognition techniques using lightweight CNN. According to the strategies applied to optimize image recognition performance, existing methods are classified into three categories: (1) model compression, (2) optimization of lightweight network, and (3) combining Transformer with lightweight network. In addition, some representative methods are tested on three commonly used datasets for performance comparison. Finally, technical challenges and future research trends in this field are discussed.
Ying Liu 0026, Jiahao Xue, Daxiang Li 0002, Weidong Zhang 0005, Tuan Kiang Chiew, Zhijie Xu
Image Vis. Comput.1
2024 HFSI-TF: Hierarchical Full-Scale Interactive Transformer Model for Object Detection in Remote Sensing Image
abstract
Transformer-based object detection models usually adopt an encoding-decoding architecture that mainly combines self-attention (SA) and multilayer perceptron (MLP). Although this architecture does not require nonmaximum suppression (NMS) and can really achieve end-to-end object detection, it also suffers from the disadvantage of insufficient multiscale object perception in the image, which leads to low accuracy in detecting small objects. Focusing on these issues, a new full-scale bidirectional interactive attention (FSBDIA) mechanism is constructed, thereby a novel hierarchical full-scale interactive transformer (HFSI-TF) model is designed for object detection in remote sensing image (RSI). First, in order to enhance the multiscale perception ability of the model, the FSBDIA mechanism is designed under the guidance of full-scale information. Then, based on FSBDIA, a hierarchical HFSI-TF encoder is constructed to interactively fuse multilayer feature maps layer by layer, thereby obtaining multiscale encoded features of RSI. Finally, a mixed cross attention (MCA) mechanism is also constructed, and an iterative decoding architecture is designed based on it to improve the accuracy of small object detection. Comparative experiments based on two benchmark datasets (i.e., DIOR and HRSC2016) show that the designed HFSI-TF model can effectively improve the accuracy of object detection in RSI, and the model we designed has superior performance compared to other state-of-the-art methods.
Daxiang Li 0002, Bingying Li, Ying Liu 0026
IEEE Geosci. Remote. Sens. Lett.3
2024 PSCLI-TF: Position-Sensitive Cross-Layer Interactive Transformer Model for Remote Sensing Image Scene Classification
abstract
In the scene classification task of remote sensing image (RSI), in order to fully perceive multi-scale local objects in the image and explore their interdependencies to mine the scene semantics of RSI, this letter designs a novel Position-Sensitive Cross-Layer Interactive Transformer (PSCLI-TF) model to improve the accuracy of RSI scene classification. Firstly, ResNet50 is utilized as the backbone to extract the multi-layer feature maps of RSI. Then, in order to enhance the model’s position sensitivity to local objects in RSI, a new Position-Sensitive Cross-Layer Interactive Attention (PSCLIA) mechanism is designed, and based on it a novel PSCLI-TF encoder is constructed to perform layer-by-layer interactive fusion on the multi-layer feature maps to obtain the multi-granularity Cross-Layer Fusion (CLF) feature of RSI. Finally, a prototype-based self-supervised loss function is constructed to alleviate the semantic gap problem of "large intra-class variance and small inter-class variance" in RSI scene classification. Comparative experimental results based on three datasets (i.e., AID, NWPU and UCM) indicate that the classification performance of the designed PSCLI-TF model is highly competitive compared to other state-of-the-art methods.
Daxiang Li 0002, Runyuan Liu, Ying Liu 0026
IEEE Geosci. Remote. Sens. Lett.4
2023 STVP: A Spatiotemporal Visual Perception Method for User-generated Content Video Quality Assessment
abstract
With the popularity and development of short video applications, the behavior of using mobile devices to shoot and share user-generated content (UGC) videos has become increasingly common. Video quality assessment (VQA) is critical in guaranteeing end-user viewing experiences. UGC-VQA is a challenging problem due to the complexity and variety of distortion types of UGC videos and the absence of reference videos. To improve the consistency of UGC-VQA results and human subjective ratings, in this paper, we propose a UGC-VQA method based on spatiotemporal visual perception (STVP). Firstly, a hierarchical feature fusion module was added to the feature extraction network to realize the fusion of low-level visual features and high-level semantic features, and obtain the quality perception features with rich visual information. Then, we use the self-attention to weight different frames to distinguish their importance. The long short-term memory (LSTM) network and the time pool are used to model long-term dependencies and temporal memory effects. Experimental results on UGC-VQA datasets show that the proposed method achieves a performance improvement of nearly 2%, and its evaluation results are more consistent with human visual perception.
Da Ai, Mingyue Lu, Ying Liu 0026
VCIP4
2023 Perceptual quantization parameter selection for crime scene investigation tool images
Yanchao Gong, Kaifang Yang, Ying Liu 0026, Keng-Pang Lim
Frontiers Comput. Sci.5
2022 A Full-Reference Image Quality Assessment Method with Saliency and Error Feature Fusion
abstract
Image quality assessment (IQA) has obtained certain achievements with the help of convolutional neural network (CNN). To promote the evaluation performance, most existing methods focus on optimizing the structure and parameters of neural networks, while some useful features of image are ignored that can easily be acquired. In this paper, we propose a saliency and error feature fusion IQA (SEFF-IQA) method. Instead of the image itself, two image features, the error between the reference image and the distorted image, and the subjective saliency of distorted image are taken as inputs of the CNN for training. The evaluation score of image quality were obtained by a conventional CNN that trained on frequently used public databases. The proposed method possesses one basic architecture of the CNN only and reduces the volume of training data remarkably compared with state-of-art approaches. Experimental results show that the proposed method is more consistent with human subjective perception than other existing deep learning-based methods.
Da Ai, Yunhong Liu, Yurong Yang, Mingyue Lu, Ying Liu 0026, Nam Ling
ISCAS5
2022 Information Entropy Augmented High Density Crowd Counting Network
abstract
The research proposes an innovated structure of the density map-based crowd counting network augmented by information entropy. The network comprises of a front-end network to extract features and a back-end network to generate density maps. In order to validate the assumption that the entropy can boost the accuracy of density map generation, a multi-scale entropy map extraction process is imported into the front-end network along with a fine-tuned convolutional feature extraction process, In the back-end network, extracted features are decoded into the density map with a multi-column dilated convolution network. Finally, the decoded density map can be mapped as the estimated counting number. Experimental results indicate that the devised network is capable of accurately estimating the count in extremely high crowd density. Compared to similar structured networks which don’t adapt entropy feature, the proposed network exhibits higher performance. This result proves the feature of information entropy is capable of enhancing the efficiency of density map-based crowd counting approaches.
Yu Hao 0002, Lingzhe Wang, Ying Liu 0026, JiuLun Fan 0001
Int. J. Semantic Web Inf. Syst.3
2022 Shoeprint Image Retrieval Based on Dual Knowledge Distillation for Public Security Internet of Things
abstract
In order to implement rapid retrieval of large-scale crime scene investigation shoeprint image (SPI) in the intelligent mobile terminal of Public Security Internet of Things (PSIoT), a novel dual knowledge distillation (DKD) network is designed by fusing spatial attention (SA) distillation and feature distillation to solve its problems of limited computing and storage capacity. First, ResNet50 was modified by adding a new SA module and hash layer as the teacher model, and a light convolutional neural network (CNN) with the same SA and hash layer is designed as the student model. Then, the attention distillation loss function is constructed to distill the SA knowledge in the teacher module to the student module to improve the ability of the convolutional layer at the front end of the student network to capture the underlying visual features of the SPI. Finally, the feature distillation loss function is constructed to distill the semantic knowledge in the teacher module to the student module to improve the ability of the student module to express high-level semantics of the SPI. we compare our method on two data sets of SPID and FID-300 with other state-of-the-art methods in the SPI retrieval domain. The experimental results show that our method can improve the SPI retrieval baseline by a large margin and better than other methods.
Daxiang Li 0002, Yang Li 0161, Ying Liu 0026
IEEE Internet Things J.3
2022 Group Bilinear CNNs for Dual-Polarized SAR Ship Classification
abstract
Ship classification from synthetic aperture radar (SAR) images tends to be a hotspot in the remote sensing community. Currently, more efforts have been made to the single-polarization (single-pol) SAR ship classification with limited performance. This letter proposes to explore the dual-polarization (dual-pol) SAR images for better ship classification. To be specific, a novel group bilinear convolutional neural network (GBCNN) model is developed to deeply extract discriminative second-order representations of ship targets from the pairwise VH and VV polarization SAR images. Particularly, the deep bilinear features are efficiently acquired by performing the bilinear pooling on sub-groups of deep feature maps derived, respectively, from the single-pol SAR images (self-bilinear pooling) and dual-pol SAR images (cross-bilinear pooling). To fully explore the polarization information, the multi-polarization fusion loss (MPFL) is constructed to train the proposed model for superior SAR ship representation learning. By extensive experiments, the proposed method can achieve an overall accuracy of 88.80% and 66.90% on the 3- and 5-category dual-pol OpenSARShip data sets, which outperform the state-of-the-art methods by at least 2.00% and 2.37%, respectively.
Jinglu He, Wenlong Chang, Ying Liu 0026, Yinghua Wang, Hongwei Liu 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Remote Sensing Image Scene Classification Model Based on Dual Knowledge Distillation
abstract
In the application of remote sensing image (RSI) scene classification, in order to solve the contradiction between the accuracy of Convolutional Neural Network (CNN) and the large amount of model parameters, a novel dual knowledge distillation (DKD) model combining dual attention (DA) and spatial structure (SS) is designed. First, new DA and SS modules are constructed and introduced into ResNet101 and light-weight CNN designed as teacher and student networks respectively. Then, in order to improve its local feature extraction and high-level semantic representation abilities for RSI by transmission the DA and SS knowledge in the teacher network to the student network, we design the corresponding DA and SS distillation losses. The comparative experimental results based on AID and NWPU-45 datasets show that when the training ratio is 20%, the accuracy of the student network after DKD is improved by 7.57% and 7.28% respectively, and in the case of fewer parameters, DKD has higher accuracy than most other methods.
Daxiang Li 0002, Yixuan Nan, Ying Liu 0026
IEEE Geosci. Remote. Sens. Lett.3
2022 Joint Processing of Spatial Resolution Enhancement and Spectral Unmixing for Hyperspectral Image
abstract
Spatial resolution enhancement and its subsequent tasks are always separated in conventional hyperspectral image (HSI) processing model. The requirement of the following task, such as unmixing, cannot be referred by spatial resolution enhancement. Moreover, errors and artifacts will also be transmitted and accumulated. In this work, we propose a joint processing method of spatial resolution enhancement and spectral unmixing for HSI (J-SRE-Un), where these two tasks are treated as constraints for each other to simultaneously achieve better performance. Experiments on both simulated and real data demonstrate the effectiveness and superiority of our method.
Ying Liu 0026, Yuquan Gan
IEEE Geosci. Remote. Sens. Lett.2
2022 3D Layout Estimation via Weakly Supervised Learning of Plane Parameters From 2D Segmentation
abstract
The task of 3D layout estimation in an indoor scene is to predict the holistic 3D structural information of the scene from an RGB image. It is costly to obtain the ground truth 3D layout, and this issue severely restricts the learning based 3D layout estimation approaches. In this paper, we present a novel weakly supervised learning framework that is able to learn the 3D layout effectively with 2D layout segmentation mask as supervision. We employ a deep neural network to predict the plane parameters and camera intrinsic parameters in the image. Based on the predicted plane instances, the 3D layout as well as the corresponding depth map and 2D segmentation can be generated. The key objectives for learning meaningful plane parameters are the label consistency of layout segmentation and depth consistency of border pixels from adjacent planes, with which the ground truth 2D layout segmentation is able to supervise the learning of the 3D layout. We further incorporate 3D geometric reasoning and prior knowledge in the learning process to ensure that the learned 3D layout is realistic and reasonable. Experimental results show that our method can produce accurate 3D layout estimates by weakly supervised learning.
Weidong Zhang 0005, Youmei Zhang, Ran Song 0001, Ying Liu 0026, Wei Zhang 0021
IEEE Trans. Image Process.4
2021 An Adaptive Feature-based Quantization Algorithm for Point Cloud Compression
abstract
To reduce over-rasterization distortion caused by global uniform quantization for static surface point cloud, an adaptive quantization coding method based on feature mining is proposed. Combining spatial position and texture feature of point clouds with level of details, the quantization increment is dynamically set according to feature priority, which can reserve the number of effective points to the maximum extent, and reduce the rasterization distortion. Experimental results show that the proposed method can effectively enhance the subjective reconstruction quality of compressed point cloud, gaining better results of rate-distortion optimization.
Da Ai, Hongying Lu, Yurong Yang, Ying Liu 0026
PCS4
2021 Tyre pattern image retrieval - current status and challenges
abstract
Tyre pattern image retrieval (TPIR) is an important tool in the investigation of criminal activities and traffic accidents. Although content-based image retrieval (CBIR) has been developed for decades with abundant results, the study on TPIR which started in the 1990s has not made much progress. The lack of large standard test datasets is a crucial shortcoming which limits the research in this field. Information presented in this paper is a result of the authors’ literature research on recent academic publications and practical field investigation in the public security and transportation sectors. The state-of-the-art technologies in the field of TPIR are surveyed in detail from two aspects of tyre patterns – their low-level spatial features and high-level semantic features. Existing algorithms are examined and their pros and cons are compared and verified through experimental results. This paper also surveys the available tyre pattern datasets used in all available literature. Finally, with the considerations on technology trends in image retrieval and application requirements in TPIR, the future research directions in this field are laid out.
Ying Liu 0026, Qiqi Liu, JiuLun Fan 0001, Jianlong Fu, Yuan Qingan, Tuan Kiang Chiew, Nam Ling
Connect. Sci.1
2021 Quantization Parameter Cascading for Surveillance Video Coding Considering All Inter Reference Frames
abstract
Video surveillance and its applications have become increasingly ubiquitous in modern daily life. In video surveillance system, video coding as a critical enabling technology determines the effective transmission and storage of surveillance videos. In order to meet the real-time or time-critical transmission requirements of video surveillance systems, the low-delay (LD) configuration of the advanced high efficiency video coding (HEVC) standard is usually used to encode surveillance videos. The coding efficiency of the LD configuration is closely related to the quantization parameter (QP) cascading technique which selects or determines the QPs for encoding. However, the quantization parameter cascading (QPC) technique currently adopted for the LD configuration in HEVC test model (i.e., HM) is not optimized since it has not taken full account of the reference dependency in coding. In this paper, an efficient QPC technique for surveillance video coding, referred to as QPC-SV, is proposed, considering all inter reference frames under the LD configuration. Experimental results demonstrate the efficacy of the proposed QPC-SV. Compared with the default configuration of QPC in the HM, the QPC-SV achieves significant rate-distortion performance gain with average BD-rates of -9.35% and -9.76% for the LDP and LDB configurations, respectively.
Yanchao Gong, Kaifang Yang, Ying Liu 0026, Keng-Pang Lim, Nam Ling, Hong Ren Wu
IEEE Trans. Image Process.3
2021 Learning wavelet coefficients for face super-resolution
abstract
Abstract Face image super-resolution imaging is an important technology which can be utilized in crime scene investigations and public security. Modern CNN-based super-resolution produces excellent results in terms of peak signal-to-noise ratio and the structural similarity index (SSIM). However, perceptual quality is generally poor, and the details of the facial features are lost. To overcome this problem, we propose a novel deep neural network to predict the super-resolution wavelet coefficients in order to obtain clearer facial images. Firstly, this paper uses prior knowledge of face images to manually emphases relevant facial features with more attention. Then, a linear low-rank convolution in the network is used. Finally, image edge features from canny detector are applied to enhance super-resolution images during training. The experimental results show that the proposed method can achieve competitive PSNR and SSIM and produces images with much higher perceptual quality.
Ying Liu 0026, Sun Dinghua, Keng-Pang Lim, Tuan Kiang Chiew, Yi Lai
Vis. Comput.1
2020 A Super-Fast Deep Network for Moving Object Detection
abstract
Deep learning methods have been actively applied to intelligent video surveillance for moving object detection in recent years and demonstrated impressive results. However, these models render superior accuracy at the cost of high computational complexity. In this work, we devised a new deep network structure that significantly improves inference speed, yet requires 10 times smaller model size and achieves 10 times reduction in floatingpoint operations as compared to existing deep learning models with tolerable accuracy loss.
Bingxin Hou, Ying Liu 0026, Nam Ling
ISCAS2
2020 Graph convolution network with node feature optimization using cross attention for few-shot learning
abstract
Graph convolution network (GCN) is an important method recently developed for few-shot learning. The adjacency matrix in GCN models is constructed based on graph node features to represent the graph node relationships, according to which, the graph network achieves message-passing inference. Therefore, the representation ability of graph node features is an important factor affecting the learning performance of GCN. This paper proposes an improved GCN model with node feature optimization using cross attention, named GCN-NFO. Leveraging on cross attention mechanism to associate the image features of support set and query set, the proposed model extracts more representative and discriminative salient region features as initialization features of graph nodes through information aggregation. Since graph network can represent the relationship between samples, the optimized graph node features transmit information through the graph network, thus implicitly enhances the similarity of intra-class samples and the dissimilarity of inter-class samples, thus enhancing the learning capability of GCN. Intensive experimental results on image classification task using different image datasets prove that GCN-NFO is an effective few-shot learning algorithm which significantly improves the classification accuracy, compared with other existing models.
Ying Liu 0026, Yanbo Lei, Sheikh Faisal Rashid
MMAsia1
2019 A Rotation Invariant HOG Descriptor for Tire Pattern Image Classification
abstract
Texture feature is important in describing tire pattern image which provides useful clue in solving crime cases and traffic accidents. In this paper, we propose a novel texture feature extraction method based on HOG (Histogram of Oriented Gradient) and dominant gradient (DG) in tire pattern images, named HOG-DG. The proposed HOG-DG is not only robust to illumination and scale changes but also is rotation-invariant. In the proposed HOG-DG, HOG features are first computed from circular local cells, and HOG features from an image are concatenated and normalized using the DG to construct the HOG-DG feature. HOG-DG is used to train a support-vector-machine (SVM) classifier for tire pattern classification. Experimental results demonstrate its outstanding performance for tire pattern description.
Ying Liu 0026, Yuxiang Ge, Qiqi Liu, Yanbo Lei, Dengsheng Zhang, Guojun Lu
ICASSP1
2019 Temporal-Layer-Motivated Lambda Domain Picture Level Rate Control for Random-Access Configuration in H.265/HEVC
abstract
Rate control is a key technique for video communication systems. The aim of rate control is to transmit the best possible quality video sequences under various restrictions, such as channel bandwidth, buffer capacity, maximum time delay allowed for a given service, and so on. The λ domain rate control technique (λ-RC) has been integrated into the latest High Efficiency Video Coding standard (H.265/HEVC) test model, due to its accurate bit estimation and high rate-distortion performance. However, it is found that the λ-RC is not the optimal choice under the random-access configuration. When the random-access configuration is used, pictures are organized into temporal layers, where pictures in different layers are of different importance in terms of prediction. In this paper, a picture level lambda domain rate control technique for the randomaccess configuration in H.265/HEVC is proposed. The influence of temporal layers is effectively considered in the proposed algorithm referred to as TL-λ-PRC. Experimental results verify that the proposed TL-λ-PRC is efficient in coding performance and accurate in bit estimation. Compared with the λ-RC with the fixed ration bit allocation which has been implemented in the test model of H.265/HEVC (HM 14.0), TL-λ-PRC achieves an average reduction of 4.10% and 3.49% for slow motion and fast motion sequences, respectively, in BD-rate (Bjøntegaard-Delta bitrate) with more accurate bit estimation. The performance of different algorithms in terms of algorithm complexity and quality fluctuation are also carefully analyzed in this contribution.
Yanchao Gong, Shuai Wan, Kaifang Yang, Hong Ren Wu, Ying Liu 0026
IEEE Trans. Circuits Syst. Video Technol.5
2019 A novel image retrieval algorithm based on transfer learning and fusion features
Ying Liu 0026, Yanan Peng, Keng-Pang Lim, Nam Ling
World Wide Web1
2018 Fast Single Image Dehazing via Positive Correlation
abstract
In this paper, we propose a fast single image dehazing method based on positive correlation. Firstly, a linear model is built to describe the positive correlation between the minimum channel of the hazy image and its corresponding depth map. Then, the transmission map and the atmospheric light are separately obtained using the created linear model. Finally, based on the traditional atmospheric scattering model, the haze-free image can be recovered with the transmission map and the atmospheric light. Experimental results on numerous hazy images demonstrate that proposed method has better performance and lower time complexity than the state-of-the-art methods.
Bingheng Li, Yi Lai, Chaoyan Wu, Ying Liu 0026
ICPR4
2018 A Graphical Simulator for Modeling Complex Crowd Behaviors
abstract
Abnormal crowd behaviors of varied real-world settings could represent or pose serious threat to public safety. The video data required for relevant analysis are often difficult to acquire due to security, privacy and data protection issues. Without large amounts of realistic crowd data, it is difficult to develop and verify crowd behavioral models, event detection techniques, and corresponding test and evaluations. This paper presented a synthetic method for generating crowd movements and tendency based on existing social and behavioral studies. Graph and tree searching algorithms as well as game engine-enabled techniques have been adopted in the study. The main outcomes of this research include a categorization model for entity-based behaviors following a linear aggregation approach; and the construction of an innovative agent-based pipeline for the synthesis of A-Star path-finding algorithm and an enhanced Social Force Model. A Spatial-Temporal Texture (STT) technique has been adopted for the evaluation of the model's effectiveness. Tests have highlighted the visual similarities between STTs extracted from the simulations and their counterparts - video recordings - from the real-world.
Yu Hao 0002, Zhijie Xu, Ying Liu 0026, Jing Wang 0033, JiuLun Fan 0001
IV3
2014 Pyramid Match Kernel and Classifier Ensemble-Based MIL Algorithm for Pornographic Images Filtering
abstract
In this paper, a novel multi-instance learning (MIL) algorithm based on pyramid match kernel (PMK) and classifier ensemble is proposed for recognizing pornographic scene from image database. First, an improved JSEG image segmentation technique is deployed for dividing every image into several regions, and regards the whole image as a "bag", the low-level visual features (i.e. color and texture) of each segmented region as "instance". As a result, the pornographic images filtering problem can be transferred into a typical MIL problem. Second, similarity between the multi-instance bags is measured by PMK method, which allows MIL problem to be solved directly by the support vector machine (SVM). Finally, many base classifiers based on PMK with different levels are constructed, and the performance weighting rule is used to dynamically determine the weights of them, so the strategy of classifier ensemble is used to improve the filtering accuracy. In a real condition image set that the ratio of normal image to pornographic image is 9:1, experimental results show that the proposed algorithm, named PMKCE-MIL, is robust, and its performance is superior to other algorithms.
Daxiang Li 0002, Jing Wang 0033, Ying Liu 0026
Int. J. Pattern Recognit. Artif. Intell.3
2014 Multiple kernel-based multi-instance learning algorithm for image classification
Daxiang Li 0002, Jing Wang 0033, Ying Liu 0026, Dianwei Wang
J. Vis. Commun. Image Represent.4
2009 A Bayesian approach integrating regional and global features for image semantic learning
abstract
In content-based image retrieval, the ldquosemantic gaprdquo between visual image features and user semantics makes it hard to predict abstract image categories from low-level features. We present a hybrid system that integrates global features (G-features) and region features (R-features) for predicting image semantics. As an intermediary between image features and categories, we introduce the notion of mid-level concepts, which enables us to predict an image's category in three steps. First, a G-prediction system uses G-features to predict the probability of each category for an image. Simultaneously, a R-prediction system analyzes R-features to identify the probabilities of mid-level concepts in that image. Finally, our hybrid H-prediction system based on a Bayesian network reconciles the predictions from both R-prediction and G-prediction to produce the final classifications. Results of experimental validations show that this hybrid system outperforms both G-prediction and R-prediction significantly.
Luong-Dong Nguyen, Ghim-Eng Yap, Ying Liu 0026, Ah-Hwee Tan, Liang-Tien Chia, Joo-Hwee Lim
ICME3
2008 Region-based image retrieval with high-level semantics using decision tree learning
Ying Liu 0026, Dengsheng Zhang, Guojun Lu
Pattern Recognit.1
2007 Integrating Semantic Templates with Decision Tree for Image Semantic Learning
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Ah-Hwee Tan
MMM (2)1
2007 A survey of content-based image retrieval with high-level semantics
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma
Pattern Recognit.1
2006 Study on texture feature extraction in region-based image retrieval system
abstract
Texture is an important feature to describe images. Though lots of work has been done for efficient texture feature extraction from rectangular images, no much effort has been made in texture feature extraction from arbitrary-shaped regions in region-based image retrieval (RBIR) system. In this paper, we present an efficient texture feature extraction algorithm for arbitrary-shaped regions. This algorithm first extends an arbitrary-shaped region into a rectangular area onto which block transformation can be applied. Based on the projection-onto-convex-sets (POCS) theory, a set of coefficients best describing the original region are finally obtained, from which texture feature of the region can be extracted. Via intensive experiments, we select a set of parameters proper for image retrieval purpose. Experimental results on real-world image database demonstrate the effectiveness of the proposed algorithm for image retrieval purpose
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma
MMM1
2005 Deriving High-Level Concepts Using Fuzzy-ID3 Decision Tree for Image Retrieval
abstract
To improve the retrieval accuracy of content-based image retrieval, an important task is to reduce the 'semantic gap' between low-level image features and the richness of human semantics. We present a region-based image retrieval system using high-level semantic concepts. The contribution of the paper is two-fold. First, salient low-level features are extracted from arbitrarily-shaped regions. Second, a fuzzy-ID3 decision tree learning method is proposed to derive association rules which map low-level image features to high-level concepts. Experimental results prove that, by reducing the 'semantic gap', the proposed system not only improves the retrieval accuracy, but also supports users in query-by-keyword.
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma
ICASSP (2)1
2005 Region-Based Image Retrieval with High-Level Semantic Color Names
abstract
Performance of traditional content-based image retrieval systems is far from user’s expectation due to the ‘semantic gap’ between low-level visual features and the richness of human semantics. In attempt to reduce the ‘semantic gap’, this paper introduces a region-based image retrieval system with high-level semantic color names. In this system, database images are segmented into color-texture homogeneous regions. For each region, we define a color name as that used in our daily life. In the retrieval process, images containing regions of same color name as that of the query are selected as candidates. These candidate images are further ranked based on their color and texture features. In this way, the system reduces the ‘semantic gap’ between numerical image features and the rich semantics in the user’s mind. Experimental results show that the proposed system provides promising retrieval results with few features used.
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma
MMM1
2004 Extracting texture features from arbitrary-shaped regions for image retrieval
abstract
Lots of work has been done in texture feature extraction for rectangular images, but not as much attention has been paid to the arbitrary-shaped regions available in region-based image retrieval (RBIR) systems. In This work, we present a texture feature extraction algorithm, based on projection onto convex sets (POCS) theory. POCS iteratively concentrates more and more energy into the selected coefficients from which texture features of an arbitrary-shaped region can be extracted. Experimental results demonstrate the effectiveness of the proposed algorithm for image retrieval purposes.
Ying Liu 0026, Xiaofang Zhou 0001, Wei-Ying Ma
ICME1
2004 Automatic Texture Segmentation for Texture-based Image Retrieval
abstract
Texture-segmentation is the crucial initial step for texture-based image retrieval. Texture is the main difficulty faced to a segmentation method. Many image segmentation algorithms either can't handle texture properly or cannot obtain texture features directly during segmentation which can be used for retrieval purpose. This paper describes an automatic texture segmentation algorithm based on a set of features derived from wavelet domain, which are effective in texture description for retrieval purpose. Simulation results show that the proposed algorithm can efficiently capture the textured regions in arbitrary images, with the features of each region extracted as well. The features of each textured region can be directly used to index image database with applications as texture-based image retrieval.
Ying Liu 0026, Xiaofang Zhou 0001
MMM1
2003 Texture segmentation based on features in wavelet domain for image retrieval
Ying Liu 0026, Xiaofang Zhou 0001
VCIP1