Kai Liu 0032

dblp:73/4566-32 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0001-6301-7756ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Efficient Optimization Criterion for Multi-View Feature Representation Learning
abstract
The training of contemporary machine learning (ML) models, particularly deep neural networks (DNNs), often relies on enormous data sources to properly tune model parameters. As a result, achieving competitive results with limited training data and computational resources has been recognized as a significant bottleneck to advance ML. To address these issues, multi-view representation learning has emerged. However, how to efficiently build multi-view learning models remains a big challenge. In this paper, a novel optimization criterion is proposed to tackle this challenge. Specifically, the proposed criterion ensures speedy and effective parameter selection, reducing the effort to reach optimal design of the model while maintaining performance. To validate the efficiency and generalizability of the presented solution, experiments were conducted on face recognition and few-shot learning for image classification using four databases of different scales. Experimental results demonstrate the superiority of the proposed approach, offering an efficient yet robust solution to the data-scarcity challenge.
Lei Gao 0001, Kai Liu 0032, Kevin Tang, Ling Guan
ISM2
2025 A discriminative multi-modal adaptation neural network model for video action recognition
Lei Gao 0001, Kai Liu 0032, Ling Guan
Neural Networks2
2025 Mathematics-Inspired Models: A Green and Interpretable Learning Paradigm for Multimedia Computing
abstract
The advances of machine learning (ML), and AI in general, have attracted unprecedented attention in intelligent multimedia computing and many other fields. However, due to the concern for sustainability and black-box nature of ML models, especially deep neural networks (DNNs), green and interpretable learnings have been extensively studied in recent years, despite suspicions on effectiveness, subjectivity of interpretability, and complexity. To address these concerns and suspicions, this article starts with a survey on recent discoveries in green learning and interpretable learning and then presents mathematics-inspired (M-I) learning models. We will demonstrate that the M-I models are green in nature with numerous interpretable properties. Finally, we present several examples in multi-view information computing on both static image-based and dynamic video-based tasks to demonstrate that the M-I methodology promises a plausible and sustainable path for natural evolution of ML, which is worth further investment in.
Lei Gao 0001, Kai Liu 0032, Ling Guan
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Towards Efficient Multi-view Representation Learning
abstract
The proliferation of deep neural networks (DNNs) has drawn unprecedented interest in the study of various contents such as image, audio, video, to name a few. However, due to the data-driven nature, the high computational requirement and slow running time are considered as Achilles’ heels of DNN-based algorithms, limiting the progress of DNNs in time-sensitive applications. Recently, distinct discriminant canonical correlation analysis network (DDCCANet), a multi-view neural network, has shown great generalizability across multiple application domains, both analytically and experimentally. However, Although the computational requirement and running time by DDCCANet are more manageable than DNN-based algorithms, they can be substantially further improved. This paper proposes two new algorithms for multi-view feature representation learning, namely incremental DDCCANet (IDDCCANet) with substantial save in computational memory and GPU-accelerated DDCCANet (GADDCCANet) with drastically accelerated running time, forming a practically significant platform for multi-view feature representation learning. To validate the power of the proposed algorithms, experiments are conducted on several data sets with different types of inputs (e.g., raw image pixels, classical and DNN-based features). Experimental results clearly show that the proposed algorithms provide promising solutions to address the two longstanding challenges.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan
ISM1
2021 A two-stream heterogeneous network for action recognition based on skeleton and RGB modalities
abstract
Recent years, skeleton based action recognition with graph convolutional network (GCN) has achieved great success. However, since skeleton data only includes human body joints coordinates, other key information on actions is missing such as the subtle motion of hands, the objects the human is interacting, leading to an unsatisfactory performance. In this respect, the RGB data can offer help to recognize actions that skeleton-based methods have limitations on. In this work, we propose a novel two-stream heterogeneous network consisting of GCN and CNN networks for action recognition. Specifically, the GCN network takes the skeletal sequence as input to exploit skeleton information. For the RGB video, the CNN model, ResNet (2+1)D, is adapted to exploit RGB information. Afterwards, the discriminant canonical correlation analysis (DCCA) method is utilized to integrate the output feature maps from the skeleton and RGB streams, resulting in improved performance. Experimental results on the large-scale dataset NTU RGB+D show that the proposed model outperforms state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISM1
2021 Integrating vertex and edge features with Graph Convolutional Networks for skeleton-based action recognition
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
Neurocomputing1
2021 A Multi-Stream Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action Recognition
abstract
Recently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN with HCRF to retain the human skeleton structure information even during the classification stage. Our model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework which takes the relative coordinate of the joints and bone direction as two static feature streams, and the temporal displacements between two consecutive frames as the dynamic feature stream. Experimental results on three challenging benchmarks (NTU RGB+D, N-UCLA, SYSU) show the superior performance of the proposed model over state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
IEEE Trans. Multim.1
2020 A Vertex-Edge Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
The Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability of exploiting the joint information from the graph structure of the skeleton data. Recently, as a strong and complementary modality for action recognition, the bone information from skeleton data has attracted more attention. However, most existing GCN-based methods extract the bone and joint features with two separate GCN networks, ignoring the dependencies between joints and bones. In this work, a Vertex-Edge Graph Convolutional Network(VE-GCN) is proposed to reveal the information across joints, bones and their relationships simultaneously. In addition, we learn the additional connections among joints and bones for various action samples besides the natural connections of the skeleton. Then we conduct the convolution operation on joints and their neighbors based on these additional connections. Moreover, the conditional random field (CRF) is utilized as the loss function to achieve improved performance. Experimental results on two large-scale datasets NTU RGB+D and NTU RGB+D 120 show that the proposed model outperforms state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISCAS1
2019 Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action Recognition
abstract
Recently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN and HCRF to retain the human skeleton structure information during the classification stage. The proposed model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework that takes the relative coordinates of the joints and bone direction as two static feature streams and the temporal displacements as the dynamic feature stream. Experimental results on two challenging benchmarks (NTU RGB+D, N-UCLA) show the superior performance of the proposed model over state-of-the-art models.
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan
ISM1