VLDB 2026 Research / reviewers in the wild / expert
Lijuan Zhou 0002
dblp:77/3457-2
· DBLP profile ↗
17ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0002-6418-6284ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attributed Graph Clustering with Dual Contrastive Regularization
Lijuan Zhou 0002, Changyong Niu |
NLPCC (3) | 1 |
| 2025 | Multi-view attributed graph clustering based on graph diffusion convolution with adaptive fusion
Lijuan Zhou 0002, Zhihong Zhang 0007 |
Expert Syst. Appl. | 1 |
| 2025 | Multi-modal and Multi-part with Skeletons and Texts for Action Recognition
Lijuan Zhou 0002, Xuri Jiao |
Expert Syst. Appl. | 1 |
| 2025 | Generative External Knowledge for Zero-shot Action Recognition
Lijuan Zhou 0002, Jianing Mao, Xinhang Xu |
Expert Syst. Appl. | 1 |
| 2024 | Node Embedding Enhancement Model Based on Joint Optimization of Clustering Distribution
Zhihong Zhang 0007, Lijuan Zhou 0002 |
ICIC (13) | 3 |
| 2024 | Dual-Adaptive Fusion Multi-View Clustering Based on Graph AutoencoderabstractThe widespread application of multi-view graph data has facilitated the development of multi-view graph clustering. Effectively learning multi-view node representations is crucial for discovering inherent patterns in complex systems. However, most existing methods struggle to handle data with both multi-attribute and multi-relation simultaneously, while both attributes and relations are essential for graph clustering. Therefore, this paper proposes a dual-adaptive fusion multi-view clustering method based on graph autoencoder. It utilizes multi-view encoders and decoders to encode and reconstruct inputs separately. Additionally, a dual-adaptive fusion module is introduced to integrate multi-view node representations. Through consistency clustering, the proposed method explores the probability distribution consistency among different views, thereby achieving consistent clustering results. Experimental results on three datasets demonstrate the effectiveness of the proposed method in clustering tasks. Changyong Niu, Lijuan Zhou 0002 |
IJCNN | 3 |
| 2024 | Multi-Modal Transformer with Skeleton and Text for Action RecognitionabstractDynamic skeleton data has been widely used for human action recognition due to its high-level semantic information and environmental robustness, represented as the 2D/3D coordinates of human joints. However, previous methods mostly utilized skeleton data only without considering the crucial role of text information in helping machines understand visual contents. This paper proposes a novel method based on multi-modal Transformer with skeleton and text (namely MMT-ST) for action recognition. The proposed method performs action captioning and recognition tasks simultaneously, which dynamically updates action recognition based on the results of action captioning. MMT-ST employs a transformer as the backbone and consists of four components: two single-modal encoders, a cross encoder, and a decoder. The single-modal encoders respectively embed skeletons and texts. The cross encoder aims to learn the underlying correlations between two modalities and further perform action recognition task through a classification head. The decoder is employed to conduct the action captioning task. Additionally, a two-stage training strategy is employed to ensure smoother model training. Extensive experiments conducted on NTU RGB+D, NTU RGB+D 120 and ETRI-Activity 3D datasets demonstrate the effectiveness of the proposed method. Lijuan Zhou 0002, Xuri Jiao |
IJCNN | 1 |
| 2024 | Adaptable Weighted Voting Fusion for Multi-modality-based Action RecognitionabstractIn action recognition tasks, voting fusion can be used to combine classification results from multiple modalities to improve recognition accuracy and robustness. This paper proposes a novel weighted voting fusion method for multi-modality-based action recognition, which includes a weight generation method and three fusion strategies based on these weights. For weight generation, action instances are first classified based on single modality to obtain prediction scores for each action class. An adaptive weight for each modality is then generated by assigning a higher value to the modality with better classification, which is used for balancing the modality contributions of different actions. In the fusion stage, three fusion strategies are proposed to apply the adaptive weights to obtain the final class label, including maximum fusion, elimination weighted voting and maximum weighted voting. Experiments conducted on three kinds of modality fusion demonstrate the effectiveness of the proposed method. Lijuan Zhou 0002, Changyong Niu |
IJCNN | 1 |
| 2024 | Static graph convolution with learned temporal and channel-wise graph topology generation for skeleton-based action recognition
Chuankun Li, Shuai Li 0005, Yanbo Gao, Lijuan Zhou 0002, Wanqing Li 0001 |
Comput. Vis. Image Underst. | 4 |
| 2023 | Attributed Multi-relational Graph Embedding Based on GCN
Zhuo Xie, Guoping Zhao, Lijuan Zhou 0002, Zhaohui Gong, Zhihong Zhang 0007 |
ICIC (2) | 4 |
| 2023 | Joint Node Representation Learning and Clustering for Attributed Graph via Graph Diffusion ConvolutionabstractIn recent years, the representation learning method based on graph convolution network has made the latest achievements in attributed graph clustering. However, these methods only deal with clustering as a downstream task, and better performance can be achieved if clustering is combined with the node learning representation process. This paper proposes a novel method of graph clustering based on graph diffusion convolution network, which jointly conducts node representation learning and clustering. The graph diffusion strategy is applied on the node attributes to assign the near-by nodes with high weights in the step of feature propagation. The output nodes after diffusion were reconstructed by a linear encoder-decoder, and also could be represented by the clusters. The joint learning is achieved by minimizing both errors of reconstruction and representation. Experiments conducted on three public datasets and three real datasets from Zhengzhou Commodity Exchange demonstrate the effectiveness of the proposed method in the task of node clustering. Lijuan Zhou 0002, Zhihong Zhang 0007 |
IJCNN | 4 |
| 2023 | Improving Class Representation for Zero-Shot Action RecognitionabstractZero-Shot Action Recognition (ZSAR) enables models to infer new action classes from previously seen data without any samples of those new classes. How an action class is represented in an understandable and processable format influences the performance in ZSAR. Semantic representations of action classes have been made in various forms, such as attributes, class labels, and text descriptions, while in video recognition, the action classes can also have visual representations in the form of images. This paper proposes a novel method by improving class representation for ZSAR. On the one hand, to improve the collection and quality of text descriptions, this paper uses ChatGPT to generate descriptions and designs conversation-based text prompts that can quickly obtain high-quality descriptions of many actions. On the other hand, to overcome the ambiguity of single-modal class representation, we propose the Image-based Description Refinement (IDR) method to obtain multimodal class representation. Specifically, action classes are represented by relevant images from the web and descriptions, and action videos are represented by spatio-temporal features and extracted objects. By training on the seen set to learn the mapping of multimodal representations for classes and videos, we can infer video classes on the unseen set from the similarity of the mapped representations. Experiments on two popular benchmarks and two elderly daily activity datasets show the effectiveness of our method. In particular, it has a significant improvement in the case of less available video samples. Lijuan Zhou 0002, Jianing Mao |
MMAsia | 1 |
| 2023 | Learning body part-based pose lexicons for semantic action recognitionabstractAbstract Semantic action recognition aims to classify actions based on the associated semantics, which can be applied in video captioning and human‐machine interaction. In this paper the problem is addressed by jointly learning multiple pose lexicons based on multiple body parts. Specifically, multiple visual pose models are learnt, and one visual pose model is associated with one body part, which characterises the likelihood of an observed video frame being generated from hidden visual poses. Moreover, multiple pose lexicon models are simultaneously learnt along with visual pose models. One pose lexicon model is associated with one body part that establishes a probabilistic mapping between the hidden visual poses and semantic poses parsed from textual instructions. To capture the temporal relations among body parts, a transition model is also learnt to measure the probability of the alignment transitioned from one position to another position. The body part‐based pose lexicon learning provides a novel method of cross‐modality semantic correlation, which can be applied in other spatial and temporal data. Action classification is finally formulated as the problem of finding the maximum posterior probability that a given multiple sequences of visual frames follow multiple sequences of semantic poses, subject to the most likely visual pose sequences and alignment sequences. Experiments were conducted on five action datasets to validate the effectiveness of the proposed method. Lijuan Zhou 0002 |
IET Comput. Vis. | 1 |
| 2020 | Jointly Learning Visual Poses and Pose Lexicon for Semantic Action RecognitionabstractA novel method for semantic action recognition through learning a pose lexicon is presented in this paper. A pose lexicon comprises a set of semantic poses, a set of visual poses, and a probabilistic mapping between the visual and semantic poses. This paper assumes that both the visual poses and mapping are hidden and proposes a method to simultaneously learn a visual pose model that estimates the likelihood of an observed video frame being generated from hidden visual poses, and a pose lexicon model establishes the probabilistic mapping between the hidden visual poses and the semantic poses parsed from textual instructions. Specifically, the proposed method consists of two-level hidden Markov models. One level represents the alignment between the visual poses and semantic poses. The other level represents a visual pose sequence, and each visual pose is modeled as a Gaussian mixture. An expectation-maximization algorithm is developed to train a pose lexicon. With the learned lexicon, action classification is formulated as a problem of finding the maximum posterior probability of a given sequence of video frames that follows a given sequence of semantic poses, constrained by the most likely visual pose and the alignment sequences. The proposed method was evaluated on MSRC-12, WorkoutSU-10, WorkoutUOW-18, Combined-15, Combined-17, and Combined-50 action datasets using cross-subject, cross-dataset, zero-shot, and seen/unseen protocols. Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Semantic action recognition by learning a pose lexicon
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
Pattern Recognit. | 1 |
| 2016 | Learning a pose lexicon for semantic action recognitionabstractThis paper presents a novel method for learning a pose lexicon comprising semantic poses defined by textual instructions and their associated visual poses defined by visual features. The proposed method simultaneously takes two input streams, semantic poses and visual pose candidates, and statistically learns a mapping between them to construct the lexicon. With the learned lexicon, action recognition can be cast as the problem of finding the maximum translation probability of a sequence of semantic poses given a stream of visual pose candidates. Experiments evaluating pre-trained and zero-shot action recognition conducted on MSRC-12 gesture and WorkoutSu-10 exercise datasets were used to verify the efficacy of the proposed method. Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona |
ICME | 1 |
| 2011 | Studies on the Automatic Recognition of Modern Chinese Conjunction Usages
Hongying Zan, Lijuan Zhou 0002, Kunli Zhang |
ICIC (1) | 2 |