Jiale Lu

dblp:310/7338 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-6963-4525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 ST-CGNet: A spatiotemporal gesture recognition network with triplet attention and dual feature fusion
Tingyu Zhou, Jiale Lu, Xingyan Zuo
Pattern Recognit.5
2024 Multi-Prototype Space Learning for Commonsense-Based Scene Graph Generation
abstract
In the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset.
Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He
AAAI4
2024 CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learning
abstract
3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive ground-truth annotations. 2) Closed-set classes training hampers the recognition of novel objects and predicates. Addressing these issues, our idea is to extract cross-modality features by CLIP from text and image data naturally related to 3D point clouds. Cross-modality features are used to train a robust 3D scene graph (3DSG)feature extractor. Specifically, we propose a novel Cross-Modality Contrastive Learning 3DSGG (CCL-3DSGG) method. Firstly, to align the text with 3DSG, the text is parsed into word level that are consistent with the 3DSG annotation. To enhance robustness during the alignment, adjectives are exchanged for different objects as negative samples. Then, to align the image with 3DSG, the camera view is treated as a positive sample and other views as negatives. Lastly, the recognition of novel object and predicate classes is achieved by calculating the cosine similarity between prompts and 3DSG features. Our rigorous experiments confirm the superior open-vocabulary capability and applicability of CCL-3DSGG in real-world contexts.
Lianggangxu Chen, Jiale Lu, Shaohui Lin, Changbo Wang, Gaoqi He
CVPR3
2024 Improving rare relation inferring for scene graph generation using bipartite graph network
Jiale Lu, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Chunhua Gu, Changbo Wang, Gaoqi He
Comput. Vis. Image Underst.1
2023 Scene Graph Generation using Depth-based Multimodal Network
abstract
Scene graph generation (SGG) provides an efficient way for scene understanding. However, it has been plagued by the inaccurate classification of relative spatial relationship and incorrect feature information aggregation from distant objects. In this paper, we innovatively introduce the depth information of objects into SGG and propose a multimodal edge-featured graph attention network (MEGA-Net). MEGA-Net primarily comprises three modules. First, the edge-aware message passing (EMP) module extracts multimodal features and fuses them as edge features in the graph network via a quadrilinear model. Multimodal features consist of depth features, visual features, spatial features, and linguistic features. The depth feature in EMP provides the relative spatial relationship among objects which prevents the tail spatial predicates from being recognized as the head predicates. Second, we propose a depth-based self-supervised graph attention (DSGAT) module to predict the correlation probability between object pairs. By encoding the depth ranking of different object pairs in 2D images, DSGAT learns more accurate directional attention to avoid unrelated neighbors. Third, we introduce a predicate aware loss (PA-Loss) to alleviate the feature redundancy problem caused by extra depth information. This is achieved by introducing semantic frequency information that reflects the priority between different types of relationships. Systematic experiments show that our method achieves state-of-the-art performance on two popular datasets, VG and VRD.
Lianggangxu Chen, Jiale Lu, Changbo Wang, Gaoqi He
ICME2
2023 Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph Generation
abstract
A scene graph generation task is largely restricted under a class imbalance. Previous methods have alleviated the class imbalance problem by incorporating commonsense information into the classification, enabling the prediction model to rectify the incorrect head class into the correct tail class. However, the results of commonsense-based models are typically overcorrected, e.g., the visually correct head class is forcibly modified into the wrong tail class. We argue that there are two principal reasons for this phenomenon. First, existing models ignore the semantic gap between commonsense knowledge and real scenes. Second, current commonsense fusion strategies propagate the neighbors in the visual-linguistic contexts without long-range correlation. To alleviate overcorrection, we formulate the commonsense-based scene graph generation task as two sub-problems: scene-induced commonsense graph generation (SI-CGG) and commonsense-inspired scene graph generation (CI-SGG). In SI-CGG module, unlike conventional methods using fixed commonsense graph, we adaptively adjust the node embeddings in a commonsense graph according to their visual appearance and configure the new reasoning edge under a specific visual context. The CI-SGG module is proposed to propagate the information from scene-induced commonsense graph back to the scene graph. It updates the representations of each node in scene graph by the aggregation of neighbourhood information at different scales. Through maximum likelihood optimisation of the logarithmic Gaussian process, the scene graph automatically adapt to the different neighbors in the visual-linguistic contexts. Systematic experiments on the Visual Genome dataset show that our full method achieves state-of-the-art performance.
Lianggangxu Chen, Jiale Lu, Youqi Song, Changbo Wang, Gaoqi He
ACM Multimedia2
2023 Prior Knowledge-driven Dynamic Scene Graph Generation with Causal Inference
abstract
The task of dynamic scene graph generation (DSGG) aims at constructing a set of frame-level scene graphs for the given video. It suffers from two kinds of spurious correlation problems. First, the spurious correlation between input object pair and predicate label is caused by the biased predicate sample distribution in dataset. Second, the spurious correlation between contextual information and predicate label arises from interference caused by background content in both the current frame and adjacent frames of the video sequence. To alleviate spurious correlations, our work is formulated into two sub-tasks: video-specific commonsense graph generation (VsCG) and causal inference (CI). VsCG module aims to alleviate the first correlation by integrating prior knowledge into prediction. Information of all the frames in current video is used to enhance the commonsense graph constructed from co-occurrence patterns of all training samples. Thus, the commonsense graph has been augmented with video-specific temporal dependencies. Then, a CI strategy with both intervention and counterfactual is used. The intervention component further eliminates the first correlation by forcing the model to consider all possible predicate categories fairly, while the counterfactual component resolves the second correlation by removing the bad effect from context. Comprehensive experiments on the Action Genome dataset show that the proposed method achieves state-of-the-art performance.
Jiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He
ACM Multimedia1
2022 DH-GCN: Saliency-Aware Complex Scene Graph Generation Using Dual-Hierarchy Graph Convolutional Network
abstract
In reality, complex scene plagues numerous scene graph generation models because realistic scene contains myriad of objects and complicated relationships. Most current methods suffer poor performance when encountering complex scenes. We find that there are two principal reasons for this phenomenon. First, the construction of graph loses sight of the hierarchy of objects. Second, there exists redundant information in feature optimization. To facilitate this issue, this paper proposes an innovative dual-hierarchy graph convolutional network (DH-GCN), which is a conceptually elegant and efficient top-down approach. In specific, DH-GCN leverages salient object detector to hierarchize objects and give gist nodes more accurate representation. Moreover, the dual-hierarchy message propagation is designed to refine the representation hierarchically and eliminate redundant information. Systematic experiments on Visual Genome dataset show the superiority of our method over strong baseline methods.
Jiale Lu, Lianggangxu Chen, Yiqing Cai, Haoyue Guan, Changhong Lu, Changbo Wang, Gaoqi He
ICME1
2022 Exploring Contextual Relationships in 3D Cloud Points by Semantic Knowledge Mining
abstract
Abstract 3D scene graph generation (SGG) aims to predict the class of objects and predicates simultaneously in one 3D point cloud scene with instance segmentation. Since the underlying semantic of 3D point clouds is spatial information, recent ideas of the 3D SGG task usually face difficulties in understanding global contextual semantic relationships and neglect the intrinsic 3D visual structures. To build the global scope of semantic relationships, we first propose two types of Semantic Clue (SC) from entity level and path level, respectively. SC can be extracted from the training set and modeled as the co‐occurrence probability between entities. Then a novel Semantic Clue aware Graph Convolution Network (SC‐GCN) is designed to explicitly model each SC of which the message is passed in their specific neighbor pattern. For constructing the interactions between the 3D visual and semantic modalities, a visual‐language transformer (VLT) module is proposed to jointly learn the correlation between 3D visual features and class label embeddings. Systematic experiments on the 3D semantic scene graph (3DSSG) dataset show that our full method achieves state‐of‐the‐art performance.
Lianggangxu Chen, Jiale Lu, Yiqing Cai, Changbo Wang, Gaoqi He
Comput. Graph. Forum2