VLDB 2026 Research / reviewers in the wild / expert
Lianggangxu Chen
dblp:303/6651
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-4131-0828ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data augmentation with attention framework for robust deepfake detection
Sardor Mamarasulov, Lianggangxu Chen, Changgu Chen, Changbo Wang |
Vis. Comput. | 2 |
| 2024 | Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationabstractIn the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset. Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He |
AAAI | 1 |
| 2024 | Kumaraswamy Wavelet for Heterophilic Scene Graph GenerationabstractGraph neural networks (GNNs) has demonstrated its capabilities in the field of scene graph generation (SGG) by updating node representations from neighboring nodes. Actually it can be viewed as a form of low-pass filter in the spatial domain, which smooths node feature representation and retains commonalities among nodes. However, spatial GNNs does not work well in the case of heterophilic SGG in which fine-grained predicates are always connected to a large number of coarse-grained predicates. Blind smoothing undermines the discriminative information of the fine-grained predicates, resulting in failure to predict them accurately. To address the heterophily, our key idea is to design tailored filters by wavelet transform from the spectral domain. First, we prove rigorously that when the heterophily on the scene graph increases, the spectral energy gradually shifts towards the high-frequency part. Inspired by this observation, we subsequently propose the Kumaraswamy Wavelet Graph Neural Network (KWGNN). KWGNN leverages complementary multi-group Kumaraswamy wavelets to cover all frequency bands. Finally, KWGNN adaptively generates band-pass filters and then integrates the filtering results to better accommodate varying levels of smoothness on the graph. Comprehensive experiments on the Visual Genome and Open Images datasets show that our method achieves state-of-the-art performance. Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
AAAI | 1 |
| 2024 | CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learningabstract3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive ground-truth annotations. 2) Closed-set classes training hampers the recognition of novel objects and predicates. Addressing these issues, our idea is to extract cross-modality features by CLIP from text and image data naturally related to 3D point clouds. Cross-modality features are used to train a robust 3D scene graph (3DSG)feature extractor. Specifically, we propose a novel Cross-Modality Contrastive Learning 3DSGG (CCL-3DSGG) method. Firstly, to align the text with 3DSG, the text is parsed into word level that are consistent with the 3DSG annotation. To enhance robustness during the alignment, adjectives are exchanged for different objects as negative samples. Then, to align the image with 3DSG, the camera view is treated as a positive sample and other views as negatives. Lastly, the recognition of novel object and predicate classes is achieved by calculating the cosine similarity between prompts and 3DSG features. Our rigorous experiments confirm the superior open-vocabulary capability and applicability of CCL-3DSGG in real-world contexts. Lianggangxu Chen, Jiale Lu, Shaohui Lin, Changbo Wang, Gaoqi He |
CVPR | 1 |
| 2024 | FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion ModelsabstractIn recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset, which diverges from user input prompts. The underlying reason behind the inaccurate generated results lies in the model's difficulty in sampling from specific intervals of the initial noise distribution corresponding to the prompt. Moreover, it is challenging to directly optimize the initial distribution, given that the diffusion process involves multiple denoising steps. In this paper, we introduce a Fine-tuning Initial Noise Distribution (FIND) framework with policy optimization, which unleashes the powerful potential of pre-trained diffusion networks by directly optimizing the initial distribution to align the generated contents with user-input prompts. To this end, we first reformulate the diffusion denoising procedure as a one-step Markov decision process and employ policy optimization to directly optimize the initial distribution. In addition, a dynamic reward calibration module is proposed to ensure training stability during optimization. Furthermore, we introduce a ratio clipping algorithm to utilize historical data for network training and prevent the optimized distribution from deviating too far from the original policy to restrain excessive optimization magnitudes. Extensive experiments demonstrate the effectiveness of our method in both text-to-image and text-to-video tasks, surpassing SOTA methods in achieving consistency between prompts and the generated content. Our method achieves 10 times faster than the SOTA approach. Changgu Chen, Libing Yang, Lianggangxu Chen, Gaoqi He, Changbo Wang, Yang Li 0041 |
ACM Multimedia | 4 |
| 2024 | Improving rare relation inferring for scene graph generation using bipartite graph network
Jiale Lu, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Chunhua Gu, Changbo Wang, Gaoqi He |
Comput. Vis. Image Underst. | 2 |
| 2023 | Explicit Invariant Feature Induced Cross-Domain Crowd CountingabstractCross-domain crowd counting has shown progressively improved performance. However, most methods fail to explicitly consider the transferability of different features between source and target domains. In this paper, we propose an innovative explicit Invariant Feature induced Cross-domain Knowledge Transformation framework to address the inconsistent domain-invariant features of different domains. The main idea is to explicitly extract domain-invariant features from both source and target domains, which builds a bridge to transfer more rich knowledge between two domains. The framework consists of three parts, global feature decoupling (GFD), relation exploration and alignment (REA), and graph-guided knowledge enhancement (GKE). In the GFD module, domain-invariant features are efficiently decoupled from domain-specific ones in two domains, which allows the model to distinguish crowds features from backgrounds in the complex scenes. In the REA module both inter-domain relation graph (Inter-RG) and intra-domain relation graph (Intra-RG) are built. Specifically, Inter-RG aggregates multi-scale domain-invariant features between two domains and further aligns local-level invariant features. Intra-RG preserves taskrelated specific information to assist the domain alignment. Furthermore, GKE strategy models the confidence of pseudolabels to further enhance the adaptability of the target domain. Various experiments show our method achieves state-of-theart performance on the standard benchmarks. Code is available at https://github.com/caiyiqing/IF-CKT. Yiqing Cai, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Changhong Lu, Changbo Wang, Gaoqi He |
AAAI | 2 |
| 2023 | Learning Local Features of Motion Chain for Human Motion Prediction
Lianggangxu Chen, Chen Li 0035, Changbo Wang, Gaoqi He |
CGI (3) | 2 |
| 2023 | Scene Graph Generation using Depth-based Multimodal NetworkabstractScene graph generation (SGG) provides an efficient way for scene understanding. However, it has been plagued by the inaccurate classification of relative spatial relationship and incorrect feature information aggregation from distant objects. In this paper, we innovatively introduce the depth information of objects into SGG and propose a multimodal edge-featured graph attention network (MEGA-Net). MEGA-Net primarily comprises three modules. First, the edge-aware message passing (EMP) module extracts multimodal features and fuses them as edge features in the graph network via a quadrilinear model. Multimodal features consist of depth features, visual features, spatial features, and linguistic features. The depth feature in EMP provides the relative spatial relationship among objects which prevents the tail spatial predicates from being recognized as the head predicates. Second, we propose a depth-based self-supervised graph attention (DSGAT) module to predict the correlation probability between object pairs. By encoding the depth ranking of different object pairs in 2D images, DSGAT learns more accurate directional attention to avoid unrelated neighbors. Third, we introduce a predicate aware loss (PA-Loss) to alleviate the feature redundancy problem caused by extra depth information. This is achieved by introducing semantic frequency information that reflects the priority between different types of relationships. Systematic experiments show that our method achieves state-of-the-art performance on two popular datasets, VG and VRD. Lianggangxu Chen, Jiale Lu, Changbo Wang, Gaoqi He |
ICME | 1 |
| 2023 | Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph GenerationabstractA scene graph generation task is largely restricted under a class imbalance. Previous methods have alleviated the class imbalance problem by incorporating commonsense information into the classification, enabling the prediction model to rectify the incorrect head class into the correct tail class. However, the results of commonsense-based models are typically overcorrected, e.g., the visually correct head class is forcibly modified into the wrong tail class. We argue that there are two principal reasons for this phenomenon. First, existing models ignore the semantic gap between commonsense knowledge and real scenes. Second, current commonsense fusion strategies propagate the neighbors in the visual-linguistic contexts without long-range correlation. To alleviate overcorrection, we formulate the commonsense-based scene graph generation task as two sub-problems: scene-induced commonsense graph generation (SI-CGG) and commonsense-inspired scene graph generation (CI-SGG). In SI-CGG module, unlike conventional methods using fixed commonsense graph, we adaptively adjust the node embeddings in a commonsense graph according to their visual appearance and configure the new reasoning edge under a specific visual context. The CI-SGG module is proposed to propagate the information from scene-induced commonsense graph back to the scene graph. It updates the representations of each node in scene graph by the aggregation of neighbourhood information at different scales. Through maximum likelihood optimisation of the logarithmic Gaussian process, the scene graph automatically adapt to the different neighbors in the visual-linguistic contexts. Systematic experiments on the Visual Genome dataset show that our full method achieves state-of-the-art performance. Lianggangxu Chen, Jiale Lu, Youqi Song, Changbo Wang, Gaoqi He |
ACM Multimedia | 1 |
| 2023 | Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceabstractThe task of dynamic scene graph generation (DSGG) aims at constructing a set of frame-level scene graphs for the given video. It suffers from two kinds of spurious correlation problems. First, the spurious correlation between input object pair and predicate label is caused by the biased predicate sample distribution in dataset. Second, the spurious correlation between contextual information and predicate label arises from interference caused by background content in both the current frame and adjacent frames of the video sequence. To alleviate spurious correlations, our work is formulated into two sub-tasks: video-specific commonsense graph generation (VsCG) and causal inference (CI). VsCG module aims to alleviate the first correlation by integrating prior knowledge into prediction. Information of all the frames in current video is used to enhance the commonsense graph constructed from co-occurrence patterns of all training samples. Thus, the commonsense graph has been augmented with video-specific temporal dependencies. Then, a CI strategy with both intervention and counterfactual is used. The intervention component further eliminates the first correlation by forcing the model to consider all possible predicate categories fairly, while the counterfactual component resolves the second correlation by removing the bad effect from context. Comprehensive experiments on the Action Genome dataset show that the proposed method achieves state-of-the-art performance. Jiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
ACM Multimedia | 2 |
| 2023 | Video-based spatio-temporal scene graph generation with efficient self-supervision tasks
Lianggangxu Chen, Yiqing Cai, Changhong Lu, Changbo Wang, Gaoqi He |
Multim. Tools Appl. | 1 |
| 2022 | DH-GCN: Saliency-Aware Complex Scene Graph Generation Using Dual-Hierarchy Graph Convolutional NetworkabstractIn reality, complex scene plagues numerous scene graph generation models because realistic scene contains myriad of objects and complicated relationships. Most current methods suffer poor performance when encountering complex scenes. We find that there are two principal reasons for this phenomenon. First, the construction of graph loses sight of the hierarchy of objects. Second, there exists redundant information in feature optimization. To facilitate this issue, this paper proposes an innovative dual-hierarchy graph convolutional network (DH-GCN), which is a conceptually elegant and efficient top-down approach. In specific, DH-GCN leverages salient object detector to hierarchize objects and give gist nodes more accurate representation. Moreover, the dual-hierarchy message propagation is designed to refine the representation hierarchically and eliminate redundant information. Systematic experiments on Visual Genome dataset show the superiority of our method over strong baseline methods. Jiale Lu, Lianggangxu Chen, Yiqing Cai, Haoyue Guan, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 2 |
| 2022 | Exploring Contextual Relationships in 3D Cloud Points by Semantic Knowledge MiningabstractAbstract 3D scene graph generation (SGG) aims to predict the class of objects and predicates simultaneously in one 3D point cloud scene with instance segmentation. Since the underlying semantic of 3D point clouds is spatial information, recent ideas of the 3D SGG task usually face difficulties in understanding global contextual semantic relationships and neglect the intrinsic 3D visual structures. To build the global scope of semantic relationships, we first propose two types of Semantic Clue (SC) from entity level and path level, respectively. SC can be extracted from the training set and modeled as the co‐occurrence probability between entities. Then a novel Semantic Clue aware Graph Convolution Network (SC‐GCN) is designed to explicitly model each SC of which the message is passed in their specific neighbor pattern. For constructing the interactions between the 3D visual and semantic modalities, a visual‐language transformer (VLT) module is proposed to jointly learn the correlation between 3D visual features and class label embeddings. Systematic experiments on the 3D semantic scene graph (3DSSG) dataset show that our full method achieves state‐of‐the‐art performance. Lianggangxu Chen, Jiale Lu, Yiqing Cai, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 1 |
| 2021 | Social-Scene-Aware Generative Adversarial Networks for Pedestrian Trajectory Prediction
Binhao Huang, Zhenwei Ma, Lianggangxu Chen, Gaoqi He |
CGI | 3 |
| 2021 | Leveraging Intra-Domain Knowledge to Strengthen Cross-Domain Crowd CountingabstractUnsupervised cross-domain counting research using synthetic datasets becomes imminent when considering the laborious labeling for supervised methods. However, the existing methods only focus on learning domain shared knowledge to narrow the gap between the source domain and target domain (inter-domain gap). Nevertheless, these methods do not consider the enormous distribution gap among the target domain data itself (intra-domain gap). In this paper, we propose a two-step domain adaptation method with multi-level feature response branches, which further uses the intra-domain knowledge to strengthen the target domain’s adaptability. Specifically, we first use different feature response branches to learn inter-domain knowledge more robustly, reducing the prediction inconsistency of different scenarios. Subsequently, the trained model is used to generate pseudo-labels for the target domain. The entire model was retrained by using pseudo-labels. Various experiments on synthetic dataset GCC and three real public datasets validate our proposed method’s availability with higher accuracy. Yiqing Cai, Lianggangxu Chen, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 2 |