VLDB 2026 Research / reviewers in the wild / expert
Rongquan Wang
dblp:227/1635
· DBLP profile ↗
22ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-3375-9561ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A cognition-inspired multimodal framework with association features and pyramid graph fusion network for personality prediction
Rongquan Wang, Xianyu Xu, Faten S. Alamri, Erik Cambria |
Expert Syst. Appl. | 1 |
| 2026 | GraphMamba: A two-stage hybrid framework for multimodal sentiment analysis
Rongquan Wang, Xinru Wu, Erik Cambria |
Knowl. Based Syst. | 1 |
| 2025 | A novel multimodal personality prediction method based on pretrained models and graph relational transformer networkabstractMultimodal personality analysis aims to identify and express human personality traits in videos. However, RNN and its variants have a limited ability to learn long-term temporal dependencies and existing methods neglect bimodal association features. Based on the fact that visual modalities play a dominant role in this task. Therefore, we propose a personality prediction method catering to learning and fusing intra-modal and intermodal feature dynamics. We first utilize pretrained models’ encoders to extract unimodal spatial scene features from videos. Then, we use xLSTM to capture sequence dependencies between different scene frames used as scene features. Meanwhile, we design a graph relational transformer network to learn longer intra-modal temporal interaction in three unimodal spatial features. Then, we calculate the similarity scores between visual and audio or text features as bimodal association features. Second, we design a multimodal attention feature fusion module to determine the contribution of each feature and aggregate these features. Finally, the MLP model is trained and used to predict scores for personality traits. Experiments on two benchmark datasets demonstrate that our method outperforms the existing methods and achieves state-of-the-art performance. Our code is available at https://github.com/RongquanWang/MP-PMGRT. Rongquan Wang, Xianyu Xu, Huimin Ma 0001 |
ICASSP | 1 |
| 2025 | DADet: Safeguarding Image Conditional Diffusion Models Against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection
Xinlong Ding, Jiawei Li 0016, Yudong Zhang 0008, Rongquan Wang, Huimin Ma 0001, Jiansheng Chen 0001 |
ICCV | 6 |
| 2025 | Puzzle-MAE: A Puzzle-Inspired Mask Autoencoder for Multi-Modal FusionabstractMost unsupervised methods in the video domain rely on simple encoder-decoder structures, often resulting in discrepancies between the features extracted from unmasked patches and those from the original patches. To address this issue, we propose a novel self-supervised learning framework, PuzzleMAE, which extracts features from both masked and unmasked patches and aligns them with original image representations to improve feature consistency. Inspired by the human ability to solve puzzles through holistic image recognition and the exploitation of spatial adjacency, we propose the Global-Local Attention Module, which effectively integrates global contextual information with local feature representations. Furthermore, we introduce 3D Relative Position Embedding and Structural Position Embedding to emulate human-like spatial and structural awareness of positional relationships during the puzzle-solving process. The effectiveness of our method is validated on two downstream tasks: the First Impression V2 and DFEW datasets. Xin Li 0034, Bochao Zou, Rongquan Wang, Huimin Ma 0001 |
ICME | 3 |
| 2025 | A multimodal personality prediction framework based on adaptive graph transformer network and multi-task learningabstractAbstract Multimodal personality analysis targets accurately detecting personality traits by incorporating related multimodal information. However, existing methods focus on unimodal features while overlooking the bimodal association features crucial for this interdisciplinary task. Therefore, we propose a multimodal personality prediction framework based on an adaptive graph transformer network and multi‐task learning. Firstly, we utilize pre‐trained models to learn specific representations from different modalities. Here, we employ pre‐trained multimodal models' encoders as the backbones of the modality‐specific extraction methods to mine unimodal features. Specifically, we introduce a novel adaptive graph transformer network to mine personality‐related bimodal association features. This network effectively learns higher‐order temporal dependencies based on relational graphs and emphasizes more significant features. Furthermore, we utilize a multimodal channel attention residual fusion module to obtain the fused features, and we propose a multimodal and unimodal joint learning regression head to learn and predict scores for personality traits. We design a multi‐task loss function to enhance the robustness and accuracy of personality prediction. Experimental results on the two benchmark datasets demonstrate the effectiveness of our framework, which outperforms the state‐of‐the‐art methods. The code is available at https://github.com/RongquanWang/PPF-AGTNMTL . Rongquan Wang, Xi-Le Zhao, Xianyu Xu |
Comput. Graph. Forum | 1 |
| 2024 | PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizesabstract3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target dataset, particularly towards object sizes. State-of-the-art unsupervised domain adaptation approaches explicitly address the variation in object sizes by appropriately scaling the source data. However, such methods require additional target domain statistics information, which contradicts the original unsupervised assumption. In this work, we present PLS, a novel unsupervised domain adaptation method for 3D object detection to overcome the object sizes bias via Pseudo-Label Sizes, which utilizes only source domain annotations. PLS alternates between generating high-quality pseudo-label sizes through the detector and model training with the pseudo-label sizes to scale and augment the source data. This iterative process enables the detector to be trained with augmented data that resembles the target domain sizes, thereby improving the performance of detector in cross-domain scenarios. Our experimental results show the outstanding performance of our PLS in various scenarios. In addition, PLS is a plug-and-play module that can be used to directly replace existing weakly-supervised scaling methods. Experimental results show that existing excellent architectures with PLS are able to achieve better performance, and making them completely unsupervised. Rongquan Wang, Xin Li 0034, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001 |
ICASSP | 2 |
| 2024 | EMo Transformer: Transformer-Based Depression Detection via Eye MovementsabstractDepressive disorder has become a prevalent psychological illness that significantly impacts individuals’ daily lives. Traditional questionnaire assessment and clinical interviews suffer from issues such as subjectivity and a high consumption of medical resources. With the advancement of artificial intelligence, there is a growing number of depression detection methods based on statistical features. However, these methods have problems of insufficient stimulus extraction and neglecting temporal information. In order to solve these problems, we propose a transformer-based model named EMo Transformer, designed for detecting depression by effectively extracting features from stimuli and combining them with eye movements. Additionally, due to challenge in collecting data from depression patients, we design a simple and effective data augmentation method to solve this challenge. Subsequently, we design an ensemble model using the models with and without data augmentation. The experimental results of accuracy 91.95% demonstrate that our method is effective. Xin Li 0034, Haizhuang Liu, Rongquan Wang, Bochao Zou, Huimin Ma 0001 |
ICME | 3 |
| 2024 | CMT: Co-training Mean-Teacher for Unsupervised Domain Adaptation on 3D Object Detection
Junbao Zhuo, Xin Li 0034, Haizhuang Liu, Rongquan Wang, Jiansheng Chen 0001, Huimin Ma 0001 |
ACM Multimedia | 5 |
| 2023 | Multi-source Information Fusion for Depression Detection
Rongquan Wang, Huiwei Wang, Huimin Ma 0001 |
PRCV (5) | 1 |
| 2022 | Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial LabelsabstractPrevious weakly-supervised methods of 3D object detection in driving scenes mainly rely on spatial labels, which provide the location, dimension, or orientation information. The annotation of 3D spatial labels is time-consuming. There also exist methods that do not require spatial labels, but their detections may fall on object parts rather than entire objects or backgrounds. In this paper, a novel cross-modal weakly-supervised 3D progressive refinement framework (WS3DPR) for 3D object detection that only needs image-level class annotations is introduced. The proposed framework consists of two stages: 1) classification refinement for potential objects localization and 2) regression refinement for spatial pseudo labels reasoning. In the first stage, a region proposal network is trained by cross-modal class knowledge transferred from 2D image to 3D point cloud and class information propagation. In the second stage, the locations, dimensions, and orientations of 3D bounding boxes are further refined with geometric reasoning based on 2D frustum and 3D region. When only image-level class labels are available, proposals with different 3D locations become overlapped in 2D, leading to the misclassification of foreground objects. Therefore, a 2D-3D semantic consistency block is proposed to disentangle different 3D proposals after projection. The overall framework progressively learns features in a coarse to fine manner. Comprehensive experiments on the KITTI3D dataset demonstrate that our method achieves competitive performance compared with previous methods with a lightweight labeling process. Haizhuang Liu, Huimin Ma 0001, Bochao Zou, Rongquan Wang, Jiansheng Chen 0001 |
ACM Multimedia | 6 |
| 2022 | Detecting protein complexes with multiple properties by an adaptive harmony search algorithmabstractBACKGROUND: Accurate identification of protein complexes in protein-protein interaction (PPI) networks is crucial for understanding the principles of cellular organization. Most computational methods ignore the fact that proteins in a protein complex have a functional similarity and are co-localized and co-expressed at the same place and time, respectively. Meanwhile, the parameters of the current methods are specified by users, so these methods cannot effectively deal with different input PPI networks. RESULT: To address these issues, this study proposes a new method called MP-AHSA to detect protein complexes with Multiple Properties (MP), and an Adaptation Harmony Search Algorithm is developed to optimize the parameters of the MP algorithm. First, a weighted PPI network is constructed using functional annotations, and multiple biological properties and the Markov cluster algorithm (MCL) are used to mine protein complex cores. Then, a fitness function is defined, and a protein complex forming strategy is designed to detect attachment proteins and form protein complexes. Next, a protein complex filtering strategy is formulated to filter out the protein complexes. Finally, an adaptation harmony search algorithm is developed to determine the MP algorithm's parameters automatically. CONCLUSIONS: Experimental results show that the proposed MP-AHSA method outperforms 14 state-of-the-art methods for identifying protein complexes. Also, the functional enrichment analyses reveal that the protein complexes identified by the MP-AHSA algorithm have significant biological relevance. Rongquan Wang, Huimin Ma 0001 |
BMC Bioinform. | 1 |
| 2022 | Correction: Detecting protein complexes with multiple properties by an adaptive harmony search algorithm
Rongquan Wang, Huimin Ma 0001 |
BMC Bioinform. | 1 |
| 2022 | Visibility of points: Mining occlusion cues for monocular 3D object detection
Huazhen Chu, Lisha Mo, Rongquan Wang, Huimin Ma 0001 |
Neurocomputing | 3 |
| 2022 | Learning representative viewpoints in 3D shape recognition
Huazhen Chu, Chao Le, Rongquan Wang, Xi Li 0010, Huimin Ma 0001 |
Vis. Comput. | 3 |
| 2021 | Device-Adaptive 2D Gaze Estimation: A Multi-Point Differential Framework
Runtong Li, Huimin Ma 0001, Rongquan Wang |
ICIG (2) | 3 |
| 2021 | Depression Detection by Combining Eye Movement with Image SemanticsabstractDepression is a common mental disorder that affects patients’ daily life. Most existing depression detection methods consume a lot of medical resources and exist at risk of subjective judgment. Therefore, we propose an objective and convenient experimental paradigm. Firstly, it selects emotional images as stimuli and records the subjects’ eye movement data. Secondly, we establish a connection between image processing and subjects’ psychological conditions analysis. Rather than some AI-based methods focus on feature engineering of recorded data, we design the saliency difference detection network and semantic segmentation network to explore the images’ deep semantic features and combine them with the subjects’ gaze pattern. Finally, we train a mental state classifier of Support Vector Machine to detect depression. The experimental results demonstrate that it achieves accuracy up to 90.06%, which outperforms previous methods. Huimin Ma 0001, Zeyu Pan, Rongquan Wang |
ICIP | 4 |
| 2021 | LiDAR-Based Symmetrical Guidance for 3D Object Detection
Huazhen Chu, Huimin Ma 0001, Haizhuang Liu, Rongquan Wang |
PRCV (4) | 4 |
| 2021 | STA-GCN: Spatio-Temporal AU Graph Convolution Network for Facial Micro-expression Recognition
Xinhui Zhao, Huimin Ma 0001, Rongquan Wang |
PRCV (1) | 3 |
| 2020 | A novel graph clustering method with a greedy heuristic search algorithm for mining protein complexes from dynamic and static PPI networks
Rongquan Wang, Guixia Liu |
Inf. Sci. | 1 |
| 2019 | Identifying protein complexes based on an edge weight algorithm and core-attachment structureabstractBACKGROUND: Protein complex identification from protein-protein interaction (PPI) networks is crucial for understanding cellular organization principles and functional mechanisms. In recent decades, numerous computational methods have been proposed to identify protein complexes. However, most of the current state-of-the-art studies still have some challenges to resolve, including their high false-positives rates, incapability of identifying overlapping complexes, lack of consideration for the inherent organization within protein complexes, and absence of some biological attachment proteins. RESULTS: In this paper, to overcome these limitations, we present a protein complex identification method based on an edge weight method and core-attachment structure (EWCA) which consists of a complex core and some sparse attachment proteins. First, we propose a new weighting method to assess the reliability of interactions. Second, we identify protein complex cores by using the structural similarity between a seed and its direct neighbors. Third, we introduce a new method to detect attachment proteins that is able to distinguish and identify peripheral proteins and overlapping proteins. Finally, we bind attachment proteins to their corresponding complex cores to form protein complexes and discard redundant protein complexes. The experimental results indicate that EWCA outperforms existing state-of-the-art methods in terms of both accuracy and p-value. Furthermore, EWCA could identify many more protein complexes with statistical significance. Additionally, EWCA could have better balance accuracy and efficiency than some state-of-the-art methods with high accuracy. CONCLUSIONS: In summary, EWCA has better performance for protein complex identification by a comprehensive comparison with twelve algorithms in terms of different evaluation metrics. The datasets and software are freely available for academic research at https://github.com/RongquanWang/EWCA . Rongquan Wang, Guixia Liu |
BMC Bioinform. | 1 |
| 2018 | Predicting overlapping protein complexes based on core-attachment and a local modularity structureabstractBACKGROUND: In recent decades, detecting protein complexes (PCs) from protein-protein interaction networks (PPINs) has been an active area of research. There are a large number of excellent graph clustering methods that work very well for identifying PCs. However, most of existing methods usually overlook the inherent core-attachment organization of PCs. Therefore, these methods have three major limitations we should concern. Firstly, many methods have ignored the importance of selecting seed, especially without considering the impact of overlapping nodes as seed nodes. Thus, there may be false predictions. Secondly, PCs are generally supposed to be dense subgraphs. However, the subgraphs with high local modularity structure usually correspond to PCs. Thirdly, a number of available methods lack handling noise mechanism, and miss some peripheral proteins. In summary, all these challenging issues are very important for predicting more biological overlapping PCs. RESULTS: In this paper, to overcome these weaknesses, we propose a clustering method by core-attachment and local modularity structure, named CALM, to detect overlapping PCs from weighted PPINs with noises. Firstly, we identify overlapping nodes and seed nodes. Secondly, for a node, we calculate the support function between a node and a cluster. In CALM, a cluster which initially consists of only a seed node, is extended by adding its direct neighboring nodes recursively according to the support function, until this cluster forms a locally optimal modularity subgraph. Thirdly, we repeat this process for the remaining seed nodes. Finally, merging and removing procedures are carried out to obtain final predicted clusters. The experimental results show that CALM outperforms other classical methods, and achieves ideal overall performance. Furthermore, CALM can match more complexes with a higher accuracy and provide a better one-to-one mapping with reference complexes in all test datasets. Additionally, CALM is robust against the high rate of noise PPIN. CONCLUSIONS: By considering core-attachment and local modularity structure, CALM could detect PCs much more effectively than some representative methods. In short, CALM could potentially identify previous undiscovered overlapping PCs with various density and high modularity. Rongquan Wang, Guixia Liu, Lingtao Su, Liyan Sun |
BMC Bioinform. | 1 |