EDBT 2026 Demo / reviewers in the wild / expert
Congbo Ma
dblp:204/8300
· DBLP profile ↗
19ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-3270-5609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LHG: LLM-enhanced and Heterogeneous Graph-induced for Unsupervised Social Event Detection
Zitai Qiu, Rongwei Xu 0001, Congbo Ma, Shan Xue 0001, Jian Yang 0001, Guanfeng Liu 0001, Quan Z. Sheng, Amin Beheshti, Jia Wu 0001 |
WWW | 3 |
| 2026 | PIGCN: Physics-Inspired Graph Convolution Networks for Heterogeneous Social Event Detection
Yongsheng Yu 0001, Congbo Ma, Zitai Qiu, Shan Xue 0001, Jian Yang 0001, Jia Wu 0001 |
WWW | 2 |
| 2026 | Real-Time Multi-Modal Social Event Detection: A New Dataset and a Key Instance-Driven, Quality-Aware Graph Neural NetworkabstractSocial event detection (SED) involves identifying and analyzing significant real-world events using data generated on social media platforms. With the rapid growth of platforms like Weibo and Twitter, users are sharing not just text but also images. However, most existing SED methods remain text-focused, limiting their ability to fully capture the complexity of real-world social dynamics. Moreover, the lack of multi-modal datasets specifically designed for SED has blocked the development of models that can effectively exploit these rich content types. To address these limitations, we introduced WEIBO2022, an extensive multi-modal SED dataset that includes both text and image data. The dataset is available in two versions: WEIBO2022-Medium, containing 25,435 entries and WEIBO2022-Large, containing 79,825 entries. In addition, we presented a novel network called the Key Instance-driven, Quality-aware Graph Neural Network (KQGNN), which features a key instance-driven library, a quality-aware learning process, and a multi-modal fusion module, enhancing its ability to detect events accurately in both offline and real-time settings. Extensive experiments showcase the exceptional performance and superiority of the proposed model, showing improvements in detection accuracy and effective prevention of catastrophic forgetting during continuous training. Yifei Han, Congbo Ma, Shan Xue 0001, Jian Yang 0001, Jia Wu 0001 |
IEEE Trans. Big Data | 2 |
| 2025 | HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMsabstract6173 Qing Li 0038, Jiahui Geng, Zongxiong Chen, Derui Zhu, Yuxia Wang 0003, Congbo Ma, Chenyang Lyu, Fakhri Karray |
ACL (1) | 6 |
| 2025 | Explicit and Implicit Data Augmentation for Social Event DetectionabstractSocial event detection involves identifying and categorizing important events from social media, which relies on labeled data, but annotation is costly and labor-intensive. To address this problem, we propose Augmentation framework for Social Event Detection (SED-Aug), a plug-and-play dual augmentation framework, which combines explicit text-based and implicit feature-space augmentation to enhance data diversity and model robustness. The explicit augmentation utilizes LLMs to enhance textual information through five diverse generation strategies. For implicit augmentation, we design five novel perturbation techniques that operate in the feature space on structural fused embeddings. These perturbations are crafted to keep the semantic and relational properties of the embeddings and make them more diverse. Specifically, SED-Aug outperforms the best baseline model by approximately 17.67% on the Twitter2012 dataset and by about 15.57% on the Twitter2018 dataset in terms of the average F1 score. Congbo Ma, Yuxia Wang 0003, Jia Wu 0001, Jian Yang 0001, Jing Du 0003, Zitai Qiu, Qing Li 0038, Hu Wang 0003, Preslav Nakov |
ACL (1) | 1 |
| 2025 | Text is All You Need: LLM-enhanced Incremental Social Event DetectionabstractSocial event detection (SED) is the task of identifying, categorizing, and tracking events from social data sources such as social media posts, news articles, and online discussions.Existing state-of-the-art (SOTA) SED models predominantly rely on graph neural networks (GNNs), which involve complex graph construction and time-consuming training processes, limiting their practicality in real-world scenarios.In this paper, we rethink the key challenge in SED: the informal expressions and abbreviations of short texts on social media platforms, which impact clustering accuracy.We propose a novel framework, LLM-enhanced Social Event Detection (LSED), which leverages the rich background knowledge of LLMs to address this challenge.Specifically, LSED utilizes LLMs to formalize and disambiguate short texts by completing abbreviations and summarizing informal expressions.Furthermore, we introduce hyperbolic space embeddings, which are more suitable for natural language sentence representations, to enhance clustering performance.Extensive experiments on two challenging realworld datasets demonstrate that LSED outperforms existing SOTA models, achieving improvements in effectiveness, efficiency, and stability.Our work highlights the potential of LLMs in SED and provides a practical solution for real-world applications.The code is available at GitHub 1 . Zitai Qiu, Congbo Ma, Jia Wu 0001, Jian Yang 0001 |
ACL (1) | 2 |
| 2025 | Rethinking Transformer-Based Multi-Document Summarization: An Empirical Investigation
Congbo Ma, Wei Zhang 0098, Dileepa Pitawela, Haojie Zhuang, Yanfeng Shu, Qing Li 0038 |
ADMA (2) | 1 |
| 2024 | Enhancing Chemistry-Domain Scientific Paper Summarization by Knowledge Graphs
Yutong Qu, Jian Yang 0001, Weitong Chen 0001, Yan Jiao, Lishan Yang 0002, Congbo Ma |
ADMA (2) | 6 |
| 2024 | Disentangling Specificity for Abstractive Multi-document Summarization
Congbo Ma, Wei Zhang 0098, Hu Wang 0005, Haojie Zhuang, Mingyu Guo 0001 |
IJCNN | 1 |
| 2024 | Distributionally-Adaptive Variational Meta Learning for Brain Graph Classification
Jing Du 0003, Guangwei Dong, Congbo Ma, Shan Xue 0001, Jia Wu 0001, Jian Yang 0001, Amin Beheshti, Quan Z. Sheng, Alexis Giral |
MICCAI (10) | 3 |
| 2024 | An Efficient Automatic Meta-Path Selection for Social Event Detection via Hyperbolic SpaceabstractSocial events reflect changes in communities, such as natural disasters and emergencies. Detection of these situations can help residents and organizations in the community avoid danger and reduce losses. The complex nature of social messages makes social event detection on social media challenging. The challenges that have a greater impact on social media detection models are as follows: (1) the amount of social media data is huge but its availability is small; (2) social media data is a tree structure and traditional Euclidean space embedding will distort embedded features; and (3) the heterogeneity of social media networks makes existing models unable to capture rich information well. To solve the above challenges, we propose a Heterogeneous Information Graph representation via Hyperbolic space combined with an Automatic Meta-path selection (GraphHAM) model, an efficient framework that automatically selects the meta-path's weight and combines hyperbolic space to learn information on social media. In particular, we apply an efficient automatic meta-path selection technique and convert the selected meta-path into a vector, thereby reducing the requisite amount of labeled data for the model. We also design a novel Hyperbolic Multi-Layer Perceptron (HMLP) to further learn the semantic and structural information of social information. Extensive experiments show that GraphHAM can achieve outstanding performance on real-world data using only 20% of the whole dataset as the training set. Our code can be found on GitHub https://github.com/ZITAIQIU/GraphHAM. Zitai Qiu, Congbo Ma, Jia Wu 0001, Jian Yang 0001 |
WWW | 2 |
| 2023 | Multi-Modal Learning with Missing Modality via Shared-Specific Feature ModellingabstractThe missing modality issue is critical but non-trivial to be solved by multi-modal models. Current methods aiming to handle the missing modality problem in multi-modal tasks, either deal with missing modalities only during evaluation or train separate models to handle specific missing modality settings. In addition, these models are designed for specific tasks, so for example, classification models are not easily adapted to segmentation tasks and vice versa. In this paper, we propose the Shared-Specific Feature Modelling (ShaSpec) method that is considerably simpler and more effective than competing approaches that address the issues above. ShaSpec is designed to take advantage of all available input modalities during training and evaluation by learning shared and specific features to better represent the input data. This is achieved from a strategy that relies on auxiliary tasks based on distribution alignment and domain classification, in addition to a residual feature fusion procedure. Also, the design simplicity of ShaSpec enables its easy adaptation to multiple tasks, such as classification and segmentation. Experiments are conducted on both medical image segmentation and computer vision classification, with results indicating that ShaSpec outperforms competing methods by a large margin. For instance, on BraTS2018, ShaSpec improves the SOTA by more than 3% for enhancing tumour, 5% for tumour core and 3% for whole tumour.11This work received funding from the Australian Government the through Medical Research Futures Fund: Primary Health Care Research Data Infrastructure Grant 2020 and from Endometriosis Australia. G.C. was supported by Australian Research Council through grant FT190100525. Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
CVPR | 3 |
| 2023 | Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
Hu Wang 0005, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
MICCAI (4) | 2 |
| 2023 | Data Hiding With Deep Learning: A Survey Unifying Digital Watermarking and SteganographyabstractThe advancement of secure communication and identity verification fields has significantly increased through the use of deep learning techniques for data hiding. By embedding information into a noise-tolerant signal, such as audio, video, or images, digital watermarking and steganography techniques can be used to protect sensitive intellectual property (IP) and enable confidential communication, ensuring that the information embedded is only accessible to authorized parties. This survey provides an overview of recent developments in deep learning techniques deployed for data hiding, categorized systematically according to model architectures and noise injection methods. In addition, potential future research directions that unite digital watermarking and steganography on software engineering to enhance security and mitigate risks are suggested and deliberated. This contribution furthers the creation of a more trustworthy digital world and advances responsible artificial intelligence (AI). Olivia Byrnes, Hu Wang 0005, Ruoxi Sun 0001, Congbo Ma, Huaming Chen, Qi Wu 0001, Minhui Xue 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2022 | Uncertainty-Aware Multi-modal Learning via Cross-Modal Random Network Prediction
Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
ECCV (37) | 4 |
| 2022 | Incorporating Linguistic Knowledge for Abstractive Multi-document Summarization
Congbo Ma, Wei Zhang 0098, Hu Wang 0005, Mingyu Guo 0001 |
PACLIC | 1 |
| 2021 | Improving Deep Learning based Multi-document Summarization through Linguistic KnowledgeabstractMulti-document summarization is one of the most important tasks in the field of Natural Language Processing (NLP) and it gains increasing attention in recent years. It aims to generate one summary across several topic-related documents. Compared with extractive summarization, abstractive summarization is more similar to human-written ones. Proposing effective and efficient abstractive multi-document summarization models is significant to the NLP community. Existing deep learning based multi-document summarization models rely on the exceptional ability of neural networks to extract distinct features. However, they have missed out important linguistic knowledge such as dependencies between words since linguistics information in texts is full of meaningful knowledge with respect to the input documents. Besides, how models automatically evaluate the quality of the summary is crucial to design a high-performance summarization model since the evaluation indicator objectively measures the effectiveness of a method. In this proposal, we bring forward two research questions and corresponding solutions for the abstractive multi-document summarization task. Congbo Ma |
SIGIR | 1 |
| 2020 | Unsupervised Representation Learning by Predicting Random DistancesabstractDeep neural networks have gained great success in a broad range of tasks due to its remarkable capability to learn semantically rich features from high-dimensional data. However, they often require large-scale labelled data to successfully learn such features, which significantly hinders their adaption in unsupervised learning tasks, such as anomaly detection and clustering, and limits their applications to critical domains where obtaining massive labelled data is prohibitively expensive. To enable unsupervised learning on those domains, in this work we propose to learn features without using any labelled data by training neural networks to predict data distances in a randomly projected space. Random mapping is a theoretically proven approach to obtain approximately preserved distances. To well predict these distances, the representation learner is optimised to learn genuine class structures that are implicitly embedded in the randomly projected space. Empirical results on 19 real-world datasets show that our learned representations substantially outperform a few state-of-the-art methods for both anomaly detection and clustering tasks. Code is available at: \url{https://git.io/RDP} Hu Wang 0005, Guansong Pang, Chunhua Shen, Congbo Ma |
IJCAI | 4 |
| 2019 | Multi-label Thoracic Disease Image Classification with Cross-Attention Networks
Congbo Ma, Hu Wang 0005, Steven C. H. Hoi |
MICCAI (6) | 1 |