EDBT 2026 Demo / reviewers in the wild / expert
Wang Gao 0002
dblp:189/1639-2
· DBLP profile ↗
17ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-9671-489XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model evaluation
domain-specific benchmark |
1.0 | 1 | 2026 | TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.0 | 1 | 2026 | TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice · ACL (1) 2026 |
High-performance computing
parallel compression |
0.3 | 1 | 2017 | Parallelization of Massive Textstream Compression Based on Compressed Sensing · ACM Trans. Inf. Syst. 2017 |
Information theory › signal processing
compressed sensing |
0.3 | 1 | 2017 | Parallelization of Massive Textstream Compression Based on Compressed Sensing · ACM Trans. Inf. Syst. 2017 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 2.0underdetermined linear system solving · 0.9compressed sensing · 0.9parallel procedures · 0.6parallel procedure · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax PracticeabstractGang Hu, Yating Chen, Haiyan Ding, Wang Gao, Huang Jiajia, Min Peng, Qianqian Xie, Kun Yue. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Gang Hu 0003, Yating Chen, Haiyan Ding, Wang Gao 0002, Min Peng 0002, Qianqian Xie, Kun Yue |
ACL (1) | 4 |
| 2026 | CodeMNER: Vision-Language Models are Better Multimodal Named Entity Recognizers via Progressive Vision-Code AlignmentabstractWith the explosive growth of multimedia content on social media, Multimodal Named Entity Recognition (MNER) has garnered significant attention. However, current paradigms predominantly rely on general Vision-Language Models (VLMs) to generate natural language responses. Such unstructured text generation struggles to precisely articulate the complex structured information inherent in MNER tasks, often resulting in outputs that lack logical rigor and explicit structural constraints. To address these limitations, we propose CodeMNER, a novel framework that reformulates MNER tasks as a multimodal code generation problem. By synthesizing executable code instead of natural language, CodeMNER leverages the inherent syntactic rigor and deterministic executability of programming languages, thereby significantly enhancing the model’s capacity for identifying and classifying named entities. Despite the evident advantages of the code generation paradigm, standard VLMs lack the joint alignment between structured code semantics and natural visual representations, making it challenging to directly establish the mapping from visual contexts to executable code. To this end, we design a progressive four-stage training pipeline, encompassing mid-training, supervised fine-tuning, reinforcement learning with verifiable rewards, and downstream adaptation. This pipeline bridges the inherent vision-code alignment gap and augments model performance on MNER. Extensive experiments across standard Twitter-2015 and Twitter-2017 datasets demonstrate that CodeMNER achieves state-of-the-art performance, surpassing existing baselines. Jiakang Yu 0001, Shizhou Huang, Xiaode Chen, Hongtao Deng, Wang Gao 0002 |
ICMR | 5 |
| 2025 | SeaFBen: A Multilingual Benchmark for Large Language Models in Southeast Asian FinanceabstractLarge language models (LLMs) excel in general financial tasks and low-resource languages, but their potential in Southeast Asia’ s multilingual financial domain remains underexplored due to cultural diversity, data scarcity, and task complexity. To address this, we introduce SeaFBen, the first open-source benchmark for Southeast Asian multilingual financial tasks. Covering 22k samples across 20 datasets in 5 major languages (Thai, Indonesian, Vietnamese, Filipino, and Malay) from highly populated countries, SeaFBen evaluates 5 key tasks: Knowledge Understanding, Investment Tendency, Credit Rating, Financial Decision-making, and Numerical Reasoning. It pioneers multilingual financial task evaluation, regional localization, and introduces 5 new datasets. Evaluating 12 LLMs reveals significant performance differences, particularly in numerical reasoning, with ChatGPT and SeaLLMs excelling while PolyLM-13B underperforms. Moreover, SeaLLMs’ evaluation reflects language and task performance biases caused by differences in underlying fine-tuning tasks. SeaFBen1is a vital resource for advancing LLM research and applications in financial domain. Gang Hu 0003, QingQing Wang, Wanlong Yu, Siqi Lv, AiJia Zhao, Wang Gao 0002 |
IJCNN | 6 |
| 2025 | AMCCL: Adaptive Multi-scale Convolution Fusion Network with Contrastive Learning for Multimodal Sentiment Analysis
Jiakang Yu 0001, Hongtao Deng, Wang Gao 0002 |
PRICAI (4) | 4 |
| 2024 | Event-centric hierarchical hyperbolic graph for multi-hop question answering over knowledge graphs
Wang Gao 0002, Wenguang Yao, Hongtao Deng |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Syntax-based argument correlation-enhanced end-to-end model for scientific relation extraction
Wang Gao 0002, Lang Zhang, Hongtao Deng |
Neurocomputing | 2 |
| 2024 | Duplicate question detection in community-based platforms via interaction networks
Wang Gao 0002, Baoping Yang |
Multim. Tools Appl. | 1 |
| 2023 | Generative non-autoregressive unsupervised keyphrase extraction with neural topic modeling
Yinxia Lou, Wang Gao 0002, Hongtao Deng |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Identifying informative tweets during a pandemic via a topic-aware neural language model
Wang Gao 0002, Lin Li 0001, Xiaohui Tao 0001 |
World Wide Web (WWW) | 1 |
| 2021 | Event Detection in Social Media via Graph Neural Network
Wang Gao 0002, Lin Li 0001, Xiaohui Tao 0001 |
WISE (1) | 1 |
| 2020 | Generation of topic evolution graphs from short text streams
Wang Gao 0002, Min Peng 0002, Hua Wang 0002, Yanchun Zhang, Weiguang Han, Gang Hu 0003, Qianqian Xie |
Neurocomputing | 1 |
| 2020 | DTC: Transfer learning for commonsense machine comprehension
Weiguang Han, Min Peng 0002, Qianqian Xie, Gang Hu 0003, Wang Gao 0002, Hua Wang 0002, Yanchun Zhang, Zuopeng Liu |
Neurocomputing | 5 |
| 2020 | Unsupervised software repositories mining and its application to code searchabstractSummary Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code‐description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code‐Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete‐time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high‐quality software repositories for various programming languages, especially for the large‐scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query‐expansion code search. Meanwhile, for SQL and C# programs, compared to the state‐of‐the‐art query‐expansion method QECK, the improvement of QECK CodeMF is 2% and 6% on Recall@10, and 4% and 14% on mean reciprocal rank, respectively. Gang Hu 0003, Min Peng 0002, Yihan Zhang 0005, Qianqian Xie, Wang Gao 0002, Mengting Yuan 0001 |
Softw. Pract. Exp. | 5 |
| 2019 | Incorporating word embeddings into topic modeling of short text
Wang Gao 0002, Min Peng 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Tian |
Knowl. Inf. Syst. | 1 |
| 2018 | Topic-Net Conversation Model
Min Peng 0002, Dian Chen 0004, Qianqian Xie, Yanchun Zhang, Hua Wang 0002, Gang Hu 0003, Wang Gao 0002, Yihan Zhang 0005 |
WISE (1) | 7 |
| 2017 | Parallelization of Massive Textstream Compression Based on Compressed SensingabstractCompressing textstreams generated by social networks can both reduce storage consumption and improve efficiency such as fast searching. However, the compression process is a challenge due to the large scale of textstreams. In this article, we propose a textstream compression framework based on compressed sensing theory and design a series of matching parallel procedures. The new approach uses a linear projection technique in the textstream compression process, achieving fast compression speed and low compression ratio. Two processes are executed by designing elaborated parallel procedures for efficient compressing and decompressing of large-scale textstreams. The decompression process is implemented for approximate solutions of underdetermined linear systems. Experimental results show that the new method can efficiently achieve the compression and decompression tasks on a large amount of text generated by social networks. Min Peng 0002, Wang Gao 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Hu 0003, Gang Tian |
ACM Trans. Inf. Syst. | 2 |
| 2017 | A probabilistic method for emerging topic tracking in Microblog stream
Min Peng 0002, Hua Wang 0002, Jinli Cao, Wang Gao 0002, Xiuzhen Zhang 0001 |
World Wide Web | 5 |