Wang Gao 0002

dblp:189/1639-2 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-9671-489XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model evaluation
domain-specific benchmark
1.012026
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice · ACL (1) 2026
High-performance computing
parallel compression
0.312017
Parallelization of Massive Textstream Compression Based on Compressed Sensing · ACM Trans. Inf. Syst. 2017
Information theory › signal processing
compressed sensing
0.312017
Parallelization of Massive Textstream Compression Based on Compressed Sensing · ACM Trans. Inf. Syst. 2017

Methods — techniques the papers use, named apart from their topics

benchmark construction · 2.0underdetermined linear system solving · 0.9compressed sensing · 0.9parallel procedures · 0.6parallel procedure · 0.3
YearPublicationVenuePosition
2026 TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
abstract
Gang Hu, Yating Chen, Haiyan Ding, Wang Gao, Huang Jiajia, Min Peng, Qianqian Xie, Kun Yue. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Gang Hu 0003, Yating Chen, Haiyan Ding, Wang Gao 0002, Min Peng 0002, Qianqian Xie, Kun Yue
ACL (1)4
2026 CodeMNER: Vision-Language Models are Better Multimodal Named Entity Recognizers via Progressive Vision-Code Alignment
abstract
With the explosive growth of multimedia content on social media, Multimodal Named Entity Recognition (MNER) has garnered significant attention. However, current paradigms predominantly rely on general Vision-Language Models (VLMs) to generate natural language responses. Such unstructured text generation struggles to precisely articulate the complex structured information inherent in MNER tasks, often resulting in outputs that lack logical rigor and explicit structural constraints. To address these limitations, we propose CodeMNER, a novel framework that reformulates MNER tasks as a multimodal code generation problem. By synthesizing executable code instead of natural language, CodeMNER leverages the inherent syntactic rigor and deterministic executability of programming languages, thereby significantly enhancing the model’s capacity for identifying and classifying named entities. Despite the evident advantages of the code generation paradigm, standard VLMs lack the joint alignment between structured code semantics and natural visual representations, making it challenging to directly establish the mapping from visual contexts to executable code. To this end, we design a progressive four-stage training pipeline, encompassing mid-training, supervised fine-tuning, reinforcement learning with verifiable rewards, and downstream adaptation. This pipeline bridges the inherent vision-code alignment gap and augments model performance on MNER. Extensive experiments across standard Twitter-2015 and Twitter-2017 datasets demonstrate that CodeMNER achieves state-of-the-art performance, surpassing existing baselines.
Jiakang Yu 0001, Shizhou Huang, Xiaode Chen, Hongtao Deng, Wang Gao 0002
ICMR5
2025 SeaFBen: A Multilingual Benchmark for Large Language Models in Southeast Asian Finance
abstract
Large language models (LLMs) excel in general financial tasks and low-resource languages, but their potential in Southeast Asia’ s multilingual financial domain remains underexplored due to cultural diversity, data scarcity, and task complexity. To address this, we introduce SeaFBen, the first open-source benchmark for Southeast Asian multilingual financial tasks. Covering 22k samples across 20 datasets in 5 major languages (Thai, Indonesian, Vietnamese, Filipino, and Malay) from highly populated countries, SeaFBen evaluates 5 key tasks: Knowledge Understanding, Investment Tendency, Credit Rating, Financial Decision-making, and Numerical Reasoning. It pioneers multilingual financial task evaluation, regional localization, and introduces 5 new datasets. Evaluating 12 LLMs reveals significant performance differences, particularly in numerical reasoning, with ChatGPT and SeaLLMs excelling while PolyLM-13B underperforms. Moreover, SeaLLMs’ evaluation reflects language and task performance biases caused by differences in underlying fine-tuning tasks. SeaFBen1is a vital resource for advancing LLM research and applications in financial domain.
Gang Hu 0003, QingQing Wang, Wanlong Yu, Siqi Lv, AiJia Zhao, Wang Gao 0002
IJCNN6
2025 AMCCL: Adaptive Multi-scale Convolution Fusion Network with Contrastive Learning for Multimodal Sentiment Analysis
Jiakang Yu 0001, Hongtao Deng, Wang Gao 0002
PRICAI (4)4
2024 Event-centric hierarchical hyperbolic graph for multi-hop question answering over knowledge graphs
Wang Gao 0002, Wenguang Yao, Hongtao Deng
Eng. Appl. Artif. Intell.2
2024 Syntax-based argument correlation-enhanced end-to-end model for scientific relation extraction
Wang Gao 0002, Lang Zhang, Hongtao Deng
Neurocomputing2
2024 Duplicate question detection in community-based platforms via interaction networks
Wang Gao 0002, Baoping Yang
Multim. Tools Appl.1
2023 Generative non-autoregressive unsupervised keyphrase extraction with neural topic modeling
Yinxia Lou, Wang Gao 0002, Hongtao Deng
Eng. Appl. Artif. Intell.4
2023 Identifying informative tweets during a pandemic via a topic-aware neural language model
Wang Gao 0002, Lin Li 0001, Xiaohui Tao 0001
World Wide Web (WWW)1
2021 Event Detection in Social Media via Graph Neural Network
Wang Gao 0002, Lin Li 0001, Xiaohui Tao 0001
WISE (1)1
2020 Generation of topic evolution graphs from short text streams
Wang Gao 0002, Min Peng 0002, Hua Wang 0002, Yanchun Zhang, Weiguang Han, Gang Hu 0003, Qianqian Xie
Neurocomputing1
2020 DTC: Transfer learning for commonsense machine comprehension
Weiguang Han, Min Peng 0002, Qianqian Xie, Gang Hu 0003, Wang Gao 0002, Hua Wang 0002, Yanchun Zhang, Zuopeng Liu
Neurocomputing5
2020 Unsupervised software repositories mining and its application to code search
abstract
Summary Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code‐description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code‐Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete‐time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high‐quality software repositories for various programming languages, especially for the large‐scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query‐expansion code search. Meanwhile, for SQL and C# programs, compared to the state‐of‐the‐art query‐expansion method QECK, the improvement of QECK CodeMF is 2% and 6% on Recall@10, and 4% and 14% on mean reciprocal rank, respectively.
Gang Hu 0003, Min Peng 0002, Yihan Zhang 0005, Qianqian Xie, Wang Gao 0002, Mengting Yuan 0001
Softw. Pract. Exp.5
2019 Incorporating word embeddings into topic modeling of short text
Wang Gao 0002, Min Peng 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Tian
Knowl. Inf. Syst.1
2018 Topic-Net Conversation Model
Min Peng 0002, Dian Chen 0004, Qianqian Xie, Yanchun Zhang, Hua Wang 0002, Gang Hu 0003, Wang Gao 0002, Yihan Zhang 0005
WISE (1)7
2017 Parallelization of Massive Textstream Compression Based on Compressed Sensing
abstract
Compressing textstreams generated by social networks can both reduce storage consumption and improve efficiency such as fast searching. However, the compression process is a challenge due to the large scale of textstreams. In this article, we propose a textstream compression framework based on compressed sensing theory and design a series of matching parallel procedures. The new approach uses a linear projection technique in the textstream compression process, achieving fast compression speed and low compression ratio. Two processes are executed by designing elaborated parallel procedures for efficient compressing and decompressing of large-scale textstreams. The decompression process is implemented for approximate solutions of underdetermined linear systems. Experimental results show that the new method can efficiently achieve the compression and decompression tasks on a large amount of text generated by social networks.
Min Peng 0002, Wang Gao 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Hu 0003, Gang Tian
ACM Trans. Inf. Syst.2
2017 A probabilistic method for emerging topic tracking in Microblog stream
Min Peng 0002, Hua Wang 0002, Jinli Cao, Wang Gao 0002, Xiuzhen Zhang 0001
World Wide Web5