Haocheng Gao

dblp:399/2872 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 46% Question answering and dialogue systems · 23% Trustworthy machine learning · 23%
Network and information security
1 paper
Digital forensics and information hiding · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
AI-generated content detection
1.012026
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images · ACL (1) 2026
Computer vision › Vision and language
multimodal benchmark
1.012026
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026
Digital forensics and information hiding › digital forensics › multimedia forensics
image forensics
1.012026
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images · ACL (1) 2026
Information retrieval
retrieval-augmented generation
0.312026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0forensic analysis · 2.0benchmark construction · 2.0multimodal large language model · 1.0
YearPublicationVenuePosition
2026 FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
abstract
We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.
Zichen Tang, Haihong E, Rongjin Li, Linwei Jia, Zhuodi Hao, Zhongjun Yang, Yuanze Li, Haolin Tian, Peizhi Zhao, Xianghe Wang, Xueyuan Lin, Ruofei Bai, Zijian Xie, Ruining Cao, Haocheng Gao
AAAI21
2026 Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
abstract
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Haolin Tian, Pengqi Sun, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Zhongjun Yang, Jintong Chen, Siying Lin
ACL (1)11
2026 AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
abstract
Bo Zhang, Tzu-Yen Ma, Zichen Tang, Junpeng Ding, Zirui Wang, Yizhuo Zhao, Peilin Gao, Zijie Xi, Zixin Ding, Haiyang Sun, Haocheng Gao, Yuan Liu, Liangjia Wang, Yiling Huang, Yujie Wang, Yuyue Zhang, Ronghui Xi, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Haihong E. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tzu-Yen Ma, Zichen Tang, Junpeng Ding, Yizhuo Zhao, Peilin Gao, Zijie Xi, Zixin Ding, Haocheng Gao, Liangjia Wang, Yuyue Zhang, Ronghui Xi, Yuanze Li, Zhongjun Yang, Haihong E
ACL (1)11
2026 The Parallel-Semantics Program Dependence Graph for Parallel Optimization
abstract
Modern shared-memory parallel programming models, such as OpenMP and Cilk, enable developers to encode a parallel execution plan within their code. Existing compilers, including Clang and GCC, directly lower or add additional compatible parallelism on top of the developers’ plan. However, when better parallel execution plans exist that are incompatible with the original plan, compilers lack the capability of disregarding it and replacing it with a better one. To address this problem, this paper introduces the parallel-semantics program dependence graph (PS-PDG), an extension of the program dependence graph (PDG) abstraction that can simultaneously represent parallel semantics derived from both the developer’s original plan and the compiler’s own analysis. To demonstrate the power of PS-PDG, this paper also introduces GINO, an LLVM-based compiler capable of optimizing parallel execution plans using PS-PDG. Through exploring, reasoning, and implementing better parallel execution plans unlocked by PS-PDG, GINO outperforms the developer’s original parallel execution plan by 46.6% at most, and by 15% on average over 56 cores across 8 benchmarks from the NAS benchmark suite.
Yian Su, Brian Homerding, Haocheng Gao, Federico Sossai, Yebin Chon, David I. August, Simone Campanoni
CGO3
2025 Characterizing and detecting Python version incompatibilities caused by inconsistent version specifications
Haocheng Gao, Wei Chen 0018, Yi Li 0008, Haoxiang Tian 0001, Dan Ye 0004
J. Syst. Softw.2