VLDB 2026 Research / reviewers in the wild / expert
Chengye Wang
dblp:71/2174 · also Cheng-Ye Wang
· DBLP profile ↗
15ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-6184-9863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 41% Vision and language · 21% Question answering and dialogue systems · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
document understanding |
1.0 | 1 | 2026 | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction · ACL (1) 2026 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
1.0 | 1 | 2026 | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | MMVU: Measuring Expert-Level Multi-Discipline Video Understanding · CVPR 2025 |
Natural language and speech › Information extraction and text analysis › fact-checking
scientific claim verification |
0.9 | 1 | 2025 | SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › fact-checking
explainable fact-checking |
0.8 | 1 | 2024 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
fact-checking |
0.8 | 1 | 2024 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation
foundation model evaluation |
0.3 | 1 | 2025 | SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering |
0.3 | 1 | 2025 | AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › document understanding
financial document analysis |
0.2 | 1 | 2024 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
error analysis · 2.6retrieval-augmented generation · 1.7benchmark construction · 1.7supervised fine-tuning · 1.0reinforcement learning with verifiable rewards · 1.0multimodal reasoning · 1.0rubric-based evaluation · 0.9reinforcement learning · 0.9expert annotation · 0.9large language model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciMDR: Advancing Scientific Multimodal Document ReasoningabstractZiyu Chen, Yilun Zhao, Chengye Wang, Rilyn R. Han, Manasi Patwardhan, Arman Cohan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilun Zhao 0001, Chengye Wang, Rilyn Han, Manasi Patwardhan 0001, Arman Cohan |
ACL (1) | 3 |
| 2026 | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX ReconstructionabstractExisting document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientific publishing.We study page-level reconstruction of scientific PDFs into compilable LaTeX and introduce TEX-OCR-Bench, a benchmark, and TEXOCR-Train, a large-scale training corpus, for this task.TEXOCR-Bench features a multi-dimensional evaluation suite that jointly assesses transcription fidelity, structural faithfulness, and endto-end compilability.Leveraging TEXOCR-Train, we train a 2B-parameter model, TEX-OCR, using supervised fine-tuning (SFT) and reinforcement learning (RL) with verifiable rewards derived from LaTeX unit tests that directly enforce compilability and referential integrity.Experiments across 21 frontier models on TEXOCR-Bench show that existing systems frequently violate key document invariants, including consistent section structure, correct float placement, and valid label-reference links, which undermines compilation reliability and downstream usability.Our analysis further reveals that RL with verifiable rewards yields consistent improvements over SFT alone, particularly on structural and compilation metrics. Chengye Wang, Zexi Kuang, Yilun Zhao 0001 |
ACL (1) | 1 |
| 2026 | MCTP: A Multi-Coupled Dynamics Trajectory Planning Scheme for Autonomous Driving in Extreme ConditionsabstractTrajectory planning is essential for ensuring the safe operation of autonomous vehicles. However, existing methods rarely consider the vehicle’s multi-coupled dynamics, including lateral-longitudinal motion coupling, tire force coupling, and lateral instability. This omission can result in infeasible trajectories, vehicle instability, or even accidents under extreme conditions. To address this challenge, this study presents a multi-coupled dynamics trajectory planning (MCTP) scheme. MCTP establishes a coupled kinematics model to accurately represent vehicle motion states and constructs a tire force representation model, which based solely on vehicle motion states, facilitating seamless integration into trajectory planning. By incorporating coupled tire force characteristics and lateral stability analysis, a set of coupled dynamic constraints is formulated to ensure trajectory feasibility and lateral stability. Additionally, a multi-objective function is designed to further optimize trajectory safety, dynamic feasibility, and lateral stability, with the optimal trajectory obtained through receding horizon optimization. Closed-loop validation on both hardware-in-the-loop and real-vehicle experimental platforms demonstrates that, MCTP generates trajectories with enhanced safety and feasibility. It also improves tracking stability margins and dynamics performance, highlighting its effectiveness in handling extreme conditions. Xuepeng Hu, Yu Zhang 0222, Chengye Wang, Shaoyang Shi, Zhenfeng Wang, Yechen Qin |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific ResearchabstractYilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan |
ACL (1) | 5 |
| 2025 | SciVer: Evaluating Foundation Models for Multimodal Scientific Claim VerificationabstractWe introduce SCIVER, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context.SCIVER consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification.To enable fine-grained evaluation, each example includes expert-annotated supporting evidence.We assess the performance of 21 state-of-the-art multimodal foundation models, including o4mini, Gemini-2.5-Flash,Llama-3.2-Vision, and Qwen2.5-VL.Our experiment reveals a substantial performance gap between these models and human experts on SCIVER.Through an in-depth analysis of retrieval-augmented generation (RAG), and human-conducted error evaluations, we identify critical limitations in current open-source models, offering key insights to advance models' comprehension and reasoning in multimodal scientific literature tasks. Chengye Wang, Yifei Shen 0006, Zexi Kuang, Arman Cohan, Yilun Zhao 0001 |
ACL (1) | 1 |
| 2025 | MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingabstractWe introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains. Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan |
CVPR | 11 |
| 2025 | UMU-Bench: Closing the Modality Gap in Multimodal Unlearning EvaluationabstractAlthough Multimodal Large Language Models (MLLMs) have advanced numerous fields, their training on extensive multimodal datasets introduces significant privacy concerns, prompting the necessity for efficient unlearning methods.However, current multimodal unlearning approaches often directly adapt techniques from unimodal contexts, largely overlooking the critical issue of modality alignment, i.e., consistently removing knowledge across both unimodal and multimodal settings. To close this gap, we introduce UMU-bench, a unified benchmark specifically targeting modality misalignment in multimodal unlearning. UMU-bench consists of a meticulously curated dataset featuring 653 individual profiles, each described with both unimodal and multimodal knowledge.Additionally, novel tasks and evaluation metrics focusing on modality alignment are introduced, facilitating a comprehensive analysis of unimodal and multimodal unlearning effectiveness. Through extensive experimentation with state-of-the-art unlearning algorithms on UMU-bench, we demonstrate prevalent modality misalignment issues in existing methods. These findings underscore the critical need for novel multimodal unlearning approaches explicitly considering modality alignment. Chengye Wang, Yuyuan Li 0001, Xiaohua Feng 0002, Chaochao Chen 0001, Jianwei Yin |
NeurIPS | 1 |
| 2025 | TLSLeaf: Unsupervised Instance Segmentation of Broadleaf Leaf Count and Area From TLS Point CloudsabstractTerrestrial laser scanning (TLS) has revolutionized tree-level measurement, enabling accurate trunk, and branch analysis, but it faces challenges in counting and measuring individual leaves in broad-leaved trees. We introduce TLSLeaf, an innovative method designed to accurately measure leaf count and area, specifically tailored for complex tree canopies. TLSLeaf addresses critical gaps in TLS-based leaf analysis through a four-step process: 1) wood-leaf separation; 2) individual leaf segmentation; 3) detection and repair of incomplete leaves; and 4) leaf counting and area measurement. TLSLeaf integrates graph theory and shortest-path algorithms for effective branch-leaf separation, followed by instance segmentation using a similarity graph. To overcome the challenge of TLS canopy occlusion and imperfect leaf scans, incomplete leaves are detected and repaired based on symmetry and concavity principles. TLSLeaf’s accuracy was validated through field scanning, in situ measurements, and destructive sampling. TLSLeaf showed a percentage error of 1.49%–8.60% for leaf counts (ranging from 201 to 4000 leaves) and achieved an$R^{2}$of 0.95 and RMSE of 5.87 cm2 for leaf areas (ranging from 9.98 to 179.77 cm2). TLSLeaf integrates semantic and instance segmentation through an unsupervised framework, presenting a “white-box” approach that ensures transparency and reproducibility. TLSLeaf represents a significant advancement in leaf-level analysis of TLS point clouds, with potential applications in enhancing tree reconstruction, 3-D radiative transfer modeling, and canopy photosynthesis. Guangpeng Fan, Ruoyoulan Wang, Chengye Wang, Jialing Zhou, Binghong Zhang, Zhiming Xin, Huijie Xiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsabstractYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, Arman Cohan. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yilun Zhao 0001, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu 0001, Xiangru Tang, Chen Zhao 0013, Arman Cohan |
EMNLP | 4 |
| 1991 | Detecting clouds and cloud shadows on aerial photographs
Chengye Wang, Liuqing Huang, Azriel Rosenfeld |
Pattern Recognit. Lett. | 1 |
| 1983 | An experiment in multispectral, multitemporal crop classification using relaxation techniques
Larry Davis 0001, Chengye Wang, Huchen Xie |
Comput. Vis. Graph. Image Process. | 2 |
| 1983 | Some experiments in relaxation image matching using corner features
Chengye Wang, Hanfang Sun, Shiro Yada, Azriel Rosenfeld |
Pattern Recognit. | 1 |
| 1983 | Multispectral image smoothing guided by global distribution of pixel valuesabstractMultispectral images can be effectively smoothed by using the global distribution of pixel values to guide a local selective averaging process. After several iterations of this process, a typical image is virtually segmented into regions of constant value while significant edges in the image are preserved. Leslie J. Kitchen, Matti Pietikäinen, Azriel Rosenfeld, Chengye Wang |
IEEE Trans. Syst. Man Cybern. | 4 |
| 1982 | An experiment in multi-spectral, multi-temporal crop classification using relaxation techniques
Larry Davis 0001, Chengye Wang, Huchen Xie |
Comput. Graph. Image Process. | 2 |
| 1982 | Two remarks on multidimensional texture analysis
David G. Morgenthaler, Chengye Wang, Azriel Rosenfeld |
Pattern Recognit. Lett. | 2 |