Chengye Wang

dblp:71/2174 · also Cheng-Ye Wang · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-6184-9863ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Information extraction and text analysis · 41% Vision and language · 21% Question answering and dialogue systems · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
document understanding
1.012026
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction · ACL (1) 2026
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
1.012026
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding · CVPR 2025
Natural language and speech › Information extraction and text analysis › fact-checking
scientific claim verification
0.912025
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025
Information retrieval
retrieval-augmented generation
0.912025
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › fact-checking
explainable fact-checking
0.812024
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024
Natural language and speech › Information extraction and text analysis
fact-checking
0.812024
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024
Machine learning › Transfer learning and domain adaptation
foundation model evaluation
0.312025
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering
0.312025
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › document understanding
financial document analysis
0.212024
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

error analysis · 2.6retrieval-augmented generation · 1.7benchmark construction · 1.7supervised fine-tuning · 1.0reinforcement learning with verifiable rewards · 1.0multimodal reasoning · 1.0rubric-based evaluation · 0.9reinforcement learning · 0.9expert annotation · 0.9large language model · 0.8
YearPublicationVenuePosition
2026 SciMDR: Advancing Scientific Multimodal Document Reasoning
abstract
Ziyu Chen, Yilun Zhao, Chengye Wang, Rilyn R. Han, Manasi Patwardhan, Arman Cohan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yilun Zhao 0001, Chengye Wang, Rilyn Han, Manasi Patwardhan 0001, Arman Cohan
ACL (1)3
2026 TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
abstract
Existing document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientific publishing.We study page-level reconstruction of scientific PDFs into compilable LaTeX and introduce TEX-OCR-Bench, a benchmark, and TEXOCR-Train, a large-scale training corpus, for this task.TEXOCR-Bench features a multi-dimensional evaluation suite that jointly assesses transcription fidelity, structural faithfulness, and endto-end compilability.Leveraging TEXOCR-Train, we train a 2B-parameter model, TEX-OCR, using supervised fine-tuning (SFT) and reinforcement learning (RL) with verifiable rewards derived from LaTeX unit tests that directly enforce compilability and referential integrity.Experiments across 21 frontier models on TEXOCR-Bench show that existing systems frequently violate key document invariants, including consistent section structure, correct float placement, and valid label-reference links, which undermines compilation reliability and downstream usability.Our analysis further reveals that RL with verifiable rewards yields consistent improvements over SFT alone, particularly on structural and compilation metrics.
Chengye Wang, Zexi Kuang, Yilun Zhao 0001
ACL (1)1
2026 MCTP: A Multi-Coupled Dynamics Trajectory Planning Scheme for Autonomous Driving in Extreme Conditions
abstract
Trajectory planning is essential for ensuring the safe operation of autonomous vehicles. However, existing methods rarely consider the vehicle’s multi-coupled dynamics, including lateral-longitudinal motion coupling, tire force coupling, and lateral instability. This omission can result in infeasible trajectories, vehicle instability, or even accidents under extreme conditions. To address this challenge, this study presents a multi-coupled dynamics trajectory planning (MCTP) scheme. MCTP establishes a coupled kinematics model to accurately represent vehicle motion states and constructs a tire force representation model, which based solely on vehicle motion states, facilitating seamless integration into trajectory planning. By incorporating coupled tire force characteristics and lateral stability analysis, a set of coupled dynamic constraints is formulated to ensure trajectory feasibility and lateral stability. Additionally, a multi-objective function is designed to further optimize trajectory safety, dynamic feasibility, and lateral stability, with the optimal trajectory obtained through receding horizon optimization. Closed-loop validation on both hardware-in-the-loop and real-vehicle experimental platforms demonstrates that, MCTP generates trajectories with enhanced safety and feasibility. It also improves tracking stability margins and dynamics performance, highlighting its effectiveness in handling extreme conditions.
Xuepeng Hu, Yu Zhang 0222, Chengye Wang, Shaoyang Shi, Zhenfeng Wang, Yechen Qin
IEEE Trans Autom. Sci. Eng.3
2025 AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
abstract
Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan
ACL (1)5
2025 SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
abstract
We introduce SCIVER, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context.SCIVER consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification.To enable fine-grained evaluation, each example includes expert-annotated supporting evidence.We assess the performance of 21 state-of-the-art multimodal foundation models, including o4mini, Gemini-2.5-Flash,Llama-3.2-Vision, and Qwen2.5-VL.Our experiment reveals a substantial performance gap between these models and human experts on SCIVER.Through an in-depth analysis of retrieval-augmented generation (RAG), and human-conducted error evaluations, we identify critical limitations in current open-source models, offering key insights to advance models' comprehension and reasoning in multimodal scientific literature tasks.
Chengye Wang, Yifei Shen 0006, Zexi Kuang, Arman Cohan, Yilun Zhao 0001
ACL (1)1
2025 MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
abstract
We introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains.
Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan
CVPR11
2025 UMU-Bench: Closing the Modality Gap in Multimodal Unlearning Evaluation
abstract
Although Multimodal Large Language Models (MLLMs) have advanced numerous fields, their training on extensive multimodal datasets introduces significant privacy concerns, prompting the necessity for efficient unlearning methods.However, current multimodal unlearning approaches often directly adapt techniques from unimodal contexts, largely overlooking the critical issue of modality alignment, i.e., consistently removing knowledge across both unimodal and multimodal settings. To close this gap, we introduce UMU-bench, a unified benchmark specifically targeting modality misalignment in multimodal unlearning. UMU-bench consists of a meticulously curated dataset featuring 653 individual profiles, each described with both unimodal and multimodal knowledge.Additionally, novel tasks and evaluation metrics focusing on modality alignment are introduced, facilitating a comprehensive analysis of unimodal and multimodal unlearning effectiveness. Through extensive experimentation with state-of-the-art unlearning algorithms on UMU-bench, we demonstrate prevalent modality misalignment issues in existing methods. These findings underscore the critical need for novel multimodal unlearning approaches explicitly considering modality alignment.
Chengye Wang, Yuyuan Li 0001, Xiaohua Feng 0002, Chaochao Chen 0001, Jianwei Yin
NeurIPS1
2025 TLSLeaf: Unsupervised Instance Segmentation of Broadleaf Leaf Count and Area From TLS Point Clouds
abstract
Terrestrial laser scanning (TLS) has revolutionized tree-level measurement, enabling accurate trunk, and branch analysis, but it faces challenges in counting and measuring individual leaves in broad-leaved trees. We introduce TLSLeaf, an innovative method designed to accurately measure leaf count and area, specifically tailored for complex tree canopies. TLSLeaf addresses critical gaps in TLS-based leaf analysis through a four-step process: 1) wood-leaf separation; 2) individual leaf segmentation; 3) detection and repair of incomplete leaves; and 4) leaf counting and area measurement. TLSLeaf integrates graph theory and shortest-path algorithms for effective branch-leaf separation, followed by instance segmentation using a similarity graph. To overcome the challenge of TLS canopy occlusion and imperfect leaf scans, incomplete leaves are detected and repaired based on symmetry and concavity principles. TLSLeaf’s accuracy was validated through field scanning, in situ measurements, and destructive sampling. TLSLeaf showed a percentage error of 1.49%–8.60% for leaf counts (ranging from 201 to 4000 leaves) and achieved an$R^{2}$of 0.95 and RMSE of 5.87 cm2 for leaf areas (ranging from 9.98 to 179.77 cm2). TLSLeaf integrates semantic and instance segmentation through an unsupervised framework, presenting a “white-box” approach that ensures transparency and reproducibility. TLSLeaf represents a significant advancement in leaf-level analysis of TLS point clouds, with potential applications in enhancing tree reconstruction, 3-D radiative transfer modeling, and canopy photosynthesis.
Guangpeng Fan, Ruoyoulan Wang, Chengye Wang, Jialing Zhou, Binghong Zhang, Zhiming Xin, Huijie Xiao
IEEE Trans. Geosci. Remote. Sens.3
2024 FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents
abstract
Yilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, Arman Cohan. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yilun Zhao 0001, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu 0001, Xiangru Tang, Chen Zhao 0013, Arman Cohan
EMNLP4
1991 Detecting clouds and cloud shadows on aerial photographs
Chengye Wang, Liuqing Huang, Azriel Rosenfeld
Pattern Recognit. Lett.1
1983 An experiment in multispectral, multitemporal crop classification using relaxation techniques
Larry Davis 0001, Chengye Wang, Huchen Xie
Comput. Vis. Graph. Image Process.2
1983 Some experiments in relaxation image matching using corner features
Chengye Wang, Hanfang Sun, Shiro Yada, Azriel Rosenfeld
Pattern Recognit.1
1983 Multispectral image smoothing guided by global distribution of pixel values
abstract
Multispectral images can be effectively smoothed by using the global distribution of pixel values to guide a local selective averaging process. After several iterations of this process, a typical image is virtually segmented into regions of constant value while significant edges in the image are preserved.
Leslie J. Kitchen, Matti Pietikäinen, Azriel Rosenfeld, Chengye Wang
IEEE Trans. Syst. Man Cybern.4
1982 An experiment in multi-spectral, multi-temporal crop classification using relaxation techniques
Larry Davis 0001, Chengye Wang, Huchen Xie
Comput. Graph. Image Process.2
1982 Two remarks on multidimensional texture analysis
David G. Morgenthaler, Chengye Wang, Azriel Rosenfeld
Pattern Recognit. Lett.2