Ketong Chen

dblp:348/5678 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 46% Information extraction and text analysis · 23% Vision and language · 23%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
document understanding
1.012026
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding · AAAI 2026
Computer vision › Vision and language
multimodal benchmark
1.012026
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding · AAAI 2026
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.712023
Characterizing Out-of-Distribution Error via Optimal Transport · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.712023
Characterizing Out-of-Distribution Error via Optimal Transport · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model
0.312026
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding · AAAI 2026
Mathematical optimization
optimal transport
0.212023
Characterizing Out-of-Distribution Error via Optimal Transport · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

pseudo-label shift · 1.3optimal transport · 1.3confidence thresholding · 1.3multi-agent pipeline · 1.0large language model · 1.0
YearPublicationVenuePosition
2026 MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
abstract
Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature simplistic layouts, and support limited tasks. Consequently, they fail to evaluate model performance for Visually Rich Document Understanding (VRDU), a critical challenge involving complex layouts and dense text. To address this, we introduce DocWeaver, a novel multi-agent pipeline that leverages Large Language Models to automatically generate a new benchmark. The result is MosaicDoc, a large-scale, bilingual (Chinese and English) resource designed to push the boundaries of VRDU. Sourced from newspapers and magazines, MosaicDoc features diverse and complex layouts (including multi-column and non-Manhattan), rich stylistic variety from 196 publishers, and comprehensive multi-task annotations (OCR, VQA, reading order, and localization). With 72K images and over 600K QA pairs, MosaicDoc serves as a definitive benchmark for the field. Our extensive evaluation of state-of-the-art models on this benchmark reveals their current limitations in handling real-world document complexity and charts a clear path for future research.
Ketong Chen
AAAI1
2023 Characterizing Out-of-Distribution Error via Optimal Transport
abstract
Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the actual error, sometimes by a large margin, which greatly impacts their applicability to real tasks. In this work, we identify *pseudo-label shift*, or the difference between the predicted and true OOD label distributions, as a key indicator of this underestimation. Based on this observation, we introduce a novel method for estimating model performance by leveraging optimal transport theory, Confidence Optimal Transport (COT), and show that it provably provides more robust error estimates in the presence of pseudo-label shift. Additionally, we introduce an empirically-motivated variant of COT, Confidence Optimal Transport with Thresholding (COTT), which applies thresholding to the individual transport costs and further improves the accuracy of COT's error estimates. We evaluate COT and COTT on a variety of standard benchmarks that induce various types of distribution shift -- synthetic, novel subpopulation, and natural -- and show that our approaches significantly outperform existing state-of-the-art methods with up to 3x lower prediction errors.
Yuzhe Lu, Yilong Qin, Runtian Zhai, Andrew Shen, Ketong Chen, Zhenlin Wang 0002, Soheil Kolouri, Simon Stepputtis, Joseph Campbell, Katia P. Sycara
NeurIPS5