VLDB 2026 Research / reviewers in the wild / expert
Yiheng Xu
dblp:29/4870
· DBLP profile ↗
20ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PipeViT: Accelerating Vision Transformers via Intra-Layer PipeliningabstractVision Transformers (ViTs) have achieved high performance across various computer vision tasks by leveraging the attention mechanism. However, the attention module in ViTs severely hindered inference performance due to its low operational intensity. Existing approaches improve ViTs efficiency through pruning, sparsity, and linearization, but at the cost of fine-tuning overhead and accuracy degradation. In this paper, we propose PipeViT, a memory-efficient and low-latency accelerator for ViTs inference. The key insight of PipeViT is to exploit intra-layer acceleration opportunities. Specifically, we first fuse the attention operations into a single operator to reduce memory access overhead. Then, we divide the input of attention into multiple tiles to reduce the on-chip memory requirement. Finally, we pipeline the tiled attention computation to improve overall throughput. Based on the optimized dataflow, we design a heterogeneous dual-core architecture for efficient pipeline execution. Furthermore, to maximize hardware utilization, the architecture can be reconfigured into a single core with higher parallelism during the execution of the feed-forward network. Experimental results show that PipeViT achieves up to $19.3 \times 1.5 \times, 2.1 \times$, and $2.0 \times$ improvements in Frames Per Second (FPS) compared to state-of-the-art accelerators, including ViTA, Auto-ViT, MEViT, and HeatViT. Additionally, PipeViT achieves up to $8.0 \times$ and $2.6 \times$ higher energy efficiency compared to CPU and GPU implementations, respectively. Xilang Zhou, Yiheng Xu, Haodong Lu 0001, Jun Yu 0010, Kun Wang 0005 |
ASP-DAC | 2 |
| 2026 | Pareto Optimization with Robust Evaluation for Noisy Subset SelectionabstractSubset selection is a fundamental problem in combinatorial optimization, which has a wide range of applications such as influence maximization and sparse regression. The goal is to select a subset of limited size from a ground set in order to maximize a given objective function. However, the evaluation of the objective function in real-world scenarios is often noisy. Previous algorithms, including the greedy algorithm and multi-objective evolutionary algorithms POSS and PONSS, either struggle in noisy environments or consume excessive computational resources. In this paper, we focus on the noisy subset selection problem with a cardinality constraint, where the evaluation of a subset is noisy. We propose a novel approach based on Pareto Optimization with Robust Evaluation for noisy subset selection (PORE), which maximizes a robust evaluation function and minimizes the subset size simultaneously. PORE can efficiently identify well-structured solutions and handle computational resources, addressing the limitations observed in PONSS. Our experiments, conducted on real-world datasets for influence maximization and sparse regression, demonstrate that PORE significantly outperforms previous methods, including the classical greedy algorithm, POSS, and PONSS. Further validation through ablation studies confirms the effectiveness of our robust evaluation function. Yiheng Xu, Danxuan Liu, Weiyong Yang, Chao Qian 0001 |
GECCO | 1 |
| 2026 | PRISM: Prior-enhanced Inference for Spatial Transcriptomic Cell Type MappingabstractMOTIVATION: Cell type annotation in spatial transcriptomics (ST) is fundamental for deciphering complex tissue organization and spatially resolved biological processes. Most existing methods perform ST cell type annotation by transferring labels from single-cell RNA-seq (scRNA) data to ST data, but typically rely on weakly constrained representations that neglect structured spatial dependencies and treat marker gene selection as an isolated preprocessing step. This renders them vulnerable to substantial domain gaps as well as platform-specific noise, resulting in unstable predictions and limited biological interpretability. RESULTS: To address these issues, we propose Prior-enhanced Inference for Spatial Transcriptomic Cell Type Mapping (PRISM), a novel three-stage framework integrating biological prior construction, pseudo-label generation, and multi-level ST refinement. First, PRISM constructs a cross-domain biological prior to explicitly extract marker genes to enforce positive biological discriminability. Next, it adopts a prior-enhanced self-training strategy, where scRNA-trained ensembles generate reliable pseudo-label candidates for ST data, serving as a robust anchor for cross-domain adaptation. Finally, the framework consolidates high-quality ensemble predictions selected via metric-guided evaluation, encodes spatial information, and optimizes the model under dual-directional biological constraints. Extensive experiments on eleven ST datasets across six platforms, two species, and multiple tissue contexts validate PRISM. Specifically, on the five labeled benchmarks, PRISM shows strong overall performance under both Accuracy and Macro-F1 evaluation across brain and non-brain tissues. Moreover, under fully label-free settings, PRISM achieves the best overall composite rank across all datasets, demonstrating strong robustness to domain shift and platform heterogeneity. AVAILABILITY AND IMPLEMENTATION: PRISM is available at https://github.com/lilab-ai4s/PRISM and https://doi.org/10.5281/zenodo.20529683. Yiheng Xu, Xuehao Wang, Congcong Ge |
Bioinform. | 1 |
| 2025 | Modality Perception Network for Multi-modal Rumor Detection
Xiangyuan Jiang, Yiheng Xu, Wei Zhou 0028 |
ICIC (1) | 3 |
| 2025 | AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsabstractGraphical User Interface (GUI) agents hold great potential for automating complex tasks across diverse digital environments, from web applications to desktop software. However, the development of such agents is hindered by the lack of high-quality, multi-step trajectory data required for effective training. Existing approaches rely on expensive and labor-intensive human annotation, making them unsustainable at scale. To address this challenge, we propose AgentTrek, a scalable data synthesis pipeline that generates high-quality web agent trajectories by leveraging web tutorials. Our method automatically gathers tutorial-like texts from the internet, transforms them into task goals with step-by-step instructions, and employs a visual-language model (VLM) agent to simulate their execution in a real digital environment. A VLM-based evaluator ensures the correctness of the generated trajectories. We demonstrate that training GUI agents with these synthesized trajectories significantly improves their grounding and planning performance over the current models. Moreover, our approach is more cost-efficient compared to traditional human annotation methods. This work underscores the potential of guided replay with web tutorials as a viable strategy for large-scale GUI agent training, paving the way for more capable and autonomous digital agents. Yiheng Xu, Dunjie Lu, Zhennan Shen, Caiming Xiong, Tao Yu 0009 |
ICLR | 1 |
| 2025 | Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionabstractAutomating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis data collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research. Yiheng Xu, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu 0009, Caiming Xiong |
ICML | 1 |
| 2025 | OpenCUA: Open Foundations for Computer-Use AgentsabstractVision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state–action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld‑Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research. Xinyuan Wang 0010, Dunjie Lu, Junlin Yang, Tianbao Xie, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Xiaochuan Li 0003, Junda Chen, Boyuan Zheng 0001, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu 0006, Jixuan Chen, Yuxiao Ye, Yipu Wang, Diyi Yang, Victor Zhong, Y. Charles, Tao Yu 0009 |
NeurIPS | 9 |
| 2025 | Scaling Computer-Use Grounding via User Interface Decomposition and SynthesisabstractGraphical user interface (GUI) grounding, the ability to map natural language instructions to specific actions on graphical user interfaces, remains a critical bottleneck in computer use agent development. Current benchmarks oversimplify grounding tasks as short referring expressions, failing to capture the complexity of real-world interactions that require software commonsense, layout understanding, and fine-grained manipulation capabilities. To address these limitations, we introduce OSWorld-G, a comprehensive benchmark comprising 564 finely annotated samples across diverse task types including text matching, element recognition, layout understanding, and precise manipulation. Additionally, we synthesize and release the largest computer use grounding dataset Jedi, which contains 4 million examples through multi-perspective decoupling of tasks. Our multi-scale models trained on Jedi demonstrate its effectiveness by outperforming existing approaches on ScreenSpot-v2, ScreenSpot-Pro, and our OSWorld-G. Furthermore, we demonstrate that improved grounding with Jedi directly enhances agentic capabilities of general foundation models on complex computer tasks with state-of-the-art performance, improving from 23% to 51% on OSWorld. Through detailed ablation studies, we identify key factors contributing to grounding performance and verify that combining specialized data for different interface elements enables compositional generalization to novel interfaces. All benchmark, data, checkpoints, and code are open-sourced and available at https://osworld-grounding.github.io. Tianbao Xie, Xiaochuan Li 0003, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang 0010, Yiheng Xu, Doyen Sahoo, Tao Yu 0009, Caiming Xiong |
NeurIPS | 11 |
| 2024 | ML-based Modeling to Predict I/O Performance on Different Storage Sub-systemsabstractParallel applications can spend a significant amount of time performing I/O on large-scale supercomputers. Fast near-compute storage accelerators called burst buffers can reduce the time a processor spends performing I/O and mitigate I/O bottlenecks. However, determining if a given application could be accelerated using burst buffers is not straightforward even for storage experts. The relationship between an application's I/O characteristics (such as I/O volume, processes involved, etc.) and the best storage sub-system for it can be complicated. As a result, adapting parallel applications to use burst buffers efficiently is a trial-and-error process. In this work, we present a Python-based tool called PrismIO that enables programmatic analysis of I/O traces. Using PrismIO, we identify performance bottlenecks when using burst buffers and parallel file systems, and explain why certain I/O patterns perform poorly. Further, we use machine learning to model the relationship between I/O characteristics and file system selections. We use IOR, an I/O benchmark to gather performance data for training the machine learning model. Our model can predict the better performing storage system for unseen IOR scenarios with an accuracy of 94.47% and for four real applications with an accuracy of 95.86%. Yiheng Xu, Pranav Sivaraman, Hariharan Devarajan, Kathryn Mohror, Abhinav Bhatele |
HiPC | 1 |
| 2024 | Lemur: Harmonizing Natural Language and Code for Language AgentsabstractWe introduce Lemur and Lemur-Chat, openly accessible language models optimized
for both natural language and coding capabilities to serve as the backbone
of versatile language agents. The evolution from language chat models to
functional language agents demands that models not only master human interaction,
reasoning, and planning but also ensure grounding in the relevant environments.
This calls for a harmonious blend of language and coding capabilities
in the models. Lemur and Lemur-Chat are proposed to address this necessity,
demonstrating balanced proficiencies in both domains, unlike existing
open-source models that tend to specialize in either. Through meticulous pretraining
using a code-intensive corpus and instruction fine-tuning on text and code
data, our models achieve state-of-the-art averaged performance across diverse
text and coding benchmarks. Comprehensive experiments demonstrate Lemur’s
superiority over existing open-source models and its proficiency across various
agent tasks involving human communication, tool usage, and interaction under
fully- and partially- observable environments. The harmonization between natural
and programming languages enables Lemur-Chat to significantly narrow the
gap with proprietary models on agent abilities, providing key insights into developing
advanced open-source agents adept at reasoning, planning, and operating
seamlessly across environments. Our model and code have been open-sourced at
https://github.com/OpenLemur/Lemur. Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu 0033, Binyuan Hui, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu 0009 |
ICLR | 1 |
| 2024 | OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsabstractAutonomous agents that accomplish complex computer tasks with minimal human interventions have the potential to transform human-computer interaction, significantly enhancing accessibility and productivity. However, existing benchmarks either lack an interactive environment or are limited to environments specific to certain applications or domains, failing to reflect the diverse and complex nature of real-world computer use, thereby limiting the scope of tasks and agent scalability. To address this issue, we introduce OSWorld, the first-of-its-kind scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems such as Ubuntu, Windows, and macOS. OSWorld can serve as a unified, integrated computer environment for assessing open-ended computer tasks that involve arbitrary applications. Building upon OSWorld, we create a benchmark of 369 computer tasks involving real web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications. Each task example is derived from real-world computer use cases and includes a detailed initial state setup configuration and a custom execution-based evaluation script for reliable, reproducible evaluation. Extensive evaluation of state-of-the-art LLM/VLM-based agents on OSWorld reveals significant deficiencies in their ability to serve as computer assistants. While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success, primarily struggling with GUI grounding and operational knowledge. Comprehensive analysis using OSWorld provides valuable insights for developing multimodal generalist agents that were not possible with previous benchmarks. Our code, environment, baseline models, and data are publicly available at this https URL. Tianbao Xie, Jixuan Chen, Xiaochuan Li 0003, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, Tao Yu 0009 |
NeurIPS | 12 |
| 2024 | Benchmarking mapping algorithms for cell-type annotating in mouse brain by integrating single-nucleus RNA-seq and Stereo-seq dataabstractLimited gene capture efficiency and spot size of spatial transcriptome (ST) data pose significant challenges in cell-type characterization. The heterogeneity and complexity of cell composition in the mammalian brain make it more challenging to accurately annotate ST data from brain. Many algorithms attempt to characterize subtypes of neuron by integrating ST data with single-nucleus RNA sequencing (snRNA-seq) or single-cell RNA sequencing. However, assessing the accuracy of these algorithms on Stereo-seq ST data remains unresolved. Here, we benchmarked 9 mapping algorithms using 10 ST datasets from four mouse brain regions in two different resolutions and 24 pseudo-ST datasets from snRNA-seq. Both actual ST data and pseudo-ST data were mapped using snRNA-seq datasets from the corresponding brain regions as reference data. After comparing the performance across different areas and resolutions of the mouse brain, we have reached the conclusion that both robust cell-type decomposition and SpatialDWLS demonstrated superior robustness and accuracy in cell-type annotation. Testing with publicly available snRNA-seq data from another sequencing platform in the cortex region further validated our conclusions. Altogether, we developed a workflow for assessing suitability of mapping algorithm that fits for ST datasets, which can improve the efficiency and accuracy of spatial data annotation. Quyuan Tao, Yiheng Xu, Youzhe He |
Briefings Bioinform. | 2 |
| 2023 | Tooth Segmentation from Cone-Beam CT Images Through Boundary Refinement
Yiheng Xu, Mingkun Zhang, Sibo Huang |
ICANN (4) | 1 |
| 2022 | MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingabstractMultimodal pre-training with text, layout, and image has made significant progress for Visually Rich Document Understanding (VRDU), especially the fixed-layout documents such as scanned document images.While, there are still a large number of digital documents where the layout information is not fixed and needs to be interactively and dynamically rendered for visualization, making existing layout-based pre-training approaches not easy to apply.In this paper, we propose MarkupLM for document understanding tasks with markup languages as the backbone, such as HTML/XMLbased documents, where text and markup information is jointly pre-trained.Experiment results show that the pre-trained MarkupLM significantly outperforms the existing strong baseline models on several document understanding tasks.The pre-trained model and code will be publicly available at https:// aka.ms/markuplm. Yiheng Xu, Lei Cui 0001, Furu Wei |
ACL (1) | 2 |
| 2022 | DiT: Self-supervised Pre-training for Document Image TransformerabstractImage Transformer has recently achieved significant progress for natural image understanding, either using supervised (ViT, DeiT, etc.) or self-supervised (BEiT, MAE, etc.) pre-training techniques. In this paper, we propose DiT, a self-supervised pre-trained Document Image Transformer model using large-scale unlabeled text images for Document AI tasks, which is essential since no supervised counterparts ever exist due to the lack of human-labeled document images. We leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR. Experiment results have illustrated that the self-supervised pre-trained DiT model achieves new state-of-the-art results on these downstream tasks, e.g. document image classification (91.11 - 92.69), document layout analysis (91.0 - 94.9), table detection (94.23 - 96.55) and text detection for OCR (93.07 - 94.29). The code and pre-trained models are publicly available at https://aka.ms/msdit. Yiheng Xu, Tengchao Lv, Lei Cui 0001, Cha Zhang, Furu Wei |
ACM Multimedia | 2 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 2 |
| 2021 | LayoutReader: Pre-training of Text and Layout for Reading Order DetectionabstractReading order detection is the cornerstone to understanding visually-rich documents (e.g., receipts and forms).Unfortunately, no existing work took advantage of advanced deep learning models because it is too laborious to annotate a large enough dataset.We observe that the reading order of WORD documents is embedded in their XML metadata; meanwhile, it is easy to convert WORD documents to PDFs or images.Therefore, in an automated manner, we construct ReadingBank, a benchmark dataset that contains reading order, text, and layout information for 500,000 document images covering a wide spectrum of document types.This first-ever large-scale dataset unleashes the power of deep neural networks for reading order detection.Specifically, our proposed LayoutReader captures the text and layout information for reading order prediction using the seq2seq model.It performs almost perfectly in reading order detection and significantly improves both open-source and commercial OCR engines in ordering text lines in their results in our experiments.The dataset and models are publicly available at https: //aka.ms/layoutreader. Zilong Wang 0002, Yiheng Xu, Lei Cui 0001, Jingbo Shang, Furu Wei |
EMNLP (1) | 2 |
| 2020 | Graph Convolutional Networks with Markov Random Field Reasoning for Social Spammer DetectionabstractThe recent growth of social networking platforms also led to the emergence of social spammers, who overwhelm legitimate users with unwanted content. The existing social spammer detection methods can be characterized into two categories: features based ones and propagation-based ones. Features based methods mainly rely on matrix factorization using tweet text features, and regularization using social graphs is incorporated. However, these methods are fully supervised and can only utilize labeled part of social graphs, which fail to work in a real-world semi-supervised setting. The propagation-based methods primarily employ Markov Random Fields (MRFs) to capture human intuitions in user following relations, which cannot take advantages of rich text features. In this paper, we propose a novel social spammer detection model based on Graph Convolutional Networks (GCNs) that operate on directed social graphs by explicitly considering three types of neighbors. Furthermore, inspired by the propagation-based methods, we propose a MRF layer with refining effects to encapsulate these human insights in social relations, which can be formulated as a RNN through mean-field approximate inference, and stack on top of GCN layers to enable end-to-end training. We evaluate our proposed method on two real-world social network datasets, and the results demonstrate that our method outperforms the state-of-the-art approaches. Defu Lian, Yiheng Xu, Le Wu 0001, Enhong Chen |
AAAI | 3 |
| 2020 | DocBank: A Benchmark Dataset for Document Layout AnalysisabstractDocument layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture.Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient.In this paper, we present DocBank, a benchmark dataset that contains 500K document pages with fine-grained tokenlevel annotations for document layout analysis.DocBank is constructed using a simple yet effective way with weak supervision from the L A T E X documents available on the arXiv.com.With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis.We build several strong baselines and manually split train/dev/test sets for evaluation.Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents.The DocBank dataset is publicly available at https: //github.com/doc-analysis/DocBank. Minghao Li 0004, Yiheng Xu, Lei Cui 0001, Shaohan Huang, Furu Wei, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 2 |
| 2020 | LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingabstractPre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the LayoutLM to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at https://aka.ms/layoutlm. Yiheng Xu, Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001 |
KDD | 1 |