VLDB 2026 Research / reviewers in the wild / expert
Yongxin Shi
dblp:359/4310
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingabstractRecent multimodal large language models (MLLMs) still struggle with long document understanding due to two fundamental challenges: information interference from abundant irrelevant content, and the quadratic computational cost of Transformer-based architectures. Existing approaches primarily fall into two categories: token compression, which sacrifices fine-grained details; and introducing external retrievers, which increase system complexity and prevent end-to-end optimization. To address these issues, we conduct an in-depth analysis and observe that MLLMs exhibit a human-like coarse-to-fine reasoning pattern: early Transformer layers attend broadly across the document, while deeper layers focus on relevant evidence pages. Motivated by this insight, we posit that the inherent evidence localization capabilities of MLLMs can be explicitly leveraged to perform retrieval during the reasoning process, facilitating efficient long document understanding. To this end, we propose URaG, a simple-yet-effective framework that Unifies Retrieval and Generation within a single MLLM. URaG introduces a lightweight cross-modal retrieval module that converts the early Transformer layers into an efficient evidence selector, identifying and preserving the most relevant pages while discarding irrelevant content. This design enables the deeper layers to concentrate computational resources on pertinent information, improving both accuracy and efficiency. Extensive experiments demonstrate that URaG achieves state-of-the-art performance while reducing computational overhead by 44-56%. Yongxin Shi, Zeyu Shan, Dezhi Peng, Zening Lin |
AAAI | 1 |
| 2026 | Physics-guided and dual attention salient object detection in sand-dust remote sensing images
Yongxin Shi, Weiyi Wei, Shengxia Gao, Linfeng Cao |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Predicting the Original Appearance of Damaged Historical DocumentsabstractHistorical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily focus on binarization, enhancement, etc., neglecting the repair of these damages. To this end, we present a new task, termed Historical Document Repair (HDR), which aims to predict the original appearance of damaged historical documents. To fill the gap in this field, we propose a large-scale dataset HDR28K and a diffusion-based network DiffHDR for historical document repair. Specifically, HDR28K contains 28,552 damaged-repaired image pairs with character-level annotations and multi-style degradations. Moreover, DiffHDR augments the vanilla diffusion framework with semantic and spatial information and a meticulously designed character perceptual loss for contextual and visual coherence. Experimental results demonstrate that the proposed DiffHDR trained on HDR28K significantly surpasses existing approaches and exhibits remarkable performance in handling real scenarios. Notably, DiffHDR can also be extended to document editing and text block generation, showcasing its high flexibility and generalization capacity. We believe this study could pioneer a new direction of document processing and contribute to the inheritance of invaluable cultures and civilizations. Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 0002, Chongyu Liu |
AAAI | 3 |
| 2025 | MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical StudiesabstractWith the rapid development of Multimodal Large Language Models (MLLMs), their potential in Chinese Classical Studies (CCS), a field which plays a vital role in preserving and promoting China’s rich cultural heritage, remains largely unexplored due to the absence of specialized benchmarks. To bridge this gap, we propose MCS-Bench, the first-of-its-kind multimodal benchmark specifically designed for CCS across multiple subdomains. MCS-Bench spans seven core subdomains (Ancient Chinese Text, Calligraphy, Painting, Oracle Bone Script, Seal, Cultural Relic, and Illustration), with a total of 45 meticulously designed tasks. Through extensive evaluation of 37 representative MLLMs, we observe that even the top-performing model (InternVL2.5-78B) achieves an average score below 50, indicating substantial room for improvement. Our analysis reveals significant performance variations across different tasks and identifies critical challenges in areas such as Optical Character Recognition (OCR) and cultural context interpretation. MCS-Bench not only establishes a standardized baseline for CCS-focused MLLM research but also provides valuable insights for advancing cultural heritage preservation and innovation in the Artificial General Intelligence (AGI) era. Data and code will be publicly available. Yang Liu 0353, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi, Kai Ding 0009 |
ACL (1) | 4 |
| 2025 | Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationabstractYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuyi Zhang 0002, Peirong Zhang 0001, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo |
ACL (1) | 5 |
| 2025 | TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningabstractChinese Classical Studies (CCS) is a pivotal gateway to ancient Chinese culture. Spanning ancient texts, illustrations, paintings, and calligraphy, CCS presents significant challenges for non-specialists due to its language and visual complexity. While Large Language Models (LLMs) have been explored to facilitate CCS, current methods primarily focus on textual analysis, overlooking the rich visual information intrinsic to classical materials. To bridge this gap, we propose TongGu-VL, a pioneering specialized MLLM designed for CCS applications. Our contributions are threefold. First, we construct CCS358K, a comprehensive multimodal instruction dataset to enhance MLLMs' CCS capabilities. Second, we propose Parameter Sensitivity-Guided Instruction Tuning (PSG-IT), a novel method that mitigates catastrophic forgetting without data replay. It effectively preserves TongGu-VL's general skills, while optimizing its CCS performance. Third, we design a Visual-Text Early Fusion (VTEF) module, which harnesses MLLMs' modality alignment to generate instruction-aware visual representations, thus improving language modeling. Extensive experimental results demonstrate that our model outperforms existing MLLMs on a broad range of CCS tasks, while maintaining general capabilities that benefit other domains beyond CCS. Our model and dataset will be publicly available. Jiahuan Cao, Yang Liu 0353, Peirong Zhang 0001, Yongxin Shi, Kai Ding 0009 |
ACM Multimedia | 4 |
| 2025 | RDFCNet: RGB-guided depth feature calibration network for RGB-D salient object detection
Weiyi Wei, Yongxin Shi, Sixuan Liu |
Neurocomputing | 2 |
| 2025 | MegaHan97K: A large-scale dataset for mega-category Chinese character recognition with over 97K categories
Yuyi Zhang 0002, Yongxin Shi, Peirong Zhang 0001, Yixin Zhao, Zhenhua Yang |
Pattern Recognit. | 2 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 5 |
| 2024 | WenMind: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Classical Literature and Language ArtsabstractLarge Language Models (LLMs) have made significant advancements across numerous domains, but their capabilities in Chinese Classical Literature and Language Arts (CCLLA) remain largely unexplored due to the limited scope and tasks of existing benchmarks. To fill this gap, we propose WenMind, a comprehensive benchmark dedicated for evaluating LLMs in CCLLA. WenMind covers the sub-domains of Ancient Prose, Ancient Poetry, and Ancient Literary Culture, comprising 4,875 question-answer pairs, spanning 42 fine-grained tasks, 3 question formats, and 2 evaluation scenarios: domain-oriented and capability-oriented. Based on WenMind, we conduct a thorough evaluation of 31 representative LLMs, including general-purpose models and ancient Chinese LLMs. The results reveal that even the best-performing model, ERNIE-4.0, only achieves a total score of 64.3, indicating significant room for improvement of LLMs in the CCLLA domain. We also provide insights into the strengths and weaknesses of different LLMs and highlight the importance of pre-training data in achieving better results.Overall, WenMind serves as a standardized and comprehensive baseline, providing valuable insights for future CCLLA research. Our benchmark and related code are available at \url{https://github.com/SCUT-DLVCLab/WenMind}. Jiahuan Cao, Yang Liu 0353, Yongxin Shi, Kai Ding 0009 |
NeurIPS | 3 |
| 2023 | M5HisDoc: A Large-scale Multi-style Chinese Historical Document Analysis BenchmarkabstractRecognizing and organizing text in correct reading order plays a crucial role in historical document analysis and preservation. While existing methods have shown promising performance, they often struggle with challenges such as diverse layouts, low image quality, style variations, and distortions. This is primarily due to the lack of consideration for these issues in the current benchmarks, which hinders the development and evaluation of historical document analysis and recognition (HDAR) methods in complex real-world scenarios. To address this gap, this paper introduces a complex multi-style Chinese historical document analysis benchmark, named M5HisDoc. The M5 indicates five properties of style, ie., Multiple layouts, Multiple document types, Multiple calligraphy styles, Multiple backgrounds, and Multiple challenges. The M5HisDoc dataset consists of two subsets, M5HisDoc-R (Regular) and M5HisDoc-H (Hard). The M5HisDoc-R subset comprises 4,000 historical document images. To ensure high-quality annotations, we meticulously perform manual annotation and triple-checking. To replicate real-world conditions for historical document analysis applications, we incorporate image rotation, distortion, and resolution reduction into M5HisDoc-R subset to form a new challenging subset named M5HisDoc-H, which contains the same number of images as M5HisDoc-R. The dataset exhibits diverse styles, significant scale variations, dense texts, and an extensive character set. We conduct benchmarking experiments on five tasks: text line detection, text line recognition, character detection, character recognition, and reading order prediction. We also conduct cross-validation with other benchmarks. Experimental results demonstrate that the M5HisDoc dataset can offer new challenges and great opportunities for future research in this field, thereby providing deep insights into the solution for HDAR. The dataset is available at https://github.com/HCIILAB/M5HisDoc. Yongxin Shi, Chongyu Liu, Dezhi Peng, Cheng Jian, Jiarong Huang |
NeurIPS | 1 |