VLDB 2026 Research / reviewers in the wild / expert
Xinyu Wang 0010
dblp:68/1277-10
· DBLP profile ↗
18ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-9082-094XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Are Large Vision Language Models Good Game Players?abstractLarge Vision Language Models (LVLMs) have demonstrated remarkable abilities in understanding and reasoning about both visual and textual information. However, existing evaluation methods for LVLMs, primarily based on benchmarks like Visual Question Answering and image captioning, often fail to capture the full scope of LVLMs' capabilities. These benchmarks are limited by issues such as inadequate assessment of detailed visual perception, data contamination, and a lack of focus on multi-turn reasoning. To address these challenges, we propose LVLM-Playground, a game-based evaluation framework designed to provide a comprehensive assessment of LVLMs' cognitive and reasoning skills in structured environments. LVLM-Playground uses a set of games to evaluate LVLMs on four core tasks: Perceiving, Question Answering, Rule Following, and End-to-End Playing, with each target task designed to assess specific abilities, including visual perception, reasoning, decision-making, etc. Based on this framework, we conduct extensive experiments that explore the limitations of current LVLMs, such as handling long structured outputs and perceiving detailed and dense elements. Code and data are publicly available at https://github.com/xinke-wang/LVLM-Playground. Xinyu Wang 0010, Bohan Zhuang, Qi Wu 0001 |
ICLR | 1 |
| 2025 | OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningabstractScoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10,000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1,500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https://github.com/Yuliang-Liu/MultimodalOCR. Zhebin Kuang, Jiajun Song, Mingxin Huang, Linghao Zhu, Qidi Luo, Xinyu Wang 0010, Hao Lu 0003, Guozhi Tang, Bin Shan, Chunhui Lin, Binghong Wu, Hao Feng 0009, Hao Liu 0003, Can Huang 0002, Jingqun Tang, Wei Chen 0088, Xiang Bai |
NeurIPS | 9 |
| 2025 | NavBench: Probing Multimodal Large Language Models for Embodied NavigationabstractMultimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3,200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge. Yanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An 0002, Yutong Xie 0001, Xinyu Wang 0010, Qi Wu 0001 |
NeurIPS | 7 |
| 2025 | Text in the dark: Extremely low-light text image enhancement
Che-Tsung Lin, Chun Chet Ng, Zhi Qin Tan, Wan Jun Nah, Xinyu Wang 0010, Jie-Long Kew, Po-Hao Hsu, Shang-Hong Lai, Chee Seng Chan, Christopher Zach |
Signal Process. Image Commun. | 5 |
| 2024 | Deciphering Oracle Bone Language with Diffusion ModelsabstractHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han, Yongge Liu, Lianwen Jin, Xiang Bai, Yuliang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haisu Guan, Huanxin Yang, Xinyu Wang 0010, Shengwei Han, Yongge Liu, Xiang Bai |
ACL (1) | 3 |
| 2024 | ModaVerse: Efficiently Transforming Modalities with LLMsabstractHumans possess the capability to comprehend diverse modalities and seamlessly transfer information between them. In this work, we introduce ModaVerse, a Multi-modal Large Language Model (MLLM) capable of comprehending and transforming content across various modalities in-cluding images, videos, and audio. Predominant MLLM frameworks have largely relied on aligning latent spaces of textual and non-textual features. This alignment process, which synchronizes a language model trained on tex-tual data with encoders and decoders trained on multimodal data, often necessitates extensive training of several projection layers in multiple stages. Inspired by LLM-as-agent methodologies, we propose a novel Input/Output (I/O) alignment mechanism that operates directly at the level of natural language. It aligns the LLM's output with the input of generative models, avoiding the complexities associated with latent feature alignments, and simplifying the multiple training stages of existing MLLMs into a single, efficient process. By conducting experiments on several benchmarks, we demonstrate that our approach at-tains comparable performance with the state of the art while achieving considerable efficiencies in data usage. The code is available at https://github.com/xinke-wang/ModaVerse. Xinyu Wang 0010, Bohan Zhuang, Qi Wu 0001 |
CVPR | 1 |
| 2024 | Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction
Pengjie Wang 0005, Kaile Zhang, Xinyu Wang 0010, Shengwei Han, Yongge Liu, Xiang Bai |
ICDAR (1) | 3 |
| 2024 | When IC meets text: Towards a rich annotated integrated circuit text dataset
Chun Chet Ng, Che-Tsung Lin, Zhi Qin Tan, Xinyu Wang 0010, Jie-Long Kew, Chee Seng Chan, Christopher Zach |
Pattern Recognit. | 4 |
| 2024 | Improving Handwritten Mathematical Expression Recognition via Similar Symbol DistinguishingabstractHandwritten mathematical expression recognition (HMER) is an essential task in the OCR community, which consists of two sub-tasks, i.e., symbol recognition and structure parsing. Modern literature treats HMER as a LaTeX sequence predicting problem that simultaneously recognizes symbols and parses the structures of MEs. Although deep learning-based HMER methods have been achieving promising results on public benchmarks, it is admitted that the misclassification error between visually similar symbols still prevents these approaches from more generalized scenes. In this paper, we try to solve this issue from three aspects. 1) We enhanced the feature extraction progress by introducing path signature features, which incorporates local writing details and global spatial information. 2) We developed a language model that uses contextual information to correct the symbols misclassified by vision-only-based recognition models. 3) We solved the misalignment problem in existing ensemble method by designing a dynamic time warping (DTW) based algorithm. By combining the above improvements, our method achieved state-of-the-art results on three CROHME benchmarks, outperforming previous methods by a large margin. Zhe Li 0046, Xinyu Wang 0010, Yichao Huang, Kai Ding 0009 |
IEEE Trans. Multim. | 2 |
| 2023 | SPTS v2: Single-Point Scene Text SpottingabstractEnd-to-end scene text spotting has made significant progress due to its intrinsic synergy between text detection and recognition. Previous methods commonly regard manual annotations such as horizontal rectangles, rotated rectangles, quadrangles, and polygons as a prerequisite, which are much more expensive than using single-point. Our new framework, SPTS v2, allows us to train high-performing text-spotting models using a single-point annotation. SPTS v2 reserves the advantage of the auto-regressive Transformer with an Instance Assignment Decoder (IAD) through sequentially predicting the center points of all text instances inside the same predicting sequence, while with a Parallel Recognition Decoder (PRD) for text recognition in parallel, which significantly reduces the requirement of the length of the sequence. These two decoders share the same parameters and are interactively connected with a simple but effective information transmission process to pass the gradient and information. Comprehensive experiments on various existing benchmark datasets demonstrate the SPTS v2 can outperform previous state-of-the-art single-point text spotters with fewer parameters while achieving 19× faster inference speed. Within the context of our SPTS v2 framework, our experiments suggest a potential preference for single-point representation in scene text spotting when compared to other representations. Such an attempt provides a significant opportunity for scene text spotting applications beyond the realms of existing paradigms. Jiaxin Zhang 0003, Dezhi Peng, Mingxin Huang, Xinyu Wang 0010, Jingqun Tang, Can Huang 0002, Dahua Lin, Chunhua Shen, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | SPTS: Single-Point Text SpottingabstractExisting scene text spotting (i.e., end-to-end text detection and recognition) methods rely on costly bounding box annotations (e.g., text-line, word-level, or character-level bounding boxes). For the first time, we demonstrate that training scene text spotting models can be achieved with an extremely low-cost annotation of a single-point for each instance. We propose an end-to-end scene text spotting method that tackles scene text spotting as a sequence prediction task. Given an image as input, we formulate the desired detection and recognition results as a sequence of discrete tokens and use an auto-regressive Transformer to predict the sequence. The proposed method is simple yet effective, which can achieve state-of-the-art results on widely used benchmarks. Most significantly, we show that the performance is not very sensitive to the positions of the point annotation, meaning that it can be much easier to be annotated or even be automatically generated than the bounding box that requires precise positions. We believe that such a pioneer attempt indicates a significant opportunity for scene text spotting applications of a much larger scale than previously possible. The code is available at https://github.com/shannanyinxiang/SPTS. Dezhi Peng, Xinyu Wang 0010, Jiaxin Zhang 0003, Mingxin Huang, Songxuan Lai, Jing Li 0036, Shenggao Zhu, Dahua Lin, Chunhua Shen, Xiang Bai |
ACM Multimedia | 2 |
| 2021 | ICDAR 2021 Competition on Integrated Circuit Text Spotting and Aesthetic Assessment
Chun Chet Ng, Akmalul Khairi Bin Nazaruddin, Yeong Khang Lee, Xinyu Wang 0010, Chee Seng Chan, Yipeng Sun, Lixin Fan |
ICDAR (4) | 4 |
| 2021 | Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
Tong He 0001, Hao Chen 0041, Xinyu Wang 0010, Canjie Luo, Shuaitao Zhang, Chunhua Shen |
Int. J. Comput. Vis. | 4 |
| 2020 | On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringabstractVisual Question Answering (VQA) methods have made incredible progress, but suffer from a failure to generalize. This is visible in the fact that they are vulnerable to learning coincidental correlations in the data rather than deeper relations between image content and ideas expressed in language. We present a dataset that takes a step towards addressing this problem in that it contains questions expressed in two languages, and an evaluation process that co-opts a well understood image-based metric to reflect the method’s ability to reason. Measuring reasoning directly encourages generalization by penalizing answers that are coincidentally correct. The dataset reflects the scene-text version of the VQA problem, and the reasoning evaluation can be seen as a text-based version of a referring expression challenge. Experiments and analyses are provided that show the value of the dataset. The dataset is available at www.est-vqa.org. Xinyu Wang 0010, Chunhua Shen, Chun Chet Ng, Canjie Luo, Chee Seng Chan, Anton van den Hengel |
CVPR | 1 |
| 2020 | Human Detection Aided by Deeply Learned Semantic MasksabstractHuman detection is one of the long-standing computer vision tasks, and it has been a cornerstone for many real-world applications, such as photo album organization, video surveillance, and autonomous driving. Benefiting from deep learning technologies, such as convolutional neural networks and modern object detectors, have been achieving much improved accuracy in generic object detection tasks. In this paper, we aim to improve deep learning-based human detection. Our main idea is to exploit semantic context information for human detection by using deep-learnt semantic features provided by semantic segmentation masks. Segmentation masks play as an attention mechanism and enforce the detectors to focus on the image regions where potential object candidates are likely to appear. Meanwhile, the extra segmentation mask channel can also guide the convolutional kernels to automatically learn more discriminative features which make it easier to distinguish the background and foreground. We implement our methods with two popular detection frameworks, i.e., faster R-CNN and SSD and experimentally analyze the effectiveness of the proposed methods. Evaluation results on the widely used MS-COCO dataset and the very recent CrowdHuman dataset are provided. Our proposed methods outperform the baseline detectors and achieve better performance on highly occluded human detection. Xinyu Wang 0010, Chunhua Shen, Shugong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Real-Time Deep Tracking via Corrective Domain AdaptationabstractVisual tracking is one of the fundamental problems in computer vision. Recently, some deep-learning-based tracking algorithms have been illustrating record-breaking performances. However, due to the high complexity of neural networks, most deep trackers suffer from low tracking speed and are, thus, impractical in many real-world applications. Some recently proposed deep trackers with smaller network structure achieve high efficiency while at the cost of significant decrease in precision. In this paper, we propose to transfer the deep feature, which is learned originally for image classification to the visual tracking domain. The domain adaptation is achieved via some “grafted” auxiliary networks, which are trained by regressing the object location in tracking frames. This adaptation improves the tracking performance significantly both on accuracy and efficiency. The yielded deep tracker is real time and also illustrates the state-of-the-art accuracies in the experiment involving two well-adopted benchmarks with more than 100 test videos. Furthermore, the adaptation is also naturally used for introducing the objectness concept into visual tracking. This removes a long-standing target ambiguity in visual tracking tasks, and we illustrate the empirical superiority of the more well-defined task. Xinyu Wang 0010, Fumin Shen, Yi Li 0025, Fatih Porikli, Mingwen Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Deep tracking with objectnessabstractVisual tracking is a fundamental problem in computer vision. However, due to the (sometimes) ambiguous target information given at the first frame, it has also been criticized as less well-posed compared with other tasks with clearly-defined targets, such as object detection and semantic segmentation. In this paper, we try to evaluate the importance of object category in visual tracking by tracking objects with known object types. The proposed algorithm, termed Deep-Track with Objectness (DTO), naturally combines the state-of-the-art deep-learning-based detectors and trackers, which essentially share a large part of the network. In DTO, a deep tracker, which is scale-fixed and sensitive to small translations tracks the object in a relative short lifespan. A deep detector, which is scale-changeable and robust to pose or illumination changes guides the deep tracker in a longer lifespan. As the deep tracker and detector share the main part of their networks, no much extra computation is imposed while the performance gain is significant. We test the proposed algorithm on two well-accepted benchmarks and on both of them, the proposed method increases the tracking accuracies remarkably compared with state-of-the-art visual trackers. Xinyu Wang 0010, Yi Li 0025, Fatih Porikli, Mingwen Wang 0001 |
ICIP | 1 |
| 2017 | Robust and real-time deep tracking via multi-scale domain adaptationabstractVisual tracking is a fundamental problem in computer vision. Recently, some deep-learning-based tracking algorithms have been achieving record-breaking performances. However, due to the high complexity of deep learning, most deep trackers suffer from low tracking speed, and thus are impractical in many real-world applications. Some new deep trackers with smaller network structure achieve high efficiency while at the cost of significant decrease on precision. In this paper, we propose to transfer the feature for image classification to the visual tracking domain via convolutional channel reductions. The channel reduction could be simply viewed as an additional convolutional layer with the specific task. It not only extracts useful information for object tracking but also significantly increases the tracking speed. To better accommodate the useful feature of the target in different scales, the adaptation filters are designed with different sizes. The yielded visual tracker is real-time and also illustrates the state-of-the-art accuracies in the experiment involving two well-adopted benchmarks with more than 100 test videos. Xinyu Wang 0010, Yi Li 0025, Fumin Shen, Fatih Porikli |
ICME | 1 |