Xiaoxin Chen 0001

dblp:17/2084-1 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LENS: Learning to Segment Anything with Unified Reinforced Reasoning
abstract
Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically ignore explicit chain-of-thought (CoT) reasoning at test time, which limits their ability to generalize to unseen prompts and domains. To address this issue, we introduce LENS, a scalable reinforcement-learning framework that jointly optimizes the reasoning process and segmentation in an end-to-end manner. We propose unified reinforcement-learning rewards that span sentence-, box-, and segment-level cues, encouraging the model to generate informative CoT rationales while refining mask quality. Using a publicly available 3-billion-parameter vision–language model, i.e., Qwen2.5-VL-3B-Instruct, LENS achieves an average cIoU of 81.2% on the RefCOCO, RefCOCO+, and RefCOCOg benchmarks, outperforming the strong fine-tuned method, i.e., GLaMM, by up to 5.6%. These results demonstrate that RL-driven CoT reasoning significantly enhances text-prompted segmentation and offers a practical path toward more generalizable Segment Anything models (SAM).
Lianghui Zhu, Bin Ouyang, Tianheng Cheng, Haocheng Shen, Longjin Ran, Xiaoxin Chen 0001, Li Yu 0003, Wenyu Liu 0001, Xinggang Wang
AAAI8
2026 EVF-SAM: Early Vision-Language Fusion for text-prompted Segment Anything Model
Tianheng Cheng, Lianghui Zhu, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
Image Vis. Comput.7
2025 Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy
abstract
In recent years, the use of large language models (LLMs) for text classification has attracted widespread attention. Despite this, the classification accuracy of LLMs has not yet universally surpassed that of smaller models. LLMs can enhance their performance in text classification through fine-tuning. However, existing data quality research based on LLMs is challenging to apply directly to solve text classification problems. To further improve the performance of LLMs in classification tasks, this paper proposes a data quality enhancement (DQE) method for text classification based on LLMs. This method starts by using a greedy algorithm to select data, dividing the dataset into sampled and unsampled subsets, and then performing fine-tuning of the LLMs using the sampled data. Subsequently, this model is used to predict the outcomes for the unsampled data, categorizing incorrectly predicted data into uncovered, difficult, and noisy data. Experimental results demonstrate that our method effectively enhances the performance of LLMs in text classification tasks and significantly improves training efficiency, saving nearly half of the training time. Our method has achieved state-of-the-art performance in several open-source classification tasks.
Caiquan Liu, Chen Sang, Xiaoxin Chen 0001
COLING6
2025 Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
abstract
Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex lay-outs. Moreover, existing fine-tuning datasets for this domain often fall short in providing the detailed contextual information for robust understanding, leading to hallucinations and limited comprehension of spatial relationships among visual elements. To address these challenges, we propose an innovative pipeline that utilizes adaptive generation of markup languages, such as Markdown, JSON, HTML, and TiKZ, to build highly structured document representations and deliver contextually-grounded responses. We intro-duce two fine-grained structured datasets: DocMark-Pile, comprising approximately 3.8M pretraining data pairs for document parsing, and DocMark-Instruct, featuring 624k fine-tuning data annotations for grounded instruction following. Extensive experiments demonstrate that our pro-posed model significantly outperforms existing state-of-the-art MLLMs across a range of visual document understanding benchmarks, facilitating advanced reasoning and comprehension capabilities in complex visual scenarios. Our code and models are released at https://github.com/Euphoria16/DocMark.
Han Xiao 0010, Yina Xie, Guanxin Tan, Ke Wang 0036, Aojun Zhou, Hao Li 0069, Hao Shao, Peng Gao 0007, Yafei Wen, Xiaoxin Chen 0001, Shuai Ren 0002, Hongsheng Li 0001
CVPR13
2025 BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
abstract
The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective and accessible deployment platform for MLLMs, enabling seamless integration into everyday tasks. However, deploying MLLMs on mobile phones presents challenges due to limitations in memory size and computational capability, making it difficult to achieve smooth and real-time processing without extensive optimization. In this paper, we present BlueLM-V-3B, an algorithm and system co-design approach specifically tailored for the efficient deployment of MLLMs on mobile platforms. To be specific, we redesign the dynamic resolution scheme adopted by mainstream MLLMs and implement system optimization for hardware-aware deployment to optimize model inference on mobile phones. BlueLM-V-3B boasts the following key highlights: (1) Small Size: BlueLM-V-3B features a language model with 2.7B parameters and a vision encoder with 400M parameters. (2) Fast Speed: BlueLM-V-3B achieves a generation speed of 24.4 token/s on the MediaTek Dimensity 9300 processor with 4-bit LLM weight quantization. (3) Strong Performance: BlueLM-V-3B has attained the highest average score of 66.1 on the OpenCompass benchmark among models with ≤ 4B parameters and surpassed a series of models with much larger parameter sizes (e.g., MiniCPM-V-2.6, InternVL2-8B).
Boheng Chen, Yina Xie, Guanxin Tan, Renshou Wu, Liuyang Bian, Zhaoxiong Wang, Yanzhou Yang, Han Xiao 0010, Aojun Zhou, Yafei Wen, Xiaoxin Chen 0001, Shuai Ren 0002, Hongsheng Li 0001
CVPR20
2025 SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?
abstract
Large Language Models (LLMs) have become integral to daily life, especially advancing as intelligent assistants through on-device deployment on smartphones.However, existing LLM evaluation benchmarks predominantly focus on objective tasks like mathematics and coding in English, which do not necessarily reflect the practical use cases of on-device LLMs in realworld mobile scenarios, especially for Chinese users.To address these gaps, we introduce SmartBench, the first benchmark designed to evaluate the capabilities of on-device LLMs in Chinese mobile contexts.We analyze functionalities provided by representative smartphone manufacturers and divide them into five categories: text summarization, text Q&A, information extraction, content creation, and notification management, further detailed into 20 specific tasks.For each task, we construct highquality datasets comprising 50 to 200 questionanswer pairs that reflect everyday mobile interactions, and we develop automated evaluation criteria tailored for these tasks.We conduct comprehensive evaluations of on-device LLMs and MLLMs using SmartBench and also assess their performance after quantized deployment on real smartphone NPUs.Our contributions provide a standardized framework for evaluating on-device LLMs in Chinese, promoting further development and optimization in this critical area.Code and data will be available at https://github.com/vivo-ai-lab/ SmartBench.
Haohao Gao, Renshou Wu, Shuai Ren 0002, Xiaoxin Chen 0001, Hongsheng Li 0001
EMNLP5
2025 GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
abstract
Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existing datasets, including limited object categories, insufficient textual diversity, and a scarcity of high-quality annotations. To mitigate these limitations, we introduce GroundingSuite, which comprises: (1) an automated data annotation framework leveraging multiple Vision-Language Model (VLM) agents; (2) a large-scale training dataset encompassing 9.56 million diverse referring expressions and their corresponding segmentations; and (3) a meticulously curated evaluation benchmark consisting of 3,800 images. The GroundingSuite training dataset facilitates substantial performance improvements, enabling models trained on it to achieve state-of-the-art results. Specifically, a cIoU of 68.9 on gRefCOCO and a gIoU of 55.3 on RefCOCOm. Moreover, the GroundingSuite annotation framework demonstrates superior efficiency compared to the current leading data annotation method, i.e., $4.5 \times$ faster than GLaMM.
Lianghui Zhu, Tianheng Cheng, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
ICCV8
2025 GenieBlue: Integrating Both Linguistic and Multimodal Capabilities for Large Language Models on Mobile Devices
Renshou Wu, Haohao Gao, Xi Chen 0072, Xue Yang 0005, Aojun Zhou, Yafei Wen, Xiaoxin Chen 0001, Shuai Ren 0002, Hongsheng Li 0001
ICCV11
2025 ControlAR: Controllable Image Generation with Autoregressive Models
abstract
Autoregressive (AR) models have reformulated image generation as next-token prediction, demonstrating remarkable potential and emerging as strong competitors to diffusion models. However, control-to-image generation, akin to ControlNet, remains largely unexplored within AR models. Although a natural approach, inspired by advancements in Large Language Models, is to tokenize control images into tokens and prefill them into the autoregressive model before decoding image tokens, it still falls short in generation quality compared to ControlNet and suffers from inefficiency. To this end, we introduce ControlAR, an efficient and effective framework for integrating spatial controls into autoregressive image generation models. Firstly, we explore control encoding for AR models and propose a lightweight control encoder to transform spatial inputs (e.g., canny edges or depth maps) into control tokens. Then ControlAR exploits the conditional decoding method to generate the next image token conditioned on the per-token fusion between control and image tokens, similar to positional encodings. Compared to prefilling tokens, using conditional decoding significantly strengthens the control capability of AR models but also maintains the model efficiency. Furthermore, the proposed ControlAR surprisingly empowers AR models with arbitrary-resolution image generation via conditional decoding and specific controls. Extensive experiments can demonstrate the controllability of the proposed ControlAR for the autoregressive control-to-image generation across diverse inputs, including edges, depths, and segmentation masks. Furthermore, both quantitative and qualitative results indicate that ControlAR surpasses previous state-of-the-art controllable diffusion models, e.g., ControlNet++.
Zongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun, Haocheng Shen, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
ICLR7
2025 Predictive Data Selection: The Data That Predicts Is the Data That Teaches
abstract
Language model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient manner. Specifically, we draw inspiration from recent findings showing that compression efficiency (i.e., normalized loss) of diverse models on certain text correlates strongly with their downstream performance, when the text domain aligns with the downstream benchmarks (Huang et al., 2024). Building on this observation, we hypothesize that data on which model losses are predictive of downstream abilities also contribute effectively to learning, which shares similar intuition with Thrush et al. (2024). To leverage this insight, we introduce predictive data selection (PreSelect), a lightweight and efficient data selection method that requires training and deploying only a fastText-based scorer. Through comprehensive experiments with 1B and 3B parameter models, we demonstrate that models trained on 30B tokens selected with PreSelect surpass the performance of the vanilla baseline trained on 300B tokens, achieving a 10x reduction in compute requirements. Furthermore, PreSelect significantly outperforms other competitive data selection baselines, such as DCLM and FineWeb-Edu on a scale of 3B models trained on 100B tokens. We open-source our trained data selection scorer along with the curated datasets at https://github.com/hkust-nlp/PreSelect.
KaShun Shum, Yuzhen Huang 0002, Hongjian Zou, Yixuan Liao, Xiaoxin Chen 0001, Qian Liu 0033, Junxian He
ICML6
2025 UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
abstract
In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectively. The reward model, UI-Genie-RM, features an image-text interleaved architecture that efficiently processes historical context and unifies action-level and task-level rewards. To support the training of UI-Genie-RM, we develop deliberately-designed data generation strategies including rule-based verification, controlled trajectory corruption, and hard negative mining. To address the second challenge, a self-improvement pipeline progressively expands solvable complex GUI tasks by enhancing both the agent and reward models through reward-guided exploration and outcome verification in dynamic environments. For training the model, we generate UI-Genie-RM-517k and UI-Genie-Agent-16k, establishing the first reward-specific dataset for GUI agents while demonstrating high-quality synthetic trajectory generation without manual annotation. Experimental results show that UI-Genie achieves state-of-the-art performance across multiple GUI agent benchmarks with three generations of data-model self-improvement. We open-source our complete framework implementation and generated datasets to facilitate further research in https://github.com/Euphoria16/UI-Genie.
Han Xiao 0010, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Lue Fan, Liuyang Bian, Shuai Ren 0002, Yafei Wen, Xiaoxin Chen 0001, Aojun Zhou, Hongsheng Li 0001
NeurIPS13
2025 Progressive Visual Prompt Learning with Contrastive Feature Re-formation
Haocheng Shen, Boheng Chen, Yixuan Liao, Xiaoxin Chen 0001, Limin Wang 0002
Int. J. Comput. Vis.6
2024 A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language Models
abstract
Due to the continuous emergence of new data, version updates have become an indispensable requirement for Large Language Models (LLMs).The training paradigms for version updates of LLMs include pre-training from scratch (PTFS) and continual pre-training (CPT).Preliminary experiments demonstrate that PTFS achieves better pre-training performance, while CPT has lower training cost.Moreover, their performance and training cost gaps widen progressively with version updates.To investigate the underlying reasons for this phenomenon, we analyze the effect of learning rate adjustments during the two stages of CPT: preparing an initialization checkpoint and continual pre-training based on this checkpoint.We find that a large learning rate in the first stage and a complete learning rate decay process in the second stage are crucial for version updates of LLMs.Hence, we propose a learning rate path switching training paradigm.Our paradigm comprises one main path, where we pre-train a LLM with the maximal learning rate, and multiple branching paths, each of which corresponds to an update of the LLM with newly-added training data.Extensive experiments demonstrate the effectiveness and generalization of our paradigm.Particularly, when training four versions of LLMs, our paradigm reduces the total training cost to 58% compared to PTFS, while maintaining comparable pretraining performance.
Jianheng Huang, Yixuan Liao, Xiaoxin Chen 0001, Junfeng Yao, Jinsong Su
EMNLP6
2024 DocReal: Robust Document Dewarping of Real-Life Images via Attention-Enhanced Control Point Prediction
abstract
Document image dewarping is a crucial task in computer vision with numerous practical applications. The control point method, as a popular image dewarping approach, has attracted attention due to its simplicity and efficiency. However, inaccurate control point prediction due to varying background noises and deformation types can result in unsatisfactory performance. To address these issues, we propose a robust document dewarping approach for real-life images, namely DocReal, which utilizes Enet to effectively remove background noise and an attention-enhanced control point (AECP) module to better capture local deformations. Moreover, we augment the training data by synthesizing 2D images with 3D deformations and additional deformation types. Our proposed method achieves state-of-the-art performance on the DocUNet benchmark and a newly proposed benchmark of 200 Chinese distorted images, exhibiting superior dewarping accuracy, OCR performance, and robustness to various types of image distortion.
Fangchen Yu, Yina Xie, Yafei Wen, Guozhi Wang, Shuai Ren 0002, Xiaoxin Chen 0001, Jianfeng Mao, Wenye Li 0001
WACV7
2023 Real-Time Image Demoiréing on Mobile Devices
Yuxin Zhang 0002, Mingbao Lin, Xunchao Li, Guozhi Wang, Fei Chao 0001, Shuai Ren 0002, Yafei Wen, Xiaoxin Chen 0001, Rongrong Ji
ICLR9
2021 Weakly-Supervised Instance Segmentation via Class-Agnostic Learning With Salient Images
abstract
Humans have a strong class-agnostic object segmentation ability and can outline boundaries of unknown objects precisely, which motivates us to propose a box-supervised class-agnostic object segmentation (BoxCaseg) based solution for weakly-supervised instance segmentation. The BoxCaseg model is jointly trained using box-supervised images and salient images in a multi-task learning manner. The fine-annotated salient images provide class-agnostic and precise object localization guidance for box-supervised images. The object masks predicted by a pretrained BoxCaseg model are refined via a novel merged and dropped strategy as proxy ground truth to train a Mask R-CNN for weakly-supervised instance segmentation. Only using 7991 salient images, the weakly-supervised Mask R-CNN is on par with fully-supervised Mask R-CNN on PASCAL VOC and significantly outperforms previous state-of-the-art box-supervised instance segmentation methods on COCO. The source code, pretrained models and datasets are available at https://github.com/hustvl/BoxCaseg.
Xinggang Wang, Jiapei Feng, Bin Hu 0020, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001
CVPR6
2021 EEM: An End-to-end Evaluation Metric for Scene Text Detection and Recognition
Jiedong Hao, Yafei Wen, Jun Gan, Shuai Ren 0002, Xiaoxin Chen 0001
ICDAR (4)7