VLDB 2026 Research / reviewers in the wild / expert
Weifeng Lin
dblp:193/7842
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Competitive multitasking with constraints relaxation for constrained multi-objective optimization
Xinyu Zhou 0002, Weifeng Lin, Long Fan, Mingwen Wang 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007 |
Int. J. Comput. Vis. | 5 |
| 2025 | A dual-population and three-stage constrained multi-objective evolutionary algorithm based on alternative evolution and degenerationabstractConstrained multi-objective optimization problems (CMOPs) are challenging due to conflicting objectives and numerous constraints. In recent years, dual-population constrained multi-objective evolutionary algorithms (DP-CMOEAs) have been shown impressive performance in solving CMOPs, in which the main population searches for the constrained Pareto front (CPF), while the auxiliary population focuses on locating the unconstrained Pareto front (UPF) and providing valuable information to the main population. However, for existing DP-CMOEAs, the way of utilizing the auxiliary population is either under-utilized or over-utilized, leaving room for further performance improvement. To tackle this, this paper proposes a novel DP-CMOEA based on alternative evolution and degeneration. The proposed algorithm utilizes the auxiliary population by alternating between evolutionary and degradation stages based on its state. In the third stage, the auxiliary population ceases to evolve, thus avoiding unnecessary computational resource waste. Experimental results on three benchmark suites and three real-world problems show that the proposed algorithm outperforms seven state-of-the-art constrained multi-objective evolutionary algorithms (CMOEAs) in terms of both convergence and diversity. Yanjun Zhu, Weifeng Lin |
CEC | 3 |
| 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You WantabstractIn this paper, we present the Draw-and-Understand framework, exploring how to integrate visual prompting understanding capabilities into Multimodal Large Language Models (MLLMs). Visual prompts allow users to interact through multi-modal instructions, enhancing the models' interactivity and fine-grained image comprehension. In this framework, we propose a general architecture adaptable to different pre-trained MLLMs, enabling it to recognize various types of visual prompts (such as points, bounding boxes, and free-form shapes) alongside language understanding. Additionally, we introduce MDVP-Instruct-Data, a multi-domain dataset featuring 1.2 million image-visual prompt-text triplets, including natural images, document images, scene text images, mobile/web screenshots, and remote sensing images. Building on this dataset, we introduce MDVP-Bench, a challenging benchmark designed to evaluate a model's ability to understand visual prompting instructions. The experimental results demonstrate that our framework can be easily and effectively applied to various MLLMs, such as SPHINX-X and LLaVA. After training with MDVP-Instruct-Data and image-level instruction datasets, our models exhibit impressive multimodal interaction capabilities and pixel-level understanding, while maintaining their image-level visual perception performance. Weifeng Lin, Ruichuan An, Peng Gao 0007, Bocheng Zou, Yulin Luo, Siyuan Huang 0004, Shanghang Zhang, Hongsheng Li 0001 |
ICLR | 1 |
| 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language InstructionsabstractThis paper presents a versatile image-to-image visual assistant, PixWizard, designed for image generation, manipulation, and translation based on free-from language instructions. To this end, we tackle a variety of vision tasks into a unified image-text-to-image generation framework and curate an Omni Pixel-to-Pixel Instruction-Tuning Dataset. By constructing detailed instruction templates in natural language, we comprehensively include a large set of diverse vision tasks such as text-to-image generation, image restoration, image grounding, dense image prediction, image editing, controllable generation, inpainting/outpainting, and more. Furthermore, we adopt Diffusion Transformers (DiT) as our foundation model and extend its capabilities with a flexible any resolution mechanism, enabling the model to dynamically process images based on the aspect ratio of the input, closely aligning with human perceptual processes. The model also incorporates structure-aware and semantic-aware guidance to facilitate effective fusion of information from the input image. Our experiments demonstrate that PixWizard not only shows impressive generative and understanding abilities for images with diverse resolutions but also exhibits generalization capabilities with unseen tasks and human instructions. Weifeng Lin, Renrui Zhang, Le Zhuo, Shitian Zhao, Siyuan Huang 0004, Junlin Xie, Peng Gao 0007, Hongsheng Li 0001 |
ICLR | 1 |
| 2025 | MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought ReasoningabstractChain-of-Thought (CoT) has widely enhanced mathematical reasoning in Large Language Models (LLMs), but it still remains challenging for extending it to multimodal domains. Existing works either adopt a similar textual reasoning for image input, or seek to interleave visual signals into mathematical CoT. However, they face three key limitations for math problem-solving: *reliance on coarse-grained box-shaped image regions, limited perception of vision encoders on math content, and dependence on external capabilities for visual modification*. In this paper, we propose **MINT-CoT**, introducing **M**athematical **IN**terleaved **T**okens for **C**hain-**o**f-**T**hought visual reasoning. MINT-CoT adaptively interleaves relevant visual tokens into textual reasoning steps via an Interleave Token, which dynamically selects visual regions of any shapes within math figures. To empower this capability, we construct the MINT-CoT dataset, containing 54K mathematical problems aligning each reasoning step with visual regions at the token level, accompanied by a rigorous data generation pipeline. We further present a three-stage MINT-CoT training strategy, progressively combining text-only CoT SFT, interleaved CoT SFT, and interleaved CoT RL, which derives our MINT-CoT-7B model. Extensive experiments demonstrate the effectiveness of our method for effective visual interleaved reasoning in mathematical domains, where MINT-CoT-7B outperforms the baseline model by +34.08% on MathVista and +28.78% on GeoQA, respectively. Xinyan Chen 0001, Renrui Zhang, Dongzhi Jiang, Aojun Zhou, Shilin Yan, Weifeng Lin, Hongsheng Li 0001 |
NeurIPS | 6 |
| 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosabstractWe present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2$-$2.4$\times$ faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding. Weifeng Lin, Ruichuan An, Tianhe Ren, Renrui Zhang, Wentao Zhang 0001, Lei Zhang 0006, Hongsheng Li 0001 |
NeurIPS | 1 |
| 2025 | UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI AgentsabstractIn this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectively. The reward model, UI-Genie-RM, features an image-text interleaved architecture that efficiently processes historical context and unifies action-level and task-level rewards. To support the training of UI-Genie-RM, we develop deliberately-designed data generation strategies including rule-based verification, controlled trajectory corruption, and hard negative mining. To address the second challenge, a self-improvement pipeline progressively expands solvable complex GUI tasks by enhancing both the agent and reward models through reward-guided exploration and outcome verification in dynamic environments. For training the model, we generate UI-Genie-RM-517k and UI-Genie-Agent-16k, establishing the first reward-specific dataset for GUI agents while demonstrating high-quality synthetic trajectory generation without manual annotation. Experimental results show that UI-Genie achieves state-of-the-art performance across multiple GUI agent benchmarks with three generations of data-model self-improvement. We open-source our complete framework implementation and generated datasets to facilitate further research in https://github.com/Euphoria16/UI-Genie. Han Xiao 0010, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Lue Fan, Liuyang Bian, Shuai Ren 0002, Yafei Wen, Xiaoxin Chen 0001, Aojun Zhou, Hongsheng Li 0001 |
NeurIPS | 5 |
| 2024 | M2SD: Multiple Mixing Self-Distillation for Few-Shot Class-Incremental LearningabstractFew-shot Class-incremental learning (FSCIL) is a challenging task in machine learning that aims to recognize new classes from a limited number of instances while preserving the ability to classify previously learned classes without retraining the entire model. This presents challenges in updating the model with new classes using limited training data, particularly in balancing acquiring new knowledge while retaining the old. We propose a novel method named Multiple Mxing Self-Distillation (M2SD) during the training phase to address these issues. Specifically, we propose a dual-branch structure that facilitates the expansion of the entire feature space to accommodate new classes. Furthermore, we introduce a feature enhancement component that can pass additional enhanced information back to the base network by self-distillation, resulting in improved classification performance upon adding new classes. After training, we discard both structures, leaving only the primary network to classify new class instances. Extensive experiments demonstrate that our approach achieves superior performance over previous state-of-the-art methods. Jinhao Lin, Weifeng Lin, Jun Huang 0004, Ronghua Luo |
AAAI | 3 |
| 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsabstractWe propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory. Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICML | 5 |
| 2023 | Building A Mobile Text Recognizer via Truncated SVD-based Knowledge Distillation-Guided NAS
Weifeng Lin, Canyu Xie, Dezhi Peng, Cong Yao, Mengchao He |
BMVC | 1 |
| 2023 | Scale-Aware Modulation Meet TransformerabstractThis paper presents a new vision Transformer, Scale-Aware Modulation Transformer (SMT), that can handle various downstream tasks efficiently by combining the convolutional network and vision Transformer. The proposed Scale-Aware Modulation (SAM) in the SMT includes two primary novel designs. Firstly, we introduce the Multi-Head Mixed Convolution (MHMC) module, which can capture multi-scale features and expand the receptive field. Secondly, we propose the Scale-Aware Aggregation (SAA) module, which is lightweight but effective, enabling information fusion across different heads. By leveraging these two modules, convolutional modulation is further enhanced. Furthermore, in contrast to prior works that utilized modulations throughout all stages to build an attention-free network, we propose an Evolutionary Hybrid Network (EHN), which can effectively simulate the shift from capturing local to global dependencies as the network becomes deeper, resulting in superior performance. Extensive experiments demonstrate that SMT significantly outperforms existing state-of-the-art models across a wide range of visual tasks. Specifically, SMT with 11.5M / 2.4GFLOPs and 32M / 7.7GFLOPs can achieve 82.2% and 84.3% top-1 accuracy on ImageNet-1K, respectively. After pretrained on ImageNet-22K in 2242resolution, it attains 87.1% and 88.1% top-1 accuracy when finetuned with resolution 2242and 3842, respectively. For object detection with Mask R-CNN, the SMT base trained with 1× and 3× schedule outperforms the Swin Transformer counterpart by 4.2 and 1.3 mAP on COCO, respectively. For semantic segmentation with UPerNet, the SMT base test at single- and multi-scale surpasses Swin by 2.0 and 1.1 mIoU respectively on the ADE20K. Our code is available at https://github.com/AFeng-x/SMT. Weifeng Lin, Jun Huang 0007 |
ICCV | 1 |
| 2021 | Constrained inverse minimum flow problems under the weighted Hamming distance
Weifeng Lin, Longcheng Liu, Anzhen Peng |
Theor. Comput. Sci. | 2 |
| 2016 | The immediate effects of Traditional Bone Setting in patients with acute low back painabstractBACKGROUND: Low back pain is common. For acute low back pain, most clinical practice guidelines agree on the use of reassurance, recommendations to stay active, brief education, paracetamol, non-steroidal anti-inflammatory drugs, spinal manipulation therapy, muscle relaxants, and weak opioids. Traditional bone setting (Chinese manipulation TBS) has been applied in cases of low back pain for years, and showed good treatment effect. However, there is lack of evidence-based research to show the immediate treatment effects of TBS on the acute low back pain patients. OBJECTIVE: To compare the treatment effect between Chinese bone-setting and non-steroidal anti-inflammatory drugs (paracetamol ) for the management of patient with acute low back pain. METHODS: Patients with acute low back pain were randomly divided into 2 groups. Patients in the experimental group received the TBS. Patients in the control group received paracetamol. The pain intensity and positive rate of straight leg raising were assessed. All the patients were evaluated two times: before and after treatment 3 hours later (i.e. immediate effect). RESULT AND CONCLUSION: The results suggest that both groups have the immediate effect of decreasing pain and the experimental group obtained a greater improvement than control group in the positive rate of straight leg raising. Weifeng Lin, Jiayou Zhao, Qiang Tian, Kiulam Chung, Zhiyong Fan, Shuhua Lai, Rusong Guo |
BIBM | 1 |