Lin Chen 0026

dblp:13/3479-26 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-5935-3877ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
abstract
Zhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma, Ziheng Wang, Yu Zeng, Zehui Chen, Lin Chen, Wenxuan Huang, Wei-Jie Xu, Yi Cao, Feng Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruiyan Han, XinYu Sun, Lin Chen 0026, Wenxuan Huang 0001, Wei-Jie Xu, Yi Cao 0006, Feng Zhao 0004
ACL (1)8
2025 Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generation
abstract
High-quality image captions are essential for improving modality alignment and visual understanding in Large Vision-Language Models (LVLMs).However, the scarcity of ultradetailed image caption data limits further advancements.This paper presents a systematic pipeline for generating high-quality, ultradetailed image captions, encompassing both pre-processing and post-processing stages.In the pre-processing stage, we classify and deduplicate images, extract visual information using expert tools, and leverage GPT-4o with structured prompts to generate initial captions.To enhance comprehensiveness, we introduce an expansion strategy based on Large Language Models (LLMs), defining eight descriptive dimensions to refine and extend captions, which serve as seed data for training a proprietary captioner model.In the post-processing stage, we incorporate human error-correction annotations and an active learning-inspired approach to refine low-quality samples.Using high-quality corrected data, we apply Direct Preference Optimization (DPO) and develop a critic-rewrite pipeline, training a sentence-level critic model to mitigate hallucinations.Experimental results demonstrate that our ultra-detailed captions significantly enhance LVLMs' perception and cognitive abilities across multiple vision-language benchmarks.The code and dataset are available at https://github.com/yuzeng0-0/UltraCaption.
Yukun Qi, Xikun Bao, Lin Chen 0026, Shiting Huang, Feng Zhao 0004
EMNLP5
2024 FreeDrag: Feature Dragging for Reliable Point-Based Image Editing
abstract
To serve the intricate and varied demands of image editing, precise and flexible manipulation in image content is indispensable. Recently, Drag-based editing methods have gained impressive performance. However, these methods predominantly center on point dragging, resulting in two noteworthy drawbacks, namely “miss tracking ”, where dif-ficulties arise in accurately tracking the predetermined han-dle points, and “ambiguous tracking”, where tracked points are potentially positioned in wrong regions that closely re-semble the handle points. To address the above issues, we propose FreeDrag, a feature dragging methodology designed to free the burden on point tracking. The Free-Drag incorporates two key designs, i.e., template feature via adaptive updating and line search with backtracking, the former improves the stability against drastic content change by elaborately controlling the feature updating scale after each dragging, while the latter alleviates the misguidance from similar points by actively restricting the search area in a line. These two technologies together contribute to a more stable semantic dragging with higher efficiency. Comprehensive experimental results substantiate that our approach significantly outperforms pre-existing methodologies, offering reliable point-based editing even in various complex scenarios.
Pengyang Ling, Lin Chen 0026, Pan Zhang 0001, Huaian Chen, Yi Jin 0002, Jinjin Zheng
CVPR2
2024 Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
abstract
In this paper, we first assess and harness various Vision Foundation Models (VFMs) in the context of Domain Generalized Semantic Segmentation (DGSS). Driven by the motivation that Leveraging Stronger pre-trained models and Fewer trainable parameters for Superior generalizability, we introduce a robust fine-tuning approach, namely “Rein”, to parameter-efficiently harness VFMs for DGSS. Built upon a set of trainable tokens, each linked to distinct instances, Rein precisely refines and forwards the feature maps from each layer to the next layer within the backbone. This process produces diverse refinements for different categories within a single image. With fewer trainable parameters, Rein efficiently fine-tunes VFMs for DGSS tasks, surprisingly surpassing full parameter fine-tuning. Extensive experiments across various settings demonstrate that Rein significantly outperforms state-of-the-art methods. Remarkably, with just an extra 1% of trainable parameters within the frozen backbone, Rein achieves a mIoU of 78.4% on the Cityscapes, without accessing any real urban-scene datasets. Code is available at https://github.com/w1oves/Rein.git.
Zhixiang Wei, Lin Chen 0026, Yi Jin 0002, Xiaoxiao Ma 0006, Pengyang Ling, Ben Wang 0005, Huaian Chen, Jinjin Zheng
CVPR2
2024 ShareGPT4V: Improving Large Multi-modal Models with Better Captions
Lin Chen 0026, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Conghui He, Jiaqi Wang 0003, Feng Zhao 0004, Dahua Lin
ECCV (17)1
2024 VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models
Haodong Duan, Junming Yang 0001, Yuxuan Qiao, Xinyu Fang, Lin Chen 0026, Yuan Liu 0025, Xiaoyi Dong, Yuhang Zang, Pan Zhang 0001, Jiaqi Wang 0003, Dahua Lin, Kai Chen 0026
ACM Multimedia5
2024 Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
abstract
Vision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is crucial for model refinement, despite the inherent difficulty due to the intertwined nature of seeing and reasoning in existing VLMs. To tackle this issue, we present Prism, an innovative framework designed to disentangle the perception and reasoning processes involved in visual question solving. Prism comprises two distinct stages: a perception stage that utilizes a VLM to extract and articulate visual information in textual form, and a reasoning stage that formulates responses based on the extracted visual information using a Large Language Model (LLM). This modular design enables the systematic comparison and assessment of both proprietary and open-source VLM for their perception and reasoning strengths. Our analytical framework provides several valuable insights, underscoring Prism's potential as a cost-effective solution for vision-language tasks. By combining a streamlined VLM focused on perception with a powerful LLM tailored for reasoning, Prism achieves superior results in general vision-language tasks while substantially cutting down on training and operational expenses. Quantitative evaluations show that Prism, when configured with a vanilla 2B LLaVA and freely accessible GPT-3.5, delivers performance on par with VLMs $10 \times$ larger on the rigorous multimodal benchmark MMStar.
Yuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang 0001, Lin Chen 0026, Songyang Zhang 0001, Jiaqi Wang 0003, Dahua Lin, Kai Chen 0026
NeurIPS5
2023 Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement
abstract
Most prior semantic segmentation methods have been developed for day-time scenes, while typically underperforming in night-time scenes due to insufficient and complicated lighting conditions. In this work, we tackle this challenge by proposing a novel night-time semantic segmentation paradigm, i.e., disentangle then parse (DTP). DTP explicitly disentangles night-time images into light-invariant reflectance and light-specific illumination components and then recognizes semantics based on their adaptive fusion. Concretely, the proposed DTP comprises two key components: 1) Instead of processing lighting-entangled features as in prior works, our Semantic-Oriented Disentanglement (SOD) framework enables the extraction of reflectance component without being impeded by lighting, allowing the network to consistently recognize the semantics under cover of varying and complicated lighting conditions. 2) Based on the observation that the illumination component can serve as a cue for some semantically confused regions, we further introduce an Illumination-Aware Parser (IAParser) to explicitly learn the correlation between semantics and lighting, and aggregate the illumination features to yield more precise predictions. Extensive experiments on the night-time segmentation task with various settings demonstrate that DTP significantly outperforms state-of-the-art methods. Furthermore, with negligible additional parameters, DTP can be directly used to benefit existing day-time methods for night-time segmentation. Code and dataset are available at https://github.com/w1oves/DTP.git.
Zhixiang Wei, Lin Chen 0026, Tao Tu 0006, Pengyang Ling, Huaian Chen, Yi Jin 0002
ICCV2
2022 Reusing the Task-specific Classifier as a Discriminator: Discriminator-free Adversarial Domain Adaptation
abstract
Adversarial learning has achieved remarkable performances for unsupervised domain adaptation (UDA). Existing adversarial UDA methods typically adopt an additional discriminator to play the min-max game with a feature extractor. However, most of these methods failed to effectively leverage the predicted discriminative information, and thus cause mode collapse for generator. In this work, we address this problem from a different perspective and design a simple yet effective adversarial paradigm in the form of a discriminator-free adversarial learning network (DALN), wherein the category classifier is reused as a discriminator, which achieves explicit domain alignment and category distinguishment through a unified objective, enabling the DALN to leverage the predicted discriminative information for sufficient feature alignment. Basically, we introduce a Nuclear-norm Wasserstein discrepancy (NWD) that has definite guidance meaning for performing discrimination. Such NWD can be coupled with the classifier to serve as a discriminator satisfying the K-Lipschitz constraint without the requirements of additional weight clipping or gradient penalty strategy. Without bells and whistles, DALN compares favorably against the existing state-of-the-art (SOTA) methods on a variety of public datasets. Moreover, as a plug-and-play technique, NWD can be directly used as a generic regularizer to benefit existing UDA algorithms. Code is available at https://github.com/xiaoachen98/DALN.
Lin Chen 0026, Huaian Chen, Zhixiang Wei, Xin Jin 0014, Xiao Tan 0004, Yi Jin 0002, Enhong Chen
CVPR1
2022 Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation
abstract
In unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bridging (DB). However, for dense prediction tasks such as domain adaptive semantic segmentation (DASS), existing solutions have mostly relied on rough style transfer and how to elegantly bridge domains is still under-explored. In this work, we resort to data mixing to establish a deliberated domain bridging (DDB) for DASS, through which the joint distributions of source and target domains are aligned and interacted with each in the intermediate space. At the heart of DDB lies a dual-path domain bridging step for generating two intermediate domains using the coarse-wise and the fine-wise data mixing techniques, alongside a cross-path knowledge distillation step for taking two complementary models trained on generated intermediate samples as ‘teachers’ to develop a superior ‘student’ in a multi-teacher distillation manner. These two optimization steps work in an alternating way and reinforce each other to give rise to DDB with strong adaptation power. Extensive experiments on adaptive segmentation tasks with different settings demonstrate that our DDB significantly outperforms state-of-the-art methods.
Lin Chen 0026, Zhixiang Wei, Xin Jin 0014, Huaian Chen, Miao Zheng, Kai Chen 0026, Yi Jin 0002
NeurIPS1