Mengcheng Lan

dblp:250/5850 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-3311-0295ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains a significant challenge.In this work, we propose a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. We first introduce image-wise semantic descriptors, a patch-aligned textual representation of segmentation masks that integrates naturally into the language modeling pipeline. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptorsby 74% and accelerating inference by $3\times$3×, without compromising performance. Building upon this, our initial framework Text4Segachieves strong segmentation performance across a wide range of vision tasks. To further improve granularity and compactness, we propose box-wise semantic descriptors, which localizes regions of interest using bounding boxes and represents region masks via structured mask tokens called semantic bricks. This leads to our refined model, Text4Seg++, which formulates segmentation as a next-brick prediction task, combining precision, scalability, and generative efficiency. Comprehensive experiments on natural and remote sensing datasets show that Text4Seg++consistently outperforms state-of-the-art models across diverse benchmarks without any task-specific fine-tuning, while remaining compatible with existing MLLM backbones. Our work highlights the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li 0001, Yiping Ke, Xudong Jiang 0001, Yingchen Yu, Yunqing Zhao, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Multi-Atlas Brain Network Classification Through Consistency Distillation and Complementary Information Fusion
abstract
Brain network analysis plays a crucial role in identifying distinctive patterns associated with neurological disorders. Functional magnetic resonance imaging (fMRI) enables the construction of brain networks by analyzing correlations in blood-oxygen-level-dependent (BOLD) signals across different brain regions, known as regions of interest (ROIs). These networks are typically constructed using atlases that parcellate the brain based on various hypotheses of functional and anatomical divisions. However, there is no standard atlas for brain network classification, leading to limitations in detecting abnormalities in disorders. Recent methods leveraging multiple atlases fail to ensure consistency across atlases and lack effective ROI-level information exchange, limiting their efficacy. To address these challenges, we propose the Atlas-Integrated Distillation and Fusion network (AIDFusion), a novel framework designed to enhance brain network classification using fMRI data. AIDFusion introduces a disentangle Transformer to filter out inconsistent atlas-specific information and distill meaningful cross-atlas connections. Additionally, it enforces subject- and population-level consistency constraints to improve cross-atlas coherence. To further enhance feature integration, AIDFusion incorporates an inter-atlas message-passing mechanism that facilitates the fusion of complementary information across brain regions. We evaluate AIDFusion on four resting-state fMRI datasets encompassing different neurological disorders. Experimental results demonstrate its superior classification performance and computational efficiency compared to state-of-the-art methods. Furthermore, a case study highlights AIDFusion's ability to extract interpretable patterns that align with established neuroscience findings, reinforcing its potential as a robust tool for multi-atlas brain network analysis.
Jiaxing Xu, Mengcheng Lan, Xia Dong, Kai He 0001, Wayne Zhang 0001, Qingtian Bian, Yiping Ke
IEEE J. Biomed. Health Informatics2
2026 BrainPrompt+: Multi-Level Brain Prompt Learning for Knowledge-Guided Neurological Disorder Identification
abstract
Accurate identification of neurological disorders such as Alzheimer's disease (AD), Parkinson's disease (PD), and Autism Spectrum Disorder (ASD) is challenging due to subtle early-stage symptoms and heterogeneous brain dynamics. Resting-state functional MRI (rs-fMRI) enables the construction of functional brain networks, where Graph Neural Networks (GNNs) have shown promise for disease classification. However, existing GNN-based methods face three key limitations: correlation-based graph construction introduces noise and negative edges; domain knowledge about brain regions is ignored; and demographic or clinical metadata are fused through simplistic encodings. To overcome these limitations, we propose BrainPrompt+, a knowledge-guided framework that integrates Large Language Models (LLMs) with multi-level natural language prompts. Five types of prompts are introduced: spectral (frequency-domain BOLD features), spatial (inter-ROI connectivity), ROI (anatomical and functional knowledge), disease (progression stages), and subject (demographic context). These prompts are encoded by a frozen LLM and incorporated into a GNN pipeline, unifying imaging, clinical, and external knowledge in a semantically enriched and interpretable manner. Experiments on three rs-fMRI datasets show that BrainPrompt+ consistently outperforms state-of-the-art baselines, achieving accuracy gains of up to 8.93%. Biomarker analysis further demonstrates that the highlighted ROIs align with established neuroscience findings, confirming the interpretability of the model. BrainPrompt+ thus establishes a flexible and generalizable paradigm for knowledge-guided brain network analysis. The source code is available at https://github.com/AngusMonroe/BrainPromptPlus.
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Yue Xun, Qika Lin, Peifan Ran, Yiping Ke, Mengling Feng
IEEE Trans. Medical Imaging5
2025 Text4Seg: Reimagining Image Segmentation as Text Generation
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74\% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Yue Zhou 0005, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ICLR1
2025 BrainOOD: Out-of-distribution Generalizable Brain Network Analysis
abstract
In neuroscience, identifying distinct patterns linked to neurological disorders, such as Alzheimer's and Autism, is critical for early diagnosis and effective intervention. Graph Neural Networks (GNNs) have shown promising in analyzing brain networks, but there are two major challenges in using GNNs: (1) distribution shifts in multi-site brain network data, leading to poor Out-of-Distribution (OOD) generalization, and (2) limited interpretability in identifying key brain regions critical to neurological disorders. Existing graph OOD methods, while effective in other domains, struggle with the unique characteristics of brain networks. To bridge these gaps, we introduce BrainOOD, a novel framework tailored for brain networks that enhances GNNs' OOD generalization and interpretability. BrainOOD framework consists of a feature selector and a structure extractor, which incorporates various auxiliary losses including an improved Graph Information Bottleneck (GIB) objective to recover causal subgraphs. By aligning structure selection across brain networks and filtering noisy features, BrainOOD offers reliable interpretations of critical brain regions. Our approach outperforms 16 existing methods and improves generalization to OOD subjects by up to 8.5%. Case studies highlight the scientific validity of the patterns extracted, which aligns with the findings in known neuroscience literature. We also propose the first OOD brain network benchmark, which provides a foundation for future research in this field. Our code is available at https://github.com/AngusMonroe/BrainOOD.
Jiaxing Xu, Yongqiang Chen 0002, Xia Dong, Mengcheng Lan, Qingtian Bian, James Cheng, Yiping Ke
ICLR4
2025 BrainPrompt: Multi-level Brain Prompt Enhancement for Neurological Condition Identification
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Xia Dong, Yiping Ke, Mengling Feng
MICCAI (12)5
2024 Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification
abstract
Understanding neurological disorder is a fundamental problem in neuroscience, which often requires the analysis of brain networks derived from functional magnetic resonance imaging (fMRI) data. Despite the prevalence of Graph Neural Networks (GNNs) and Graph Transformers in various domains, applying them to brain networks faces challenges. Specifically, the datasets are severely impacted by the noises caused by distribution shifts across sub- populations and the neglect of node identities, both obstruct the identification of disease-specific patterns. To tackle these challenges, we propose Contrasformer, a novel contrastive brain network Transformer. It generates a prior-knowledge-enhanced contrast graph to address the distribution shifts across sub-populations by a two-stream attention mechanism. A cross attention with identity embedding highlights the identity of nodes, and three auxiliary losses ensure group consistency. Evaluated on 4 functional brain network datasets over 4 different diseases, Contrasformer outperforms the state-of-the-art methods for brain networks by achieving up to 10.8% improvement in accuracy, which demonstrates its efficacy in neurological disorder identification. Case studies illustrate its interpretability, especially in the context of neuroscience. This paper provides a solution for analyzing brain networks, offering valuable insights into neurological disorders. Our code is available at https://github.com/AngusMonroe/Contrasformer.
Jiaxing Xu, Kai He 0001, Mengcheng Lan, Qingtian Bian, Wei Li 0231, Tieying Li, Yiping Ke, Miao Qiao
CIKM3
2024 ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ECCV (47)1
2024 ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ECCV (68)1
2024 Learning to Discover Knowledge: A Weakly-Supervised Partial Domain Adaptation Approach
abstract
Domain adaptation has shown appealing performance by leveraging knowledge from a source domain with rich annotations. However, for a specific target task, it is cumbersome to collect related and high-quality source domains. In real-world scenarios, large-scale datasets corrupted with noisy labels are easy to collect, stimulating a great demand for automatic recognition in a generalized setting, i.e., weakly-supervised partial domain adaptation (WS-PDA), which transfers a classifier from a large source domain with noises in labels to a small unlabeled target domain. As such, the key issues of WS-PDA are: 1) how to sufficiently discover the knowledge from the noisy labeled source domain and the unlabeled target domain, and 2) how to successfully adapt the knowledge across domains. In this paper, we propose a simple yet effective domain adaptation approach, termed as self-paced transfer classifier learning (SP-TCL), to address the above issues, which could be regarded as a well-performing baseline for several generalized domain adaptation tasks. The proposed model is established upon the self-paced learning scheme, seeking a preferable classifier for the target domain. Specifically, SP-TCL learns to discover faithful knowledge via a carefully designed prudent loss function and simultaneously adapts the learned knowledge to the target domain by iteratively excluding source examples from training under the self-paced fashion. Extensive evaluations on several benchmark datasets demonstrate that SP-TCL significantly outperforms state-of-the-art approaches on several generalized domain adaptation tasks. Code is available at https://github.com/mc-lan/SP-TCL.
Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu
IEEE Trans. Image Process.1
2023 MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video Prediction
abstract
The mainstream of the existing approaches for video prediction builds up their models based on a Single-In-Single-Out (SISO) architecture, which takes the current frame as input to predict the next frame in a recursive manner. This way often leads to severe performance degradation when they try to extrapolate a longer period of future, thus limiting the practical use of the prediction model. Alternatively, a Multi-In-Multi-Out (MIMO) architecture that outputs all the future frames at one shot naturally breaks the recursive manner and therefore prevents error accumulation. However, only a few MIMO models for video prediction are proposed and they only achieve inferior performance due to the date. The real strength of the MIMO model in this area is not well noticed and is largely under-explored. Motivated by that, we conduct a comprehensive investigation in this paper to thoroughly exploit how far a simple MIMO architecture can go. Surprisingly, our empirical studies reveal that a simple MIMO model can outperform the state-of-the-art work with a large margin much more than expected, especially in dealing with long-term error accumulation. After exploring a number of ways and designs, we propose a new MIMO architecture based on extending the pure Transformer with local spatio-temporal blocks and a new multi-output decoder, namely MIMO-VP, to establish a new standard in video prediction. We evaluate our model in four highly competitive benchmarks. Extensive experiments show that our model wins 1st place on all the benchmarks with remarkable performance gains and surpasses the best SISO model in all aspects including efficiency, quantity, and quality. A dramatic error reduction is achieved when predicting 10 frames on Moving MNIST and Weather datasets respectively. We believe our model can serve as a new baseline to facilitate the future research of video prediction tasks. The code will be released.
Shuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen, Xunlai Chen, Xiaoguang Han 0001, Shuguang Cui
AAAI2
2023 SmooSeg: Smoothness Prior for Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation is a challenging task that segments images into semantic groups without manual annotation. Prior works have primarily focused on leveraging prior knowledge of semantic consistency or priori concepts from self-supervised learning methods, which often overlook the coherence property of image segments. In this paper, we demonstrate that the smoothness prior, asserting that close features in a metric space share the same semantics, can significantly simplify segmentation by casting unsupervised semantic segmentation as an energy minimization problem. Under this paradigm, we propose a novel approach called SmooSeg that harnesses self-supervised learning methods to model the closeness relationships among observations as smoothness signals. To effectively discover coherent semantic segments, we introduce a novel smoothness loss that promotes piecewise smoothness within segments while preserving discontinuities across different segments. Additionally, to further enhance segmentation quality, we design an asymmetric teacher-student style predictor that generates smoothly updated pseudo labels, facilitating an optimal fit between observations and labeling outputs. Thanks to the rich supervision cues of the smoothness prior, our SmooSeg significantly outperforms STEGO in terms of pixel accuracy on three datasets: COCOStuff (+14.9\%), Cityscapes (+13.0\%), and Potsdam-3 (+5.7\%).
Mengcheng Lan, Xinjiang Wang, Yiping Ke, Jiaxing Xu, Litong Feng, Wayne Zhang 0001
NeurIPS1
2023 Dual-Level Adaptive and Discriminative Knowledge Transfer for Cross-Domain Recognition
abstract
Unsupervised domain adaptation is an appealing technique to learn robust classifiers for unlabeled target domain by borrowing knowledge from well-established source domain. However, previous works mainly suffer from two limitations: 1) the classifier trained on labeled source data may be prone to overfitting the source distribution, lowering its performance on the target domain; 2) the adaptation process will be misled by conditional distribution matching using hard pseudo labels of target samples. This paper presents a Dual-Level Adaptive and Discriminative (DLAD) classifier learning framework, in which transfer classifier and distribution adaptation can be mutually beneficial for effective knowledge transfer. Specifically, we aim to achieve a domain-level adaptive classifier by considering structural risk minimization (SRM) on both domains and performing weighted distribution adaptation, which facilitates joint classifier learning in a semi-supervised manner. To further achieve a class-level discriminative classifier, we explicitly leverage unlabeled target data to promote classifier learning based on class probabilities, which refines the decision boundary to be more discriminative for unlabeled target data. To the best of our knowledge, DLAD is the first attempt to consider the principle of SRM on the target domain, which significantly boosts the discriminative power of transfer classifier and yields a tighter generalization bound. Experimental evaluations on several standard cross-domain datasets show that DLAD significantly outperforms other competitive methods.
Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Ligang Liu 0001
IEEE Trans. Multim.2
2022 Group Correspondence: A Statistical Perspective for Incomplete Multi-View Clustering Augmentation
abstract
Cross-view consistency is the fundamental property of multiview clustering. However, in incomplete multi-view scenarios, existing methods can only pursue consistency through the paired data while ignoring the information in unpaired data. In this paper, we show a new insight from the data pattern and provide a novel perspective to incorporate unpaired data for consistency maximization by mining group correspondence. We first formulate cross-view consistency in a statistical perspective to by-pass the strict demand of instance correspondence, and then propose a technique to construct corresponding groups across views to enhance the objective of consistency maximization. Our proposal can be used as a universal plug-in to augment existing approaches. We test the efficacy and generality of our proposal by adapting it to two base methods as augmentations and comparing the augmented models against the original ones and other baselines. Experiment results demonstrate the effectiveness of our proposal and validate the value of our insight.
Tianyou Liang, Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu
ICME3
2022 Generalized Multi-View Collaborative Subspace Clustering
abstract
In real-world applications, complete or incomplete multi-view data are common, which leads to the problem of generalized multi-view clustering. Recently, researchers attempt to learn the latent representation in the common subspace from heterogeneous data, which usually suffers from feature degeneration. Moreover, there are limited efforts on simultaneously revealing the underlying subspace structure and exploring the complementary information from incomplete multiple views. In this paper, we introduce a novel Generalized Multi-view Collaborative Subspace Clustering (GMCSC) framework to address the above issues, in which consensus subspace structure of all views and embedding subspaces for each view are jointly learned to benefit each other. Specifically, we develop a novel collaborative subspace learning strategy based on self-representation learning, which provides a brand-new way of pursuing the complete subspace structure directly from multi-view data. Furthermore, we explore complementary information by enforcing the consistency across different views and preserving the view-specific information of each view, which can alleviate the problem of feature degeneration and enhance the reasonability of using a consensus representation for multiple views. Experimental results on six benchmark datasets demonstrate that the proposed method can significantly outperform the state-of-the-art algorithms.
Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu
IEEE Trans. Circuits Syst. Video Technol.1
2022 Multiview Consensus Structure Discovery
abstract
Multiview subspace learning has attracted much attention due to the efficacy of exploring the information on multiview features. Most existing methods perform data reconstruction on the original feature space and thus are vulnerable to noisy data. In this article, we propose a novel multiview subspace learning method, called multiview consensus structure discovery (MvCSD). Specifically, we learn the low-dimensional subspaces corresponding to different views and simultaneously pursue the structure consensus over subspace clustering for multiple views. In such a way, latent subspaces from different views regularize each other toward a common consensus that reveals the underlying cluster structure. Compared to existing methods, MvCSD leverages the consensus structure derived from the subspaces of diverse views to better exploit the intrinsic complementary information that well reflects the essence of data. Accordingly, the proposed MvCSD is capable of producing a more robust and accurate representation structure which is crucial for multiview subspace learning. The proposed method can be optimized effectively, with theoretical convergence guarantee, by alternatively iterating the argument Lagrangian multiplier algorithm and the eigendecomposition. Extensive experiments on diverse datasets demonstrate the advantages of our method over the state-of-the-art methods.
Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu
IEEE Trans. Cybern.2
2021 Coupled Knowledge Transfer for Visual Data Recognition
abstract
Transfer learning aims to learn an effective classifier for unlabeled target data by borrowing knowledge from well-labeled source data. However, most existing work has emphasized on learning domain invariant features to reduce the distribution discrepancy, which may suffer from the negative transfer problem caused by structure inconsistencies or distribution outliers. To address this challenge, in this paper, we propose a novel transfer learning approach, which seamlessly integrates domain invariant feature learning, discriminative structure preservation and sample reweighting into a unified learning model. Specifically, we attempt to learn domain invariant features by jointly adapting the marginal and conditional distributions. To transfer discriminative knowledge inferred from data, we enforce the structure consistency between the original feature space and the latent feature space. Furthermore, to enhance the robustness of our model, an efficient and more generalized sample reweighting strategy is developed to assign target predictions with different levels of confidence. The key advantage over previous methods is that our model can adaptively select pivot samples in target domain and retain the properties of discriminative structures underlying data domains, which enables coupled knowledge transfer during the learning process. Experimental results on several benchmark datasets have verified the superiority of the proposed method over other state-of-the-art algorithms.
Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu
IEEE Trans. Circuits Syst. Video Technol.2
2020 Constrained Discriminative Projection Learning for Image Classification
abstract
Projection learning is widely used in extracting discriminative features for classification. Although numerous methods have already been proposed for this goal, they barely explore the label information during projection learning and fail to obtain satisfactory performance. Besides, many existing methods can learn only a limited number of projections for feature extraction which may degrade the performance in recognition. To address these problems, we propose a novel constrained discriminative projection learning (CDPL) method for image classification. Specifically, CDPL can be formulated as a joint optimization problem over subspace learning and classification. The proposed method incorporates the low-rank constraint to learn a robust subspace which can be used as a bridge to seamlessly connect the original visual features and objective outputs. A regression function is adopted to explicitly exploit the class label information so as to enhance the discriminability of subspace. Unlike existing methods, we use two matrices to perform feature learning and regression, respectively, such that the proposed approach can obtain more projections and achieve superior performance in classification tasks. The experiments on several datasets show clearly the advantages of our method against other state-of-the-art methods.
Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Dapeng Tao
IEEE Trans. Image Process.2