VLDB 2026 Research / reviewers in the wild / expert
Munan Ning
dblp:214/9635
· DBLP profile ↗
26ranked-venue papers
5as first author
21since 2021 · last 2026
0009-0005-3418-085XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-Based Large Language ModelsabstractVideo-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving artificial general intelligence, a truly intelligent Video-LLM model should not only see and understand the surroundings, but also possess human-level commonsense, and make well-informed decisions for users. To guide the development of such a model, the establishment of a robust and comprehensive evaluation system becomes crucial. To this end, this paper proposes Video-Bench, a new comprehensive benchmark along with a toolkit specifically designed for evaluating Video-LLMs. The benchmark comprises 10 meticulously crafted tasks, evaluating the capabilities of Video-LLMs across three distinct levels: video-exclusive understanding, prior knowledge-based question-answering, and comprehension and decision-making. In addition, we introduce an automatic toolkit tailored to process model outputs for various tasks, facilitating the calculation of metrics and conveniently generating final scores. We evaluate 9 representative Video-LLMs using Video-Bench. The findings reveal that current Video-LLMs still fall considerably short of achieving human-like comprehension and analysis of real-world video, and offer valuable insights for future research directions. The benchmark and toolkit are available at https://github.com/PKU-YuanGroup/Video-Bench. Munan Ning, Yujia Xie, Bin Lin 0014, Jiaxi Cui, Lu Yuan 0001, Dongdong Chen 0001, Li Yuan 0007 |
Comput. Vis. Media | 1 |
| 2026 | MoE-LLaVA: Mixture of Experts for Large Vision-Language ModelsabstractRecently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA. Bin Lin 0014, Zhenyu Tang 0004, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin 0001, Munan Ning, Jiebo Luo 0001, Li Yuan 0007 |
IEEE Trans. Multim. | 8 |
| 2025 | UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationabstractMultimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation methods face limitations due to the significant human workload required to design Q&A pairs for visual images, which inherently restricts the scale and scope of evaluations. Although automated MLLM-as-judge approaches attempt to reduce the human workload through automatic evaluations, they often introduce biases. To address these problems, we propose an Unsupervised Peer review MLLM Evaluation framework. It utilizes only image data, allowing models to automatically generate questions and conduct peer review assessments of answers from other models, effectively alleviating the reliance on human workload. Additionally, we introduce the vision-language scoring system to mitigate the bias issues, which focuses on three aspects: (i) response correctness; (ii) visual understanding and reasoning; and (iii) image-text correlation. Experimental results demonstrate that UPME achieves a Pearson correlation of 0.944 with human evaluations on the MMstar dataset and 0.814 on the ScienceQA dataset, indicating that our framework closely aligns with human-designed benchmarks and inherent human preferences. Qihui Zhang, Munan Ning, Zheyuan Liu 0012, Yue Huang 0001, Yanbo Wang 0005, Jiayi Ye, Yibing Song, Li Yuan 0007 |
CVPR | 2 |
| 2025 | OpenDUN: To Discover Unknown Number of Visual CategoriesabstractOpen-Set methods have relaxed the underlying assumption made by most image recognition studies that all samples in the test and training datasets belong to the same classes by considering only a part of classes are known in training dataset. However, most of these approaches require a known or predefined number of novel classes, which is often not the case in real applications. In this study, we aim at a more difficult but practical scenario, where the number of novel classes is unknown. By merging the unlabeled samples into clusters instead of directly assigning categorical labels to them, the proposed end-to-end framework can simultaneously estimate the number of novel classes and learn the appropriate division of unlabeled samples. In addition, a cluster scattering strategy is introduced such that the erroneous merging can be alleviated. Comprehensive experiments on three benchmark datasets are conducted to demonstrate the superiority of the proposed method in both the estimation of the novel class number and the classification of unlabeled samples. Sik Chit Wu, Munan Ning, Dong Wei 0004, Yefeng Zheng 0001, Donghuan Lu, Li Yuan 0007 |
ICME | 2 |
| 2025 | CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepabstractCurrent text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes.
Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled from layout planning, making it difficult to refine the layout during synthesis.
We present CoT-Diff, a framework that brings step-by-step CoT-style reasoning into T2I generation by tightly integrating Multimodal Large Language Model (MLLM)-driven 3D layout planning with the diffusion process.
CoT-Diff enables layout-aware reasoning inline within a single diffusion round: at each denoising step, the MLLM evaluates intermediate predictions, dynamically updates the 3D scene layout, and continuously guides the generation process.
The updated layout is converted into semantic conditions and depth maps, which are fused into the diffusion model via a condition-aware attention mechanism, enabling precise spatial control and semantic injection.
Experiments on 3D Scene benchmarks show that CoT-Diff significantly improves spatial alignment and compositional fidelity, and outperforms the state-of-the-art method by 34.7% in complex scene spatial accuracy, thereby validating the effectiveness of this entangled generation paradigm. Zheyuan Liu 0012, Munan Ning, Qihui Zhang, Yiwei Yang 0007, Yibing Song, Fan Wang 0019, Li Yuan 0007 |
NeurIPS | 2 |
| 2024 | Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting
Junwu Zhang, Zhenyu Tang 0004, Yatian Pang, Xinhua Cheng, Peng Jin 0001, Yida Wei, Munan Ning, Li Yuan 0007 |
ECCV (25) | 8 |
| 2024 | Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionabstractLarge Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding.Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models.However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers.In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other.As a result, Video-LLaVA outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively.Additionally, our Video-LLaVA also achieves superior performances on a broad range of 9 image benchmarks.Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos.We aim for this work to provide modest insights into the multi-modal inputs for the LLM. Bin Lin 0014, Jiaxi Cui, Munan Ning, Peng Jin 0001, Li Yuan 0007 |
EMNLP | 5 |
| 2024 | Temporal Contrastive Learning for Spiking Neural Networks
Haonan Qiu, Zeyin Song, Yanqi Chen, Munan Ning, Wei Fang 0006, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
ICANN (10) | 4 |
| 2024 | LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentabstractThe video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N ≥ 3) beyond vision and language. We thus propose LanguageBind, taking the language as the bind across different modalities because the language modality is well-explored and contains rich semantics. Specifically, we freeze the language encoder acquired by VL pretraining and then train encoders for other modalities with contrastive learning. As a result, all modalities are mapped to a shared feature space, implementing multi-modal semantic alignment. While LanguageBind ensures that we can extend VL modalities to N modalities, we also need a high-quality dataset with alignment data pairs centered on language. We thus propose VIDAL-10M with 10 Million data with Video, Infrared, Depth, Audio and their corresponding Language. In our VIDAL-10M, all videos are from short video platforms with complete semantics rather than truncated segments from long videos, and all the video, depth, infrared, and audio modalities are aligned to their textual descriptions. LanguageBind has achieved superior performance on a wide range of 15 benchmarks covering video, audio, depth, and infrared. Moreover, multiple experiments have provided evidence for the effectiveness of LanguageBind in achieving indirect alignment and complementarity among diverse modalities. Bin Lin 0014, Munan Ning, Jiaxi Cui, Hongfa Wang, Yatian Pang, Junwu Zhang, Caiwan Zhang, Zhifeng Li 0001, Wei Liu 0005, Li Yuan 0007 |
ICLR | 3 |
| 2024 | Self-architectural knowledge distillation for spiking neural networks
Haonan Qiu, Munan Ning, Zeyin Song, Wei Fang 0006, Yanqi Chen, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
Neural Networks | 2 |
| 2024 | An Organ-Aware Diagnosis Framework for Radiology Report GenerationabstractRadiology report generation (RRG) is crucial to save the valuable time of radiologists in drafting the report, therefore increasing their work efficiency. Compared to typical methods that directly transfer image captioning technologies to RRG, our approach incorporates organ-wise priors into the report generation. Specifically, in this paper, we propose Organ-aware Diagnosis (OaD) to generate diagnostic reports containing descriptions of each physiological organ. During training, we first develop a task distillation (TD) module to extract organ-level descriptions from reports. We then introduce an organ-aware report generation module that, for one thing, provides a specific description for each organ, and for another, simulates clinical situations to provide short descriptions for normal cases. Furthermore, we design an auto-balance mask loss to ensure balanced training for normal/abnormal descriptions and various organs simultaneously. Being intuitively reasonable and practically simple, our OaD outperforms SOTA alternatives by large margins on commonly used IU-Xray and MIMIC-CXR datasets, as evidenced by a 3.4% BLEU-1 improvement on MIMIC-CXR and 2.0% BLEU-2 improvement on IU-Xray. Pengchong Qiao, Lin Wang 0026, Munan Ning, Li Yuan 0007, Yefeng Zheng 0001, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Unsupervised Domain Adaptation for Medical Image Segmentation by Disentanglement Learning and Self-TrainingabstractUnsupervised domain adaption (UDA), which aims to enhance the segmentation performance of deep models on unlabeled data, has recently drawn much attention. In this paper, we propose a novel UDA method (namely DLaST) for medical image segmentation via disentanglement learning and self-training. Disentanglement learning factorizes an image into domain-invariant anatomy and domain-specific modality components. To make the best of disentanglement learning, we propose a novel shape constraint to boost the adaptation performance. The self-training strategy further adaptively improves the segmentation performance of the model for the target domain through adversarial learning and pseudo label, which implicitly facilitates feature alignment in the anatomy space. Experimental results demonstrate that the proposed method outperforms the state-of-the-art UDA methods for medical image segmentation on three public datasets, i.e., a cardiac dataset, an abdominal dataset and a brain dataset. The code will be released soon. Qingsong Xie, Yuexiang Li, Nanjun He, Munan Ning, Kai Ma 0002, Guoxing Wang, Yong Lian 0001, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | A Model-Agnostic Framework for Universal Anomaly Detection of Multi-organ and Multi-modal Images
Donghuan Lu, Munan Ning, Liansheng Wang 0002, Dong Wei 0004, Yefeng Zheng 0001 |
MICCAI (3) | 3 |
| 2023 | MIL-ViT: A multiple instance vision transformer for fundus image classification
Qi Bi, Xu Sun 0006, Kai Ma 0002, Cheng Bian, Munan Ning, Nanjun He, Yawen Huang, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation SegmentationabstractUnsupervised domain adaption has been widely adopted in tasks with scarce annotated data. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data, leading to inferior performance. To address this issue, we first propose to introduce active sample selection to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, both source and target domains can be better characterized as multimodal distributions, in which way more complementary and informative samples are selected from the target domain. With only a little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, achieving a large performance gain. In addition, a powerful semi-supervised domain adaptation strategy is proposed to alleviate the long-tail distribution problem and further improve the segmentation performance. Extensive experiments are conducted on public datasets, and the results demonstrate that the proposed approach outperforms state-of-the-art methods by large margins and achieves similar performance to the fully-supervised upperbound, i.e., 71.4% mIoU on GTA5 and 71.8% mIoU on SYNTHIA. The effectiveness of each component is also verified by thorough ablation studies. Munan Ning, Donghuan Lu, Yujia Xie, Dongdong Chen 0001, Dong Wei 0004, Yefeng Zheng 0001, Yonghong Tian 0001, Shuicheng Yan, Li Yuan 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation
Xinyu Shi 0003, Dong Wei 0004, Yu Zhang 0185, Donghuan Lu, Munan Ning, Jiashun Chen, Kai Ma 0002, Yefeng Zheng 0001 |
ECCV (20) | 5 |
| 2022 | Deformer: Towards Displacement Field Learning for Unsupervised Medical Image Registration
Jiashun Chen, Donghuan Lu, Yu Zhang 0185, Dong Wei 0004, Munan Ning, Xinyu Shi 0003, Zhe Xu 0012, Yefeng Zheng 0001 |
MICCAI (6) | 5 |
| 2022 | An Inclusive Task-Aware Framework for Radiology Report Generation
Lin Wang 0026, Munan Ning, Donghuan Lu, Dong Wei 0004, Yefeng Zheng 0001, Jie Chen 0001 |
MICCAI (8) | 2 |
| 2022 | Multiscale Unsupervised Retinal Edema Area Segmentation in OCT Images
Wenguang Yuan, Donghuan Lu, Dong Wei 0004, Munan Ning, Yefeng Zheng 0001 |
MICCAI (2) | 4 |
| 2021 | Multi-Anchor Active Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaption has proven to be an effective approach for alleviating the intensive workload of manual annotation by aligning the synthetic source-domain data and the real-world target-domain samples. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data. To this end, we firstly propose to introduce a novel multi-anchor based active learning strategy to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, the source domain can be better characterized as a multimodal distribution, thus more representative and complimentary samples are selected from the target domain. With little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, resulting in a large performance gain. The multi-anchor strategy is additionally employed to model the target-distribution. By regularizing the latent representation of the target samples compact around multiple anchors through a novel soft alignment loss, more precise segmentation can be achieved. Extensive experiments are conducted on public datasets to demonstrate that the proposed approach outperforms state-of-the-art methods significantly, along with thorough ablation study to verify the effectiveness of each component. The code will be released soon at https://github.com/munanning/MADA. Munan Ning, Donghuan Lu, Dong Wei 0004, Cheng Bian, Chenglang Yuan, Kai Ma 0002, Yefeng Zheng 0001 |
ICCV | 1 |
| 2021 | MIL-VT: Multiple Instance Learning Enhanced Vision Transformer for Fundus Image Classification
Kai Ma 0002, Qi Bi, Cheng Bian, Munan Ning, Nanjun He, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001 |
MICCAI (8) | 5 |
| 2020 | Hierarchical Clustering With Hard-Batch Triplet Loss for Person Re-IdentificationabstractFor clustering-guided fully unsupervised person reidentification (re-ID) methods, the quality of pseudo labels generated by clustering directly decides the model performance. In order to improve the quality of pseudo labels in existing methods, we propose the HCT method which combines hierarchical clustering with hard-batch triplet loss. The key idea of HCT is to make full use of the similarity among samples in the target dataset through hierarchical clustering, reduce the influence of hard examples through hard-batch triplet loss, so as to generate high quality pseudo labels and improve model performance. Specifically, (1) we use hierarchical clustering to generate pseudo labels, (2) we use PK sampling in each iteration to generate a new dataset for training, (3) we conduct training with hard-batch triplet loss and evaluate model performance in each iteration. We evaluate our model on Market-1501 and DukeMTMC-reID. Results show that HCT achieves 56.4% mAP on Market-1501 and 50.7% mAP on DukeMTMC-reID which surpasses state-of-the-arts a lot in fully unsupervised re-ID and even better than most unsupervised domain adaptation (UDA) methods which use the labeled source dataset. Code will be released soon on https://github.com/zengkaiwei/HCT Kaiwei Zeng, Munan Ning, Yang Guo 0003 |
CVPR | 2 |
| 2020 | A Macro-Micro Weakly-Supervised Framework for AS-OCT Tissue Segmentation
Munan Ning, Cheng Bian, Donghuan Lu, Chenglang Yuan, Yang Guo 0003, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 1 |
| 2020 | Energy clustering for unsupervised person re-identification
Kaiwei Zeng, Munan Ning, Yang Guo 0003 |
Image Vis. Comput. | 2 |
| 2020 | Deviation based clustering for unsupervised person re-identification
Munan Ning, Kaiwei Zeng, Yang Guo 0003 |
Pattern Recognit. Lett. | 1 |
| 2017 | A joint multi-scale convolutional network for fully automatic segmentation of the left ventricleabstractLeft ventricle (LV) segmentation is crucial for quantitative analysis of the cardiac contractile function. In this paper, we propose a joint multi-scale convolutional neural network to fully automatically segment the LV. Our method adopts two kinds of multi-scale features of cardiac magnetic resonance (CMR) images, including multi-scale features directly extracted from CMR images with different scales and multi-scale features constructed by intermediate layers of standard CNN architecture. We take advantage of these two strategies and fuse their prediction results to produce more accurate segmentation results. Qualitative results demonstrate the effectiveness and robustness of our method, and quantitative evaluation indicates our method achieves LV segmentation with higher accuracy than state-of-the-art approaches. Qianqian Tong 0001, Xiangyun Liao, Mianlun Zheng, Weixu Zhu, Guian Zhang, Munan Ning |
ICIP | 7 |