EDBT 2026 Demo / reviewers in the wild / expert
Minheng Ni
dblp:263/9969
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0001-6483-0650ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FineRAG: Fine-grained Retrieval-Augmented Text-to-Image GenerationabstractRecent advancements in text-to-image generation, notably the series of Stable Diffusion methods, have enabled the production of diverse, high-quality photo-realistic images. Nevertheless, these techniques still exhibit limitations in terms of knowledge access. Retrieval-augmented image generation is a straightforward way to tackle this problem. Current studies primarily utilize coarse-grained retrievers, employing initial prompts as search queries for knowledge retrieval. This approach, however, is ineffective in accessing valuable knowledge in long-tail text-to-image generation scenarios. To alleviate this problem, we introduce FineRAG, a fine-grained model that systematically breaks down the retrieval-augmented image generation task into four critical stages: query decomposition, candidate selection, retrieval-augmented diffusion, and self-reflection. Experimental results on both general and long-tailed benchmarks show that our proposed method significantly reduces the noise associated with retrieval-augmented image generation and performs better in complex, open-world scenarios. Huaying Yuan, Ziliang Zhao 0001, Shuting Wang 0002, Shitao Xiao, Minheng Ni, Zheng Liu 0011, Zhicheng Dou |
COLING | 5 |
| 2025 | Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts ReasoningabstractAs large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent large models exhibit observable performance limitations on ambiguous instructions, where weak reasoning abilities of disambiguation can lead to catastrophic errors. To address this issue, this paper proposes Visual-O1, a multi-modal multi-turn chain-of-thought reasoning framework. It simulates human multi-modal multi-turn reasoning, providing instantial experience for highly intelligent models or empirical experience for generally intelligent models to understand ambiguous instructions. Unlike traditional methods that require models to possess high intelligence to understand long texts or perform lengthy complex reasoning, our framework does not notably increase computational overhead and is more general and effective, even for generally intelligent models. Experiments show that our method not only enhances the performance of models of different intelligence levels on ambiguous instructions but also improves their performance on general datasets. Our work highlights the potential of artificial intelligence to work like humans in real-world scenarios with uncertainty and ambiguity. We release our data and code at https://github.com/kodenii/Visual-O1. Minheng Ni, Yutao Fan, Lei Zhang 0006, Wangmeng Zuo |
ICLR | 1 |
| 2025 | MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMabstractMultimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant issue and hinders the diagnosis of multimodal reasoning failures within MLLMs. To address this, we propose the MIRAGE benchmark, which isolates reasoning hallucinations by constructing questions where input images are correctly perceived by MLLMs yet reasoning errors persist. MIRAGE introduces multi-granular evaluation metrics: accuracy, factuality, and LLMs hallucination score for hallucination quantification. Our analysis reveals strong correlations between question types and specific hallucination patterns, particularly systematic failures of MLLMs in spatial reasoning involving complex relationships (\emph{e.g.}, complex geometric patterns across images). This highlights a critical limitation in the reasoning capabilities of current MLLMs and provides targeted insights for hallucination mitigation on specific types. To address these challenges, we propose Logos, a method that combines curriculum reinforcement fine-tuning to encourage models to generate logic-consistent reasoning chains by stepwise reducing learning difficulty, and collaborative hint inference to reduce reasoning complexity. Logos establishes a baseline on MIRAGE, and reduces the logical hallucinations in original base models. Link: \url{https://bit.ly/25mirage}. Bowen Dong 0001, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, Lei Zhang 0006 |
NeurIPS | 2 |
| 2025 | Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement FinetuningabstractRecent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains challenging due to inherent limitations in text-only CoT, such as visual hallucinations and insufficient multimodal integration. In this paper, we introduce Point-RFT, a multimodal reasoning framework explicitly designed to leverage visually grounded CoT reasoning for visual document understanding. Our approach consists of two stages: First, we conduct format finetuning using a curated dataset of 71K diverse visual reasoning problems, each annotated with detailed, step-by-step rationales explicitly grounded to corresponding visual elements. Second, we employ reinforcement finetuning targeting visual document understanding. On ChartQA, our approach improves accuracy from 70.88% (format-finetuned baseline) to 90.04%, surpassing the 83.92% accuracy achieved by reinforcement finetuning relying solely on text-based CoT. The result shows that our grounded CoT is more effective for multimodal reasoning compared with the text-only CoT. Moreover, Point-RFT exhibits superior generalization capability across several out-of-domain visual document reasoning benchmarks, including CharXiv, PlotQA, IconQA, TabMWP, etc., and highlights its potential in complex real-world scenarios. Minheng Ni, Zhengyuan Yang, Chung-Ching Lin, Wangmeng Zuo |
NeurIPS | 1 |
| 2024 | ORES: Open-Vocabulary Responsible Visual SynthesisabstractAvoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visual synthesis tends to be diverse, depending on the region, context, and usage scenarios. In this work, we formalize a new task, Open-vocabulary Responsible Visual Synthesis (ORES), where the synthesis model is able to avoid forbidden visual concepts while allowing users to input any desired content. To address this problem, we present a Two-stage Intervention (TIN) framework. By introducing 1) rewriting with learnable instruction through a large-scale language model (LLM) and 2) synthesizing with prompt intervention on a diffusion synthesis model, it can effectively synthesize images avoiding any concepts but following the user's query as much as possible. To evaluate on ORES, we provide a publicly available dataset, baseline models, and benchmark. Experimental results demonstrate the effectiveness of our method in reducing risks of image generation. Our work highlights the potential of LLMs in responsible visual synthesis. Our code and dataset is public available in https://github.com/kodenii/ORES. Minheng Ni, Chenfei Wu, Xiaodong Wang 0023, Shengming Yin, Zicheng Liu 0001, Nan Duan 0001 |
AAAI | 1 |
| 2024 | Responsible Visual Editing
Minheng Ni, Yeli Shen, Lei Zhang 0006, Wangmeng Zuo |
ECCV (22) | 1 |
| 2024 | Multi-Attentional Distance for Zero-Shot Classification with Text-to-Image Diffusion ModelabstractText-to-image diffusion models have demonstrated rich visual-linguistic capability. However, existing image classification methods based on diffusion models simply choose the best-predicted noise, not exploiting the relationships between visual elements and text adequately. To this end, we propose a novel Multi-attentional Distance Classifier (MDC) by exploring some beneficial information in diffusion models. Specifically, MDC joints self- and cross-attention maps to model the semantic and structural distances of the latent variables of images under different category conditions, measuring the relevance between images and categories. With two types of distances integrated, we can classify image's category with minimum distance. We evaluate MDC on CIFAR-10, STL-10, and CIFAR-100 datasets under the zero-shot setting and it shows MDC achieves superior performance to prior works. Further experiments prove that by introducing attention in the diffusion process, MDC can discover key semantic and structure information of categories among images. Codes are publicly available at https://github.com/Carlofkl/MDC. Kailai Feng, Minheng Ni, Jiaxiu Jiang, Zhilu Zhang 0001, Wangmeng Zuo |
ICME | 2 |
| 2024 | StrokeNUWA - Tokenizing Strokes for Vector Graphic SynthesisabstractTo leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model’s ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation "stroke" tokens on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a $94\times$ speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%. Zecheng Tang, Chenfei Wu, Minheng Ni, Shengming Yin, Zhengyuan Yang, Zicheng Liu 0001, Nan Duan 0001 |
ICML | 4 |
| 2023 | NUWA-XL: Diffusion over Diffusion for eXtremely Long Video GenerationabstractShengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shengming Yin, Chenfei Wu, Huan Yang 0005, Xiaodong Wang 0023, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Jianlong Fu, Ming Gong 0001, Zicheng Liu 0001, Houqiang Li, Nan Duan 0001 |
ACL (1) | 6 |
| 2023 | NÜWA-LIP: Language-guided Image Inpainting with Defect-free VQGANabstractLanguage-guided image inpainting aims to fill the defective regions of an image under the guidance of text while keeping the non-defective regions unchanged. However, directly encoding the defective images is prone to have an adverse effect on the non-defective regions, giving rise to distorted structures on non-defective parts. To better adapt the text guidance to the inpainting task, this paper proposes NÜWA-LIP, which involves defect-free VQGAN (DF-VQGAN) and a multi-perspective sequence-to-sequence module (MP-S2S). To be specific, DF-VQGAN introduces relative estimation to carefully control the receptive spreading, as well as symmetrical connections to protect structure details unchanged. For harmoniously embedding text guidance into the locally defective regions, MP-S2S is employed by aggregating the complementary perspectives from low-level pixels, high-level tokens as well as the text description. Experiments show that our DF-VQGAN effectively aids the inpainting process while avoiding unexpected changes in non-defective regions. Results on three open-domain benchmarks demonstrate the superior performance of our method against state-of-the-arts. Our code, datasets, and model will be made publicly available11https://github.com/kodenii/NUWA-LIP. Minheng Ni, Xiaoming Li 0002, Wangmeng Zuo |
CVPR | 1 |
| 2023 | ImaginaryNet: Learning Object Detectors without Real Images and Annotations
Minheng Ni, Zitong Huang, Kailai Feng, Wangmeng Zuo |
ICLR | 1 |
| 2023 | Learning 3D Photography Videos via Self-supervised Diffusion on Single Imagesabstract3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded regions. The inpainting model plays a crucial role in rendering quality, but it is normally trained on out-of-domain data. To reduce the training and inference gap, we propose a novel self-supervised diffusion model as the inpainting module. Given a single input image, we automatically construct a training pair of the masked occluded image and the ground-truth image with random cycle rendering. The constructed training samples are closely aligned to the testing instances, without the need for data annotation. To make full use of the masked images, we designed a Masked Enhanced Block (MEB), which can be easily plugged into the UNet and enhance the semantic conditions. Towards real-world animation, we present a novel task: out-animation, which extends the space and time of input objects. Extensive experiments on real datasets show that our method achieves competitive results with existing SOTA methods. Xiaodong Wang 0023, Chenfei Wu, Shengming Yin, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
IJCAI | 4 |
| 2022 | Multi-domain Spoken Language Understanding Using Domain- and Task-aware ParameterizationabstractSpoken language understanding (SLU) has been addressed as a supervised learning problem, where a set of training data is available for each domain. However, annotating data for a new domain can be both financially costly and non-scalable. One existing approach solves the problem by conducting multi-domain learning where parameters are shared for joint training across domains, which is domain-agnostic and task-agnostic . In the article, we propose to improve the parameterization of this method by using domain-specific and task-specific model parameters for fine-grained knowledge representation and transfer. Experiments on five domains show that our model is more effective for multi-domain SLU and obtain the best results. In addition, we show its transferability when adapting to a new domain with little data, outperforming the prior best model by 12.4%. Finally, we explore the strong pre-trained model in our framework and find that the contributions from our framework do not fully overlap with contextualized word representations (RoBERTa). Libo Qin 0001, Fuxuan Wei, Minheng Ni, Yue Zhang 0004, Wanxiang Che, Yangming Li, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn a dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers’ intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately. The dialog context information (contextual information) and the mutual interaction information are two key factors that contribute to the two related tasks. Unfortunately, none of the existing approaches consider the two important sources of information simultaneously. In this paper, we propose a Co-Interactive Graph Attention Network (Co-GAT) to jointly perform the two tasks. The core module is a proposed co-interactive graph interaction layer where a cross-utterances connection and a cross-tasks connection are constructed and iteratively updated with each other, achieving to consider the two types of information simultaneously. Experimental results on two public datasets show that our model successfully captures the two sources of information and achieve the state-of-the-art performance. In addition, we find that the contributions from the contextual and mutual interaction information do not fully overlap with contextualized word representations (BERT, Roberta, XLNet). Libo Qin 0001, Zhouyang Li, Wanxiang Che, Minheng Ni, Ting Liu 0001 |
AAAI | 4 |
| 2021 | M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-TrainingabstractWe present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universal representations that can map objects occurred in different modalities or texts expressed in different languages into a common semantic space. In addition, to explicitly encourage fine-grained alignment between images and non-English languages, we also propose Multimodal Code-switched Training (MCT) to combine monolingual pre-training and multimodal pre-training via a code-switch strategy. Experiments are performed on the multilingual image retrieval task across two benchmark datasets, including MSCOCO and Multi30K. M3P can achieve comparable results for English and new state-of-the-art results for non-English languages. Minheng Ni, Haoyang Huang, Edward Dong Bo Cui, Taroon Bharti, Dongdong Zhang 0001, Nan Duan 0001 |
CVPR | 1 |
| 2021 | Knowing Where to Leverage: Context-Aware Graph Convolutional Network With an Adaptive Fusion Layer for Contextual Spoken Language UnderstandingabstractSpoken language understanding (SLU) systems aim to understand users’ utterance, which is a key component of task-oriented dialogue systems. In this paper, we focus on improving the contextual SLU. The contextual SLU systems mainly focus on how to effectively incorporate dialog context information (contextual information). The existing approaches all use the same contextual information to guide slot filling at all tokens, which may inject the irrelevant information and result in ambiguity. To tackle this problem, we propose a context-aware graph convolutional network (GCN) with an adaptive fusion layer for contextual SLU. The context-aware GCN is proposed to automatically aggregate the contextual information, which frees our model from the manually designed heuristic aggregation function. Meanwhile, an adaptive fusion layer is applied at each token to dynamically incorporate relevant contextual information, which achieves a fine-grained contextual information transfer to guide the token-level slot filling. Experiments on the Simulated Dialog Dataset show that our model achieves state-of-the-art performance and outperforms other previous methods by a large margin (+3.67% on Sim-R, +4.18% on Sim-M and +3.75% on Overall dataset). In addition, we explore and analyze the pre-trained model (i.e., BERT) in our framework. We show that incorporating BERT brings a large improvement in low-resource setting. Libo Qin 0001, Wanxiang Che, Minheng Ni, Yangming Li, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers' intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately (Kim and Kim 2018). Most of the existing systems either treat them as separate tasks or just jointly model the two tasks by sharing parameters in an implicit way without explicitly modeling mutual interaction and relation. To address this problem, we propose a Deep Co-Interactive Relation Network (DCR-Net) to explicitly consider the cross-impact and model the interaction between the two tasks by introducing a co-interactive relation layer. In addition, the proposed relation layer can be stacked to gradually capture mutual knowledge with multiple steps of interaction. Especially, we thoroughly study different relation layers and their effects. Experimental results on two public datasets (Mastodon and Dailydialog) show that our model outperforms the state-of-the-art joint model by 4.3% and 3.4% in terms of F1 score on dialog act recognition task, 5.7% and 12.4% on sentiment classification respectively. Comprehensive analysis empirically verifies the effectiveness of explicitly modeling the relation between the two tasks and the multi-steps interaction mechanism. Finally, we employ the Bidirectional Encoder Representation from Transformer (BERT) in our framework, which can further boost our performance in both tasks. Libo Qin 0001, Wanxiang Che, Yangming Li, Minheng Ni, Ting Liu 0001 |
AAAI | 4 |
| 2020 | CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLPabstractMulti-lingual contextualized embeddings, such as multilingual-BERT (mBERT), have shown success in a variety of zero-shot cross-lingual tasks. However, these models are limited by having inconsistent contextualized representations of subwords across different languages. Existing work addresses this issue by bilingual projection and fine-tuning technique. We propose a data augmentation framework to generate multi-lingual code-switching data to fine-tune mBERT, which encourages model to align representations from source and multiple target languages once by mixing their context information. Compared with the existing work, our method does not rely on bilingual sentences for training, and requires only one training process for multiple target languages. Experimental results on five tasks with 19 languages show that our method leads to significantly improved performances for all the tasks compared with mBERT. Libo Qin 0001, Minheng Ni, Yue Zhang 0004, Wanxiang Che |
IJCAI | 2 |