Mingjie Xu

dblp:136/2952 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSecurity and privacy · 2 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding boxes to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap raises uncertainty about whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and utilize them to solve problems. To address this limitation, we introduce VP-Bench, aiming to assess MLLMs’ capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models’ ability to perceive VPs in natural scenes, utilizing 100K visualized prompts spanning 8 shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 21 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL-2.5 and Qwen2.5-VL). In addition, we conduct a comprehensive analysis of the factors influencing VP understanding, such as attribute variations and model scale. VP-Bench establishes a new reference framework for studying MLLMs’ ability to comprehend and resolve grounded referring questions.
Mingjie Xu, Jinpeng Chen 0003, Yuzhi Zhao, Jason Chun Lok Li, Zekang Du, Mengyang Wu, Kun Li 0015, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei
AAAI1
2026 Gene-guided multimodal data fusion for cancer patient survival analysis
Mingjie Xu, Zongbao Yang, Ruxin Wang 0001, Hao Zhang 0079
Neurocomputing1
2026 RSTFA: Efficient Training-Free Human-Preference Alignment via Rejection Sampling for Text-to-Image Diffusion Models
abstract
Given a text-to-image diffusion model pretrained on large-scale text-image pairs, can we align the model with human pReferences without further fine-tuning? In this paper, we analyze the effect of alignment tuning in diffusion models by comparing the diffusion denoising trajectory between base and aligned models. Our findings reveal that alignment tuning primarily affects superficial stylistic aspects during denoising, rather than fundamental content, suggesting superficial alignment behaviors. Based on this discovery, we introduce a novel, training-free alignment approach (RSTFA) that leverages rejection sampling at specific stylistic timesteps, ensuring human preference alignment without fine-tuning or heavy inference overhead. We provide a theoretical analysis and derive a bias bound for our rejection-sampling alignment scheme. Empirically, we show that RSTFA better preserves sample diversity than reinforcement-learning-based tuning methods. Extensive experiments on Pick-a-Pic, COCO, HPD V2, and PartiPrompts show that our method not only achieves superior alignment with human preferences compared to state-of-the-art methods, but also reduces computational demands, establishing efficient, human-centered diffusion model alignment.
Hongzheng Yang, Jason Chun Lok Li, Kun Li 0015, Wenao Ma, Mingjie Xu, Yuzhi Zhao, Lai-Man Po
IEEE Trans. Image Process.5
2025 ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
abstract
Controversial contents largely inundate the Internet, infringing various cultural norms and child protection standards. Traditional Image Content Moderation (ICM) models fall short in producing precise moderation decisions for diverse standards, while recent multimodal large language models (MLLMs), when adopted to general rule-based ICM, often produce classification and explanation results that are inconsistent with human moderators. Aiming at flexible, explainable, and accurate ICM, we design a novel rule-based dataset generation pipeline, decomposing concise human-defined rules and leveraging well-designed multi-stage prompts to enrich short explicit image annotations. Our ICM-Instruct dataset includes detailed moderation explanation and moderation Q-A pairs. Built upon it, we create our ICM-Assistant model in the framework of rule-based ICM, making it readily applicable in real practice. Our ICM-Assistant model demonstrates exceptional performance and flexibility. Specifically, it significantly outperforms existing approaches on various sources, improving both the moderation classification (36.8% on average) and moderation explanation quality (26.6% on average) consistently over existing MLLMs. Caution: Content includes offensive language or images.
Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Qinbin Li, Guang-Neng Hu, Shengchao Qin, Chi-Wing Fu
AAAI4
2025 LLaVA-SpaceSGG: Visual Instruct Tuning for Open-Vocabulary Scene Graph Generation with Enhanced Spatial Relations
abstract
Scene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks. However, existing SGG models often overlook essential spatial relationships and struggle with generalization in open-vocabulary contexts. To address these limitations, we propose LLaVA-SpaceSGG, a multimodal large language model (MLLM) designed for open-vocabulary SGG with enhanced spatial relation modeling. To train it, we collect the SGG instruction-tuning dataset, named SpaceSGG. This dataset is constructed by combining publicly available datasets and synthesizing data using open-source models within our data construction pipeline. It combines object locations, object relations, and depth information, resulting in three data formats: spatial SGG description, question-answering, and conversation. To enhance the transfer of MLLMs' inherent capabilities to the SGG task, we introduce a two-stage training paradigm. Experiments show that LLaVA-SpaceSGG outperforms other open-vocabulary SGG methods, boosting recall by 8.6% and mean recall by 28.4% compared to the baseline. Our codebase, dataset, and trained models are publicly accessible on GitHub at the following URL: https://github.com/Endlinc/LLaVA-SpaceSGG.
Mingjie Xu, Mengyang Wu, Yuzhi Zhao, Jason Chun Lok Li, Weifeng Ou
WACV1
2024 Gaze from Origin: Learning for Generalized Gaze Estimation by Embedding the Gaze Frontalization Process
abstract
Gaze estimation aims to accurately estimate the direction or position at which a person is looking. With the development of deep learning techniques, a number of gaze estimation methods have been proposed and achieved state-of-the-art performance. However, these methods are limited to within-dataset settings, whose performance drops when tested on unseen datasets. We argue that this is caused by infinite and continuous gaze labels. To alleviate this problem, we propose using gaze frontalization as an auxiliary task to constrain gaze estimation. Based on this, we propose a novel gaze domain generalization framework named Gaze Frontalization-based Auxiliary Learning (GFAL) Framework which embeds the gaze frontalization process, i.e., guiding the feature so that the eyeball can rotate and look at the front (camera), without any target domain information during training. Experimental results show that our proposed framework is able to achieve state-of-the-art performance on gaze domain generalization task, which is competitive with or even superior to the SOTA gaze unsupervised domain adaptation methods.
Mingjie Xu, Feng Lu 0005
AAAI1
2024 DCMAI: A Dynamical Cross-Modal Alignment Interaction Framework for Document Key Information Extraction
abstract
Document key information extraction (DKIE) is a crucial topic that aims at automatically comprehending documents with complex formats and layouts (invoices, business insurance, etc.). While pre-trained approaches have shown high performance on many DKIE tasks, they suffer from three major challenges. First of all, these approaches ignore the ambiguity resulting from similar text representations before cross-modal interaction. Secondly, they do not consider cross-modal representation alignment before cross-modal interaction. Finally, self-attention layers in cross-modal interaction incur significant computing costs, making it hard to perform joint representation learning from all negative samples. To address these issues, we propose a Dynamical Cross-Modal Alignment Interaction framework (DCMAI). To be more specific, 1) A prior knowledge-guided module is designed to adaptively mine fine-grained visual information to disambiguate similar text representations. 2) A crossover alignment loss is formulated to align cross-modal representations before cross-modal interaction. 3) A hierarchical interaction sampling scheme is introduced to obtain a small but efficient subset of cross-modal negative samples, and a contrastive loss is employed to improve joint representation learning. Comprehensive experiments show that the proposed DCMAI achieves state-of-the-art performance than competitive baselines on several public downstream benchmarks. Code will be open to the public.
Yonghong Song, Yongbiao Deng, Kangkang Xie, Mingjie Xu, Haijun Ren
IEEE Trans. Circuits Syst. Video Technol.5
2023 Learning a Generalized Gaze Estimator from Gaze-Consistent Feature
abstract
Gaze estimator computes the gaze direction based on face images. Most existing gaze estimation methods perform well under within-dataset settings, but can not generalize to unseen domains. In particular, the ground-truth labels in unseen domain are often unavailable. In this paper, we propose a new domain generalization method based on gaze-consistent features. Our idea is to consider the gaze-irrelevant factors as unfavorable interference and disturb the training data against them, so that the model cannot fit to these gaze-irrelevant factors, instead, only fits to the gaze-consistent features. To this end, we first disturb the training data via adversarial attack or data augmentation based on the gaze-irrelevant factors, i.e., identity, expression, illumination and tone. Then we extract the gaze-consistent features by aligning the gaze features from disturbed data with non-disturbed gaze features. Experimental results show that our proposed method achieves state-of-the-art performance on gaze domain generalization task. Furthermore, our proposed method also improves domain adaption performance on gaze estimation. Our work provides new insight on gaze domain generalization task.
Mingjie Xu, Haofei Wang 0001, Feng Lu 0005
AAAI1
2023 Joint Low-light Enhancement and Super Resolution with Image Underexposure Level Guidance
Mingjie Xu, Chaoqun Zhuang, Feifan Lv, Feng Lu 0005
BMVC1
2022 Adaptive Sparse Self-attention for Object Detection
abstract
Object detection is a fundamental task for computer vision. The majority of prior methods employ global contextual information to enhance features representation. However, we argue that prior methods suffer from the superfluous features. To address the above mentioned problems, we explore the sparsity on object detection tasks in two dimensions. Specifically, a sparse spatial attention is proposed to capture global sparse long-range relationship into local features adaptively, where a learnable channel-wise mask is obtained to reduce the superfluous channels. Meanwhile, a sparse channel attention module is used to enhance the representation of each channel by introducing sparse global semantics. Experiments demonstrate that our proposed method outperforms comparative methods on the commonly used benchmark dataset, i.e., MS-COCO. The ablative experiments show that the sparsity the effectiveness in feature extraction and bounding boxes selection.
Mingjie Xu, Yonghong Song, Kangkang Xie, Jiaxi Mu
IJCNN1
2022 RIBDetector: an RFC-guided Inconsistency Bug Detecting Approach for Protocol Implementations
abstract
The implementations of network protocols must comply with rules described in their Request For Comments (RFC) Standards. Developers' misunderstanding or negligence of RFCs may bring in inconsistency bugs, which could further cause incorrect behaviors, interoperability issues, or critical security implications. Detecting such bugs is difficult as they usually result in silent erroneous effect. Prior work on RFC-directed inconsistency bug detection usually deal with a certain protocol or ad-hoc properties in RFCs. In this paper, we present RIBDetector, an approach focusing on statically and efficiently locating inconsistency bugs that could be triggered by hand-crafted network packets in protocol implementations. Given an implementation, its corresponding RFCs and a user-provided configuration file, our approach automatically extracts rules about packet format, state transition and error handling from RFCs into a uniform format which dictates condition checks that must be performed before taking particular operations. Then we leverage common programming conventions to identify corresponding locations of the conditions and operations in implementations and use a light-weight predominator-based algorithm to detect violations of RFC rules. We implemented a prototype of RIBDetector and demonstrated its efficacy by applying it on 14 implementations of 5 network protocols. For implementations varying in size from 1.5 to 141.3 KLOC, RIBDetector consumes 17.57 seconds on average to finish its analysis. We have detected 23 new inconsistency bugs, 6 of which are confirmed and fixed by developers.
Jingting Chen, Feng Li 0045, Mingjie Xu, Wei Huo 0005
SANER3
2020 ELAID: detecting integer-Overflow-to-Buffer-Overflow vulnerabilities by light-weight and accurate static analysis
abstract
Abstract The Integer-Overflow-to-Buffer-Overflow (IO2BO) vulnerability has been widely exploited by attackers to cause severe damages to computer systems. Automatically identifying this kind of vulnerability is critical for software security. Despite many works have been done to mitigate integer overflow, existing tools either report large number of false positives or introduce unacceptable time consumption. To address this problem, in this article we present a static analysis framework. It first constructs an inter-procedural call graph and utilizes taint analysis to accurately identify potential IO2BO vulnerabilities. Then it uses a light-weight method to further filter out false positives. Specifically, it generates constraints representing the conditions under which a potential IO2BO vulnerability can be triggered, and feeds the constraints to SMT solver to decide their satisfiability. We have implemented a prototype system ELAID based on LLVM, and evaluated it on 228 programs of the NIST’s SAMATE Juliet test suite and 14 known IO2BO vulnerabilities in real world. The experiment results show that our system can effectively and efficiently detect all known IO2BO vulnerabilities.
Mingjie Xu, Feng Li 0045, Wei Huo 0005
Cybersecur.2
2018 A Light-Weight and Accurate Method of Static Integer-Overflow-to-Buffer-Overflow Vulnerability Detection
Mingjie Xu, Feng Li 0045, Wei Huo 0005, Xinhua Li, Qingjia Huang
Inscrypt1
2015 Deterministic Detection of Cloning Attacks for Anonymous RFID Systems
abstract
Cloning attacks seriously impede the security of radio-frequency identification (RFID) applications. This paper tackles deterministic clone detection for anonymous RFID systems without tag identifiers (IDs) as a priori. Existing clone detection protocols either cannot apply to anonymous RFID systems due to necessitating the knowledge of tag IDs or achieve only probabilistic detection with a few clones tolerated. This paper proposes three protocols—BASE, DeClone, and DeClone+—toward fast and deterministic clone detection for large anonymous RFID systems. BASE leverages the observation that clone tags make tag cardinality exceed ID cardinality. DeClone is built on a recent finding that clone tags cause collisions that are hardly reconciled through rearbitration. For DeClone to achieve detection certainty, this paper designs breadth first tree traversal toward quickly verifying unreconciled collisions and hence the cloning attack. DeClone+ further incorporates optimization techniques that promise faster clone detection when clone ratio is relatively high. The performance of the proposed protocols is validated through analysis and simulation. This paper also suggests feasible extensions to enrich their applicability to distributed design.
Kai Bu, Mingjie Xu, Xuan Liu 0001, Jiaqing Luo, Shigeng Zhang, Minyu Weng
IEEE Trans. Ind. Informatics2
2014 Toward Fast and Deterministic Clone Detection for Large Anonymous RFID Systems
abstract
Cloning attacks seriously impede the security of Radio-Frequency Identification (RFID) applications. In this paper, we tackle deterministic clone detection for anonymous RFID systems without tag identifiers (IDs) as a priori. Existing clone detection protocols either cannot apply to anonymous RFID systems due to necessitating the knowledge of tag IDs or achieve only probabilistic detection with a few clones tolerated. We propose two protocols, BASE and DeClone, toward fast and deterministic clone detection for large anonymous RFID systems. BASE leverages the observation that clone tags make tag cardinality exceed ID cardinality. DeClone is built on a recent finding that clone tags cause collisions that are hardly reconciled through re-arbitration. For DeClone to achieve detection certainty, we design breadth first tree traversal toward quickly verifying unreconciled collisions and hence the cloning attack. We validate their detection performance through analysis and simulation. The results show that BASE delivers faster detection for small systems while DeClone for large ones especially when clone ratio increases.
Kai Bu, Mingjie Xu, Xuan Liu 0001, Jiaqing Luo, Shigeng Zhang
MASS2
2004 Experiments on Remote Sensing image cube and its OLAP
abstract
OLAP can answer questions such as 'what next' and 'what if'. OLAP, which always needs the support of data warehouse, is a complement to data mining. In the early steps of data mining progress, OLAP tools might be helpful in tasks such as exploring the data sets, locating more important variables and unwanted data, which can help users to understand the source data and quicken the data mining progress. In order to find data or information quickly in great volumes of Remote Sensing (RS) image data sets, for one thing, metadata bases need to be built; for another thing, research and development are required to deal with spatial OLAP application servers, which can manipulate spatial data warehouse containing RS images. Thematic images such as TM2, TM4, land cover, transportation, slope and the result image of ERDAS IMAGINE Expert Classifier, etc., have been inputted into MS Access tables. On grounds of these tables and the relational multi-dimensional data model, dimension tables and the fact table have been generated, and furthermore a RS image cube structure has been constructed. After the pre-computation and materialization, a RS image cube has been created. Experiments of OLAP on this cube have been carried out. Owning to the pre-computation and materialization, queries on the cube can be carried out with no delay. As the experiments show, if data warehouse and OLAP are adopted, not only different factors and their concept hierarchies can be used conveniently, but also queries speed up.
Mingjie Xu
IGARSS1
2004 Research on remote sensing image data mining prototype system and the RSIDMM-DTM
abstract
With mass production and widespread application of remote sensing (RS) image, the management of RS data and its processing theories, techniques and algorithms need a new breakthrough. Different levels of knowledge from very large volume of RS data will be applied to RS image classification, so as to improve the efficiency and accuracy of RS image analysis and to establish an intelligent GIS based on RS images. Difficulties in RS images data mining are listed And a prototype system of RS image data mining is designed. Besides experiments are made with RSIDMM-DTM, RS image data mining classification model, based on the Microsoft decision tree mining algorithm. By comparison with ERDAS IMAGINE Expert Classifier's effects, the experiments show that the RSIDMM-DTM takes spatial relationships and other contextual information into account. In addition, the acquisition, presentation and application of knowledge are highly automatic, and the classification is rather accurate.
Mingjie Xu, Lun Wu
IGARSS1