Mingjie Li 0004

dblp:48/10103-4 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-3565-1180ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 77% Graph data management · 15% Indexing and storage engines · 8%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 43% Image and video processing · 38% Geometric modeling and processing · 19%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
2.442025
Graph-based Approximate Nearest Neighbor Search by Deep Reinforcement Routing · ACM Multimedia 2025
Deep Learning for Approximate Nearest Neighbour Search: A Survey and Future Directions · IEEE Trans. Knowl. Data Eng. 2023
Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement · IEEE Trans. Knowl. Data Eng. 2020
Information retrieval › similarity search
nearest neighbor search
1.122023
Deep Learning for Approximate Nearest Neighbour Search: A Survey and Future Directions · IEEE Trans. Knowl. Data Eng. 2023
Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement · IEEE Trans. Knowl. Data Eng. 2020
Machine learning › Generative modeling › generative adversarial network › image-to-image translation
multi-domain image translation
1.012026
Enhancing Cross-Domain Correspondence for Unsupervised Image-to-Image Translation · IEEE Trans. Multim. 2026
Visual content generation and editing
image-to-image translation
1.012026
Enhancing Cross-Domain Correspondence for Unsupervised Image-to-Image Translation · IEEE Trans. Multim. 2026
Visual content generation and editing › image-to-image translation
unpaired image translation
1.012026
Enhancing Cross-Domain Correspondence for Unsupervised Image-to-Image Translation · IEEE Trans. Multim. 2026
Graph data management
graph indexing
0.912025
Graph-based Approximate Nearest Neighbor Search by Deep Reinforcement Routing · ACM Multimedia 2025
Image and video processing › super-resolution › image super-resolution
blind super-resolution
0.912025
Image Super-Resolution With Taylor Expansion Approximation and Large Field Reception · IEEE Trans. Multim. 2025
Image and video processing › super-resolution
image super-resolution
0.912025
Image Super-Resolution With Taylor Expansion Approximation and Large Field Reception · IEEE Trans. Multim. 2025
Geometric modeling and processing
self-similarity
0.912025
Image Super-Resolution With Taylor Expansion Approximation and Large Field Reception · IEEE Trans. Multim. 2025
Information retrieval › similarity search › nearest neighbor search
high-dimensional nearest neighbor search
0.412020
Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement · IEEE Trans. Knowl. Data Eng. 2020
Indexing and storage engines
learned index
0.412020
I/O Efficient Approximate Nearest Neighbour Search based on Learned Functions · ICDE 2020
Information retrieval
similarity search
0.412020
I/O Efficient Approximate Nearest Neighbour Search based on Learned Functions · ICDE 2020
Performance modeling and evaluation
benchmarking
0.112020
Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement · IEEE Trans. Knowl. Data Eng. 2020

Methods — techniques the papers use, named apart from their topics

visual perceptual guidance · 2.0semantic perceptual matching · 2.0multi-level style embedding · 2.0CLIP · 2.0taylor expansion approximation · 0.9multi-scale large field reception · 0.9graph routing · 0.9deep reinforcement learning · 0.9recall analysis · 0.9experimental evaluation · 0.9learning to search · 0.7learning to index · 0.7deep learning · 0.7learned hashing functions · 0.4data-sensitive indexing · 0.4
YearPublicationVenuePosition
2026 SemiDDM-weather: A semi-supervised learning framework for all-in-one adverse weather removal
Fang Long, Wenkang Su 0001, Mingjie Li 0004, Yuan-Gen Wang, Xiaochun Cao
Neural Networks5
2026 Enhancing Cross-Domain Correspondence for Unsupervised Image-to-Image Translation
abstract
UNsupervised Image-to-image Translation (UNIT) aims to translate images across visual domains without paired training data, which has been widely used in style transfer, image processing, game design, etc. However, ensuring the correspondence (e.g., target category, pose, or head orientation) between generated images and inputs remains a formidable challenge. To this end, we present a new scheme, named EC-UNIT, which comprises three innovative designs aiming to Enhance cross domain Correspondence for UNIT. Specifically, 1) we propose Multi-level Style Embedding to extract multi-level style features for fusion while imposing our newly designed Hierarchical Consistency Constraints on both the content and style features (MSE&HCC), aiming to retain more style representations and facilitate feature disentanglement; 2) we develop Semantic Perceptual Matching (SPM) to minimize the semantic distribution discrepancy between the generated image and the input image by leveraging the multimodal model CLIP, dedicated to enhancing semantic consistency; 3) considering that previous works have struggled to control the image translation using pixel-level visual consistency constraints, we design Visual Perceptual Guidance (VPG) to reduce the perceptual distance between the generated image and the style input in VGG feature space, devoted to enhancing visual perceptual correspondence, thereby preventing the generation of unrealistic image details. Extensive experiments demonstrate that our EC-UNIT is more stable and outperforms current SOTA competitors in terms of image quality and diversity as well as both content and style consistency.
Binxin Lai, Wenkang Su 0001, Yuying Liang, Yuan-Gen Wang, Mingjie Li 0004, Jiantao Zhou 0001
IEEE Trans. Multim.5
2025 Graph-based Approximate Nearest Neighbor Search by Deep Reinforcement Routing
abstract
We focus on the approximate nearest neighbor search (ANNS) in high dimensional space, which is a fundamental technique in computer vision and multimedia database. Among the ANNS solutions, graph-based approaches achieve excellent performance by executing a routing algorithm on a proximity graph to retrieve the nearest neighbors. However, most of their routing strategies are heuristic-based greedy routing, leading to suboptimal search results with large number of hops. In this paper, we propose a novel routing paradigm on graphs for ANNS problem by deep reinforcement learning. We design a reinforcement model to learn the routing policy by making use of both graph global and local topology information. A hops-optimized reward mechanism is devised to enable the model to be more efficient and effective. The final searching algorithm with the learned model is able to find the nearest neighbors without any backtracking in a small number of hops. Comprehensive experiments on real-world datasets demonstrate the superiorities of the proposed method over the state-of-the-art ANNS approaches.
Mingjie Li 0004, Dian Ouyang, Ying Zhang 0001, Wei Wang 0011
ACM Multimedia1
2025 Black-box adversarial attacks against image quality assessment models
Yu Ran, Aoxiang Zhang, Mingjie Li 0004, Weixuan Tang 0004, Yuan-Gen Wang
Expert Syst. Appl.3
2025 Adaptive Multi-Lens Phase Modulation for Scale-Aware Privacy-Preserving Human Pose Recognition
abstract
Recently, optical privacy protection has emerged as a promising approach for safeguarding visual privacy at the physical acquisition stage. However, existing methods often face a trade‐off between privacy strength and human pose recognition accuracy, particularly in long‐range and multi‐scale scenarios. To address this challenge, we propose a novel adaptive optical privacy‐preserving framework that integrates a learnable optical modulation system with a human pose recognition network. The core of our method lies in a sparse‐weighted multi‐lens model, where a lightweight multilayer perceptron (MLP) predicts a sparse set of coefficients to linearly combine predefined lens phase profiles based on facial region geometry. This enables dynamic control over the point spread function (PSF), adapting the degree of image degradation to subject scale in real time. Additionally, we introduce a privacy‐aware loss function that selectively reduces facial localization accuracy while preserving body pose information. Extensive experiments on MSCOCO and FLIC datasets demonstrate that the proposed method achieves a favorable balance between privacy protection and pose estimation, outperforming previous optical‐ and software‐based baselines.
Weilong Peng, Quanwei Deng, Mingjie Li 0004, Yangtao Wang, Yan Wang 0022, Lisheng Fan, Meie Fang
IET Softw.3
2025 Image Super-Resolution With Taylor Expansion Approximation and Large Field Reception
abstract
Self-similarity techniques are booming in blind super-resolution (SR) due to accurate estimation of the degradation types involved in low-resolution images. However, high-dimensional matrix multiplication within self-similarity computation prohibitively consumes massive computational costs. We find that the high-dimensional attention map is derived from the matrix multiplication between query and key, followed by a softmax function. This softmax makes the matrix multiplication inseparable, posing a great challenge in simplifying computational complexity. To address this issue, we first propose a second-order Taylor expansion approximation (STEA) to separate the matrix multiplication of query and key, resulting in the complexity reduction from$\mathcal {O}(N^{2})$to$\mathcal {O}(N)$. Then, we design a multi-scale large field reception (MLFR) to compensate for the performance degradation caused by STEA. Finally, we apply these two core designs to laboratory and real-world scenarios by constructing LabNet and RealNet, respectively. Extensive experimental results tested on five synthetic datasets demonstrate that our LabNet sets a new benchmark in qualitative and quantitative evaluations. Tested on the real-world dataset, our RealNet achieves superior visual quality over existing methods. Ablation studies further verify the contributions of STEA and MLFR towards both LabNet and RealNet frameworks. Codes are available athttps://github.com/GZHU-DVL/STEA-MLFR.
Jiancong Feng, Yuan-Gen Wang, Mingjie Li 0004, Fengchuang Xing
IEEE Trans. Multim.3
2024 Cross-Shaped Adversarial Patch Attack
abstract
Recent studies have shown that deep learning-based classifiers are vulnerable to malicious inputs, i.e., adversarial examples. A practical solution is to construct a perceptible but localized perturbation called patch, making the well-trained models misclassified. However, most existing patch-based adversarial attacks focus on designing patches with localized rectangles, squares, or grids, ignoring the effect of the non-local patch. In this paper, we propose a novel cross-shaped patch attack paradigm (CSPA), a simple yet efficient and effective adversarial attack in Black-box scenarios. Specifically, the cross-shaped patch consists of two line segments intersected and perpendicular to each other at the midpoint. These two line segments are designed to be sufficiently thin and long to reach the four corners of the input image nearly. Thus, the patch has a globalized perturbation capacity while preserving its continuousness. The content and location of cross-shaped patch are then iteratively optimized by a carefully contrived random search-based algorithm to maximize this global property. Comprehensive experiments are conducted on four benchmark datasets against various victim networks. The results show that the proposed CSPA outperforms the existing patch-based attacks regarding both attack success rate and query efficiency by a large margin. Specifically, compared with the baselines, CSPA increases the success rate by up to 20% on ImageNet and reaches 100% on the CIFAR-100 and CIFAR-10 datasets. Meanwhile, CSPA reduces the average number of queries by up to 7 times. Even for the white-box attack scenario, our designed cross-shaped patch can still be applicable, achieving state-of-the-art performance.
Yu Ran, Mingjie Li 0004, Lin-Cheng Li, Yuan-Gen Wang, Jin Li 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 Deep Learning for Approximate Nearest Neighbour Search: A Survey and Future Directions
abstract
Approximate nearest neighbour search (ANNS) in high-dimensional space is an essential and fundamental operation in many applications from many domains such as multimedia database, information retrieval and computer vision. With the rapidly growing volume of data and the dramatically increasing demands of users, traditional heuristic-based ANNS solutions have been facing great challenges in terms of both efficiency and accuracy. Inspired by the recent successes of deep learning in many fields, substantial efforts have been devoted to applying deep learning techniques to ANNS for learning to index and learning to search, resulting in numerous algorithms that achieve state-of-the-art performance compared with conventional methods. In this survey paper, we comprehensively review the different types of deep learning-based ANNS methods according to two learning paradigms:learning to indexandlearning to search. We provide a comprehensive overview and analysis of these methods in a systematic manner. Based on the overview, we point out thatend-to-end learningwill be a new and promising research direction for deep learning-based ANNS, i.e., applying deep learning techniques to jointly learn the indexing and searching together, such that the underlying knowledge learned from data can directly contribute to the final searching performance. Finally, we conduct experiments and provide general performance analyses for the representative deep learning-based ANNS algorithms.
Mingjie Li 0004, Yuan-Gen Wang, Peng Zhang 0057, Hanpin Wang, Lisheng Fan, Enxia Li, Wei Wang 0011
IEEE Trans. Knowl. Data Eng.1
2020 I/O Efficient Approximate Nearest Neighbour Search based on Learned Functions
abstract
Approximate nearest neighbour search (ANNS) in high dimensional space is a fundamental problem in many applications, such as multimedia database, computer vision and information retrieval. Among many solutions, data-sensitive hashing-based methods are effective to this problem, yet few of them are designed for external storage scenarios and hence do not optimized for I/O efficiency during the query processing. In this paper, we introduce a novel data-sensitive indexing and query processing framework for ANNS with an emphasis on optimizing the I/O efficiency, especially, the sequential I/Os. The proposed index consists of several lists of point IDs, ordered by values that are obtained by learned hashing (i.e., mapping) functions on each corresponding data point. The functions are learned from the data and approximately preserve the order in the high-dimensional space. We consider two instantiations of the functions (linear and non-linear), both learned from the data with novel objective functions. We also develop an I/O efficient ANNS framework based on the index. Comprehensive experiments on six benchmark datasets show that our proposed methods with learned index structure perform much better than the state-of-the-art external memory-based ANNS methods in terms of I/O efficiency and accuracy.
Mingjie Li 0004, Ying Zhang 0001, Yifang Sun, Wei Wang 0011, Ivor W. Tsang, Xuemin Lin 0001
ICDE1
2020 Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement
abstract
Nearest neighbor search is a fundamental and essential operation in applications from many domains, such as databases, machine learning, multimedia, and computer vision. Because exact searching results are not efficient for a high-dimensional space, a lot of efforts have turned to approximate nearest neighbor search. Although many algorithms have been continuously proposed in the literature each year, there is no comprehensive evaluation and analysis of their performance. In this paper, we conduct a comprehensive experimental evaluation of many state-of-the-art methods for approximate nearest neighbor search. Our study (1) is cross-disciplinary (i.e., including 19 algorithms in different domains, and from practitioners) and (2) has evaluated a diverse range of settings, including 20 datasets, several evaluation metrics, and different query workloads. The experimental results are carefully reported and analyzed to understand the performance results. Furthermore, we propose a new method that achieves both high query efficiency and high recall empirically on majority of the datasets under a wide range of settings.
Ying Zhang 0001, Yifang Sun, Wei Wang 0011, Mingjie Li 0004, Wenjie Zhang 0001, Xuemin Lin 0001
IEEE Trans. Knowl. Data Eng.5
2018 An Efficient Exact Nearest Neighbor Search by Compounded Embedding
Mingjie Li 0004, Ying Zhang 0001, Yifang Sun, Wei Wang 0011, Ivor W. Tsang, Xuemin Lin 0001
DASFAA (1)1