Mengjuan Li

dblp:122/4387 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SSC-Join: An Efficient Syntactic-Semantic Collaboration Based Set Semantic Similarity Join Algorithm
Lianyin Jia, Chengchen Zeng, Mengjuan Li, Suprio Ray, Jiaman Ding, Xiuxing Li
ICDE3
2026 Efficient Query Region Expansion and Decomposition based Spatial Range Query Algorithm
abstract
Spatial range queries play a crucial role in spatial information retrieval. Existing Z-order curve based algorithms suffer from accessing a large number of invalid points outside the query region. To address this challenge, we design a simple yet efficient Z-order curve based learned index, ZPI. Building upon ZPI, we propose a novel spatial range query algorithm, ZPI-RQ. ZPI-RQ leverages efficient decomposition mechanism to address the invalid points issue. To avoid decomposing the query region into a large number of overly small blocks, a query region expansion strategy is further introduced to align each query border with a m-order dividing lines. Experimental results show that ZPI-RQ significantly outperforms state-of-the-art algorithms in query efficiency, achieving a 3.6× improvement over traditional Z-order curve-based algorithms while accessing only 1% invalid points.
Lianyin Jia, Rongjin Wang, Yingbin Su, Suprio Ray, Mengjuan Li, Jiaman Ding
SIGIR5
2026 An Asymmetric-difference based Document Exact Similarity Search Algorithm
Lianyin Jia, Yongxue Zhao, Mengjuan Li, Xiuxing Li, Jiaman Ding
SIGIR3
2026 HMSNet: Hilbert curve enhanced Mamba for real-time semantic segmentation
Lianyin Jia, Aoxiang Gao, Mengjuan Li, Xiaodong Fu, Haihe Zhou, Jiaman Ding
Pattern Recognit.3
2025 Can LLMs only talk? Experimental studies on task scheduling with Large Language Models
abstract
Large Language Models (LLMs) have emerged as a disruptive technology for Natural Language Processing (NLP), achieving success in NLP-related generative applications. However, the potential capability of LLMs in other domains remains largely unexplored. To explore the potential of task scheduling with LLMs, we model a typical task scheduling scenario in cloud computing and transfer scheduling problems as natural language prompts. Afterward, the knowledge and reasoning abilities of LLMs are enabled to generate scheduling decisions. Six well-known and open-source LLMs are integrated into our framework to perform experimental studies, and the results are evaluated from multiple perspectives and compared with each other. Besides, traditional heuristic algorithms and a basic Reinforcement Learning (RL) method are all performed for comparison. Our results demonstrate: 1) compared to most heuristic methods, the decisions made by LLMs achieve better scheduling performance; 2) compared to the basic RL method, LLMs exhibit better generalization on various workload patterns; 3) the larger parameter size of the LLMs has, the better scheduling performance it achieves. To the best of our knowledge, our experimental study is the first exploration to apply LLMs in task scheduling. Our findings highlight the promising potential of LLMs as a novel approach to task scheduling, offering new avenues for research and practice.
Mengjuan Li, Zhengguang Chen, Huan Zhou 0006, Yingwen Chen 0001, Baokang Zhao, Xue Ouyang 0003, Jinshu Su
ICCCN1
2025 A Length Enhanced B+-Tree Based Index for Efficient Set Similarity Query
abstract
Set Similarity Query (SSQ) is widely applied in various fields. The existing B+-tree-based SSQ approaches fail to fully exploit length filtering and require calculating similarity bounds in a node-wise manner, leading to low efficiency. To address these issues, we propose LeB, a novel length-enhanced B+-tree index, whose keys integrate set lengths and bucket mapping, enabling the direct pruning of sets that do not meet the length requirements. Building upon LeB, we present an efficient algorithm, LeBQ, which leverages length filtering and symmetric difference allocation to determine the key bounds for a query, enabling the key bounds computation only once for each query$Q$and avoiding costly similarity bounds computation in a node-wise manner. Efficient key filtering strategies are proposed to prune sets that cannot be similar, significantly reducing the number of candidates. Based on LeBQ, LeBQ+ further reduces the number of candidates by introducing length-independent key bounds. Experimental results on four real datasets demonstrate that LeBQ+ has a higher node access efficiency and accesses only 3.08% to 27.47% nodes compared to the existing B+-tree-based SSQ algorithm. LeBQ+is up to 99.8 × faster than the state-of-the-art algorithms.
Lianyin Jia, Shiqi Luo, Jiaman Ding, Suprio Ray, Mengjuan Li, Xiuxing Li
ICDE5
2025 3SRank: An Effective Method for Measuring Semantic Similarity Strength
abstract
The measurement of semantic similarity between words is crucial in industrial big data analysis and many other fields. However, most existing work focuses on measuring the semantic similarity between two words, other than individual words. In calculating the semantic similarity strength(3S) of a word w, a naive algorithm requires accumulating the semantic similarity between w and all other words, which is inefficient. To address this issue, we designed a Compact Concept Taxonomy Tree (CCTT) and a novel method for measuring 3S—3SRank.3SRank only accumulates the similarity scores of nodes with a similarity to w greater than a specified threshold θ, and converts the problem of finding nodes that meet the threshold into a problem of finding nodes within specific depth bounds. This significantly reduces the cost of 3S calculation. Extensive experimental results demonstrate that 3SRank can effectively measure the 3S of words and achieve a good balance between measurement accuracy and algorithm execution time. When threshold θ = 0.3, 3SRank can achieve 71.25% accuracy while using only 16% of the time required to accumulate all word similarities.
Aoxiang Gao, Lianyin Jia, Mengjuan Li, Runxin Li, Haihe Zhou, Xinming Xing
INDIN3
2025 Exploring contrastive learning and CLIP for improving image clustering
Mengjuan Li, Wenming Cao 0002, Zhiwen Yu 0002, Hangjun Che
Inf. Sci.1
2025 Efficient group based Hilbert encoding and decoding algorithms
Lianyin Jia, Songyu Wang, Shaowen Sun, Jiaman Ding, Mengjuan Li, Jinguo You, Shaojie Qiao
Pattern Recognit.5
2023 ATSA: An Adaptive Tree Seed Algorithm based on double-layer framework with tree migration and seed intelligent generation
Jianhua Jiang, Mengjuan Li, Taibo Chen
Knowl. Based Syst.3
2022 The Extreme Counts: Modeling the Performance Uncertainty of Cloud Resources with Extreme Value Theory
Mengjuan Li, Jinshu Su, Hongyun Liu, Zhiming Zhao, Xue Ouyang 0003, Huan Zhou 0006
ICSOC1
2015 Fast T-overlap query algorithms using graphics processor units and its applications in web data query
Mengjuan Li, Lianyin Jia, Jinguo You, Jianqing Xi, HaiFei Qin
World Wide Web1
2012 ETI: an efficient index for set similarity queries
Lianyin Jia, Jianqing Xi, Mengjuan Li, Decheng Miao
Frontiers Comput. Sci.3