EDBT 2026 Demo / reviewers in the wild / expert
Lianyin Jia
dblp:201/1796
· DBLP profile ↗
22ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-0269-9017ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSC-Join: An Efficient Syntactic-Semantic Collaboration Based Set Semantic Similarity Join Algorithm
Lianyin Jia, Chengchen Zeng, Mengjuan Li, Suprio Ray, Jiaman Ding, Xiuxing Li |
ICDE | 1 |
| 2026 | Efficient Query Region Expansion and Decomposition based Spatial Range Query AlgorithmabstractSpatial range queries play a crucial role in spatial information retrieval. Existing Z-order curve based algorithms suffer from accessing a large number of invalid points outside the query region. To address this challenge, we design a simple yet efficient Z-order curve based learned index, ZPI. Building upon ZPI, we propose a novel spatial range query algorithm, ZPI-RQ. ZPI-RQ leverages efficient decomposition mechanism to address the invalid points issue. To avoid decomposing the query region into a large number of overly small blocks, a query region expansion strategy is further introduced to align each query border with a m-order dividing lines. Experimental results show that ZPI-RQ significantly outperforms state-of-the-art algorithms in query efficiency, achieving a 3.6× improvement over traditional Z-order curve-based algorithms while accessing only 1% invalid points. Lianyin Jia, Rongjin Wang, Yingbin Su, Suprio Ray, Mengjuan Li, Jiaman Ding |
SIGIR | 1 |
| 2026 | An Asymmetric-difference based Document Exact Similarity Search Algorithm
Lianyin Jia, Yongxue Zhao, Mengjuan Li, Xiuxing Li, Jiaman Ding |
SIGIR | 1 |
| 2026 | A single domain generalization fault diagnosis method based on multi-scale style enhancement and causal contribution alignmentabstractIn recent years, single-domain generalization(SDG) fault diagnosis has become a prominent research focus in intelligent fault diagnosis due to its ability to generalize to previously unseen target domains based solely on a single source domain. The primary aim of domain generalization is to identify the intrinsic invariances underlying diverse data distributions, which have been found to be closely related to causality. While most existing fault diagnosis methods based on causal inference emphasize the invariance of causal features across domains, this study considers a stronger form of stability—namely, the cross domain consistency of features’ causal contributions to fault labels. Accordingly, a novel fault diagnosis method is proposed, which integrates Multi-scale Style Enhancement (MSSE) with Causal Contribution Alignment (CCA) to achieve SDG. First, to make up for the lack of data diversity in the source domain, domain shifts are simulated and diverse pseudo-domain samples are generated using a MSSE module. Second, causal contributions of features to diagnostic labels are quantified through causal attribution. Finally, the alignment of causal contributions of features between source and pseudo domains is enforced through contrastive learning and domain adversarial training, thereby promoting stable and cross domain invariant causal representations. Comprehensive experimental evaluations on two benchmark datasets verify that the proposed method consistently outperforms existing fault diagnosis approaches. Jiaman Ding, Jiachen Luo, Lianyin Jia, Hongbin Wang 0002, Xiaodong Fu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | HMSNet: Hilbert curve enhanced Mamba for real-time semantic segmentation
Lianyin Jia, Aoxiang Gao, Mengjuan Li, Xiaodong Fu, Haihe Zhou, Jiaman Ding |
Pattern Recognit. | 1 |
| 2026 | Parameter broadcasting on participant graph for federated heterogeneous graph learning
Juncheng Pu, Xiaodong Fu, Li Liu 0032, Lianyin Jia |
J. Supercomput. | 4 |
| 2026 | HQT-TI: An Efficient Hilbert Curve Based Index for Spatial Keyword QueriesabstractThis paper introduces HQT-TI, a novel indexing method designed to improve the efficiency of spatial keyword queries. HQT-TI consists of two main components: a Hilbert QuadTree (HQT) based spatial index and a Trie-Inverted index (TI) combined textual index. HQT integrates the Hilbert curve with a Quadtree, establishing a direct relationship between the two. TI combines a trie and inverted index to minimize the intersection cost associated with long lists, thus improving the speed of keyword queries. The HQT based Spatial Query algorithm (HQT-SQ) reduces overlap checks and limits irrelevant object retrieval by employing query drill-down and depth first search with limited breadth expansion in spatial queries. Meanwhile, the Segment List Intersection based Keyword Query algorithm (SLI-KQ), built on TI, efficiently handles segment list intersections for keyword queries. The combination of HQT-SQ and SLI-KQ results in HS-SK, a highly efficient spatial keyword query algorithm. Extensive experimental results demonstrate that HS-SKQ outperforms SFC-Quad by up to two orders of magnitude, achieving up to a 5.46× speedup over the best existing competitors, making it a promising solution for large-scale spatial keyword query processing. Lianyin Jia, Yongwang Miao, Suprio Ray, Jiaman Ding, Xiaodong Fu, Xiuxing Li |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | A Length Enhanced B+-Tree Based Index for Efficient Set Similarity QueryabstractSet Similarity Query (SSQ) is widely applied in various fields. The existing B+-tree-based SSQ approaches fail to fully exploit length filtering and require calculating similarity bounds in a node-wise manner, leading to low efficiency. To address these issues, we propose LeB, a novel length-enhanced B+-tree index, whose keys integrate set lengths and bucket mapping, enabling the direct pruning of sets that do not meet the length requirements. Building upon LeB, we present an efficient algorithm, LeBQ, which leverages length filtering and symmetric difference allocation to determine the key bounds for a query, enabling the key bounds computation only once for each query$Q$and avoiding costly similarity bounds computation in a node-wise manner. Efficient key filtering strategies are proposed to prune sets that cannot be similar, significantly reducing the number of candidates. Based on LeBQ, LeBQ+ further reduces the number of candidates by introducing length-independent key bounds. Experimental results on four real datasets demonstrate that LeBQ+ has a higher node access efficiency and accesses only 3.08% to 27.47% nodes compared to the existing B+-tree-based SSQ algorithm. LeBQ+is up to 99.8 × faster than the state-of-the-art algorithms. Lianyin Jia, Shiqi Luo, Jiaman Ding, Suprio Ray, Mengjuan Li, Xiuxing Li |
ICDE | 1 |
| 2025 | Dynamic Reputation Measurement of Online Services for Maximizing User Group Satisfaction
Hedan Zheng, Xiaodong Fu, Li Liu 0032, Jiaman Ding, Lianyin Jia |
ICSOC (1) | 6 |
| 2025 | 3SRank: An Effective Method for Measuring Semantic Similarity StrengthabstractThe measurement of semantic similarity between words is crucial in industrial big data analysis and many other fields. However, most existing work focuses on measuring the semantic similarity between two words, other than individual words. In calculating the semantic similarity strength(3S) of a word w, a naive algorithm requires accumulating the semantic similarity between w and all other words, which is inefficient. To address this issue, we designed a Compact Concept Taxonomy Tree (CCTT) and a novel method for measuring 3S—3SRank.3SRank only accumulates the similarity scores of nodes with a similarity to w greater than a specified threshold θ, and converts the problem of finding nodes that meet the threshold into a problem of finding nodes within specific depth bounds. This significantly reduces the cost of 3S calculation. Extensive experimental results demonstrate that 3SRank can effectively measure the 3S of words and achieve a good balance between measurement accuracy and algorithm execution time. When threshold θ = 0.3, 3SRank can achieve 71.25% accuracy while using only 16% of the time required to accumulate all word similarities. Aoxiang Gao, Lianyin Jia, Mengjuan Li, Runxin Li, Haihe Zhou, Xinming Xing |
INDIN | 2 |
| 2025 | Multi-domains personalized local differential privacy frequency estimation mechanism for utility optimizationabstractLocal Differential Privacy (LDP) has garnered considerable attention in recent years because it does not rely on trusted third parties and has low interactivity and high operational efficiency. However, current LDP frequency estimation mechanisms aggregate data using different privacy budgets within the same domain of attribute values, overlooking the aggregation requirements across different domains of attribute values. This limits the potential for enhancing the data utility under fixed privacy budgets and meeting user preferences in multiple domains of attribute values and privacy budgets. To address this issue, we define a Multi-Domains Personalized Local Differential Privacy (MDPLDP) model that allows users to freely choose domains of attribute values and privacy budgets according to their privacy preferences. Furthermore, based on the MDPLDP model, two new frequency estimation mechanisms are proposed: MDPLDP-Generalized Randomized Response and MDPLDP-basic Randomized Aggregatable Privacy-Preserving Ordinal Response. These mechanisms support cross-domains data aggregation and optimize data utility by adjusting the domains of attribute values and increasing privacy budgets. Theoretical analysis reveals that these new mechanisms have lower estimation errors than the traditional LDP mechanisms. Experiments on real and synthetic datasets demonstrate that the proposed mechanisms effectively reduce estimation errors and enhance the utility of data-frequency estimation. Xiaodong Fu, Li Liu 0032, Jiaman Ding, Wei Peng 0004, Lianyin Jia |
Comput. Secur. | 6 |
| 2025 | Semi-supervised multi-label feature selection combining nonlinear manifold structure and minimizing group sparse redundant correlation
Runxin Li, Xiaowu Li, Guofeng Shu, Lianyin Jia, Zhenhong Shang |
Expert Syst. Appl. | 5 |
| 2025 | Efficient group based Hilbert encoding and decoding algorithms
Lianyin Jia, Songyu Wang, Shaowen Sun, Jiaman Ding, Mengjuan Li, Jinguo You, Shaojie Qiao |
Pattern Recognit. | 1 |
| 2024 | Multi-label feature selection based on minimizing feature redundancy of mutual information
Gaozhi Zhou, Runxin Li, Zhenhong Shang, Xiaowu Li, Lianyin Jia |
Neurocomputing | 5 |
| 2024 | Noisy feature decomposition-based multi-label learning with missing labels
Jiaman Ding, Lianyin Jia, Xiaodong Fu |
Inf. Sci. | 3 |
| 2024 | Semi-supervised multi-label dimensionality reduction learning based on minimizing redundant correlation of specific and common features
Runxin Li, Gaozhi Zhou, Xiaowu Li, Lianyin Jia, Zhenhong Shang |
Knowl. Based Syst. | 4 |
| 2019 | A Parallel Uncertain Frequent Itemset Mining Algorithm with SparkabstractFrequent Itemset Mining (FIM) from large-scale databases has emerged as an important problem in the data mining and knowledge discovery research community. However, FIM suffers from three important limitations with the rapidly expanding of big data in all domains. First, it assumes that all items have the same importance. Second, it ignores the fact that data collected in a real-life environment is often inaccurate. Third, it is also a data-intensive and computation-intensive process which makes the FIM algorithm very time-consuming over large datasets. To address these issues, we propose a Parallel uncertain frequent itemset mining algorithm with spark (Pufim). Pufim firstly expresses item uncertainty by considering both the probability and weight, and calculates the maximum probability weight value of 1-items. Next, a distributed Pufim-tree structure is designed inspiring by FP-Tree for reducing the times of scanning the databases. Each node of Pufim-tree stores an item and its maximum probability weight value. Finally, experiments on publicly available UCI datasets demonstrate that Pufim achieves more prominent results than other related approaches across various metrics. In addition, the empirical study also shows Pufim has a good scalability. Jiaman Ding, Lianyin Jia, Jinguo You |
PDCAT | 4 |
| 2019 | Drug repositioning based on individual bi-random walks on a heterogeneous networkabstractBACKGROUND: Traditional drug research and development is high cost, time-consuming and risky. Computationally identifying new indications for existing drugs, referred as drug repositioning, greatly reduces the cost and attracts ever-increasing research interests. Many network-based methods have been proposed for drug repositioning and most of them apply random walk on a heterogeneous network consisted with disease and drug nodes. However, these methods generally adopt the same walk-length for all nodes, and ignore the different contributions of different nodes. RESULTS: In this study, we propose a drug repositioning approach based on individual bi-random walks (DR-IBRW) on the heterogeneous network. DR-IBRW firstly quantifies the individual work-length of random walks for each node based on the network topology and knowledge that similar drugs tend to be associated with similar diseases. To account for the inner structural difference of the heterogeneous network, it performs bi-random walks with the quantified walk-lengths, and thus to identify new indications for approved drugs. Empirical study on public datasets shows that DR-IBRW achieves a much better drug repositioning performance than other related competitive methods. CONCLUSIONS: Using individual random walk-lengths for different nodes of heterogeneous network indeed boosts the repositioning performance. DR-IBRW can be easily generalized to prioritize links between nodes of a network. Yuehui Wang, Maozu Guo 0001, Yazhou Ren 0001, Lianyin Jia, Guoxian Yu |
BMC Bioinform. | 4 |
| 2016 | A Dynamic Migration Method for Big Data in CloudabstractBig data applications store data sets through sharing data center under the Cloud computing environment, but the need of data set in big data applications is dynamic change over time. In face of multiple data centers, such applications meet new challenges in data migration which mainly include how to how to reduce the number of network access, how to reduce the overall time consumption, and how to improve the efficiency by the time of balancing the global load in the migration process. Facing these challenges, we first build the problem model and descript the dynamic migration method, then solve the global time consumption of data migration, the number of network access and global load balancing these three parameters. Finally, do the cloud computing simulation experiment under the Cloudsim experiment platform. The result shows that the proposed method makes the task completion time reduced by 10% and the data transmission time accounts for the roportion of the total time is reduced. When the amount of data sets is increase, the proportion can reduces to 50% or less. Network access number lower than Zipf and reached stable, in global load, the variance of the node's store space closed to zero. Jiaman Ding, Sichen Wang, Lianyin Jia |
PDCAT | 4 |
| 2015 | Fast T-overlap query algorithms using graphics processor units and its applications in web data query
Mengjuan Li, Lianyin Jia, Jinguo You, Jianqing Xi, HaiFei Qin |
World Wide Web | 2 |
| 2012 | ETI: an efficient index for set similarity queries
Lianyin Jia, Jianqing Xi, Mengjuan Li, Decheng Miao |
Frontiers Comput. Sci. | 1 |
| 2010 | Double Table Switch: An Efficient Partitioning Algorithm for Bottom-Up Computation of Data Cubes
Jinguo You, Lianyin Jia, Qingsong Huang, Jianqing Xi |
ADMA (2) | 2 |