EDBT 2026 Demo / reviewers in the wild / expert
Shengfei Shi
dblp:67/3479
· DBLP profile ↗
27ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-2932-4618ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 61% Memory systems · 30% Distributed systems · 8% | |
| Databases, data mining, and information retrieval
3 papers |
Indexing and storage engines · 89% Database system architecture and tuning · 4% Query processing and optimization · 4% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines › spatial index
r-tree |
0.9 | 1 | 2025 | Hybrid DRAM-NVM R-Trees with Consistency Guarantee · ICDE 2025 |
Indexing and storage engines
spatial index |
0.9 | 1 | 2025 | Hybrid DRAM-NVM R-Trees with Consistency Guarantee · ICDE 2025 |
Storage systems
crash consistency |
0.9 | 1 | 2025 | Hybrid DRAM-NVM R-Trees with Consistency Guarantee · ICDE 2025 |
Memory systems
non-volatile memory |
0.9 | 1 | 2025 | Hybrid DRAM-NVM R-Trees with Consistency Guarantee · ICDE 2025 |
Storage systems
storage reliability |
0.9 | 1 | 2025 | Hybrid DRAM-NVM R-Trees with Consistency Guarantee · ICDE 2025 |
Internet of things and sensor networks › wireless sensor network
sensor network simulation |
0.1 | 1 | 2010 | Data-enriched simulation of data management applications for wireless sensor networks · SenSys 2010 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2007 | InfiniteDB: a pc-cluster based parallel massive database management system · SIGMOD Conference 2007 |
Query processing and optimization
parallel query processing |
0.1 | 1 | 2007 | InfiniteDB: a pc-cluster based parallel massive database management system · SIGMOD Conference 2007 |
Distributed systems › distributed communication
data dissemination |
0.1 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Distributed systems
peer-to-peer systems |
0.1 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Distributed systems › peer-to-peer systems › overlay networks
structured overlay |
0.1 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Audio and music processing › music information retrieval
lyrics analysis |
0.0 | 1 | 2010 | Structure-aware music resizing using lyrics · WWW 2010 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2010 | Data-enriched simulation of data management applications for wireless sensor networks · SenSys 2010 |
Information retrieval › similarity search
approximate similarity search |
0.0 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Data integration and cleaning
data warehouse |
0.0 | 1 | 2007 | InfiniteDB: a pc-cluster based parallel massive database management system · SIGMOD Conference 2007 |
Information retrieval
similarity search |
0.0 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Wireless networking
mobile ad hoc networks |
0.0 | 1 | 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007 |
Methods — techniques the papers use, named apart from their topics
persistence operations · 1.7hilbert curve · 1.7statistical device emulation · 0.2dataset integration · 0.2wavelet transform · 0.2k-means clustering · 0.2structure-aware compression · 0.1data declustering · 0.1coordinator-wrapper · 0.1adaptive query optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GCTKG: Group center for keyword search over knowledge graphs
Xiaoxiao Xie, Shengfei Shi, Chao Yi |
Inf. Sci. | 2 |
| 2025 | Hybrid DRAM-NVM R-Trees with Consistency GuaranteeabstractThe non-volatile memory (NVM) with DRAM-like performance and disk-like persistency has attracted considerable attention in a variety of index structures, including hash table, B-Tree and R-Tree. However, existing NVM-optimized consistent R-Tree is still suboptimal because its single level system neglects the potential boost that DRAM can bring. In this paper, we first propose a hybrid DRAM-NVM consistent R-Tree (HR-Tree), which separately stores internal nodes in DRAM and leaf nodes in NVM. To avoid inconsistency, HR-Tree uses several auxiliary flag bits and pointers to record the process of writes to NVM and employs persistence operations to strictly control the order of writes to NVM. To reduce DRAM consumption, which mainly depends on the metadata size of a leaf node, we present a shared byte strategy to abolish restrictions on metadata size while still keeping HR-Tree consistency. Next, for further shortening search time, we propose an alternative Hilbert-curve-based hybrid R-Tree (HHR-Tree). It has better search efficiency yet leads to insertion performance degradation. Contrary to in-place update in HR-Tree, HHR-Tree applies out-of-place mechanism to enforce data consistency. We conduct comprehensive evaluations on Intel Optane DC Persistent Memory. The proposed HR-Tree outperforms FBR-Tree in terms of insertion, deletion and search throughput while HHR-Tree exhibits a significant improvement for search performance by sacrificing insertion efficiency. Chengyou Shen, Shengfei Shi, Hong Gao 0001, Yaofeng Tu |
ICDE | 4 |
| 2023 | SAT: sampling acceleration tree for adaptive database repartition
Xiaoxiao Xie, Shengfei Shi, Hongzhi Wang 0001, Mohan Li |
World Wide Web (WWW) | 2 |
| 2022 | Partial multi-label learning via specific label disambiguation
Shengfei Shi, Hongzhi Wang 0001 |
Knowl. Based Syst. | 2 |
| 2021 | Column concept determination based on multiple evidencesabstractSummary Tables on the web provide rich information. To make sufficient usage of web tables, the semantics of columns should be identified correctly. The absence, misspelling, and abbreviation in column names bring the challenges in column semantics identification. Facing this challenge, we extract multiple features including keywords, concepts, and structure from the content in the column. Thus, we could identify the column semantics by matching these multiple features. For the extraction and matching with these features, we propose efficient algorithms. Experimental results on real data sets show that our solution achieves high performance. Xianxi An, Sihan You, Zeguang Lu, Shengfei Shi |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | Semi-supervised multi-label feature selection with adaptive structure learning and manifold learning
Sitao Lv, Shengfei Shi, Hongzhi Wang 0001 |
Knowl. Based Syst. | 2 |
| 2019 | FreshJoin: An Efficient and Adaptive Algorithm for Set Containment JoinabstractAbstract This paper revisits set containment join (SCJ) problem, which uses the subset relationship (i.e., $$\subseteq$$ ⊆ ) as condition to join set-valued attributes of two relations and has many fundamental applications in commercial and scientific fields. Existing in-memory algorithms for SCJ are either signature-based or prefix-tree-based. The former incurs high CPU cost because of the enumeration of signatures, while the latter incurs high space cost because of the storage of prefix trees. This paper proposes a new adaptive parameter-free in-memory algorithm, named as frequency-hashjoin or $${\mathsf {FreshJoin}}$$ FreshJoin in short, to evaluate SCJ efficiently. $${\mathsf {FreshJoin}}$$ FreshJoin builds a flat index on-the-fly to record three kinds of signatures (i.e., two least frequent elements and a hash signature whose length is determined adaptively by the frequencies of elements in the universe set). The index consists of two sparse inverted indices and two arrays which record hash signatures of all sets in each relation. The index is well organized such that $${\mathsf {FreshJoin}}$$ FreshJoin can avoid enumerating hash signatures. The rationality of this design is explained. And, the time and space cost of the proposed algorithm, which provide a rule to choose $${\mathsf {FreshJoin}}$$ FreshJoin from existing algorithms, are analyzed. Experiments on 16 real-life datasets show that $${\mathsf {FreshJoin}}$$ FreshJoin usually reduces more than 50% of space cost while remains as competitive as the state-of-the-art algorithms in running time. Jizhou Luo, Wei Zhang 0017, Shengfei Shi, Hong Gao 0001, Jianzhong Li 0001, Shouxu Jiang |
Data Sci. Eng. | 3 |
| 2018 | A gray-box performance model for Apache Spark
Zemin Chao, Shengfei Shi, Hong Gao 0001, Jizhou Luo, Hongzhi Wang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | O2iJoin: An Efficient Index-Based Algorithm for Overlap Interval Join
Jizhou Luo, Shengfei Shi, Hongzhi Wang 0001, Jianzhong Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2018 | Data management on new processors: A survey
Hongzhi Wang 0001, Shengfei Shi, Jianzhong Li 0001, Hong Gao 0001 |
Parallel Comput. | 4 |
| 2017 | FrepJoin: an efficient partition-based algorithm for edit similarity joinabstractString similarity join (SSJ) is essential for many applications where near-duplicate objects need to be found. This paper targets SSJ with edit distance constraints. The existing algorithms usually adopt the filter-andrefine framework. They cannot catch the dissimilarity between string subsets, and do not fully exploit the statistics such as the frequencies of characters. We investigate to develop a partition-based algorithm by using such statistics. The frequency vectors are used to partition datasets into data chunks with dissimilarity between them being caught easily. A novel algorithm is designed to accelerate SSJ via the partitioned data. A new filter is proposed to leverage the statistics to avoid computing edit distances for a noticeable proportion of candidate pairs which survive the existing filters. Our algorithm outperforms alternative methods notably on real datasets. Jizhou Luo, Shengfei Shi, Hongzhi Wang 0001, Jianzhong Li 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2014 | Multi-Hierarchies: Accurately Computing Realtime Statistical Measures on Data Streams
Penghe Qi, Shengfei Shi |
WASA | 2 |
| 2011 | Efficient Computation of Measurements of Correlated Patterns in Uncertain Data
Lisi Chen 0001, Shengfei Shi, Jing Lv |
ADMA (1) | 2 |
| 2010 | Data-enriched simulation of data management applications for wireless sensor networksabstractSimulation is an essential means for evaluating WSN applications. As many WSN applications embrace in-network data processing functionalities, more sophisticated simulation tools with data-enriched test case scenarios, such as extensive environment input and accurate hardware models, are necessary. This work devises TOSSIM DB, a data-enriched WSN simulation framework, allowing unmodified applications to be evaluated as if virtually deployed in a third-party environment. Notable features of TOSSIM DB include i) integrating third-party sensor datasets as virtual test scenarios, ii) emulating the statistical properties of sensor devices, and iii) seamless integration with the TinyOS tool-chain. Yu Liu 0002, Jianzhong Li 0001, Hong Gao 0001, Shengfei Shi |
SenSys | 4 |
| 2010 | Structure-aware music resizing using lyricsabstractWorld wide web provides plenty of multimedia resources for creating rich media web applications. However, the collected music and other media resources always mismatch in the metric of time length. Existent music resizing approaches suffer from perceptual artifacts which degrade the performance of resized music. In this paper, a novel structure-aware music resizing approach is proposed. Through lyrics analysis, our approach can compress different parts of a music piece in variant compression rates. Experimental results show that the proposed method can effectively generate resized songs with good quality. Zhang Liu 0004, Chaokun Wang, Jianmin Wang 0001, Shengfei Shi |
WWW | 5 |
| 2008 | Reliable and Fast Detection of Gradual Events in Wireless Sensor Networks
Liping Peng, Hong Gao 0001, Jianzhong Li 0001, Shengfei Shi, Boduo Li |
WASA | 4 |
| 2007 | Unsupervised Outlier Detection in Sensor Networks Using Aggregation Tree
Kejia Zhang 0001, Shengfei Shi, Hong Gao 0001, Jianzhong Li 0001 |
ADMA | 2 |
| 2007 | Clustering wavelets to speed-up data dissemination in structured P2P MANETsabstractThis paper introduces a fast data dissemination method for structured peer-to-peer networks. The work is motivated on one side by the increase in non-volatile memory available on mobile devices and, on the other side, by observed behavioral patterns of the users. We envision a scenario where users come together for short periods of time (e.g. public transport, conference sessions) and wish to be able to share large collections of data. With hundreds and even thousands of data, items stored on small devices, content publication is simply too energy and time consuming. By indexing summary information obtained by a combination of multi-resolution analysis and k-means, our method (Hyper-Ad) is able to cut down the overall construction time of an overlay network such as CAN by an order of magnitude, as well as provide fast approximate similarity search on such a network. The results of our extensive experimental studies confirm that Hyper-M is both energy and time efficient, and provides good precision and recall. Mihai Lupu, Jianzhong Li 0001, Beng Chin Ooi, Shengfei Shi |
ICDE | 4 |
| 2007 | MuSQL: A Music Structured Query Language
Chaokun Wang, Jianmin Wang 0001, Jianzhong Li 0001, Jia-Guang Sun 0001, Shengfei Shi |
MMM (2) | 5 |
| 2007 | InfiniteDB: a pc-cluster based parallel massive database management systemabstractThis paper describes a PC-cluster based parallel DBMS, InfiniteDB, developed by the authors. InfiniteDB aims at efficiently storing and processing of massive databases in response to the rapidly growing in database size and the need of high performance analyzing of massive databases. It supports the parallelisms of intra-query, inter-query, intra-operation, inter-operation and pipelining. It provides effective strategies for processing massive databases including the multiple data declustering methods, the declustering-aware algorithms for the execution of relational operations and other database operations, and the adaptive query optimization method. It also provides the functions of parallel data warehousing and data mining, the coordinator-wrapper mechanism to support the integration of heterogeneous information resources on the Internet, and the fault tolerant and resilient infrastructures. It has been used in many applications and has proved quite effective for storing and processing massive databases in practice. Jianzhong Li 0001, Hong Gao 0001, Jizhou Luo, Shengfei Shi, Wei Zhang 0017 |
SIGMOD Conference | 4 |
| 2007 | AbIx: An Approach to Content-Based Approximate Query Processing in Peer-to-Peer Data Systems
Chaokun Wang, Jianmin Wang 0001, Jia-Guang Sun 0001, Shengfei Shi, Hong Gao 0001 |
J. Comput. Sci. Technol. | 4 |
| 2006 | N-gram inverted index structures on music data for theme mining and content-based information retrieval
Chaokun Wang, Jianzhong Li 0001, Shengfei Shi |
Pattern Recognit. Lett. | 3 |
| 2004 | Cell Abstract Indices for Content-Based Approximate Query Processing in Structured Peer-to-Peer Data Systems
Chaokun Wang, Jianzhong Li 0001, Shengfei Shi |
APWeb | 3 |
| 2004 | A Music Data Model and its ApplicationabstractA music data model, its query language and its application are proposed in this paper. Firstly, a music data model and its algebraic operations are given, which can be used to describe and manipulate musical data efficiently. Secondly, a structured query language on the model is proposed, which can be used to define and manage musical data. Finally, a digital music library, one of the applications of this model, is presented, which can be used to retrieve musical information, especially against musical instruments. Chaokun Wang, Jianzhong Li 0001, Shengfei Shi |
MMM | 3 |
| 2004 | A Query-Aware Routing Algorithm in Sensor Networks
Jianzhong Li 0001, Shengfei Shi |
NPC | 3 |
| 2004 | TS-Cache: A Novel Caching Strategy to Manage Data in MANET Database
Shengfei Shi, Jianzhong Li 0001, Chaokun Wang |
WAIM | 1 |
| 2002 | Using PR-Tree and HPIR to Manage Coherence of Semantic Cache for Location Dependent Data in Mobile Database
Shengfei Shi, Jianzhong Li 0001, Chaokun Wang |
WAIM | 1 |