EDBT 2026 Demo / reviewers in the wild / expert
Feng Ye 0004
dblp:70/6154-4
· DBLP profile ↗
12ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-0005-2073ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Research on the Digital Twin Polder Area System Driven by Integrating the Xin'anjiang Model and the N-BEATS ModelabstractDigital twins are propelling the next generation of the industrial revolution and serve as a key technology in enabling intelligent water conservancy. However, due to the diversity of objects within water conservancy scenarios and the complexity of related factors, research and application of digital twins in the field of water conservancy remain immature. There are still significant challenges in constructing fine‐grained, high‐fidelity digital twin for water conservancy objects and their corresponding scenarios. In this context, taking polder areas as research subjects, a digital twin polder area system is proposed, which includes the data representation of the main elements in the polder area; based on 3D engine Unity, the modeling and rendering of the polder area’s terrain, water body, water conservancy projects, and different weather conditions are achieved; the Xin’anjiang model, N‐BEATS model, and the feature engineering model we proposed are integrated to predict water level and flow rate, thereby driving the visual scenario to simulate the extent of the impact of waterlogging at different future moments. Based on satellite imagery data, actual water level data, and flow rate data from a polder area in the Lixiahe river network of Jiangsu Province in China, we measure the efficiency of scene rendering and the accuracy of the prediction models. The results show that the performance and model accuracy of the digital twin polder area system meet the practical requirements. It is more comprehensive by comparing with other works, which can be used as a reference for the construction of a digital twin system or scenario in the water conservancy field. Feng Ye 0004, Zishuo Jin |
Int. J. Intell. Syst. | 1 |
| 2024 | LLM-Based Digital Twin Water Conservancy Knowledge Graph ConstructionabstractIn the study, a novel method for extracting domain-specific knowledge in digital twin water conservancy construction is proposed, addressing the challenges of interdisciplinary complexity and the limitations in existing knowledge extraction models' understanding capabilities of domain-specific knowledge. Utilizing a large language model, the novel method incorporates the domain knowledge of digital twin water conservancy into a local model framework. It employs prompt-based fine-tuning and leverages the semantic understanding and generative capabilities of the model to enhance knowledge extraction precision. To optimize entity extraction, innovative heterogeneous entity alignment strategies are introduced. The efficacy of the proposed method is validated through comparative and ablation studies conducted on a water domain corpus. Results indicate a significant improvement over conventional models, with F1 scores for entity and relationship extraction at 88.63% and 84.46%, respectively, and an entity extraction precision rate of 90.11%. Ablation experiments further demonstrate that our method outperforms the baseline large language model, enhancing the F1 scores for entity and relation extraction by 5.5 and 3.2 percentage points, respectively. By adjusting the amount of domain knowledge injected, the influence of the amount of domain text on the knowledge extraction effect is explored. Feng Ye 0004 |
ICTAI | 2 |
| 2024 | Digital Twin System for River Network Based on Graph Neural NetworkabstractRiver network management presents a significant challenge in the field of water conservancy, and the subsequent advancements in digital technology offer a promising avenue for addressing this challenge.However, the utilization of cutting-edge technologies like digital twin and deep learning within the water conservancy domain remains largely experimental, hampered by the absence of applicable models, authentic scenarios, and viable solutions.To address this, we propose a river network prediction and response system based on digital twin.This system leverages an improved graph attention network to discern spatial correlations within the detection data of individual sites, integrating bidirectional long short-term memory neural network for predictive analysis.Concurrently, we have developed a suite of dynamic visualization tools designed to streamline real-time situation analysis for users.Through the establishment of an interactive digital twin river network system, we furnish a fresh strategic framework for addressing water conservancy challenges. Zishuo Jin, Feng Ye 0004 |
SEKE | 2 |
| 2023 | UrbanFloodKG: An Urban Flood Knowledge Graph System for Risk AssessmentabstractIncreasing numbers of people live in flood-prone areas worldwide. With continued development, urban flood will become more frequent, which has caused casualties and property damage. Researchers have been dedicating to urban flood risk assessments in recent years. However, current research is still facing the challenges of multi-modal data fusion and knowledge representation of urban flood events. Therefore, in this paper, we propose an Urban Flood Knowledge Graph (UrbanFloodKG) system that enables KG to support urban flood risk assessment. The system consists of data layer, graph layer, algorithm layer, and application layer, which implements knowledge extraction and storage functions, integrates knowledge representation learning models and graph neural network models to support link prediction and node classification tasks. We conduct model comparison experiments on link prediction and node classification tasks based on urban flood event data from Guangzhou, and demonstrate the effectiveness of the models used. Our experiments prove that the accuracy of risk assessment can reach 91% when using GEN, which provides a a promising research direction for urban flood risk assessment. Yu Wang 0297, Feng Ye 0004, Binquan Li, Gaoyang Jin, Fengsheng Li |
CIKM | 2 |
| 2023 | Workload-Aware Performance Tuning for Multimodel Databases Based on Deep Reinforcement LearningabstractCurrently, multimodel databases are widely used in modern applications, but the default configuration often fails to achieve the best performance. How to efficiently manage and tune the performance of multimodel databases is still a problem. Therefore, in this study, we present a configuration parameter tuning tool MMDTune+ for ArangoDB. First, the selection of configuration parameters is based on the random forest algorithm for feature selection. Second, a workload‐aware mechanism is based on k‐means++ and the Pearson correlation coefficient to detect workload changes and match the empirical knowledge of historically similar workloads. Finally, the ArangoDB configuration parameters are optimized based on the improved TD3 algorithm. The experimental results show that MMDTune+ can recommend higher‐quality configuration parameters for ArangoDB compared to OtterTune and CDBTune in different scenarios. Feng Ye 0004, Nadia Nedjah |
Int. J. Intell. Syst. | 2 |
| 2023 | A Benchmark for Performance Evaluation of a Multi-Model Database vs. Polyglot PersistenceabstractAs the need for handling data from various sources becomes crucial for making optimal decisions, managing multi-model data has become a key area of research. Currently, it is challenging to strike a balance between two methods: polyglot persistence and multi-model databases. Moreover, existing studies suggest that current benchmarks are not completely suitable for comparing these two methods, whether in terms of test datasets, workloads, or metrics. To address this issue, the authors introduce MDBench, an end-to-end benchmark tool. Based on the multi-model dataset and proposed workloads, the experiments reveal that ArangoDB is superior at insertion operations of graph data, while the polyglot persistence instance is better at handling the deletion operations of document data. When it comes to multi-thread and associated queries to multiple tables, the polyglot persistence outperforms ArangoDB in both execution time and resource usage. However, ArangoDB has the edge over MongoDB and Neo4j regarding reliability and availability. Feng Ye 0004, Xinjun Sheng, Nadia Nedjah |
J. Database Manag. | 1 |
| 2023 | Parameters tuning of multi-model database based on deep reinforcement learningabstractAbstract As we all know, the performance of database management system is directly linked to a vast array of knobs, which control various aspects of system operation, ranging from memory and thread counts settings to I/O optimization. Improper settings of configuration parameters are shown to have detrimental effects on performance, reliability and availability of the overall database management system. This is also true for multi-model databases, which use a single platform to support multiple data models. Existing approaches for automatic DBMS knobs tuning are not directly applicable to multi-model databases due to the diversity of multi-model database instances and workloads. Firstly, in cloud environment, they have difficulty adapting to changing environments and diverse workloads. Secondly, they rely on large-scale high-quality training samples that are difficult to obtain. Finally, they focus primarily on throughput metrics, ignoring tuning requirements for resource utilization. Therefore, in this paper, we propose a multi-model database configuration parameters tuning solution named MMDTune. It selects influential parameters, recommends the optimal configurations in a high-dimensional continuous space. For different workloads, the TD3 algorithm is improved to generate reasonable parameter adjustment plans according to the internal state of the multi-model databases. We conduct extensive experiments under 5 different workloads on real cloud databases to evaluate MMDTune. Experimental results show that MMDTune adapts well to a new hardware environment or workloads, and significantly outperforms the representative tuning tools, such as OtterTune, CDBTune. Feng Ye 0004, Nadia Nedjah |
J. Intell. Inf. Syst. | 1 |
| 2019 | Research on Index Mechanism of HBase Based on Coprocessor for Sensor DataabstractIn order to provide effective management of big data, almost two hundred different NoSQL stores have been developed, among which HBase is one of the best known. When performing data queries, the native HBase supports primary key indexes well. For non-primary key data, only the full table scan can be used, which greatly reduces the multi-condition query speed of HBase. Some additional indexing techniques have been presented to support querying on non-key indexing for HBase. For large-scale sensor data, this research field still requires effective design paradigm, in-depth experimentations and practical implementations. Therefore, we propose and implement a secondary memory index mechanism of HBase based on coprocessor for sensor data. The index mechanism can be automatically updated according to the change of the HBase table. Meanwhile, the index mechanism is persisted and maintained in memory, which can greatly improve the retrieval speed of the index data. By using real sensor datasets with different distribution patterns, experimental results show that the condition retrieval speed of the mechanism is greatly improved compared with the original HBase. In addition, compared with the secondary index mechanism based on Solr and Hibase, the performance of our proposed solution is also improved. Feng Ye 0004, Songjie Zhu, Yuansheng Lou, Yong Chen 0009, Qian Huang 0008 |
COMPSAC (1) | 1 |
| 2019 | Research of Benchmarking and Selection for TSDB
Feng Ye 0004, Songjie Zhu, Yong Chen 0009 |
ICA3PP (2) | 1 |
| 2019 | Hydrological stream data pipeline framework based on IoTDB
Yuansheng Lou, Feng Ye 0004, Yong Chen 0009 |
Serv. Oriented Comput. Appl. | 3 |
| 2018 | The Research of a Lightweight Distributed Crawling SystemabstractNowadays, information on the Internet is growing at an explosive rate. The ability of the stand-alone web crawling system has come to its bottleneck, so more and more companies turn to distributed web crawling techniques. However, existing distributed web crawling systems have some shortcomings. Thread management modules for solving thread synchronization and resource competition are usually designed by using pure multithread asynchronous methods, but the execution of this kind of modules observably reduces the performance. Moreover, the deduplication algorithms lead to low efficiency in dealing with large data sets or the problem of occupying large storage space. To solve the problems mentioned above, this paper proposes a lightweight and practical distributed crawling system, which combines Docker and distributed computing techniques. It can make full use of the computing resources of the cluster and improve the efficiency of the crawling system effectively. Taking the data of Netease news page as an example, the experimental results show that the distributed crawler proposed has higher execution efficiency. Feng Ye 0004, Zongfei Jing, Qian Huang 0008, Yong Chen 0009 |
SERA | 1 |
| 2017 | Research on Anomaly Pattern Detection in Hydrological Time SeriesabstractThe abnormal patterns in hydrological time series play an important role in the analysis and decision-making. Aiming at the problems that the amount of hydrological data is large and there is a lot of “noise” in this data, which lead to the high time complexity of traditional anomaly detection algorithm, we propose anomaly pattern detection based on density for hydrological time series. Firstly, this method makes a piecewise linear representation of the sequence through the important feature points, then extracts the slope, length and mean of the pattern, and maps them to the three-dimensional space. Finally, it calculates the local outlier factor of each pattern. The selection of important feature points and parameters in the algorithm are discussed and verified by the actual data which are historical water level of Jin-niu mountain reservoir. Experimental results show that the algorithm has low complexity and it has full mining results, which can meet the requirements of large-scale time series. Jianshu Sun, Yuansheng Lou, Feng Ye 0004 |
WISA | 3 |