Wen Nie

dblp:128/3731 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SSCard: Substring Cardinality Estimation using Suffix Tree-Guided Learned FM-Index
abstract
Accurate cardinality estimation of substring queries, which are commonly expressed using the SQL LIKE predicate, is crucial for query optimization in database systems. While both rule-based methods and machine learning-based methods have been developed to optimize various aspects of cardinality estimation, their absence of error bounds may result in substantial estimation errors, leading to suboptimal execution plans. In this paper, we propose SSCard, a novel S ub S tring Card inality estimator that leverages a space-efficient FM-Index into flexible database applications. SSCard first extends the FM-Index to support multiple strings naturally, and then organizes the FM-index using a pruned suffix tree. The suffix tree structure enables precise cardinality estimation for short patterns and achieves high compression via a pushup operation, especially on a large alphabet with skewed character distributions. Furthermore, SSCard incorporates a spline interpolation method with an error bound to balance space usage and estimation accuracy. Additional innovations include a bidirectional estimation algorithm and incremental update strategies. Extensive experimental results in five real-life datasets show that SSCard outperforms both traditional methods and recent learning-based methods, which achieves an average reduction of 20% in the average q-error, 80% in the maximum q-error, and 50% in the construction time, compared with second-best approaches.
Yirui Zhan, Wen Nie, Jun Gao 0003
Proc. ACM Manag. Data2
2025 GaussDB-Vector: A Large-Scale Persistent Real-Time Vector Database for LLM Applications
abstract
Vector databases are widely used as a fundamental tool for addressing the weaknesses of large language model (LLM) applications, specifically hallucinations and the high cost of inference. However, existing vector databases either cater to niche applications with low-latency in-memory search, or offer sophisticated data management capabilities but at the cost of low performance. To address these limitations, we propose GaussDB-Vector, a high-performance, real-time persistent vector database that excels in low-latency scalable search, real-time inserts and deletes, high availability, large-scale distributed search, and hybrid scalar-vector filtered search capabilities. These features are primarily achieved through an innovative storage architecture designed for a graph-based vector index, optimized for I/O operations and adaptable across various dataset sizes and dimensions, complemented by novel buffering strategies to further reduce I/O burdens. GaussDB-Vector supports product quantization, parallel search, and hardware acceleration via SIMD, GPUs, and NPUs in order to further accelerate queries. Experimental results show that GaussDB-Vector outperforms competitive baselines by a factor of 1 to 5 times.
Guoliang Li 0001, Ji Sun 0001, James Pan, Yongqing Xie, Ruicheng Liu, Wen Nie
Proc. VLDB Endow.7
2024 GaussML: An End-to-End In-Database Machine Learning System
abstract
In-database machine learning (In-DB ML) is appealing to database users with security and privacy concerns, as it avoids copying data out of the database to a separate machine learning system. The common way to implement in-DB ML is the ML-as-UDF approach, which utilizes the User-Defined Functions (UDFs) within SQL to implement the ML training and prediction. However, UDFs may introduce security risks with vulnerable code, and suffer from performance problems, as constrained by data access and execution patterns of SQL query operators. To address these limitations, we propose a new in-database machine learning system, namely GaussML, which provides an end-to-end machine-learning ability with native SQL interface. To support ML training/inference within SQL query, GaussML directly integrates typical ML operators into the query engine without UDFs. GaussML also introduces an ML-aware cardinality and cost estimator to optimize the SQL+ML query plan. Moreover, GaussML leverages Single Instruction Multiple Data (SIMD) and data prefetching techniques to accelerate the ML operators for training. We have implemented a series of algorithms inside GaussML in openGauss database. Compared to the state-of-the-art in-DB ML systems like Apache MADlib, our GaussML achieves 2-6× speed-up in extensive experiments.
Guoliang Li 0001, Ji Sun 0001, Lijie Xu, Shifu Li, Wen Nie
ICDE6
2022 A Deep Reinforcement Learning-Based Framework for PolSAR Imagery Classification
abstract
The deep convolutional neural network (CNN) has been extensively applied to polarimetric synthetic radar (PolSAR) imagery classification. However, its success is greatly dependent on numerous labeled samples for revealing and modeling the characteristics of different targets, thus remaining a challenge in maintaining high accuracy in limited sample cases. To address this issue, a deep reinforcement learning (RL)-based PolSAR image classification framework, named deep Q-fully CNN (DQFCN), is proposed in this article. In this framework, two ways are adopted to boost the classification performance while reducing training samples. On the information utilization hand, the spatial neighboring information and polarimetric decomposition information of PolSAR data are both extracted to enrich the feature representation of the sample. Meanwhile, the 3-D CNN architecture is adopted to learn the spatial-polarimetric jointed characteristics simultaneously. On the model learning hand, two RL learning strategies are employed to promote classification performance. The first one is learning from scratch, which does not use any label information as prior knowledge but learns from its self-generated experience. Learning from pretraining is the second strategy in which the networks are sequentially trained from labeled samples and experience data to reduce the time cost. As far as we know, it is the first time that an RL-based fully CNN has been proposed for PolSAR image classification. Experiments on three benchmark datasets prove the effectiveness of the proposed framework, suggesting that the two adopted strategies achieve boosted performance in all experiments, particularly in a limited sample size.
Wen Nie, Kui Huang, Jie Yang 0040, Pingxiang Li
IEEE Trans. Geosci. Remote. Sens.1