Expert Fit Tier Activity
Papers they would take
Guanhao Hou
99.99+
junior + senior
5 since 20215 in total
RNSG: A Range-Aware Graph Index for Efficient Range-Filtered Approximate Nearest Neighbor Search
BBC: Improving Large-𝑘 Approximate Nearest Neighbor Search with a Bucket-based Result Collector
CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data
CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates
+1 more
Jianyang Gao
99.99+
junior + senior
6 since 20216 in total
Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search
JHQ: Johnson-Lindenstrauss Enhanced Hierarchical Quantization for High-Dimensional Approximate Nearest Neighbor Search
HEXA: A Disjoint-Subgraph-Based Indexing Framework for Approximate Nearest Neighbor Search at Billion Scale
ANNiE: A Learned Query Cost Estimator for Graph-Based Approximate Nearest Neighbor Search
+1 more
in pool
Xinjun Yang
99.99+
junior + senior
15 since 202115 in total
How to Write to SSDs
Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP Systems
BtrLog: Low-Latency Logging for Cloud Database Systems
SunStorm: Geographically distributed transactions over Aurora-style systems
+1 more
in pool
Manos Athanassoulis
99.99+
senior
29 since 202151 in total
ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic Workloads
Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
Operation-Aware Hybrid Locking for Modern In-Memory Indexes
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
+1 more
in pool
Qiangqiang Dai
99.99+
junior + senior
21 since 202125 in total
Revisiting the Maximum Defective Clique Problem: Faster Branching and a Tighter Upper Bound
Aggregating maximal cliques in real-world graphs
Efficient Locally h-Clique Densest Subgraph Discovery via Divide-and-Conquer
Scalable Approximate Biclique Counting over Large Bipartite Graphs
+1 more
Zhuoyue Zhao 0001
99.99+
senior
7 since 202114 in total
Sampling-based Predictive Database Buffer Management
Storing and Indexing Multiple Tables by Interesting Orderings: For Efficient Joins, Groupings, and Updates in Relational Databases
High-Performance DBMSs with io_uring: When and How to Use It
AQD: Online Adaptive Query Dispatcher for HTAP Databases
+1 more
Baoqing Cai
99.99+
junior
2 since 20212 in total
DOT: Dynamic Knob Selection and Online Sampling for Automated Database Tuning
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Libra: One-Shot Parameter Sensitivity Estimation for Transfer Learning in Database Performance Prediction
Redbench: Workload Synthesis From Cloud Traces
+1 more
Bobbi W. Yogatama
99.99+
junior
5 since 20215 in total
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
GPU Acceleration of SQL Analytics on Compressed Data
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
+1 more
Junyong Yang
99.99+
junior
3 since 20213 in total
Effective Durable Community Search in Large Temporal Graph
Efficient Temporal Subgraph Management: A New Interval Index
Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph Representation
Finding Time-Proximity Communities in Temporal Heterogeneous Information Networks
+1 more
Xi He 0001
99.99+
senior
16 since 202130 in total
Fast and Private Max-Sum Diversification
Secure Multi-Party Sampling over Joins
Doppio: Communication-Efficient and Secure Multi-Party Shuffle Differential Privacy
Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
+1 more
Yingfan Liu
99.99+
senior
17 since 202123 in total
PAIL: Efficient kNN Search on Set-Valued Attributes
A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
BBC: Improving Large-𝑘 Approximate Nearest Neighbor Search with a Bucket-based Result Collector
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
+1 more
Per-Åke Larson
99.99+
senior
4 since 202188 in total
Swan: Hybrid MVCC Management for Efficient Transaction Processing in LSM-Tree-Based Key-Value Stores
Tux: Efficient Drop-in Networking for Database Systems
Tidehunter: Large-Value Storage With Minimal Data Relocation
Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability
+1 more
in pool
Lijun Chang
99.99+
junior + senior
33 since 2021108 in total
Maximum Defective Biclique Search in Large Bipartite Graphs
A Practical Sublinear Approximation for Group Steiner Tree
MDS-FSM: Coverage-Based Frequent Subgraph Mining in Single Graphs
Anchored Maximum Communities over Large Directed Graphs
+1 more
Yinan Li 0009
99.99+
junior
6 since 202121 in total
Bridging the Indexing Gap in Fused GPU Query Engines
Vodka: Rethink Benchmarking Philosophy in HTAP Systems
Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression
Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines
+1 more
Zhijie Zhang 0004
99.99
junior + senior
7 since 20217 in total
CEMR: An Effective Subgraph Matching Algorithm with Redundant Extension Elimination
X-Wim: Massive Parallelization of Weighted Matching in Bipartite Graphs
Aquila: A High-Concurrency System for Incremental Graph Query
Mix & Match: Subgraph Matching for Absolute Coverage
+1 more
Samuele Langhi
99.99+
junior
5 since 20216 in total
Fugue: Online Elasticity for Distributed Stateful Stream Processing
Continuous Query for Top-K Maximal Sum Intervals over Streaming Data
Incremental Stream Query Deployment under Continuous Infrastructure Changes in the Cloud-Edge Continuum
Toward Temporal Attribution Analytics in Dataflows
+1 more
Qing Wang 0031
99.99+
senior
30 since 202134 in total
SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous Memory
Terark-DS: A High-Performance and Storage-Efficient Key-Value Separation Storage Engine on Disaggregated Storage
Shard: A Scalable and Resize-optimized Hash Index on Disaggregated Memory
Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services
+1 more
in pool
Xi Zhao 0006
99.99+
junior + senior
17 since 202119 in total
Elastic Index Selection for Label-Hybrid AKNN Search
RED-ANNS: An RDMA-Enabled Distributed Framework for Graph-Based Approximate Nearest Neighbor Search
GAS: A Lightweight Framework for Filtered Search over Wide-table Vectors
RNSG: A Range-Aware Graph Index for Efficient Range-Filtered Approximate Nearest Neighbor Search
+1 more
in pool
Surajit Chaudhuri
dblp:c/SurajitChaudhuri ·
DBLP ↗
99.99+
senior
38 since 2021227 in total
BaCon: Efficient Batch Processing of Counting Queries
Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset Queries
Toward Drift-Aware Database Benchmarking
OBELISK: Efficient Offline Query Planning with Bayesian Optimization-Informed Language Model Reasoning
+1 more
Niv Dayan
99.99+
junior + senior
12 since 202125 in total
STEM2: A Fast and Space-efficient Data Structure for Exact Multi-Set Membership Query
I/O Optimizations in Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration
Rethinking Learned Index and LSM-tree Integration
CounterSnake: A lossless and generalized compression framework for diverse sketches
+1 more
Makoto Imamura
99.98
junior
7 since 202117 in total
MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean Distance
KDSelector: A Framework of Knowledge-Enhanced and Data-Efficient Selector Learning for Anomaly Detection Model Selection in Time Series
CLaP - State Detection from Time Series
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
+1 more
Youri Kaminsky
99.99
junior
5 since 20215 in total
Discovering Approximate Denial Constraints in Large Databases
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
Storage-Centric Relation Design via High-Quality Approximate Functional Dependencies
Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database Construction
+1 more
in pool
Qiange Wang
99.99+
junior + senior
21 since 202122 in total
Efficient GNN Training on Giant Graphs with Collective Batching and Scheduling
FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism
ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks
Resource-Efficient FirmCore Decomposition on Billion-scale Multilayer Graphs
+1 more
Azim Afroozeh
99.99+
junior
5 since 20215 in total
DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point Data
QStore: Quantization-Aware Compressed Model Storage
Morphing-based Compression for Data-centric ML Pipelines
Accelerating String-Heavy Queries with LLM Token Tables
+1 more
Semih Salihoglu
99.99+
senior
17 since 202138 in total
TVA: A Version-aware Temporal Graph Storage System for Real-time Analytics
Dinkel: State-Aware and Granular Framework for Validating Graph Databases
TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
Worst-Case Optimal BGPs on Temporal Graphs
+1 more
Marko Kabic
99.99+
junior
4 since 20215 in total
dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data Processing
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics
+1 more
Mengzhao Wang 0001
99.99+
junior + senior
8 since 20218 in total
Harmonizing Efficiency and Accuracy in Filtered Vector Search
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
An Experimental Evaluation of Hybrid Querying on Vectors
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
+1 more
Qintian Guo
99.98
senior
10 since 202113 in total
Theoretically and Practically Efficient Resistance Distance Computation on Large Graphs
Augmenting Social Influence of Uncertain Seeds via Probabilistic Link Insertion
AGIS: Fast Approximate Graph Pattern Mining with Structure-Informed Sampling
Lower-Bound Distance Queries under Partial Information
+1 more
in pool
Thomas Neumann 0001
99.99+
senior
46 since 2021160 in total
Chipmink: Efficient Delta Identification for Massive Object Graphs
Index Intersection for High-Dimensional Range Queries
Ken: An Execution Engine for Unstructured Database Systems
Window Function Optimization: Co-Evaluation and Other Techniques
+1 more
Ishtiyaque Ahmad
99.99+
junior + senior
9 since 202110 in total
Verifiable Authenticated Data Structure (V-ADS) for Analytic Queries
PrivMDC: Leveraging Multi-Dimensional Correlations to Answer Differentially Private Range Queries
Secure Join Operations in Multi-Identifier Databases: Performance and Practicality
Enabling Index-free Adjacency in Oblivious Graph Processing with Delayed Duplications
+1 more
in pool
Tianzheng Wang 0001
99.99+
senior
19 since 202146 in total
Demystifying and Improving Lazy Promotion in Cache Eviction
How Much Can RocksDB Chew? Achieving Near-Zero Write Stalls with Sustainable RocksDB
Operation-Aware Hybrid Locking for Modern In-Memory Indexes
SunStorm: Geographically distributed transactions over Aurora-style systems
+1 more
Tonghui Ren
99.99+
junior + senior
9 since 20219 in total
Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases
+1 more
in pool
Vivek R. Narasayya
dblp:n/VivekRNarasayya ·
DBLP ↗
99.99+
senior
20 since 202190 in total
Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
Hybrid Mixed Integer Linear Programming for Large-Scale Join Order Optimisation
QDBO: A Real-time Quantum-augmented Database System Optimizer
+1 more
Matin Najafi
99.99
senior
6 since 20217 in total
Efficient Partition-based Approaches for Diversified Top-k Subgraph Matching
gMatch: Fine-Grained and Hardware-Efficient Subgraph Matching on GPUs
Characterizing Parallel Subgraph Matching Performance: A Systematic Study of Interactions, Scalability, and Enumeration
Subgraph Enumeration: Beyond Tree Decomposition
+1 more
in pool
Matteo Interlandi
99.99+
senior
26 since 202157 in total
NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning Applications
TPCx-AI under the Microscope: A Benchmarking Debt Analysis
SEMA: A High-performance System for LLM-based Semantic Query Processing
+1 more
in pool
Ke Yi 0001
99.99
senior
45 since 2021159 in total
On Fair Epsilon Net and Geometric Hitting Set
Towards Efficient Random-Order Enumeration for Join Queries
Unbiased Binning for Fairness-aware Attribute Representation
Efficient and Secure Range Counting over Distributed Geographic Data with Query Range Protection
+1 more
Amélie Gheerbrant
99.99
junior
6 since 202114 in total
Repairing Property Graphs under PG-Constraints
Computing Why-Provenance for Property Graph Queries
A Unified Query Planning Framework for Conjunctive Regular Path Queries
Structural Normalization of Property Graphs
+1 more
Daichi Amagata
99.99+
junior
43 since 202170 in total
RT-RkNN: Reverse k Nearest Neighbor Queries as a Graphics Ray Casting Problem
FB*: A Compact Index for Efficient and Exact Density-based Clustering
Scalable Grid-based Computation of Kendall's Tau Correlation
PAIL: Efficient kNN Search on Set-Valued Attributes
+1 more
Jingzhou Fu
99.99+
senior
17 since 202117 in total
I-Rex: An Interactive Debugger for SQL
FlowLog: Efficient and Extensible Datalog via Incrementality
DBAIOps: A Reasoning LLM-Enhanced Database Operation and Maintenance System using Knowledge Graphs
Fast Verification of Strong Database Isolation
+1 more
Gengrui Zhang 0001
99.95
junior
7 since 20218 in total
Orca: Flexible Quorums Meet Dynamic Quorums
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
Remora: Scale-out Deterministic Execution for Smart Contracts
FairDAG: Consensus Fairness over Multi-Proposer Causal Design
+1 more
in pool
Hongchao Qin
99.99
junior + senior
35 since 202140 in total
TIMEST: Temporal Information Motif Estimator Using Sampling Trees
Effective Durable Community Search in Large Temporal Graph
Efficient Locally h-Clique Densest Subgraph Discovery via Divide-and-Conquer
Aggregating maximal cliques in real-world graphs
+1 more
Raghu Ramakrishnan 0001
dblp:r/RaghuRamakrishnan ·
DBLP ↗
99.99+
senior
6 since 2021182 in total
Interoperable ACID Transactions for Open Table Formats
IncreQueryFusion: On-demand Data Fusion Framework in Dynamic Data Lakes
ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling
Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
+1 more
Hyungsoo Jung 0001
99.99+
senior
9 since 202135 in total
How to Write to SSDs
Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP Systems
BtrLog: Low-Latency Logging for Cloud Database Systems
High-Performance DBMSs with io_uring: When and How to Use It
+1 more
Xiaomeng Yi
99.99+
senior
6 since 20219 in total
Aker: Density-Aware Approximate Caching for Vector Search
CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data
A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
+1 more
Xiang Song 0003
99.99
senior
22 since 202123 in total
NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments
Towards A Generalizable and Expressive Graph Neural Network for Graph-Level Tasks with Theoretical Guarantees
Scalable GNN Explanations with Distributed Shapley Values
UniTG: A Unified System for Efficient and Seamless Textual Graph Learning
+1 more
Jianshun Zhang
99.99+
junior
7 since 20217 in total
LiBox: A Learned Index as an Array to Minimize Last-Mile Search
Succinct and Fast Tiny Pointer Hash Tables
ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic Workloads
Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
+1 more
Chuan Lei
99.99+
senior
15 since 202136 in total
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
Schuyler: Self-Supervised Clustering of Tables in Relational Databases
Relational Deep Dive: Error-Aware Queries Over Unstructured Data
+1 more
Binyang Dai
99.99+
junior + senior
6 since 20216 in total
One Join Order Does Not Fit All: Reducing Intermediate Results with Per-Split Query Plans
MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
TablePuppet: Towards a Generic Framework for Learning over Relational Tables
Love-at-First-Sight: First Answers Without the Awkward Silence in Big Knowledge Graphs
+1 more
Ana Klimovic
99.94
senior
31 since 202144 in total
Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
APEROL: Adaptive Parallel Edge-to-Cloud Runtime Optimization for Layered Workflow Execution
+1 more
S. Venkatesh 0001
dblp:v/SrinivasanVenkatesh ·
DBLP ↗
99.98
senior
26 since 202187 in total
Counting HyperGraphlets via Color Coding: a Quadratic Barrier and How to Break It
Efficient GPU-Accelerated Adaptive Minimum Cost Seed Selection
Theoretically and Practically Efficient Resistance Distance Computation on Large Graphs
TRIM: An Efficient Framework for Exact Eccentricity Computation on Large-Scale Graphs
+1 more
Trinabh Gupta
99.99
senior
10 since 202119 in total
Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data Analytics
SACK: Shielding Dynamic Attribute-based Access Control in Persistent Key-Value Stores
Learned Static Function Data Structures
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
+1 more
Fotis Psallidas
99.99+
senior
7 since 202117 in total
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
Document-to-Database: Extraction Meets Relational Semantics
Streaming Validation of JSON Documents Against Schemas
Blaze: Compiling JSON Schema for 10x Faster Validation
+1 more
in pool
Calisto Zuzarte
99.99+
senior
14 since 202157 in total
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
LIO: A lightweight and interpretable query optimizer based on an evolutionary forest
Libra: One-Shot Parameter Sensitivity Estimation for Transfer Learning in Database Performance Prediction
AXE: A Task Decomposition Approach to Learned LSM Tuning
+1 more
Franco Maria Nardini
99.99
senior
49 since 202194 in total
ConANN: Conformal Approximate Nearest Neighbor Search
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
FedAugment: Table Augmentation Search over Decentralized Data Repositories
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
+1 more
in pool
Anastasia Ailamaki
dblp:a/AnastassiaAilamaki ·
DBLP ↗
99.99+
senior
38 since 2021216 in total
Sampling-based Predictive Database Buffer Management
Terabyte-Scale Analytics in the Blink of an Eye
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
AQD: Online Adaptive Query Dispatcher for HTAP Databases
+1 more
Binhang Yuan
99.99
senior
36 since 202140 in total
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
Efficient Cooperation-Aware Key and Value Management for LLM Inference
Unified Static–Dynamic Pruning for Efficient LLM Inference
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
+1 more
Lihui Liu
99.99
junior + senior
28 since 202132 in total
A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs
Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-Based Interpolation
Noisy Interactive Graph Search: An Uncertainty-Based Approach with Online Modeling of Latent Expertise and Difficulty
Sankofa: Online Query-adaptive Dynamic Graph Summaries
+1 more
Yuankai Fan
99.99+
junior + senior
10 since 202110 in total
SQL-Exchange: Transforming SQL Queries Across Domains
SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation
SemBench: A Benchmark for Semantic Query Processing Engines
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
+1 more
Xiaodong Zhang 0001
99.99+
senior
16 since 2021182 in total
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization
I/O Optimizations in Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration
GPU Acceleration of SQL Analytics on Compressed Data
Bridging the Indexing Gap in Fused GPU Query Engines
+1 more
Qian Xu 0021
99.99+
junior
3 since 20213 in total
Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search
RED-ANNS: An RDMA-Enabled Distributed Framework for Graph-Based Approximate Nearest Neighbor Search
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
ANNiE: A Learned Query Cost Estimator for Graph-Based Approximate Nearest Neighbor Search
+1 more
in pool
Xiangyao Yu
99.99+
senior
36 since 202160 in total
Global Hash Tables Strike Back! An Analysis of Parallel GROUP BY Aggregation
Swan: Hybrid MVCC Management for Efficient Transaction Processing in LSM-Tree-Based Key-Value Stores
Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized Coordination
Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services
+1 more
Sergei Vassilvitskii
99.98
senior
30 since 2021113 in total
Efficient Banzhaf-Based Data Valuation for $k$-Nearest Neighbors Classification
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
CaSh: Shapley Value Computation with Cache Optimization
Highly-Efficient Large-Scale k-means with Individual Fairness
+1 more
Penghang Liu
99.99
junior
4 since 20215 in total
TIMEST: Temporal Information Motif Estimator Using Sampling Trees
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph Representation
Efficient Temporal Edge-Core Maintenance in Streaming Graphs
+1 more
Xin Ai 0006
99.99+
junior + senior
10 since 202110 in total
Efficient GNN Training on Giant Graphs with Collective Batching and Scheduling
PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness Avoidance
ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks
FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism
+1 more
Yue Pang 0001
99.99+
junior + senior
9 since 20219 in total
Aquila: A High-Concurrency System for Incremental Graph Query
Nav-Index: A High-Performance, Adaptive Index for Shortest Path Queries in RDBMS
Efficient Temporal Subgraph Management: A New Interval Index
Mix & Match: Subgraph Matching for Absolute Coverage
+1 more
Xiaofan Li 0004
99.99
junior + senior
6 since 20217 in total
Revisiting the Maximum Defective Clique Problem: Faster Branching and a Tighter Upper Bound
Maximum Defective Biclique Search in Large Bipartite Graphs
CREST: Approximate k-Clique Counting in Real-World Networks via Refinement of Star-Based Sample Space
Efficient Hyper-truss Decomposition over Hypergraphs
+1 more
Linshan Qiu
99.99+
junior
3 since 20214 in total
X-Wim: Massive Parallelization of Weighted Matching in Bipartite Graphs
Characterizing Parallel Subgraph Matching Performance: A Systematic Study of Interactions, Scalability, and Enumeration
CEMR: An Effective Subgraph Matching Algorithm with Redundant Extension Elimination
gMatch: Fine-Grained and Hardware-Efficient Subgraph Matching on GPUs
+1 more
Tarique Siddiqui
99.99+
junior + senior
8 since 202115 in total
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines
BaCon: Efficient Batch Processing of Counting Queries
Redbench: Workload Synthesis From Cloud Traces
+1 more
Badrish Chandramouli
99.99
senior
20 since 202168 in total
Fugue: Online Elasticity for Distributed Stateful Stream Processing
How Much Can RocksDB Chew? Achieving Near-Zero Write Stalls with Sustainable RocksDB
Tidehunter: Large-Value Storage With Minimal Data Relocation
Demystifying and Improving Lazy Promotion in Cache Eviction
+1 more
Lei Liang 0002
99.99
senior
26 since 202126 in total
Graph Transformers for Query Plan Representation: Potentials and Challenges
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
+1 more
Yang Liu 0018
99.98
senior
88 since 2021128 in total
Efficient Task Assignment for Multi-Workerset Crowdsourcing with Time and Expense Considerations
Stress-Testing ML Pipelines with Adversarial Data Corruption
Auditing for Demographic Bias in Opaque Rankings
Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness
+1 more
Yuan Qiu 0002
99.98
junior + senior
7 since 20218 in total
PrivSTD: Differentially Private Spatio-temporal Trajectory Density Data Publication
Measuring Database Unfairness via Dependency Quantification Under Differential Privacy
Algorithmic Data Minimization for Machine Learning over Internet-of-Things Data Streams
Doppio: Communication-Efficient and Secure Multi-Party Shuffle Differential Privacy
+1 more
Davood Rafiei
99.99+
senior
20 since 202152 in total
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Unveiling Challenges for LLMs in Enterprise Data Engineering
Human-Centered Exploration of Table Unionability
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
+1 more
Eduardo H. M. Pena
99.99
junior + senior
6 since 202113 in total
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
Storage-Centric Relation Design via High-Quality Approximate Functional Dependencies
SHARP: Shared State Reduction for Efficient Matching of Sequential Patterns
Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset Queries
+1 more
Bowen Zhang 0012
99.99
junior
9 since 20219 in total
SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous Memory
Terark-DS: A High-Performance and Storage-Efficient Key-Value Separation Storage Engine on Disaggregated Storage
LiBox: A Learned Index as an Array to Minimize Last-Mile Search
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
+1 more
Gaurav Tarlok Kakkar
99.99+
junior + senior
6 since 20216 in total
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
SEMA: A High-performance System for LLM-based Semantic Query Processing
Ken: An Execution Engine for Unstructured Database Systems
OBELISK: Efficient Offline Query Planning with Bayesian Optimization-Informed Language Model Reasoning
+1 more
Florian Kerschbaum
99.99
senior
53 since 2021133 in total
Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data Analytics
Composition for Pufferfish Privacy
Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
Secure Join Operations in Multi-Identifier Databases: Performance and Practicality
+1 more
Amir Shaikhha
99.99
senior
24 since 202140 in total
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
Window Function Optimization: Co-Evaluation and Other Techniques
FlowLog: Efficient and Extensible Datalog via Incrementality
+1 more
Zhitao Shen
99.99
senior
8 since 202113 in total
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point Data
Revisiting Filtered ANN Benchmarks: A Hardness-Controlled Benchmark Generator for Realistic Evaluation
JHQ: Johnson-Lindenstrauss Enhanced Hierarchical Quantization for High-Dimensional Approximate Nearest Neighbor Search
+1 more
Qizhen Zhang 0001
99.98
junior + senior
13 since 202120 in total
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
CloudGlide: Deconstructing the Landscape of Cloud-Based Analytics
LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics
CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort
+1 more
in pool
Yuchen Li 0001
99.98
senior
42 since 202173 in total
AGIS: Fast Approximate Graph Pattern Mining with Structure-Informed Sampling
Efficient GPU-Accelerated Adaptive Minimum Cost Seed Selection
MDS-FSM: Coverage-Based Frequent Subgraph Mining in Single Graphs
GPU-Accelerated 𝜂-threshold Decomposition for Uncertain Graphs
+1 more
Jinsheng Ba
99.99+
junior
8 since 20218 in total
Dinkel: State-Aware and Granular Framework for Validating Graph Databases
Testing Graph Databases via Transformations Between Fixed-Length and Variable-Length Queries
DBAIOps: A Reasoning LLM-Enhanced Database Operation and Maintenance System using Knowledge Graphs
I-Rex: An Interactive Debugger for SQL
+1 more
Daniel Kocher
99.99
senior
6 since 20217 in total
Near-Duplicate Text Alignment under Weighted Jaccard Similarity
FB*: A Compact Index for Efficient and Exact Density-based Clustering
Index Intersection for High-Dimensional Range Queries
An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks
+1 more
Yitong Song 0001
99.99+
junior + senior
7 since 20217 in total
Elastic Index Selection for Label-Hybrid AKNN Search
Sparse Neighborhood Graph-Based Approximate Nearest Neighbor Search Revisited: Theoretical Analysis and Optimization
GAS: A Lightweight Framework for Filtered Search over Wide-table Vectors
RNSG: A Range-Aware Graph Index for Efficient Range-Filtered Approximate Nearest Neighbor Search
+1 more
in pool
Tim Kraska
99.99+
senior
39 since 2021126 in total
Robust Predicate Transfer with Dynamic Execution
Chipmink: Efficient Delta Identification for Massive Object Graphs
NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents
+1 more
in pool
Zoi Kaoudi
99.99
senior
17 since 202141 in total
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
Decisionhouse: Prescriptive Analytics in the Data Stack
Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning Applications
TablePuppet: Towards a Generic Framework for Learning over Relational Tables
+1 more
George Katsogiannis-Meimarakis
99.99+
junior
7 since 20217 in total
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
Database Views as Explanations for Relational Deep Learning
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
+1 more
Mohammed Saeed 0002
99.99+
junior + senior
7 since 20219 in total
Bolt-on, Verifiable Provenance for LLM-Powered Data Processing
SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text Generation
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
Multi-Objective Agentic Rewrites for Unstructured Data Processing
+1 more
Qinggang Zhang
99.97
junior
24 since 202124 in total
In-depth Analysis of Graph-based RAG in a Unified Framework
LLMs as Stratification Signals for KG Accuracy Evaluation
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
Can we trust LLM Self-Explanations for Entity Resolution?
+1 more
Xiao Hu 0005
99.99+
junior + senior
30 since 202138 in total
Towards Efficient Random-Order Enumeration for Join Queries
STEM2: A Fast and Space-efficient Data Structure for Exact Multi-Set Membership Query
Hybrid Mixed Integer Linear Programming for Large-Scale Join Order Optimisation
QDBO: A Real-time Quantum-augmented Database System Optimizer
+1 more
Da Zheng 0004
99.99+
junior + senior
25 since 202135 in total
ConRAD: Conformal Risk-Aware Neural Databases
NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments
Scalable GNN Explanations with Distributed Shapley Values
FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph Training
+1 more
Kouki Kawabata
99.95
senior
9 since 202111 in total
MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean Distance
CLaP - State Detection from Time Series
MH-GIN: Multi-scale Heterogeneous Graph-based Imputation Network for AIS Data
KDSelector: A Framework of Knowledge-Enhanced and Data-Efficient Selector Learning for Anomaly Detection Model Selection in Time Series
+1 more
Zhaoyan Sun
99.99+
junior + senior
6 since 20216 in total
ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling
Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting
LIO: A lightweight and interpretable query optimizer based on an evolutionary forest
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL
+1 more
Xiang Li 0156
99.97
junior
9 since 202110 in total
HarborMaster: Rollback Detection for Trusted Distributed Computing
SACK: Shielding Dynamic Attribute-based Access Control in Persistent Key-Value Stores
Verifiable Authenticated Data Structure (V-ADS) for Analytic Queries
Succinct and Fast Tiny Pointer Hash Tables
+1 more
in pool
M. Tamer Özsu
99.99+
senior
29 since 2021179 in total
TVA: A Version-aware Temporal Graph Storage System for Real-time Analytics
The Data World Is Not Flat: Efficient Factorized Execution for Relational Systems
One Join Order Does Not Fit All: Reducing Intermediate Results with Per-Split Query Plans
Love-at-First-Sight: First Answers Without the Awkward Silence in Big Knowledge Graphs
+1 more
Hengyi Cai
99.95
senior
30 since 202136 in total
Data-efficient Online Training for Direct Alignment in LLMs
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
+1 more
in pool
Asterios Katsifodimos
99.98
senior
30 since 202152 in total
Interoperable ACID Transactions for Open Table Formats
Toward Temporal Attribution Analytics in Dataflows
Programmable Dataflows: Abstraction and Programming Model for Data Sharing
Streaming Validation of JSON Documents Against Schemas
+1 more
Shaleen Deep
99.99+
junior + senior
20 since 202128 in total
A Unified Query Planning Framework for Conjunctive Regular Path Queries
Efficient Query Repair for Aggregate Constraints
Discovering Approximate Denial Constraints in Large Databases
Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
+1 more
Longlong Lin
99.94
junior + senior
36 since 202138 in total
Efficient Partition-based Approaches for Diversified Top-k Subgraph Matching
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
CRAFT: Corpus Relatedness Analysis Using Fourier Transforms
Counting HyperGraphlets via Color Coding: a Quadratic Barrier and How to Break It
+1 more
in pool
Kwanghyun Park 0001
99.99+
senior
15 since 202123 in total
TPCx-AI under the Microscope: A Benchmarking Debt Analysis
Morphing-based Compression for Data-centric ML Pipelines
One Pass to Parse Them All: Fused Parallel CSV Processing
QStore: Quantization-Aware Compressed Model Storage
+1 more
Rasmus Pagh
99.99
senior
25 since 2021119 in total
On Fair Epsilon Net and Geometric Hitting Set
Learned Static Function Data Structures
Efficient Banzhaf-Based Data Valuation for $k$-Nearest Neighbors Classification
Revisiting Filtered ANN Benchmarks: A Hardness-Controlled Benchmark Generator for Realistic Evaluation
+1 more
Connor Henderson
99.99
junior
2 since 20212 in total
Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache Flink
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
DOT: Dynamic Knob Selection and Online Sampling for Automated Database Tuning
+1 more
in pool
Yeye He
99.99+
senior
20 since 202146 in total
Relational Deep Dive: Error-Aware Queries Over Unstructured Data
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
Schuyler: Self-Supervised Clustering of Tables in Relational Databases
Document-to-Database: Extraction Meets Relational Semantics
+1 more
Amelie Chi Zhou
99.91
senior
26 since 202149 in total
Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
Incremental Stream Query Deployment under Continuous Infrastructure Changes in the Cloud-Edge Continuum
A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric
+1 more
in pool
Man Lung Yiu
99.99+
senior
21 since 2021136 in total
RT-RkNN: Reverse k Nearest Neighbor Queries as a Graphics Ray Casting Problem
BBC: Improving Large-𝑘 Approximate Nearest Neighbor Search with a Bucket-based Result Collector
HEXA: A Disjoint-Subgraph-Based Indexing Framework for Approximate Nearest Neighbor Search at Billion Scale
Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search
+1 more
Haoyang Li 0015
99.99+
junior + senior
13 since 202113 in total
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
SemBench: A Benchmark for Semantic Query Processing Engines
+1 more
Renzo Angles
99.98
junior
7 since 202115 in total
Structural Normalization of Property Graphs
Computing Why-Provenance for Property Graph Queries
TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs
+1 more
Jiang Xiao 0001
99.92
senior
45 since 202181 in total
Remora: Scale-out Deterministic Execution for Smart Contracts
FairDAG: Consensus Fairness over Multi-Proposer Causal Design
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
Orca: Flexible Quorums Meet Dynamic Quorums
+1 more
Xiangpeng Hao
99.99+
junior + senior
6 since 20219 in total
Operation-Aware Hybrid Locking for Modern In-Memory Indexes
How to Write to SSDs
Dynamic read & write optimization with TurtleKV
Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP Systems
+1 more
Chao Zhang 0034
99.99+
junior + senior
14 since 202116 in total
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
Vodka: Rethink Benchmarking Philosophy in HTAP Systems
AQD: Online Adaptive Query Dispatcher for HTAP Databases
Sampling-based Predictive Database Buffer Management
+1 more
Guangyi Zhang 0001
99.96
junior + senior
12 since 202114 in total
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
CaSh: Shapley Value Computation with Cache Optimization
Augmenting Social Influence of Uncertain Seeds via Probabilistic Link Insertion
Fast and Private Max-Sum Diversification
+1 more
Jimmy Lin
99.99
senior
107 since 2021278 in total
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
FedAugment: Table Augmentation Search over Decentralized Data Repositories
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
+1 more
Nils Boeschen
99.98
junior + senior
5 since 20215 in total
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
GPU Acceleration of SQL Analytics on Compressed Data
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
Bridging the Indexing Gap in Fused GPU Query Engines
+1 more
Zhiyong Wu 0010
99.99+
senior
22 since 202122 in total
Fast Verification of Strong Database Isolation
Pisco: An Isolation Bug Case Reduction and Deduplication Framework
Blaze: Compiling JSON Schema for 10x Faster Validation
Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database Construction
+1 more
Qingshuai Feng
99.91
junior + senior
6 since 20216 in total
Sankofa: Online Query-adaptive Dynamic Graph Summaries
Noisy Interactive Graph Search: An Uncertainty-Based Approach with Online Modeling of Latent Expertise and Difficulty
KAFY: An Extensible and Scalable Transformers-Based System for Trajectory Data Analysis
MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented Generation
+1 more
Yong Li 0045
99.98
senior
17 since 202120 in total
PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness Avoidance
Unified Static–Dynamic Pruning for Efficient LLM Inference
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
+1 more
Junghoon Kim 0007
99.99+
junior + senior
16 since 202118 in total
Efficient Hyper-truss Decomposition over Hypergraphs
Anchored Maximum Communities over Large Directed Graphs
Aggregating maximal cliques in real-world graphs
Efficient Locally h-Clique Densest Subgraph Discovery via Divide-and-Conquer
+1 more
Shreya Shankar
99.99+
junior + senior
16 since 202118 in total
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
+1 more
in pool
Jianguo Wang 0001
99.99+
senior
26 since 202139 in total
Aker: Density-Aware Approximate Caching for Vector Search
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
An Experimental Evaluation of Hybrid Querying on Vectors
I/O Optimizations in Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration
+1 more
in pool
Yifan Zhu 0002
99.99+
junior + senior
14 since 202116 in total
CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data
A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
ANNiE: A Learned Query Cost Estimator for Graph-Based Approximate Nearest Neighbor Search
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
+1 more
Kaisong Huang
99.99+
junior + senior
8 since 20218 in total
BtrLog: Low-Latency Logging for Cloud Database Systems
Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized Coordination
Tux: Efficient Drop-in Networking for Database Systems
Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services
+1 more
in pool
Ashwin Machanavajjhala
dblp:m/AMachanavajjhala ·
DBLP ↗
99.95
senior
16 since 202185 in total
Enabling Index-free Adjacency in Oblivious Graph Processing with Delayed Duplications
Secure Multi-Party Sampling over Joins
Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
PrivSTD: Differentially Private Spatio-temporal Trajectory Density Data Publication
+1 more
Tim Gubner
99.99+
junior
4 since 20217 in total
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
Terabyte-Scale Analytics in the Blink of an Eye
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization
dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data Processing
+1 more
in pool
Stratos Idreos
99.99+
senior
21 since 202181 in total
Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
Rethinking Learned Index and LSM-tree Integration
Tidehunter: Large-Value Storage With Minimal Data Relocation
+1 more
Junjie Xing
99.99+
junior
5 since 20216 in total
Human-Centered Exploration of Table Unionability
Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA
Unveiling Challenges for LLMs in Enterprise Data Engineering
TATA: An Efficient Framework for Task Transfer in Query Plan Representation
+1 more
Vassilis N. Ioannidis
99.99
senior
25 since 202129 in total
Towards A Generalizable and Expressive Graph Neural Network for Graph-Level Tasks with Theoretical Guarantees
Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-Based Interpolation
Graph Transformers for Query Plan Representation: Potentials and Challenges
CRAFT: Corpus Relatedness Analysis Using Fourier Transforms
+1 more
in pool
Kaiqiang Yu
99.99
junior + senior
22 since 202127 in total
Maximum Defective Biclique Search in Large Bipartite Graphs
A Practical Sublinear Approximation for Group Steiner Tree
Theoretically and Practically Efficient Resistance Distance Computation on Large Graphs
CREST: Approximate k-Clique Counting in Real-World Networks via Refinement of Star-Based Sample Space
+1 more
Brian Kroth
99.99+
senior
6 since 202110 in total
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
AXE: A Task Decomposition Approach to Learned LSM Tuning
Toward Drift-Aware Database Benchmarking
Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents
+1 more
Odysseas Papapetrou
dblp:p/OdysseasPapapetrou ·
DBLP ↗
99.99
senior
10 since 202135 in total
Continuous Query for Top-K Maximal Sum Intervals over Streaming Data
PAIL: Efficient kNN Search on Set-Valued Attributes
Optimal Approximate Matrix Multiplication over Sliding Windows
Near-Duplicate Text Alignment under Weighted Jaccard Similarity
+1 more
Ilie Sarpe
99.99
junior
7 since 20217 in total
TIMEST: Temporal Information Motif Estimator Using Sampling Trees
Efficient Temporal Edge-Core Maintenance in Streaming Graphs
Finding Time-Proximity Communities in Temporal Heterogeneous Information Networks
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
+1 more
Matteo Paganelli
99.98
junior + senior
15 since 202120 in total
PINE: Extracting Correlated Token Pairs for Explainable Entity Matching
Database Views as Explanations for Relational Deep Learning
ALER: An Active Learning Hybrid System for Efficient Entity Resolution
Multi-Objective Agentic Rewrites for Unstructured Data Processing
+1 more
Krishnaram Kenthapadi
99.94
junior + senior
18 since 202159 in total
Auditing for Demographic Bias in Opaque Rankings
Stress-Testing ML Pipelines with Adversarial Data Corruption
Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness
Efficient Task Assignment for Multi-Workerset Crowdsourcing with Time and Expense Considerations
+1 more
Sebastian Schmidl
99.95
junior + senior
6 since 20216 in total
MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean Distance
KDSelector: A Framework of Knowledge-Enhanced and Data-Efficient Selector Learning for Anomaly Detection Model Selection in Time Series
Finding Non-Redundant Simpson's Paradox in Multidimensional Data
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
+1 more
Hao Chen 0080
99.99
senior
12 since 202115 in total
SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous Memory
Terark-DS: A High-Performance and Storage-Efficient Key-Value Separation Storage Engine on Disaggregated Storage
Swan: Hybrid MVCC Management for Efficient Transaction Processing in LSM-Tree-Based Key-Value Stores
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
+1 more
in pool
Longbin Lai
99.99+
senior
19 since 202127 in total
CEMR: An Effective Subgraph Matching Algorithm with Redundant Extension Elimination
Aquila: A High-Concurrency System for Incremental Graph Query
Nav-Index: A High-Performance, Adaptive Index for Shortest Path Queries in RDBMS
gMatch: Fine-Grained and Hardware-Efficient Subgraph Matching on GPUs
+1 more
Dingwen Tao
99.99
senior
80 since 2021113 in total
PILOT-C: Physics-Informed Low-Distortion Optimal Trajectory Compression
DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point Data
QStore: Quantization-Aware Compressed Model Storage
Morphing-based Compression for Data-centric ML Pipelines
+1 more
in pool
Kaustubh Beedkar
99.99
senior
14 since 202119 in total
Programmable Dataflows: Abstraction and Programming Model for Data Sharing
Fugue: Online Elasticity for Distributed Stateful Stream Processing
SHARP: Shared State Reduction for Efficient Matching of Sequential Patterns
IncreQueryFusion: On-demand Data Fusion Framework in Dynamic Data Lakes
+1 more
Maximilian Kuschewski
99.98
junior
6 since 20216 in total
Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression
LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
+1 more
in pool
Immanuel Trummer
99.99+
junior + senior
34 since 202165 in total
Exploring Exploratory Querying
OBELISK: Efficient Offline Query Planning with Bayesian Optimization-Informed Language Model Reasoning
An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks
ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling
+1 more
in pool
Shixuan Sun
99.99+
senior
38 since 202147 in total
Efficient GPU-Accelerated Local Subgraph Counting
X-Wim: Massive Parallelization of Weighted Matching in Bipartite Graphs
TVA: A Version-aware Temporal Graph Storage System for Real-time Analytics
Resource-Efficient FirmCore Decomposition on Billion-scale Multilayer Graphs
+1 more
Stefan Grafberger
99.99+
junior + senior
10 since 202111 in total
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
CAPS: Cost-Aware ML Pipeline Selection
Decisionhouse: Prescriptive Analytics in the Data Stack
Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning Applications
+1 more
Lei Hou 0001
99.94
senior
96 since 2021127 in total
LLMs as Stratification Signals for KG Accuracy Evaluation
Can we trust LLM Self-Explanations for Entity Resolution?
In-depth Analysis of Graph-based RAG in a Unified Framework
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs
+1 more
Baotong Lu
99.99
senior
8 since 202110 in total
Efficient Cooperation-Aware Key and Value Management for LLM Inference
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
LiBox: A Learned Index as an Array to Minimize Last-Mile Search
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
+1 more
in pool
Meihao Liao
99.96
junior + senior
18 since 202118 in total
Scalable Approximate Biclique Counting over Large Bipartite Graphs
Lower-Bound Distance Queries under Partial Information
Revisiting the Maximum Defective Clique Problem: Faster Branching and a Tighter Upper Bound
Sparse Neighborhood Graph-Based Approximate Nearest Neighbor Search Revisited: Theoretical Analysis and Optimization
+1 more
Johes Bater
99.99
senior
8 since 202112 in total
Measuring Database Unfairness via Dependency Quantification Under Differential Privacy
Secure Join Operations in Multi-Identifier Databases: Performance and Practicality
A Workload-Aware Encrypted Index for Efficient Privacy-Preserving Range Queries
PrivMDC: Leveraging Multi-Dimensional Correlations to Answer Differentially Private Range Queries
+1 more
Yannis Papakonstantinou
dblp:p/YPapakonstantinou ·
DBLP ↗
99.99+
senior
9 since 202192 in total
Ken: An Execution Engine for Unstructured Database Systems
SEMA: A High-performance System for LLM-based Semantic Query Processing
Robust Predicate Transfer with Dynamic Execution
Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting
+1 more
Kyle Deeds
99.99+
junior + senior
9 since 202110 in total
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
Storage-Centric Relation Design via High-Quality Approximate Functional Dependencies
QDBO: A Real-time Quantum-augmented Database System Optimizer
Window Function Optimization: Co-Evaluation and Other Techniques
+1 more
in pool
Xiaohui Yu 0001
99.97
senior
44 since 2021114 in total
Craw: A Unified and Efficient Querying Framework for Large-Scale Video Datasets
KAFY: An Extensible and Scalable Transformers-Based System for Trajectory Data Analysis
MH-GIN: Multi-scale Heterogeneous Graph-based Imputation Network for AIS Data
GAS: A Lightweight Framework for Filtered Search over Wide-table Vectors
+1 more
Nikolay Yakovets
99.99
senior
14 since 202126 in total
Efficient Temporal Subgraph Management: A New Interval Index
Worst-Case Optimal BGPs on Temporal Graphs
Subgraph Enumeration: Beyond Tree Decomposition
Computing Why-Provenance for Property Graph Queries
+1 more
Eleni Zapridou
99.95
junior
4 since 20215 in total
Incremental Stream Query Deployment under Continuous Infrastructure Changes in the Cloud-Edge Continuum
Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache Flink
APEROL: Adaptive Parallel Edge-to-Cloud Runtime Optimization for Layered Workflow Execution
How Much Can RocksDB Chew? Achieving Near-Zero Write Stalls with Sustainable RocksDB
+1 more
Kurt Stockinger
99.99
senior
14 since 202151 in total
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases
+1 more
in pool
Peter Boncz
99.99
senior
27 since 2021108 in total
One Pass to Parse Them All: Fused Parallel CSV Processing
The Data World Is Not Flat: Efficient Factorized Execution for Relational Systems
Chipmink: Efficient Delta Identification for Massive Object Graphs
Index Intersection for High-Dimensional Range Queries
+1 more
Graham Cormode
99.98
junior + senior
36 since 2021180 in total
On Fair Epsilon Net and Geometric Hitting Set
Unbiased Binning for Fairness-aware Attribute Representation
Composition for Pufferfish Privacy
Efficient Banzhaf-Based Data Valuation for $k$-Nearest Neighbors Classification
+1 more
Philipp Skavantzos
99.99
junior
7 since 20217 in total
Structural Normalization of Property Graphs
Repairing Property Graphs under PG-Constraints
Discovering Approximate Denial Constraints in Large Databases
Dinkel: State-Aware and Granular Framework for Validating Graph Databases
+1 more
in pool
Faisal Nawab
99.90
senior
31 since 202162 in total
HarborMaster: Rollback Detection for Trusted Distributed Computing
Remora: Scale-out Deterministic Execution for Smart Contracts
Meerkat: Scalable, Network-Aware Failure Recovery for the Internet of Things
Orca: Flexible Quorums Meet Dynamic Quorums
+1 more
in pool
Avrilia Floratou
99.99
senior
12 since 202130 in total
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL
Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Streaming Validation of JSON Documents Against Schemas
+1 more
Ningyi Liao
99.98
junior + senior
12 since 202112 in total
Efficient GNN Training on Giant Graphs with Collective Batching and Scheduling
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
Towards A Generalizable and Expressive Graph Neural Network for Graph-Level Tasks with Theoretical Guarantees
Scalable GNN Explanations with Distributed Shapley Values
+1 more
Yanguo Peng
99.97
senior
26 since 202131 in total
Verifiable Authenticated Data Structure (V-ADS) for Analytic Queries
Efficient and Secure Range Counting over Distributed Geographic Data with Query Range Protection
Enabling Index-free Adjacency in Oblivious Graph Processing with Delayed Duplications
SACK: Shielding Dynamic Attribute-based Access Control in Persistent Key-Value Stores
+1 more
Dan Olteanu
99.99+
senior
23 since 202187 in total
MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
A Unified Query Planning Framework for Conjunctive Regular Path Queries
TablePuppet: Towards a Generic Framework for Learning over Relational Tables
One Join Order Does Not Fit All: Reducing Intermediate Results with Per-Split Query Plans
+1 more
Ilias Azizi
99.99
junior
2 since 20212 in total
JHQ: Johnson-Lindenstrauss Enhanced Hierarchical Quantization for High-Dimensional Approximate Nearest Neighbor Search
Harmonizing Efficiency and Accuracy in Filtered Vector Search
CRAFT: Corpus Relatedness Analysis Using Fourier Transforms
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
+1 more
Manuel Rigger
99.99
senior
26 since 202135 in total
FlowLog: Efficient and Extensible Datalog via Incrementality
Pisco: An Isolation Bug Case Reduction and Deduplication Framework
Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
Fast Verification of Strong Database Isolation
+1 more
Kaihao Ma
99.81
junior
8 since 20218 in total
ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
+1 more
in pool
Laks V. S. Lakshmanan
99.99
senior
41 since 2021208 in total
Augmenting Social Influence of Uncertain Seeds via Probabilistic Link Insertion
Efficient GPU-Accelerated Adaptive Minimum Cost Seed Selection
Noisy Interactive Graph Search: An Uncertainty-Based Approach with Online Modeling of Latent Expertise and Difficulty
Mix & Match: Subgraph Matching for Absolute Coverage
+1 more
in pool
Jiannan Wang 0001
99.99
senior
18 since 202159 in total
Relational Deep Dive: Error-Aware Queries Over Unstructured Data
Love-at-First-Sight: First Answers Without the Awkward Silence in Big Knowledge Graphs
Efficient Query Repair for Aggregate Constraints
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
+1 more
Daniel Ulrich Schmitt
99.98
junior
3 since 20213 in total
FB*: A Compact Index for Efficient and Exact Density-based Clustering
ConANN: Conformal Approximate Nearest Neighbor Search
Scalable Grid-based Computation of Kendall's Tau Correlation
Schuyler: Self-Supervised Clustering of Tables in Relational Databases
+1 more
Yihang Zheng
99.99+
junior
4 since 20214 in total
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
LIO: A lightweight and interpretable query optimizer based on an evolutionary forest
SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
+1 more
Linsen Li 0001
99.77
senior
5 since 20215 in total
FutureLight: An Efficient Future Traffic Data-Driven Reinforcement Learning Framework for Traffic Signal Controls
PILOT-C: Physics-Informed Low-Distortion Optimal Trajectory Compression
MH-GIN: Multi-scale Heterogeneous Graph-based Imputation Network for AIS Data
PrivSTD: Differentially Private Spatio-temporal Trajectory Density Data Publication
+1 more
in pool
Xiu Tang
99.99+
junior + senior
27 since 202127 in total
LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
SQL-Exchange: Transforming SQL Queries Across Domains
TATA: An Efficient Framework for Task Transfer in Query Plan Representation
+1 more
Xinyu Ma 0001
99.98
junior + senior
27 since 202127 in total
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
In-depth Analysis of Graph-based RAG in a Unified Framework
+1 more
Zequn Sun 0001
99.99
senior
27 since 202136 in total
A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs
Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-Based Interpolation
ALER: An Active Learning Hybrid System for Efficient Entity Resolution
Database Views as Explanations for Relational Deep Learning
+1 more
Jian Sha
99.98
senior
7 since 20217 in total
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
Unified Static–Dynamic Pruning for Efficient LLM Inference
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
CAPS: Cost-Aware ML Pipeline Selection
+1 more
in pool
Aditya G. Parameswaran
99.99+
senior
37 since 2021110 in total
Exploring Exploratory Querying
Document-to-Database: Extraction Meets Relational Semantics
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
+1 more
in pool
Viktor Leis
99.98
junior + senior
52 since 202185 in total
Succinct and Fast Tiny Pointer Hash Tables
A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric
Dynamic read & write optimization with TurtleKV
SunStorm: Geographically distributed transactions over Aurora-style systems
+1 more
Yongchao Liu 0004
99.99
senior
33 since 202161 in total
FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph Training
GPU-Accelerated 𝜂-threshold Decomposition for Uncertain Graphs
UniTG: A Unified System for Efficient and Seamless Textual Graph Learning
ConRAD: Conformal Risk-Aware Neural Databases
+1 more
Kavitha Srinivas
99.98
senior
23 since 202146 in total
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
Multi-Objective Agentic Rewrites for Unstructured Data Processing
+1 more
in pool
Matei Zaharia
99.96
senior
66 since 2021133 in total
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
Efficient Cooperation-Aware Key and Value Management for LLM Inference
FedAugment: Table Augmentation Search over Decentralized Data Repositories
+1 more
Vahab S. Mirrokni
dblp:m/VahabSMirrokni ·
DBLP ↗
99.87
senior
100 since 2021273 in total
Highly-Efficient Large-Scale k-means with Individual Fairness
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
Efficient Task Assignment for Multi-Workerset Crowdsourcing with Time and Expense Considerations
CaSh: Shapley Value Computation with Cache Optimization
+1 more
Yue Zeng 0004
99.95
junior
4 since 20214 in total
Effective Durable Community Search in Large Temporal Graph
Finding Non-Redundant Simpson's Paradox in Multidimensional Data
Finding Time-Proximity Communities in Temporal Heterogeneous Information Networks
Anchored Maximum Communities over Large Directed Graphs
+1 more
in pool
Dong Deng 0001
99.99
senior
20 since 202155 in total
SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
BBC: Improving Large-𝑘 Approximate Nearest Neighbor Search with a Bucket-based Result Collector
An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks
Aker: Density-Aware Approximate Caching for Vector Search
+1 more
Heena Nagda
99.82
junior
5 since 20216 in total
FairDAG: Consensus Fairness over Multi-Proposer Causal Design
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
HarborMaster: Rollback Detection for Trusted Distributed Computing
Meerkat: Scalable, Network-Aware Failure Recovery for the Internet of Things
+1 more
Tarikul Islam Papon
99.99+
junior + senior
11 since 202112 in total
How to Write to SSDs
ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic Workloads
Tux: Efficient Drop-in Networking for Database Systems
Operation-Aware Hybrid Locking for Modern In-Memory Indexes
+1 more
Yuchen Zhuang
99.90
junior
26 since 202128 in total
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs
Data-efficient Online Training for Direct Alignment in LLMs
LLMs as Stratification Signals for KG Accuracy Evaluation
+1 more
in pool
Rathijit Sen
99.97
senior
26 since 202138 in total
CloudGlide: Deconstructing the Landscape of Cloud-Based Analytics
RayDB: Building Databases with Ray Tracing Cores
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
Bridging the Indexing Gap in Fused GPU Query Engines
+1 more
Lin Ma 0006
99.99+
senior
10 since 202119 in total
DOT: Dynamic Knob Selection and Online Sampling for Automated Database Tuning
Sampling-based Predictive Database Buffer Management
High-Performance DBMSs with io_uring: When and How to Use It
Toward Drift-Aware Database Benchmarking
+1 more
in pool
Huanchen Zhang
99.99+
senior
45 since 202153 in total
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
Storing and Indexing Multiple Tables by Interesting Orderings: For Efficient Joins, Groupings, and Updates in Relational Databases
I/O Optimizations in Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration
+1 more
in pool
Fangyuan Zhang 0001
99.99+
junior + senior
15 since 202115 in total
CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data
A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
PAIL: Efficient kNN Search on Set-Valued Attributes
ANNiE: A Learned Query Cost Estimator for Graph-Based Approximate Nearest Neighbor Search
+1 more
Mourad Ouzzani
99.97
senior
11 since 2021101 in total
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
Unveiling Challenges for LLMs in Enterprise Data Engineering
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
Bolt-on, Verifiable Provenance for LLM-Powered Data Processing
+1 more
in pool
Rong-Hua Li 0001
99.99
senior
111 since 2021165 in total
Aggregating maximal cliques in real-world graphs
TRIM: An Efficient Framework for Exact Eccentricity Computation on Large-Scale Graphs
MDS-FSM: Coverage-Based Frequent Subgraph Mining in Single Graphs
Scalable Approximate Biclique Counting over Large Bipartite Graphs
+1 more
Yancan Mao
99.93
junior
12 since 202112 in total
Fugue: Online Elasticity for Distributed Stateful Stream Processing
Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache Flink
Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
APEROL: Adaptive Parallel Edge-to-Cloud Runtime Optimization for Layered Workflow Execution
+1 more
Nico Lässig
99.93
junior + senior
6 since 20216 in total
Auditing for Demographic Bias in Opaque Rankings
Stress-Testing ML Pipelines with Adversarial Data Corruption
Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness
Unbiased Binning for Fairness-aware Attribute Representation
+1 more
in pool
Sheng Wang 0007
99.99+
senior
25 since 202135 in total
Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search
RED-ANNS: An RDMA-Enabled Distributed Framework for Graph-Based Approximate Nearest Neighbor Search
CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates
Elastic Index Selection for Label-Hybrid AKNN Search
+1 more
Holger Pirk
99.99+
senior
7 since 202125 in total
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization
Global Hash Tables Strike Back! An Analysis of Parallel GROUP BY Aggregation
Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability
Robust Predicate Transfer with Dynamic Execution
+1 more
David Pujol
99.99
junior
5 since 20216 in total
Fast and Private Max-Sum Diversification
Doppio: Communication-Efficient and Secure Multi-Party Shuffle Differential Privacy
Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
Secure Multi-Party Sampling over Joins
+1 more
in pool
Ryan Marcus
99.99+
senior
21 since 202136 in total
AXE: A Task Decomposition Approach to Learned LSM Tuning
BaCon: Efficient Batch Processing of Counting Queries
Libra: One-Shot Parameter Sensitivity Estimation for Transfer Learning in Database Performance Prediction
Rethinking Learned Index and LSM-tree Integration
+1 more
in pool
Byron Choi
99.95
senior
46 since 2021135 in total
Efficient Temporal Subgraph Management: A New Interval Index
TVA: A Version-aware Temporal Graph Storage System for Real-time Analytics
Nav-Index: A High-Performance, Adaptive Index for Shortest Path Queries in RDBMS
Craw: A Unified and Efficient Querying Framework for Large-Scale Video Datasets
+1 more
Arthur Bernhardt
99.98
junior + senior
14 since 202116 in total
dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data Processing
Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
Terark-DS: A High-Performance and Storage-Efficient Key-Value Separation Storage Engine on Disaggregated Storage
+1 more
Adeel Aslam
99.99
junior
8 since 20219 in total
Continuous Query for Top-K Maximal Sum Intervals over Streaming Data
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset Queries
Vodka: Rethink Benchmarking Philosophy in HTAP Systems
+1 more
Zhonggen Li
99.99+
junior
4 since 20214 in total
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks
Characterizing Parallel Subgraph Matching Performance: A Systematic Study of Interactions, Scalability, and Enumeration
+1 more
Yizhang He
99.99
junior + senior
12 since 202112 in total
Maximum Defective Biclique Search in Large Bipartite Graphs
Theoretically and Practically Efficient Resistance Distance Computation on Large Graphs
Lower-Bound Distance Queries under Partial Information
CREST: Approximate k-Clique Counting in Real-World Networks via Refinement of Star-Based Sample Space
+1 more
Raunak Shah
99.96
junior
3 since 20214 in total
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
PILOT-C: Physics-Informed Low-Distortion Optimal Trajectory Compression
Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression
CounterSnake: A lossless and generalized compression framework for diverse sketches
+1 more
Meihao Fan
99.92
junior
3 since 20213 in total
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation
Can we trust LLM Self-Explanations for Entity Resolution?
SQL-Exchange: Transforming SQL Queries Across Domains
+1 more
Jyoti Leeka
99.99+
senior
6 since 202111 in total
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Redbench: Workload Synthesis From Cloud Traces
LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics
Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting
+1 more
Qiyan Li 0002
99.99
junior + senior
7 since 20218 in total
CEMR: An Effective Subgraph Matching Algorithm with Redundant Extension Elimination
Efficient Hyper-truss Decomposition over Hypergraphs
A Practical Sublinear Approximation for Group Steiner Tree
Mix & Match: Subgraph Matching for Absolute Coverage
+1 more
Rob Johnson 0001
99.99+
senior
18 since 202164 in total
Demystifying and Improving Lazy Promotion in Cache Eviction
Tidehunter: Large-Value Storage With Minimal Data Relocation
CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort
STEM2: A Fast and Space-efficient Data Structure for Exact Multi-Set Membership Query
+1 more
Zhuoxing Zhang
99.99+
junior + senior
8 since 20219 in total
Storage-Centric Relation Design via High-Quality Approximate Functional Dependencies
Discovering Approximate Denial Constraints in Large Databases
DBAIOps: A Reasoning LLM-Enhanced Database Operation and Maintenance System using Knowledge Graphs
Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database Construction
+1 more
Qiyao Luo
99.99
junior + senior
12 since 202112 in total
Measuring Database Unfairness via Dependency Quantification Under Differential Privacy
PrivMDC: Leveraging Multi-Dimensional Correlations to Answer Differentially Private Range Queries
Secure Join Operations in Multi-Identifier Databases: Performance and Practicality
Efficient and Secure Range Counting over Distributed Geographic Data with Query Range Protection
+1 more
in pool
Michael J. Carey 0001
99.99+
junior + senior
22 since 2021184 in total
Index Intersection for High-Dimensional Range Queries
SEMA: A High-performance System for LLM-based Semantic Query Processing
Chipmink: Efficient Delta Identification for Massive Object Graphs
Ken: An Execution Engine for Unstructured Database Systems
+1 more
Stefania Dumbrava
99.99+
senior
12 since 202115 in total
Dinkel: State-Aware and Granular Framework for Validating Graph Databases
Computing Why-Provenance for Property Graph Queries
Testing Graph Databases via Transformations Between Fixed-Length and Variable-Length Queries
TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
+1 more
Chengying Huan
99.99
junior + senior
20 since 202121 in total
Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph Representation
FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism
FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph Training
Efficient GPU-Accelerated Local Subgraph Counting
+1 more
Hongtai Cao
99.99
junior + senior
7 since 20219 in total
Efficient Partition-based Approaches for Diversified Top-k Subgraph Matching
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
Subgraph Enumeration: Beyond Tree Decomposition
X-Wim: Massive Parallelization of Weighted Matching in Bipartite Graphs
+1 more
Parimarjan Negi
99.99+
junior
6 since 202111 in total
QDBO: A Real-time Quantum-augmented Database System Optimizer
TablePuppet: Towards a Generic Framework for Learning over Relational Tables
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning Applications
+1 more
Mirek Riedewald
dblp:r/MirekRiedewald ·
DBLP ↗
99.97
senior
14 since 202161 in total
Hybrid Mixed Integer Linear Programming for Large-Scale Join Order Optimisation
Efficient Query Repair for Aggregate Constraints
Finding Non-Redundant Simpson's Paradox in Multidimensional Data
A Unified Query Planning Framework for Conjunctive Regular Path Queries
+1 more
Torsten Grust
99.99
senior
9 since 202149 in total
I-Rex: An Interactive Debugger for SQL
Window Function Optimization: Co-Evaluation and Other Techniques
Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
FlowLog: Efficient and Extensible Datalog via Incrementality
+1 more
Rubao Lee
99.99
senior
12 since 202149 in total
GPU Acceleration of SQL Analytics on Compressed Data
Scalable Grid-based Computation of Kendall's Tau Correlation
Shard: A Scalable and Resize-optimized Hash Index on Disaggregated Memory
Accelerating String-Heavy Queries with LLM Token Tables
+1 more
Karla Saur
99.99
senior
5 since 202110 in total
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
TPCx-AI under the Microscope: A Benchmarking Debt Analysis
NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Decisionhouse: Prescriptive Analytics in the Data Stack
+1 more
Yunxiang Su
99.94
junior + senior
5 since 20215 in total
MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean Distance
KDSelector: A Framework of Knowledge-Enhanced and Data-Efficient Selector Learning for Anomaly Detection Model Selection in Time Series
Optimal Approximate Matrix Multiplication over Sliding Windows
FB*: A Compact Index for Efficient and Exact Density-based Clustering
+1 more
Xiaoyao Zhong
99.98
junior + senior
4 since 20214 in total
Sparse Neighborhood Graph-Based Approximate Nearest Neighbor Search Revisited: Theoretical Analysis and Optimization
JHQ: Johnson-Lindenstrauss Enhanced Hierarchical Quantization for High-Dimensional Approximate Nearest Neighbor Search
An Experimental Evaluation of Hybrid Querying on Vectors
Harmonizing Efficiency and Accuracy in Filtered Vector Search
+1 more
Victor Giannakouris
99.99
junior
2 since 20216 in total
Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
+1 more
Kyle Luoma
99.99
junior
2 since 20212 in total
Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
+1 more
Chenghong Wang
99.97
junior
16 since 202125 in total
Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data Analytics
A Workload-Aware Encrypted Index for Efficient Privacy-Preserving Range Queries
Enabling Index-free Adjacency in Oblivious Graph Processing with Delayed Duplications
Learned Static Function Data Structures
+1 more
Rui Xue 0006
99.99
junior
5 since 20215 in total
Efficient GNN Training on Giant Graphs with Collective Batching and Scheduling
Towards A Generalizable and Expressive Graph Neural Network for Graph-Level Tasks with Theoretical Guarantees
NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments
UniTG: A Unified System for Efficient and Seamless Textual Graph Learning
+1 more
Zachary G. Ives
99.99
senior
13 since 202174 in total
MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
Relational Deep Dive: Error-Aware Queries Over Unstructured Data
Toward Temporal Attribution Analytics in Dataflows
Love-at-First-Sight: First Answers Without the Awkward Silence in Big Knowledge Graphs
+1 more
Yiming Zhang 0003
99.91
senior
78 since 2021135 in total
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
+1 more
Zhao Zhang 0009
99.80
senior
19 since 202136 in total
Remora: Scale-out Deterministic Execution for Smart Contracts
FairDAG: Consensus Fairness over Multi-Proposer Causal Design
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
Orca: Flexible Quorums Meet Dynamic Quorums
+1 more
in pool
Stefanie Scherzinger
dblp:s/StefanieScherzinger ·
DBLP ↗
99.98
senior
23 since 202141 in total
Streaming Validation of JSON Documents Against Schemas
Blaze: Compiling JSON Schema for 10x Faster Validation
Repairing Property Graphs under PG-Constraints
Document-to-Database: Extraction Meets Relational Semantics
+1 more
Yuxiang Wang 0001
99.99
junior + senior
24 since 202133 in total
A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs
Noisy Interactive Graph Search: An Uncertainty-Based Approach with Online Modeling of Latent Expertise and Difficulty
Sankofa: Online Query-adaptive Dynamic Graph Summaries
GAS: A Lightweight Framework for Filtered Search over Wide-table Vectors
+1 more
in pool
Ju Fan
99.99+
senior
49 since 202191 in total
SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation
TATA: An Efficient Framework for Task Transfer in Query Plan Representation
Schuyler: Self-Supervised Clustering of Tables in Relational Databases
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
+1 more
Kai Han 0003
99.93
senior
29 since 202174 in total
Augmenting Social Influence of Uncertain Seeds via Probabilistic Link Insertion
Efficient GPU-Accelerated Adaptive Minimum Cost Seed Selection
Efficient Task Assignment for Multi-Workerset Crowdsourcing with Time and Expense Considerations
CaSh: Shapley Value Computation with Cache Optimization
+1 more
Roman Heinrich
99.93
junior + senior
7 since 20217 in total
Incremental Stream Query Deployment under Continuous Infrastructure Changes in the Cloud-Edge Continuum
LIO: A lightweight and interpretable query optimizer based on an evolutionary forest
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
Morphing-based Compression for Data-centric ML Pipelines
+1 more
Elias Bassani
99.99
junior
9 since 20219 in total
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
FedAugment: Table Augmentation Search over Decentralized Data Repositories
SemBench: A Benchmark for Semantic Query Processing Engines
+1 more
Yao Zhao 0011
99.97
junior
6 since 20216 in total
Efficient Cooperation-Aware Key and Value Management for LLM Inference
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Unified Static–Dynamic Pruning for Efficient LLM Inference
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
+1 more
Anton Tsitsulin
99.94
senior
12 since 202116 in total
Counting HyperGraphlets via Color Coding: a Quadratic Barrier and How to Break It
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-Based Interpolation
CRAFT: Corpus Relatedness Analysis Using Fourier Transforms
+1 more
Guanming Xiong
99.98
junior
7 since 20217 in total
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
In-depth Analysis of Graph-based RAG in a Unified Framework
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
+1 more
in pool
Paolo Papotti
99.99+
junior + senior
33 since 202197 in total
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Human-Centered Exploration of Table Unionability
Database Views as Explanations for Relational Deep Learning
Multi-Objective Agentic Rewrites for Unstructured Data Processing
+1 more
Arun Kumar 0001
99.97
senior
19 since 202150 in total
LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
CAPS: Cost-Aware ML Pipeline Selection
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
QStore: Quantization-Aware Compressed Model Storage
+1 more
Dong Xie 0001
99.87
senior
12 since 202124 in total
Fast Verification of Strong Database Isolation
Towards Efficient Random-Order Enumeration for Join Queries
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
Revisiting Filtered ANN Benchmarks: A Hardness-Controlled Benchmark Generator for Realistic Evaluation
+1 more
Michael Färber 0001
99.96
senior
38 since 202158 in total
Can we trust LLM Self-Explanations for Entity Resolution?
SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text Generation
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
PINE: Extracting Correlated Token Pairs for Explainable Entity Matching
+1 more
Sijie Ruan
99.85
junior + senior
41 since 202157 in total
MH-GIN: Multi-scale Heterogeneous Graph-based Imputation Network for AIS Data
KAFY: An Extensible and Scalable Transformers-Based System for Trajectory Data Analysis
Highly-Efficient Large-Scale k-means with Individual Fairness
PrivSTD: Differentially Private Spatio-temporal Trajectory Density Data Publication
+1 more
Guanting Dong 0001
99.94
junior
42 since 202142 in total
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs
Data-efficient Online Training for Direct Alignment in LLMs
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?
+1 more
Yangfan Jiang 0001
99.85
junior
9 since 20219 in total
Algorithmic Data Minimization for Machine Learning over Internet-of-Things Data Streams
Stress-Testing ML Pipelines with Adversarial Data Corruption
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
Auditing for Demographic Bias in Opaque Rankings
+1 more
Jeffrey Tao
99.92
junior
8 since 20218 in total
Pisco: An Isolation Bug Case Reduction and Deduplication Framework
Exploring Exploratory Querying
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation
+1 more
in pool
Nikolaus Augsten
99.69
senior
13 since 202144 in total
Near-Duplicate Text Alignment under Weighted Jaccard Similarity
An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks
Craw: A Unified and Efficient Querying Framework for Large-Scale Video Datasets
BBC: Improving Large-𝑘 Approximate Nearest Neighbor Search with a Bucket-based Result Collector
+1 more
Cheng Li 0001
99.73
senior
52 since 202168 in total
PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness Avoidance
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-Design
+1 more
in pool
Hua Fan 0002
99.84
junior + senior
11 since 202119 in total
Meerkat: Scalable, Network-Aware Failure Recovery for the Internet of Things
Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP Systems
SunStorm: Geographically distributed transactions over Aurora-style systems
How to Write to SSDs
+1 more
Alexander van Renen
99.99+
junior + senior
9 since 202117 in total
Dynamic read & write optimization with TurtleKV
RayDB: Building Databases with Ray Tracing Cores
ArceKV: Towards Workload-driven LSM-compactions for Key-Value Store Under Dynamic Workloads
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
+1 more
Chaoji Zuo
99.99+
junior + senior
7 since 20218 in total
RNSG: A Range-Aware Graph Index for Efficient Range-Filtered Approximate Nearest Neighbor Search
Aker: Density-Aware Approximate Caching for Vector Search
PAIL: Efficient kNN Search on Set-Valued Attributes
CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving Data
+1 more
in pool
Fan Zhang 0036
99.99
senior
39 since 202158 in total
Anchored Maximum Communities over Large Directed Graphs
Effective Durable Community Search in Large Temporal Graph
Aggregating maximal cliques in real-world graphs
Efficient Locally h-Clique Densest Subgraph Discovery via Divide-and-Conquer
+1 more
Wan Shen Lim
99.99+
senior
11 since 202113 in total
DOT: Dynamic Knob Selection and Online Sampling for Automated Database Tuning
High-Performance DBMSs with io_uring: When and How to Use It
Tux: Efficient Drop-in Networking for Database Systems
Toward Drift-Aware Database Benchmarking
+1 more
in pool
Alfons Kemper
99.99+
senior
18 since 2021183 in total
Storing and Indexing Multiple Tables by Interesting Orderings: For Efficient Joins, Groupings, and Updates in Relational Databases
Operation-Aware Hybrid Locking for Modern In-Memory Indexes
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization
Swan: Hybrid MVCC Management for Efficient Transaction Processing in LSM-Tree-Based Key-Value Stores
+1 more
in pool
Yaoshu Wang
99.99+
senior
24 since 202130 in total
ALER: An Active Learning Hybrid System for Efficient Entity Resolution
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
Relational Deep Dive: Error-Aware Queries Over Unstructured Data
Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
+1 more
in pool
Cheng Chen 0008
99.99
junior + senior
14 since 202134 in total
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
I/O Optimizations in Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search
+1 more
in pool
Tobias Ziegler 0001
99.99+
junior + senior
19 since 202125 in total
Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized Coordination
Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services
dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data Processing
Terark-DS: A High-Performance and Storage-Efficient Key-Value Separation Storage Engine on Disaggregated Storage
+1 more
Changji Li
99.99+
senior
6 since 202110 in total
A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
Aquila: A High-Concurrency System for Incremental Graph Query
TVA: A Version-aware Temporal Graph Storage System for Real-time Analytics
ANNiE: A Learned Query Cost Estimator for Graph-Based Approximate Nearest Neighbor Search
+1 more
Srinivas Karthik
99.99+
senior
5 since 20218 in total
Vodka: Rethink Benchmarking Philosophy in HTAP Systems
AQD: Online Adaptive Query Dispatcher for HTAP Databases
Redbench: Workload Synthesis From Cloud Traces
An Experimental Evaluation of Hybrid Querying on Vectors
+1 more
Alex Thomo
99.99
senior
30 since 202187 in total
TRIM: An Efficient Framework for Exact Eccentricity Computation on Large-Scale Graphs
Theoretically and Practically Efficient Resistance Distance Computation on Large Graphs
Efficient Hyper-truss Decomposition over Hypergraphs
Scalable Approximate Biclique Counting over Large Bipartite Graphs
+1 more
Geoffrey X. Yu
99.99+
junior
4 since 20215 in total
Libra: One-Shot Parameter Sensitivity Estimation for Transfer Learning in Database Performance Prediction
AXE: A Task Decomposition Approach to Learned LSM Tuning
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines
+1 more
in pool
Dong Wen 0001
99.99
senior
49 since 202160 in total
Efficient Temporal Edge-Core Maintenance in Streaming Graphs
CEMR: An Effective Subgraph Matching Algorithm with Redundant Extension Elimination
MDS-FSM: Coverage-Based Frequent Subgraph Mining in Single Graphs
Nav-Index: A High-Performance, Adaptive Index for Shortest Path Queries in RDBMS
+1 more
Jennie Rogers
99.99
senior
6 since 202121 in total
Secure Multi-Party Sampling over Joins
Verifiable Authenticated Data Structure (V-ADS) for Analytic Queries
Doppio: Communication-Efficient and Secure Multi-Party Shuffle Differential Privacy
Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
+1 more
Zhiwen Zhang 0004
99.82
junior
16 since 202116 in total
FutureLight: An Efficient Future Traffic Data-Driven Reinforcement Learning Framework for Traffic Signal Controls
MH-GIN: Multi-scale Heterogeneous Graph-based Imputation Network for AIS Data
KAFY: An Extensible and Scalable Transformers-Based System for Trajectory Data Analysis
PrivSTD: Differentially Private Spatio-temporal Trajectory Density Data Publication
in pool
Nikos Mamoulis
99.99+
senior
38 since 2021247 in total
Continuous Query for Top-K Maximal Sum Intervals over Streaming Data
CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates
RT-RkNN: Reverse k Nearest Neighbor Queries as a Graphics Ray Casting Problem
Efficient Temporal Subgraph Management: A New Interval Index
+1 more
Tobias Maltenberger
99.99
junior + senior
5 since 20215 in total
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort
Scalable Grid-based Computation of Kendall's Tau Correlation
GPU Acceleration of SQL Analytics on Compressed Data
+1 more
Youyou Lu
99.99
senior
60 since 2021109 in total
SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous Memory
Shard: A Scalable and Resize-optimized Hash Index on Disaggregated Memory
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
Demystifying and Improving Lazy Promotion in Cache Eviction
+1 more
Jesús Camacho-Rodríguez
99.99+
senior
11 since 202119 in total
Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability
BaCon: Efficient Batch Processing of Counting Queries
Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting
Ken: An Execution Engine for Unstructured Database Systems
+1 more
in pool
Chenhao Ma 0001
99.99
junior + senior
58 since 202160 in total
AGIS: Fast Approximate Graph Pattern Mining with Structure-Informed Sampling
Finding Time-Proximity Communities in Temporal Heterogeneous Information Networks
TIMEST: Temporal Information Motif Estimator Using Sampling Trees
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
+1 more
Qiyu Liu
99.99
junior + senior
15 since 202123 in total
Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search
Sparse Neighborhood Graph-Based Approximate Nearest Neighbor Search Revisited: Theoretical Analysis and Optimization
JHQ: Johnson-Lindenstrauss Enhanced Hierarchical Quantization for High-Dimensional Approximate Nearest Neighbor Search
QBAT: Model-based Query Budget Autotuner for Clustering-based Approximate Nearest Neighbor Search
+1 more
Ioannis Mytilinis
99.98
senior
8 since 202117 in total
Fugue: Online Elasticity for Distributed Stateful Stream Processing
Incremental Stream Query Deployment under Continuous Infrastructure Changes in the Cloud-Edge Continuum
APEROL: Adaptive Parallel Edge-to-Cloud Runtime Optimization for Layered Workflow Execution
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
+1 more
in pool
Sebastian Link
99.99+
senior
33 since 2021131 in total
Storage-Centric Relation Design via High-Quality Approximate Functional Dependencies
Discovering Approximate Denial Constraints in Large Databases
Structural Normalization of Property Graphs
Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database Construction
+1 more
Ruiqi Xu 0002
99.99+
senior
6 since 202111 in total
Characterizing Parallel Subgraph Matching Performance: A Systematic Study of Interactions, Scalability, and Enumeration
X-Wim: Massive Parallelization of Weighted Matching in Bipartite Graphs
gMatch: Fine-Grained and Hardware-Efficient Subgraph Matching on GPUs
Resource-Efficient FirmCore Decomposition on Billion-scale Multilayer Graphs
+1 more
in pool
Xuanhe Zhou
99.99+
junior + senior
36 since 202138 in total
OBELISK: Efficient Offline Query Planning with Bayesian Optimization-Informed Language Model Reasoning
ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling
NeurIDA: Dynamic Modeling for Effective In-Database Analytics
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
+1 more
Junfeng Liu 0001
99.99
junior + senior
5 since 20215 in total
LiBox: A Learned Index as an Array to Minimize Last-Mile Search
Tidehunter: Large-Value Storage With Minimal Data Relocation
CounterSnake: A lossless and generalized compression framework for diverse sketches
STEM2: A Fast and Space-efficient Data Structure for Exact Multi-Set Membership Query
+1 more
Dixin Tang
99.99+
junior + senior
14 since 202124 in total
SHARP: Shared State Reduction for Efficient Matching of Sequential Patterns
Chipmink: Efficient Delta Identification for Massive Object Graphs
Window Function Optimization: Co-Evaluation and Other Techniques
Interoperable ACID Transactions for Open Table Formats
+1 more
Yuchao Tao
99.96
senior
7 since 202110 in total
Fast and Private Max-Sum Diversification
PrivMDC: Leveraging Multi-Dimensional Correlations to Answer Differentially Private Range Queries
Secure Join Operations in Multi-Identifier Databases: Performance and Practicality
Unbiased Binning for Fairness-aware Attribute Representation
+1 more
Bing Tong
99.99
junior + senior
5 since 20215 in total
Testing Graph Databases via Transformations Between Fixed-Length and Variable-Length Queries
TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
Dinkel: State-Aware and Granular Framework for Validating Graph Databases
The Data World Is Not Flat: Efficient Factorized Execution for Relational Systems
+1 more
Nikolaos Tziavelis
99.99
junior
8 since 202112 in total
Hybrid Mixed Integer Linear Programming for Large-Scale Join Order Optimisation
Towards Efficient Random-Order Enumeration for Join Queries
On Fair Epsilon Net and Geometric Hitting Set
Efficient Query Repair for Aggregate Constraints
+1 more
Shihong Gao
99.99
junior + senior
5 since 20216 in total
PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness Avoidance
FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph Training
Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph Representation
NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments
+1 more
Chunwei Liu
99.99
junior + senior
14 since 202117 in total
Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression
Accelerating String-Heavy Queries with LLM Token Tables
IncreQueryFusion: On-demand Data Fusion Framework in Dynamic Data Lakes
QStore: Quantization-Aware Compressed Model Storage
+1 more
Shantanu Sharma 0001
99.98
senior
22 since 202141 in total
Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data Analytics
Efficient and Secure Range Counting over Distributed Geographic Data with Query Range Protection
A Workload-Aware Encrypted Index for Efficient Privacy-Preserving Range Queries
SACK: Shielding Dynamic Attribute-based Access Control in Persistent Key-Value Stores
+1 more
K. Venkatesh Emani
99.99
junior
4 since 20219 in total
Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL
Blaze: Compiling JSON Schema for 10x Faster Validation
Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents
+1 more
in pool
Wenfei Fan
99.99
junior + senior
43 since 2021198 in total
Subgraph Enumeration: Beyond Tree Decomposition
Repairing Property Graphs under PG-Constraints
A Unified Query Planning Framework for Conjunctive Regular Path Queries
Love-at-First-Sight: First Answers Without the Awkward Silence in Big Knowledge Graphs
+1 more
in pool
Jorge-Arnulfo Quiané-Ruiz
99.99
senior
21 since 202168 in total
QDBO: A Real-time Quantum-augmented Database System Optimizer
Decisionhouse: Prescriptive Analytics in the Data Stack
TablePuppet: Towards a Generic Framework for Learning over Relational Tables
One Join Order Does Not Fit All: Reducing Intermediate Results with Per-Split Query Plans
+1 more
in pool
Kai Wang 0037
99.94
junior + senior
58 since 202164 in total
Revisiting the Maximum Defective Clique Problem: Faster Branching and a Tighter Upper Bound
Efficient Partition-based Approaches for Diversified Top-k Subgraph Matching
Counting HyperGraphlets via Color Coding: a Quadratic Barrier and How to Break It
Efficient GPU-Accelerated Adaptive Minimum Cost Seed Selection
+1 more
Donatello Santoro
99.99
senior
5 since 202121 in total
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
+1 more
Bryan Perozzi
99.99
senior
21 since 202146 in total
Towards A Generalizable and Expressive Graph Neural Network for Graph-Level Tasks with Theoretical Guarantees
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
Scalable GNN Explanations with Distributed Shapley Values
UniTG: A Unified System for Efficient and Seamless Textual Graph Learning
+1 more
Yuxiang Guo 0003
99.99
junior
4 since 20214 in total
MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
GAS: A Lightweight Framework for Filtered Search over Wide-table Vectors
Harmonizing Efficiency and Accuracy in Filtered Vector Search
ConANN: Conformal Approximate Nearest Neighbor Search
+1 more
Noura Alghamdi
99.86
junior
3 since 20216 in total
MS-Index: Fast Top-k Subsequence Search for Multivariate Time Series under Euclidean Distance
KDSelector: A Framework of Knowledge-Enhanced and Data-Efficient Selector Learning for Anomaly Detection Model Selection in Time Series
Optimal Approximate Matrix Multiplication over Sliding Windows
Near-Duplicate Text Alignment under Weighted Jaccard Similarity
+1 more
Chuheng Zhang
99.75
senior
19 since 202122 in total
FutureLight: An Efficient Future Traffic Data-Driven Reinforcement Learning Framework for Traffic Signal Controls
Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache Flink
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
Efficient Task Assignment for Multi-Workerset Crowdsourcing with Time and Expense Considerations
Sujaya Maiyya
99.92
junior + senior
11 since 202118 in total
Orca: Flexible Quorums Meet Dynamic Quorums
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
Remora: Scale-out Deterministic Execution for Smart Contracts
Fast Verification of Strong Database Isolation
+1 more
Yan Zhang 0117
99.99
senior
25 since 202126 in total
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
In-depth Analysis of Graph-based RAG in a Unified Framework
MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented Generation
+1 more
Torsten Hoefler
99.87
senior
150 since 2021311 in total
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
Morphing-based Compression for Data-centric ML Pipelines
+1 more
Wenjing Wang 0005
99.99
senior
6 since 20216 in total
Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
+1 more
Samson Zhou
99.96
senior
54 since 202175 in total
Efficient Banzhaf-Based Data Valuation for $k$-Nearest Neighbors Classification
Learned Static Function Data Structures
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
Algorithmic Data Minimization for Machine Learning over Internet-of-Things Data Streams
+1 more
Vincent Corvinelli
99.99
senior
10 since 202112 in total
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
Graph Transformers for Query Plan Representation: Potentials and Challenges
+1 more
Wei Hu 0007
99.99
senior
55 since 202197 in total
A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs
PINE: Extracting Correlated Token Pairs for Explainable Entity Matching
ALER: An Active Learning Hybrid System for Efficient Entity Resolution
Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-Based Interpolation
+1 more
Yida Wang 0003
99.97
senior
21 since 202129 in total
Unified Static–Dynamic Pruning for Efficient LLM Inference
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
CAPS: Cost-Aware ML Pipeline Selection
+1 more
in pool
Nan Tang 0001
99.99
senior
68 since 2021147 in total
SemBench: A Benchmark for Semantic Query Processing Engines
SQL-Exchange: Transforming SQL Queries Across Domains
Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA
TATA: An Efficient Framework for Task Transfer in Query Plan Representation
+1 more
in pool
Andreas Kipf
99.94
senior
18 since 202134 in total
One Pass to Parse Them All: Fused Parallel CSV Processing
FB*: A Compact Index for Efficient and Exact Density-based Clustering
Streaming Validation of JSON Documents Against Schemas
A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric
+1 more
Andrew Yates
99.96
senior
70 since 2021108 in total
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
CRAFT: Corpus Relatedness Analysis Using Fourier Transforms
+1 more
Zhengjie Miao
99.98
senior
13 since 202121 in total
Document-to-Database: Extraction Meets Relational Semantics
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
Human-Centered Exploration of Table Unionability
OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
+1 more
in pool
Daniel Kang 0001
99.99
junior + senior
23 since 202132 in total
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
Bolt-on, Verifiable Provenance for LLM-Powered Data Processing
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
Multi-Objective Agentic Rewrites for Unstructured Data Processing
+1 more
Xiaozhi Wang
99.95
senior
31 since 202135 in total
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs
Can we trust LLM Self-Explanations for Entity Resolution?
Data-efficient Online Training for Direct Alignment in LLMs
LLMs as Stratification Signals for KG Accuracy Evaluation
+1 more
in pool
Abolfazl Asudeh
99.85
senior
32 since 202158 in total
CaSh: Shapley Value Computation with Cache Optimization
Highly-Efficient Large-Scale k-means with Individual Fairness
Auditing for Demographic Bias in Opaque Rankings
Programmable Dataflows: Abstraction and Programming Model for Data Sharing
+1 more
Qian Sun 0005
99.99+
junior + senior
4 since 20214 in total
FutureLight: An Efficient Future Traffic Data-Driven Reinforcement Learning Framework for Traffic Signal Controls
Katia Papakonstantinopoulou
99.92
senior
10 since 202117 in total
DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point Data
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
PILOT-C: Physics-Informed Low-Distortion Optimal Trajectory Compression
in pool
Xiang Lian 0001
99.73
senior
29 since 2021110 in total
GPU-Accelerated 𝜂-threshold Decomposition for Uncertain Graphs
Noisy Interactive Graph Search: An Uncertainty-Based Approach with Online Modeling of Latent Expertise and Difficulty
Craw: A Unified and Efficient Querying Framework for Large-Scale Video Datasets
Tanu Malik
99.66
senior
19 since 202149 in total
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
HarborMaster: Rollback Detection for Trusted Distributed Computing
Pisco: An Isolation Bug Case Reduction and Deduplication Framework
Steven Euijong Whang
dblp:w/StevenEuijongWhang ·
DBLP ↗
99.99
senior
21 since 202149 in total
Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness
Stress-Testing ML Pipelines with Adversarial Data Corruption
Albert Y. Zomaya
99.80
senior
185 since 2021679 in total
Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
Meerkat: Scalable, Network-Aware Failure Recovery for the Internet of Things
Zhaoze Sun
99.83
senior
5 since 20215 in total
An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks
SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text Generation
Yu Mei 0002
99.99+
senior
5 since 20215 in total
FutureLight: An Efficient Future Traffic Data-Driven Reinforcement Learning Framework for Traffic Signal Controls