VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0097
dblp:10/4661-97
· DBLP profile ↗
16ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-6921-4926ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceabstractLarge language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services. Chenxu Niu 0001, Wei Zhang 0097, Jie Li 0057, Tongyang Wang, Xi Wang 0009, Yong Chen 0001 |
AAAI | 2 |
| 2025 | Distributed Metadata Querying on HPC SystemsabstractEfficient execution of range queries and exact queries on metadata is critical for many scientific workflows on large-scale high-performance computing (HPC) systems. Range queries in a distributed setting pose a challenge, as typical partitioning methods such as hash-based, list-based, and range-based partitioning do not work efficiently. There are also challenges in balancing the workload using locality-sensitive partitioning methods, which are capable of routing range queries to a subset of nodes to achieve efficiency. In this paper, we study prefix-hash tree and range-hash partitioning methods that fit the requirements of efficient range queries in distributed HPC systems. We also proposed a novel load-balancing method inspired by the Gossip protocol, which has similar complexity for most operations while keeping the coefficient of variation (CV) less than 7% across diverse metadata workload distributions, i.e., normal, uniform, and exponential. Our proposed methods outperform the commonly used Key-Value store, RocksDB, for exact queries by up to$15 \times$and for range queries by up to$120,000 \times$. Suben Kumer Saha, Houjan Tang, Wei Zhang 0097, Surendra Byna |
HiPC | 3 |
| 2025 | ICEAGE: Intelligent Contextual Exploration and Answer Generation Engine for Scientific Data Discovery
Chenxu Niu 0001, Wei Zhang 0097, Mert Side, Yong Chen 0001 |
SSDBM | 2 |
| 2024 | Evaluating Performance Trade-offs of Caching Strategies for AI-Powered Querying SystemsabstractWith the rapid growth of accumulated data from various scientific domains, traditional data management systems face challenges in supporting complicated queries, such as pattern search, on massive amounts of data. To serve sophisticated queries through capturing precise features from data, recent data management systems seek to use artificial intelligence (AI) within the querying process. However, the characteristic of AI inference workflow within the querying process, such as intensive computation and expensive requirements for computing resources, becomes a bottleneck of the AI-powered query systems.In this paper, we provide a generalization of AI inference workflow in the context of AI-powered data discovery and we introduce three different caching strategies corresponding to each stage in the AI inference workflow. We provide in-depth performance evaluation on the impact of these caching strategies through a series of strong scaling experiments. Our experimental results show that the AI-powered data querying performance can be significantly improved by applying different caching strategies. Hyunju Oh, Wei Zhang 0097, Christopher D. Rickett, Sreenivas R. Sukumar 0001, Surendra Byna |
IEEE Big Data | 2 |
| 2024 | IDIOMS: Index-powered Distributed Object-centric Metadata Search for Scientific Data ManagementabstractAffix-oriented metadata search is one of the essential fuzzy search capabilities that allow users to find data of interest in their voluminous data set with incomplete query conditions. With the recent transition towards object-centric data management systems in the science community, there is a paramount need for the support of such features in distributed settings. However, existing metadata search solutions either do not support efficient affix-oriented metadata search or do not suit well in a distributed setting of object-centric data management systems. To bridge this gap, we introduce IDIOMS, a metadata search solution underpinned by a distributed metadata index, meticulously designed to enable high-performance affix-oriented metadata search for parallel object-centric storage. One of the standout features of IDIOMS is its efficiency in supporting four distinct types of highly demanded metadata queries. Furthermore, IDIOMS is flexibly catering to both independent and collective metadata search operations. Our experimental comparisons with SoMeta, a state-of-the-art metadata query method, demonstrate more than 400× performance boost for independent queries and up to 300× performance improvements for collective queries, while keeping a small index management overhead. Wei Zhang 0097, Houjun Tang, Surendra Byna |
CCGrid | 1 |
| 2023 | PSQS: Parallel Semantic Querying Service for Self-describing File FormatsabstractFinding relevant datasets can be a time-consuming and challenging task, especially for self-describing file formats. Current solutions use either exact or partial keyword matching approaches to extract and process metadata queries, but they fail to capture semantic relationships between the metadata content and query keywords. To address this challenge, we introduce PSQS, a novel parallel semantic search method for self-describing files. The method leverages parallel processing and kv2vec semantic similarity measures to retrieve semantically relevant data efficiently. Our evaluation against existing metadata search solutions shows that PSQS offers a new, efficient and effective semantic search functionality for various fields where large self-describing files are used, such as scientific data management, leading to more accurate and efficient data retrieval. Chenxu Niu 0001, Wei Zhang 0097, Surendra Byna, Yong Chen 0001 |
IEEE Big Data | 2 |
| 2021 | Exploiting user activeness for data retention in HPC systemsabstractHPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology. Wei Zhang 0097, Surendra Byna, Hyogi Sim, Sankeun Lee 0001, Sudharshan S. Vazhkudai, Yong Chen 0001 |
SC | 1 |
| 2019 | Exploring Metadata Search Essentials for Scientific Data ManagementabstractScientific experiments and observations store massive amounts of data in various scientific file formats. Metadata, which describes the characteristics of the data, is commonly used to sift through massive datasets in order to locate data of interest to scientists. Several indexing data structures (such as hash tables, trie, self-balancing search trees, sparse array, etc.) have been developed as part of efforts to provide an efficient method for locating target data. However, efficient determination of an indexing data structure remains unclear in the context of scientific data management, due to the lack of investigation on metadata, metadata queries, and corresponding data structures. In this study, we perform a systematic study of the metadata search essentials in the context of scientific data management. We study a real-world astronomy observation dataset and explore the characteristics of the metadata in the dataset. We also study possible metadata queries based on the discovery of the metadata characteristics and evaluate different data structures for various types of metadata attributes. Our evaluation on real-world dataset suggests that trie is a suitable data structure when prefix/suffix query is required, otherwise hash table should be used. We conclude our study with a summary of our findings. These findings provide a guideline and offers insights in developing metadata indexing methodologies for scientific applications. Wei Zhang 0097, Surendra Byna, Chenxu Niu 0001, Yong Chen 0001 |
HiPC | 1 |
| 2019 | MIQS: metadata indexing and querying service for self-describing file formatsabstractScientific applications often store datasets in self-describing data file formats, such as HDF5 and netCDF. Regrettably, to efficiently search the metadata within these files remains challenging due to the sheer size of the datasets. Existing solutions extract the metadata and store it in external database management systems (DBMS) to locate desired data. However, this practice introduces significant overhead and complexity in extraction and querying. In this research, we propose a novel Metadata Indexing and Querying Service (MIQS), which removes the external DBMS and utilizes in-memory index to achieve efficient metadata searching. MIQS follows the self-contained data management paradigm and provides portable and schema-free metadata indexing and querying functionalities for self-describing file formats. We have evaluated MIQS with the state-of-the-art MongoDB-based metadata indexing solution. MIQS achieved up to 99% time reduction in index construction and up to 172kx search performance improvement with up to 75% reduction in memory footprint. Wei Zhang 0097, Surendra Byna, Houjun Tang, Brody Williams, Yong Chen 0001 |
SC | 1 |
| 2019 | Improving Nighttime Light Imagery With Location-Based Social Media DataabstractLocation-based social media have been extensively utilized in the concept of “social sensing” to exploit dynamic information about human activities, yet joint uses of social sensing and remote sensing images are underdeveloped at present. In this paper, the close relationship between the number of Twitter users and brightness of nighttime lights (NTL) over the contiguous United States is calculated and geotagged tweets are then used to upsample a stable light image for 2013. An associated outcome of the upsampling process is the solution of two major problems existing in the NTL image, pixel saturation, and blooming effects. Compared with the original stable light image, digital number (DN) values of the upsampled stable light image have larger correlation coefficients with gridded population (0.47 versus 0.09) and DN values of the new generation NTL image product (0.56 versus 0.52), i.e., the Visible Infrared Imaging Radiometer Suite day/night band image composite. In addition, total personal incomes of states are disaggregated to each pixel in proportion to the DN value of the pixel in the NTL images and then aggregate by counties. Personal incomes distributed by the upsampled NTL image are closer to the official demographic data than those distributed by the original stable light image. All of these results explore the potential of geotagged tweets to improve the quality of NTL images for more accurately estimating or mapping socioeconomic factors. Naizhuo Zhao, Wei Zhang 0097, Eric L. Samson, Yong Chen 0001, Guofeng Cao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Managing Rich Metadata in High-Performance Computing Systems Using a Graph ModelabstractHigh-performance computing (HPC) systems generate huge amounts of metadata about different entities such as jobs, users, and files. Existing systems can efficiently record and manage part of these metadata, mainly the POSIX metadata of data files (e.g., file size, name, and permissions mode). But another important set of metadata, referred to as “rich” metadata in this study, which record not only wider range of entities (e.g., running processes and jobs) but also more complex relationships between them, are mostly missing in current HPC systems. Yet such rich metadata are critical for supporting many advanced data management functions such as identifying data sources and parameters behind a given result; auditing data usage; or understanding details about how inputs are transformed into outputs. To uniformly and efficiently manage the rich metadata generated in HPC systems, We propose to utilize a graph model in this study. We identify the key challenges of implementing such a graph-based HPC rich metadata management system and present GraphMeta, a graph-based rich metadata management system designed and optimized for HPC platforms, to tackle these challenges. Extensive evaluations on both synthetic and real HPC metadata workloads show its advantages in both performance and scalability compared with existing solutions. Dong Dai 0001, Yong Chen 0001, Philip H. Carns, John Jenkins, Wei Zhang 0097, Robert B. Ross |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | DART: distributed adaptive radix tree for efficient affix-based keyword search on HPC systemsabstractAffix-based search is a fundamental functionality for storage systems. It allows users to find desired datasets, where attributes of a dataset match an affix. While building inverted index to facilitate efficient affix-based keyword search is a common practice for standalone databases and for desktop file systems, building local indexes or adopting indexing techniques used in a standalone data store is insufficient for high-performance computing (HPC) systems due to the massive amount of data and distributed nature of the storage devices within a system. In this paper, we propose Distributed Adaptive Radix Tree (DART), to address the challenge of distributed affix-based keyword search on HPC systems. This trie-based approach is scalable in achieving efficient affix-based search and alleviating imbalanced keyword distribution and excessive requests on keywords at scale. Our evaluation at different scales shows that, comparing with the "full string hashing" use case of the most popular distributed indexing technique - Distributed Hash Table (DHT), DART achieves up to 55× better throughput with prefix search and with suffix search, while achieving comparable throughput with exact and infix searches. Also, comparing to the "initial hashing" use case of DHT, DART maintains a balanced keyword distribution on distributed nodes and alleviates excessive query workload against popular keywords. Wei Zhang 0097, Houjun Tang, Surendra Byna, Yong Chen 0001 |
PACT | 1 |
| 2018 | AKIN: A Streaming Graph Partitioning Algorithm for Distributed Graph Storage SystemsabstractMany graph-related applications face the challenge of managing excessive and ever-growing graph data in a distributed environment. Therefore, it is necessary to consider a graph partitioning algorithm to distribute graph data onto multiple machines as the data comes in. Balancing data distribution and minimizing edge-cut ratio are two basic pursuits of the graph partitioning problem. While achieving balanced partitions for streaming graphs is easy, existing graph partitioning algorithms either fail to work on streaming workloads, or leave edge-cut ratio to be further improved. Our research aims to provide a better solution that fits the need of streaming graph partitioning in a distributed system, which further reduces the edge-cut ratio while maintaining rough balance among all partitions. We exploit the similarity measure on the degree of vertices to gather structuralrelated vertices in the same partition as much as possible, this reduces the edge-cut ratio even further as compared to the state-of-the-art streaming graph partitioning algorithm - FENNEL. Our evaluation shows that our streaming graph partitioning algorithm is able to achieve better partitioning quality in terms of edge-cut ratio (up to 20% reduction as compared to FENNEL) while maintaining decent balance between all partitions, and such improvement applies to various real-life graphs. Wei Zhang 0097, Yong Chen 0001, Dong Dai 0001 |
CCGrid | 1 |
| 2017 | IOGP: An Incremental Online Graph Partitioning Algorithm for Distributed Graph DatabasesabstractGraphs have become increasingly important in many applications and domains such as querying relationships in social networks or managing rich metadata generated in scientific computing. Many of these use cases require high-performance distributed graph databases for serving continuous updates from clients and, at the same time, answering complex queries regarding the current graph. These operations in graph databases, also referred to as online transaction processing (OLTP) operations, have specific design and implementation requirements for graph partitioning algorithms. In this research, we argue it is necessary to consider the connectivity and the vertex degree changes during graph partitioning. Based on this idea, we designed an Incremental Online Graph Partitioning (IOGP) algorithm that responds accordingly to the incremental changes of vertex degree. IOGP helps achieve better locality, generate balanced partitions, and increase the parallelism for accessing high-degree vertices of the graph. Over both real-world and synthetic graphs, IOGP demonstrates as much as 2x better query performance with a less than 10% overhead when compared against state-of-the-art graph partitioning algorithms. Dong Dai 0001, Wei Zhang 0097, Yong Chen 0001 |
HPDC | 2 |
| 2017 | POSTER: IOGP: An Incremental Online Graph Partitioning for Large-Scale Distributed Graph DatabasesabstractLarge-scale graphs are becoming critical in various domains such as social network, scientific application, knowledge discovery, and even system software, etc. Many of those use cases require large-scale high-performance graph databases, which are designed for serving continuous updates from the clients, and at the same time, answering complex queries towards the current graph in an on-line manner. Those operations in graph databases, also referred as OLTP (online transaction processing) operations, need specific design and implementation in graph partitioning algorithms. In this study, we designed an incremental online graph partitioning (IOGP), optimized for OLTP workloads. It is designed to achieve better locality, generate balanced partitions, and increase the parallelism for accessing hotspots of the graph. Our evaluation results on both real world and synthetic graphs in both simulation and real system confirm a better performance on graph queries (as much as 2X) with small overheads during graph insertion (less than 10%). Dong Dai 0001, Wei Zhang 0097, Yong Chen 0001 |
PPoPP | 2 |
| 2016 | GraphMeta: A Graph-Based Engine for Managing Large-Scale HPC Rich MetadataabstractHigh-performance computing (HPC) systems face increasingly critical metadata management challenges, especially in the approaching exascale era. These challenges arise not only from exploding metadata volumes but also from increasingly diverse metadata, which contains data provenance and user-defined attributes in addition to traditional POSIX metadata. This "rich" metadata is critical to support many advanced data management functionality such as data auditing and validation. In our prior work, we presented a graph-based model that could be a promising solution to uniformly manage such rich metadata because of its flexibility and generality. At the same time, however, graph-based rich metadata management introduces significant challenges. In this study, we first identify the challenges presented by the underlying infrastructure in supporting scalable, high-performance rich metadata management. To tackle these challenges, we then present GraphMeta, a graph-based engine designed for managing large-scale rich metadata. We also utilize a series of optimizations designed for rich metadata graphs. We evaluate GraphMeta with both synthetic and real HPC metadata workloads and compare it with other approaches. The results show that its advantages in terms of rich metadata management in HPC systems, including better performance and scalability compared with existing solutions. Dong Dai 0001, Yong Chen 0001, Philip H. Carns, John Jenkins, Wei Zhang 0097, Robert B. Ross |
CLUSTER | 5 |