Qing Zheng

dblp:35/9182 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Lessons from Profiling and Optimizing Placement in AMR Codes
abstract
Block-structured Adaptive Mesh Refinement (AMR), while essential for improving efficiency in large-scale irregular and dynamic simulations, poses unique optimization challenges. Previous work has identified load imbalance and synchronization overhead as key obstacles to performance, but the deep understanding of complex runtime behavior needed to systematically address them remains elusive. In this paper, we integrate telemetry collection, analysis, and intervention to bridge this understanding gap. Establishing reliable, actionable telemetry required systematic tuning to eliminate cross-stack performance anomalies. Leveraging this foundation we design CPLX, a tunable placement policy balancing compute load and communication locality, improving runtime by up to$\mathbf{2 1. 6 \%}$over optimized baselines. Our experience highlights the empirical nature of placement optimization, requiring theoretical models to be grounded in observed runtime behavior.
Ankush Jain, Chuck Cranor, Qing Zheng, Dominic Manno, George Amvrosiadis, Gary Grider
CLUSTER3
2025 Digital twin updating method of railway vehicle bogies based on hybrid whale sea-horse optimization
Guofu Ding, Qing Zheng, Qinghua Du
Adv. Eng. Informatics3
2025 Manufacturing service recommendation method based on knowledge graph and graph convolutional network
Qing Zheng, Tingfeng Guo, Guofu Ding, Haizhu Zhang, Kai Zhang 0051
Adv. Eng. Informatics1
2025 Unsupervised fault detection with multi-source anomaly sensitivity enhancing convolutional autoencoder for high-speed train bogie bearings
Kai Zhang 0051, Qing Zheng, Guofu Ding, Haizhu Zhang
Expert Syst. Appl.3
2025 A generalized network with domain invariance and specificity representation for bearing remaining useful life prediction under unknown conditions
Qing Zheng, Pengtao Teng, Kai Zhang 0051, Guofu Ding, Xuwei Lai, Zhaocheng Yuan
Knowl. Based Syst.1
2025 An Open-Set Faults Diagnosis Method for Bogie Mechanical Transmission Components Based on Multi-Channel Feature-Enhanced Placeholder Learning
abstract
It is crucial to recognize unknown faults for the intelligent fault diagnosis of bogie mechanical transmission components. This is essential to ensure the safety of trains in long-term service. However, the unknown faults are hardly represented accurately because of the influence of vibration signals multipath propagation. This results in a challenge to efficiently construct open-set decision boundaries (OSDB). An open-set fault diagnosis (OSFD) method based on multichannel feature-enhanced placeholder learning is proposed to improve diagnostic reliability. It utilizes the multichannel vibration signals collected by sensors located at different positions to construct a placeholder learning network based on the association of multichannel features. It can effectively improve the association of unknown class presentations and boost the effect of unknown fault identification. Besides, the placeholder learning network based on dummy classifier enhancement is proposed to strengthen the representations of known classes and ensure the generalization of OSDB. The proposed approach is validated with case studies constructed by high-speed train bogie fault datasets and the data challenge public datasets of PHM-Beijing 2024. The results demonstrate that the proposed method can enhance the effectiveness of OSFD.
Kai Zhang 0051, Qing Zheng, Guofu Ding, Jiaohao Ma, Haizhu Zhang
IEEE Trans. Ind. Informatics3
2024 CARP: Range Query-Optimized Indexing for Streaming Data
abstract
Ingestion of data generated by high-performance scientific applications continues to stress available storage resources. Efficient range-based analyses on this data can be enabled by reordering it on attributes of interest, but require expensive post-processing sorts to realize the query benefits of reordering. In-situ indexing techniques, while write-efficient, are orders of magnitude slower at range queries than sorted indices. Range queries are necessary for analyzing continuous physical attributes and tracking phenomena such as energy bands and wave fronts. We present CARP, a scalable data partitioner for range queries that reorders data in-situ as it is streamed to storage during application I/O. Motivated by our findings that real application distributions tend to be highly skewed and dynamic, CARP dynamically discovers and adapts its data partitions to track these characteristics. As a result, CARP can approximate the query performance of a sort without any ingestion overhead, making it $5 \times$ faster than prior work.
Ankush Jain, Chuck Cranor, Qing Zheng, Bradley W. Settlemyer, George Amvrosiadis, Gary Grider
SC3
2024 Iterative updating of digital twin for equipment: Progress, challenges, and trends
Guofu Ding, Qing Zheng, Sheng Feng Qin
Adv. Eng. Informatics3
2024 Quantitative evaluation of crowd intelligence innovation system health: An ecosystem perspective
Qing Zheng, Wei Guo 0032, Guofu Ding, Haizhu Zhang, Zhong-Lin Fu, Sheng Feng Qin
Adv. Eng. Informatics1
2023 KV-CSD: A Hardware-Accelerated Key-Value Store for Data-Intensive Applications
abstract
Popular software key-value stores such as LevelDB and RocksDB are often tailored for efficient writing. Yet, they tend to also perform well on read operations. This is because while data is initially stored in a format that favors writes, it is later transformed by the DB in the background into a format that better accommodates reads. Write-optimized key-value stores can still block writes. This happens when those background workers cannot keep up with the foreground insertion workload.This paper advocates for a hardware-accelerated key-value store, enabling performance-critical operations, like background data reorganization and queries, to execute directly on storage instead of a host as existing key-value stores do. This better hides background work latency, prevents it from blocking foreground writes, and improves overall I/O efficiency. Our prototype, called KV-CSD, is a key-value based computational storage device consisting of an NVMe SSD and a System-on-a-Chip (SoC) that implements an ordered key-value store atop the SSD. Through offloaded processing, KV-CSD streamlines data insertion, reduces host-device data movement for both background data reorganization and query processing, and shows up to 10.6× lower write times and up to 7.4× faster queries compared to the current state-of-the-art software key-value stores on a real scientific dataset.
Inhyuk Park, Qing Zheng, Dominic Manno, Soonyeal Yang, Jason Lee 0004, David Bonnie, Bradley W. Settlemyer, Youngjae Kim 0001, Woosuk Chung, Gary Grider
CLUSTER2
2023 Crowdsourcing Workers Recommendation Algorithm Based on Improved Deep Matrix Factorization
abstract
In the current research of recommending the workers to the demanders, most scholars only model and analyze task information and workers preferences while ignoring demander preferences. To solve this problem, this paper proposed an improved deep matrix factorization algorithm considering both workers preference and demanders preference, and then applied it to the workers recommendation of the crowdsourcing platform. First, this algorithm constructs user preferences based on tags, using ratings of tags as the initial weights. Second, this algorithm uses Word2vec to get tag vectors, multiply it with the weight, and then average all tag vectors to get the user vector. Third, this algorithm uses the cosine similarity to calculate the user similarity matrix. Last, the user similarity matrix is incorporated into the deep matrix factorization algorithm as auxiliary information. The experimental dataset comes from the Zhubajie.com. The results show that this algorithm is better than the baseline models in both MAE and RMSE.
Yuanmeng Tian, Shuying Wang, Qing Zheng
CSCWD4
2023 Population evolution analysis in collective intelligence design ecosystem
abstract
The Collective Intelligent Design Ecosystem is a dynamic ecosystem founded on an online design platform that leverages collective intelligence to support the creation of novel products. The system's primary components are its users and designers. Maintaining the system's sustainability requires expanding the scale of the designer and user populations as it evolves to stabilize. However, the unity of ecological interactions between various populations is fragmented in contemporary studies of population-scale evolution, and the parameterization of evolutionary models is illogical. To overcome this gap, this research provides a population evolution model of collective intelligent design incorporating participants' intra- and interspecific ecological connections. The model's validity is verified by the evolutionary simulation of 110 designers and 5990 users of China's largest collective intelligence design platform, the Zhubajie platform, and illuminating conclusions are in turn drawn from this simulation. First, the designer's influence on the user is greater than the user's impact on the designer. Second, keeping current members engaged is more crucial to the system's viability than luring in new ones. Third, fostering collaboration among designers while retaining user competitiveness can promote system growth. Fourth, decreasing the reliance between particular designers and users might hasten the system's evolution.
Zhong-Lin Fu, Lei Wang 0189, Wei Guo 0032, Qing Zheng, Li-Wen Shi
Adv. Eng. Informatics4
2023 Intelligent failure localization and maintenance of network based on reliability
Qing Zheng, Fang-Ming Shao
J. Supercomput.1
2023 KVRangeDB: Range Queries for a Hash-based Key-Value Device
abstract
Key–value (KV) software has proven useful to a wide variety of applications including analytics, time-series databases, and distributed file systems. To satisfy the requirements of diverse workloads, KV stores have been carefully tailored to best match the performance characteristics of underlying solid-state block devices. Emerging KV storage device is a promising technology for both simplifying the KV software stack and improving the performance of persistent storage-based applications. However, while providing fast, predictable put and get operations, existing KV storage devices do not natively support range queries that are critical to all three types of applications described above. In this article, we present KVRangeDB, a software layer that enables processing range queries for existing hash-based KV solid-state disks (KVSSDs). As an effort to adapt to the performance characteristics of emerging KVSSDs, KVRangeDB implements log-structured merge tree key index that reduces compaction I/O, merges keys when possible, and provides separate caches for indexes and values. We evaluated the KVRangeDB under a set of representative workloads, and compared its performance with two existing database solutions: a Rocksdb variant ported to work with the KVSSD, and Wisckey, a key–value database that is carefully tuned for conventional block devices. On filesystem aging workloads, KVRangeDB outperforms Wisckey by 23.7× in terms of throughput and reduce CPU usage and external write amplifications by 14.3× and 9.8×, respectively.
Qing Zheng, Jason Lee 0004, Bradley W. Settlemyer, Fei Wen 0003, A. L. Narasimha Reddy, Paul Gratz
ACM Trans. Storage2
2022 GUFI: Fast, Secure File System Metadata Search for Both Privileged and Unprivileged Users
abstract
Modern High-Performance Computing (HPC) data centers routinely store massive data sets resulting in millions of directories and billions of files. To efficiently search and sift through these files and directories we present the Grand Unified File Index (GUFI), a novel file system metadata index that enables both privileged and regular users to rapidly locate and characterize data sets of interest. GUFI uses a hierarchical index that preserves file access permissions such that the index can be securely accessed by users while still enabling efficient, advanced analysis of storage system usage by cluster administrators. Compared with the current state-of-the-art indexing for file system metadata, GUFI is able to provide speedups of 1.5× to 230× for queries executed by administrators on a real production file system namespace. Queries executed by users, which typically cannot rely on cluster-wide indexing, see even greater speedups using GUFI.
Dominic Manno, Jason Lee 0004, Prajwal Challa, Qing Zheng, David Bonnie, Gary Grider, Bradley W. Settlemyer
SC4
2021 DeltaFS: a scalable no-ground-truth filesystem for massively-parallel computing
abstract
High-Performance Computing (HPC) is known for its use of massive concurrency. But it can be challenging for a parallel filesystem's control plane to utilize cores when every client process must globally synchronize and serialize its metadata mutations with those of other clients. We present DeltaFS, a new paradigm for distributed filesystem metadata. DeltaFS allows jobs to self-commit their namespace changes to logs, avoiding the cost of global synchronization. Followup jobs selectively merge logs produced by previous jobs as needed, a principle we term No Ground Truth which allows for efficient data sharing. By avoiding unnecessary synchronization of metadata operations, DeltaFS improves metadata operation throughput up to 98X leveraging parallelism on the nodes where job processes run. This speedup grows as job size increases. DeltaFS enables efficient inter-job communication, reducing overall workflow runtime by significantly improving client metadata operation latency up to 49X and resource usage up to 52X.
Qing Zheng, Chuck Cranor, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
SC1
2020 Mochi: Composing Data Services for High-Performance Computing Environments
Robert B. Ross, George Amvrosiadis, Philip H. Carns, Chuck Cranor, Matthieu Dorier, Kevin Harms, Gregory R. Ganger, Garth A. Gibson, Samuel K. Gutierrez, Robert Latham, Robert W. Robey, Dana Robinson, Bradley W. Settlemyer, Galen M. Shipman, Shane Snyder, Jérome Soumagne, Qing Zheng
J. Comput. Sci. Technol.17
2020 Streaming Data Reorganization at Scale with DeltaFS Indexed Massive Directories
abstract
Complex storage stacks providing data compression, indexing, and analytics help leverage the massive amounts of data generated today to derive insights. It is challenging to perform this computation, however, while fully utilizing the underlying storage media. This is because, while storage servers with large core counts are widely available, single-core performance and memory bandwidth per core grow slower than the core count per die. Computational storage offers a promising solution to this problem by utilizing dedicated compute resources along the storage processing path. We present DeltaFS Indexed Massive Directories (IMDs), a new approach to computational storage. DeltaFS IMDs harvest available (i.e., not dedicated) compute, memory, and network resources on the compute nodes of an application to perform computation on data. We demonstrate the efficiency of DeltaFS IMDs by using them to dynamically reorganize the output of a real-world simulation application across 131,072 CPU cores. DeltaFS IMDs speed up reads by 1,740× while only slightly slowing down the writing of data during simulation I/O for in situ data processing.
Qing Zheng, Chuck Cranor, Ankush Jain, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
ACM Trans. Storage1
2019 Compact Filters for Fast Online Data Partitioning
abstract
We are approaching a point in time when it will be infeasible to catalog and query data after it has been generated. This trend has fueled research on in-situ data processing (i.e. operating on data as it is streamed to storage). One important example of this approach is in-situ data indexing. Prior work has shown the feasibility of indexing at scale as a two-step process. First, one partitions data by key across the CPU cores of a parallel job. Then each core indexes its subset as data is persisted. Online partitioning requires transferring data over the network so that it can be indexed and stored by the core responsible for the data. This approach is becoming increasingly costly as new computing platforms emphasize parallelism instead of individual core performance that is crucial for communication libraries and systems software in general. In addition to indexing, scalable online data partitioning is also useful in other contexts such as load balancing and efficient compression. We present FilterKV, an efficient data management scheme for fast online data partitioning of key-value (KV) pairs. FilterKV reduces the total amount of data sent over the network and to storage. We achieve this by: (a) partitioning pointers to KV pairs instead of the KV pairs themselves and (b) using a compact format to represent and store KV pointers. Results from LANL show that FilterKV can reduce total write slowdown (including partitioning overhead) by up to 3x across 4096 CPU cores.
Qing Zheng, Chuck Cranor, Ankush Jain, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
CLUSTER1
2018 Scaling embedded in-situ indexing with deltaFS
Qing Zheng, Chuck Cranor, Danhao Guo, Gregory R. Ganger, George Amvrosiadis, Garth A. Gibson, Bradley W. Settlemyer, Gary Grider
SC1
2017 SlimDB: A Space-Efficient Key-Value Storage Engine For Semi-Sorted Data
abstract
Modern key-value stores often use write-optimized indexes and compact in-memory indexes to speed up read and write performance. One popular write-optimized index is the Log-structured merge-tree (LSM-tree) which provides indexed access to write-intensive data. It has been increasingly used as a storage backbone for many services, including file system metadata management, graph processing engines, and machine learning feature storage engines. Existing LSM-tree implementations often exhibit high write amplifications caused by compaction, and lack optimizations to maximize read performance on solid-state disks. The goal of this paper is to explore techniques that leverage common workload characteristics shared by many systems using key-value stores to reduce the read/write amplification overhead typically associated with general-purpose LSM-tree implementations. Our experiments show that by applying these design techniques, our new implementation of a key-value store, SlimDB, can be two to three times faster, use less memory to cache metadata indices, and show lower tail latency in read operations compared to popular LSM-tree implementations such as LevelDB and RocksDB.
Kai Ren 0001, Qing Zheng, Joy Arulraj, Garth A. Gibson
Proc. VLDB Endow.2
2015 ShardFS vs. IndexFS: replication vs. caching strategies for distributed metadata management in cloud storage systems
abstract
The rapid growth of cloud storage systems calls for fast and scalable namespace processing. While few commercial file systems offer anything better than federating individually non-scalable namespace servers, a recent academic file system, IndexFS, demonstrates scalable namespace processing based on client caching of directory entries and permissions (directory lookup state) with no per-client state in servers. In this paper we explore explicit replication of directory lookup state in all servers as an alternative to caching this information in all clients. Both eliminate most repeated RPCs to different servers in order to resolve hierarchical permission tests. Our realization for server replicated directory lookup state, ShardFS, employs a novel file system specific hybrid optimistic and pessimistic concurrency control favoring single object transactions over distributed transactions. Our experimentation suggests that if directory lookup state mutation is a fixed fraction of operations (strong scaling for metadata), server replication does not scale as well as client caching, but if directory lookup state mutation is proportional to the number of jobs, not the number of processes per job, (weak scaling for metadata), then server replication can scale more linearly than client caching and provide lower 70 percentile response times as well.
Kai Ren 0001, Qing Zheng, Garth A. Gibson
SoCC3
2014 IndexFS: Scaling File System Metadata Performance with Stateless Caching and Bulk Insertion
abstract
The growing size of modern storage systems is expected to exceed billions of objects, making metadata scalability critical to overall performance. Many existing distributed file systems only focus on providing highly parallel fast access to file data, and lack a scalable metadata service. In this paper, we introduce a middleware design called Index FS that adds support to existing file systems such as PVFS, Lustre, and HDFS for scalable high-performance operations on metadata and small files. Index FS uses a table-based architecture that incrementally partitions the namespace on a per-directory basis, preserving server and disk locality for small directories. An optimized log-structured layout is used to store metadata and small files efficiently. We also propose two client-based storm free caching techniques: bulk namespace insertion for creation intensive workloads such as N-N check pointing, and stateless consistent metadata caching for hot spot mitigation. By combining these techniques, we have demonstrated Index FS scaled to 128 metadata servers. Experiments show our out-of-core metadata throughput out-performing existing solutions such as PVFS, Lustre, and HDFS by 50% to two orders of magnitude.
Kai Ren 0001, Qing Zheng, Swapnil Patil 0001, Garth A. Gibson
SC2
2013 Curriculum integration for the ECE undergraduate core courses in electronics
abstract
This paper discusses the work in progress to restructure the Electronics curriculum in the Department of Electrical and Computer Engineering (ECE) in order to improve the system integration learning experience gained by the undergraduate students. The goals of restructuring the ECE Electronics curriculum are as follows: a) train and prepare students to design and analyze complex electronic systems first at the subsystem and system level before teaching and learning electronics at the component level; and b) strengthen the infrastructure for the system integration learning experience with other courses such as power electronics through the use of integrated projects developed for the Electronics curriculum. To realize these goals, the curriculum of Electronics I and Electronics II are redesigned. Electronics I is designed to focus on the design and analysis of electronic circuits, devices, and processes at the system and subsystem level. Electronics II is designed to focus on the study, operation, and analysis of electronic circuits, devices, and processes at the component level. Centralized projects are selected as platforms to allow students to develop the skills in designing and analyzing electronic systems. The students' performance and survey show that the Electronics curriculum restructure has a positive impact on students' learning.
Qing Zheng, Pengtao Lin, Fong Mak, Ramakrishnan Sundaram
FIE1
2013 COSBench: cloud object storage benchmark
abstract
With object storage systems being increasingly recognized as a preferred way to expose one's storage infrastructure to the web, the past few years have witnessed an explosion in the acceptance of these systems. Unfortunately, the proliferation of available solutions and the complexity of each individual one, coupled with a lack of dedicated workload, makes it very challenging for one to evaluate and tune the performance of different systems. To help address this problem, we present the Cloud Object Storage Benchmark (COSBench). It is a benchmark tool that we have developed at Intel with the goal of facilitating both performance comparison and system optimization of these systems. In this paper, we describe the design and implementation of this tool, focusing on its extensibility and scalability. In addition, we discuss how people can use this tool to perform system characterization and how the latter can facilitate system comparison and optimization. To demonstrate the value of our tool, we report the results of our experiments conducted on two Swift setups we built in our lab. We also share some of our experiences in turning our setups to achieve higher performance.
Qing Zheng, Haopeng Chen, Yaguang Wang, Jiangang Duan
ICPE1
2012 COSBench: A Benchmark Tool for Cloud Object Storage Services
abstract
With object storage services becoming increasingly accepted as replacements for traditional file or block systems, it is important to effectively measure the performance of these services. Thus people can compare different solutions or tune their systems for better performance. However, little has been reported on this specific topic as yet. To address this problem, we present COSBench (Cloud Object Storage Benchmark), a benchmark tool that we are currently working on in Intel for cloud object storage services. In addition, in this paper, we also share the results of the experiments we have performed so far.
Qing Zheng, Haopeng Chen, Yaguang Wang, Jiangang Duan, Zhiteng Huang
IEEE CLOUD1
2011 Work in progress - Design and delivery of the graduate course on electronic systems integration
abstract
This paper discusses the approach to improve the preparation of graduate students in Electrical Engineering for projects in the electronics industry by introducing the graduate course: Electronic Systems Design and Integration. The course emphasizes the understanding and skills necessary for students to achieve competencies at the system and sub-system level of electronic project design, test, and validation. Basic and advanced electronic circuits are studied and modeled in terms of their input-output characteristics. In this course, the circuit simulation software, the printed circuit board design software, printed circuit board maker and its related software are introduced to students. Students are required to complete several subsystem level and system level projects by employing the above software and hardware to design, build, test, and validate the final products. Through this procedure, students learn the required software and hardware for the electronic systems design and become familiar with the process to manufacture printed circuit board based electronic circuits. The purpose of offering this course is to (a) enable graduate students to apply what they learn in electronics courses during undergraduate studies to design complex circuits for electronic systems integration, and (b) to prepare students for electronic systems design and integration projects in industry.
Qing Zheng, Ramakrishnan Sundaram, Fong Mak
FIE1
2002 Multiobjective optimization of current waveforms for switched reluctance motors by genetic algorithm
abstract
In this paper a genetic algorithm (GA) is employed to determine the desired current waveforms for switched reluctance motors (SRM) through generating appropriate reference phase torques for a given desired torque using the torque sharing function (TSF). The objective is to yield smoother phase current waveforms in general, and achieve minimum phase current variations in particular. This problem is formulated into a multiobjective optimization task with certain constraints. Due to the highly nonlinear relationship between the SRM torque and current, this optimization task is an NP-hard problem. To deal with the difficulty, the problem is further coded so that a GA can be applied to facilitate the search of global minimum. Simulation results verify the effectiveness of the proposed method.
Jianxin Xu 0001, Sanjib Kumar Panda, Qing Zheng
IEEE Congress on Evolutionary Computation3