EDBT 2026 Demo / reviewers in the wild / expert
Ziqiang Feng
dblp:02/5202
· DBLP profile ↗
12ranked-venue papers
5as first author
2since 2021 · last 2023
0000-0002-0787-2448ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Computer networks · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Query processing and optimization · 52% Indexing and storage engines · 29% Data stream processing · 16% | |
| Computer networks
1 paper |
Edge and fog computing · 77% Internet of things and sensor networks · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Edge and fog computing
distributed learning |
0.7 | 1 | 2023 | Low-Bandwidth Self-Improving Transmission of Rare Training Data · MobiCom 2023 |
Indexing and storage engines › column store
main-memory column store |
0.3 | 2 | 2016 | Fast Multi-Column Sorting in Main-Memory Column-Stores · SIGMOD Conference 2016 Accelerating aggregation using intra-cycle parallelism · ICDE 2015 |
Query processing and optimization
sorting |
0.2 | 1 | 2016 | Fast Multi-Column Sorting in Main-Memory Column-Stores · SIGMOD Conference 2016 |
Query processing and optimization
aggregation |
0.2 | 1 | 2015 | Accelerating aggregation using intra-cycle parallelism · ICDE 2015 |
Indexing and storage engines › in-memory storage
in-memory data layout |
0.2 | 1 | 2015 | ByteSlice: Pushing the Envelop of Main Memory Data Processing with a New Storage Layout · SIGMOD Conference 2015 |
Cloud and datacenter computing
database-as-a-service |
0.2 | 1 | 2015 | Thrifty: Offering Parallel Database as a Service using the Shared-Process Approach · SIGMOD Conference 2015 |
Cloud and datacenter computing
multi-tenancy |
0.2 | 1 | 2015 | Thrifty: Offering Parallel Database as a Service using the Shared-Process Approach · SIGMOD Conference 2015 |
Internet of things and sensor networks › data dissemination
sensor data transmission |
0.2 | 1 | 2023 | Low-Bandwidth Self-Improving Transmission of Rare Training Data · MobiCom 2023 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2015 | Thrifty: Offering Parallel Database as a Service using the Shared-Process Approach · SIGMOD Conference 2015 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.7semi-supervised learning · 0.7few-shot learning · 0.7diversity sampling · 0.7active learning · 0.7SIMD · 0.5tenant placement · 0.4query routing · 0.4pattern template matching · 0.3iceberg cell computation · 0.3code massaging · 0.2intra-cycle parallelism · 0.2bit-parallel algorithms · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Low-Bandwidth Self-Improving Transmission of Rare Training DataabstractA severe bandwidth mismatch between incoming sensor data rate and wireless backhaul bandwidth often exists on unmanned probes when collecting new training data for machine learning (ML). To overcome this mismatch, we describe a self-improving ML-based transmission system called Hawk. Starting from a weak model that is trained on just a few examples, it seamlessly pipelines semi-supervised learning, active learning, and transfer learning, with asynchronous bandwidth-sensitive data transmission to a distant human for labeling. When a significant number of true positives (TPs) have been labeled, Hawk trains an improved model to replace the old model. This iterative workflow, called Live Learning, continues until a sufficient number of TPs have been collected. For very rare events on challenging datasets, and bandwidths as low as 12 kbps, a team of 7 probes using Hawk discovers up to 87% of the TPs that could have been discovered via full preview, transmission and labeling of all mission data. Hawk also uses diversity sampling and few-shot learning. Shilpa Anna George, Haithem Turki, Ziqiang Feng, Deva Ramanan, Padmanabhan Pillai, Mahadev Satyanarayanan |
MobiCom | 3 |
| 2022 | ByteStore: Hybrid Layouts for Main-Memory Column StoresabstractThe performance of main memory column stores highly depends on the scan and lookup operations on the base column layouts. Existing column-stores adopt a homogeneous column layout, leading to sub-optimal performance on real workloads since different columns possess different data characteristics. In this paper, we propose ByteStore, a column store that uses different storage layouts and corresponding encoding methods for different columns. Extensive experiments show that ByteStore outperforms homogeneous storage engines by up to 5.2×. Ziqiang Feng, Eric Lo 0001, Hailin Qin |
IEEE Big Data | 2 |
| 2017 | Live Synthesis of Vehicle-Sourced Data Over 4G LTEabstractAccurate, up-to-date maps of transient traffic and hazards are invaluable to drivers, city managers, and the emerging class of self-driving vehicles. We present LiveMap, a scalable, automated system for acquiring, curating, and disseminating detailed, continually-updated road conditions in a region. LiveMap leverages in-vehicle cameras, sensors, and processors to crowd-source hazard detection without human intervention. We build a real-time simulation framework that allows a mix of real and simulated components to be tested together at scale. We demonstrate that LiveMap can work well at city scales within the limits of today's cellular network bandwidth. We also show the feasibility of accurate, in-vehicle, computer-vision-based hazard detection. Wenlu Hu, Ziqiang Feng, Jan Harkes, Padmanabhan Pillai, Mahadev Satyanarayanan |
MSWiM | 2 |
| 2017 | Efficient Pattern-Based Aggregation on Sequence DataabstractA Sequence OLAP(S-OLAP) system provides a platform on which pattern-based aggregate (PBA) queries on a sequence database are evaluated. In its simplest form, a PBA query consists of a pattern template T and an aggregate function F. A pattern template is a sequence of variables, each is defined over a domain. Each variable is instantiated with all possible values in its corresponding domain to derive all possible patterns of the template. Sequences are grouped based on the patterns they possess. The answer to a PBA query is a sequence cuboid (s-cuboid), which is a multidimensional array of cells. Each cell is associated with a pattern instantiated from the query's pattern template. The value of each s-cuboid cell is obtained by applying the aggregate function F to the set of data sequences that belong to that cell. Since a pattern template can involve many variables and can be arbitrarily long, the induced s-cuboid for a PBA query can be huge. For most analytical tasks, however, only iceberg cells with very large aggregate values are of interest. This paper proposes an efficient approach to identifying and evaluating iceberg cells of s-cuboids. Experimental results show that our algorithms are orders of magnitude faster than existing approaches. Zhian He, Petrie Wong, Ben Kao, Eric Lo 0001, Reynold Cheng, Ziqiang Feng |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | Competitive Distributed Spectrum Access in QoS-Constrained Cognitive Radio Networks
Ziqiang Feng, Ian J. Wassell |
GLOBECOM | 1 |
| 2016 | Dynamic power control and optimization scheme for QoS-constrained cooperative wireless sensor networksabstractCooperative transmission can significantly reduce the power consumption associated with long distance transmission in wireless sensor networks (WSNs). In this paper, we analyze the optimal power consumption of cluster-based multi-hop transmission for a cooperative WSN. With specific Quality of Service (QoS) constraints on delay and channel capacity, we show that the power optimization problem of the whole network has no closed-form solution in a slow flat Rayleigh fading environment. Thus we propose a dynamic power control and optimization (DPCO) scheme that can jointly determine the optimal number of cooperative sensors and their transmission power. We further propose a channel approximation algorithm that can significantly reduce the computational complexity of the DPCO scheme. Ziqiang Feng, Ian J. Wassell |
ICC | 1 |
| 2016 | Fast Multi-Column Sorting in Main-Memory Column-StoresabstractSorting is a crucial operation that could be used to implement SQL operators such as GROUP BY, ORDER BY, and SQL:2003 PARTITION BY. Queries with multiple attributes in those clauses are common in real workloads. When executing queries of that kind, state-of-the-art main-memory column-stores require one round of sorting per input column. With the advent of recent fast scans and denormalization techniques, that kind of multi-column sorting could become a bottleneck. In this paper, we propose a new technique called "code massaging", which manipulates the bits across the columns so that the overall sorting time can be reduced by eliminating some rounds of sorting and/or by improving the degree of SIMD data level parallelism. Empirical results show that a main-memory column-store with code massaging can achieve speedup of up to 4.7X, 4.7X, 4X, and 3.2X on TPC-H, TPC-H skew, TPC-DS, and real workload, respectively. Wenjian Xu, Ziqiang Feng, Eric Lo 0001 |
SIGMOD Conference | 2 |
| 2015 | TLB misses: The Missing Issue of Adaptive Radix Tree?abstractEfficient main-memory index structures are crucial to main-memory database systems. Adaptive Radix Tree (ART) is the most recent in-memory index structure. ART is designed to avoid cache miss, leverage SIMD data parallelism, minimize branch mis-prediction, and have small memory footprint. When an in-memory index structure like ART has significantly few cache misses and branch mis-predictions, it is natural to question whether misses in Translation Lookaside Buffer (TLB) matters. In this paper, we try to confirm whether this is the case and if the answer is positive, what are the measures that we can take to alleviate that and how effective they are. Petrie Wong, Ziqiang Feng, Wenjian Xu, Eric Lo 0001, Ben Kao |
DaMoN | 2 |
| 2015 | Accelerating aggregation using intra-cycle parallelismabstractModern CPUs have word width of 64 bits but real data values are usually represented using bits fewer than a CPU word. This underutilization of CPU at register level has motivated the recent development of bit-parallel algorithms that carry out data processing operations (e.g., filter scan) on CPU words packed with data values (e.g., 8 data values are packed into one 64-bit word). Bit-parallel algorithms fully unleash the intra-cycle parallelism of modern CPUs and they are especially attractive to main-memory column stores whose goal is to process data at the speed of the “bare metal”. Main-memory column stores generally focus on analytical queries, where aggregation is a common operation. Current bit-parallel algorithms, however, have not covered aggregation yet. In this paper, we present a suite of bit-parallel algorithms to accelerate all standard aggregation operations: SUM, MIN, MAX, AVG, MEDIAN, COUNT. The algorithms are designed to fully leverage the intra-cycle parallelism in CPU cores when aggregating words of packed values. Experimental evaluation shows that our bit-parallel aggregation algorithms exhibit significant performance benefits compared with non-bit-parallel methods. Ziqiang Feng, Eric Lo 0001 |
ICDE | 1 |
| 2015 | ByteSlice: Pushing the Envelop of Main Memory Data Processing with a New Storage LayoutabstractScan and lookup are two core operations in main memory column stores. A scan operation scans a column and returns a result bit vector that indicates which records satisfy a filter. Once a column scan is completed, the result bit vector is converted into a list of record numbers, which is then used to look up values from other columns of interest for a query. Recently there are several in-memory data layout proposals that aim to improve the performance of in-memory data processing. However, these solutions all stand at either end of a trade-off --- each is either good in lookup performance or good in scan performance, but not both. In this paper we present ByteSlice, a new main memory storage layout that supports both highly efficient scans and lookups. ByteSlice is a byte-level columnar layout that fully leverages SIMD data-parallelism. Micro-benchmark experiments show that ByteSlice achieves a data scan speed at less than 0.5 processor cycle per column value --- a new limit of main memory data scan, without sacrificing lookup performance. Our experiments on TPC-H data and real data show that ByteSlice offers significant performance improvement over all state-of-the-art approaches. Ziqiang Feng, Eric Lo 0001, Ben Kao, Wenjian Xu |
SIGMOD Conference | 1 |
| 2015 | Thrifty: Offering Parallel Database as a Service using the Shared-Process ApproachabstractRecently, Amazon has announced Redshift, a Parallel-Database-as-a Service (PDaaS). Redshift adopts the "virtual cluster" approach to implement multitenancy, which has the merit of hard isolation among tenants (i.e., tenants do not interfere even when sharing resources). However, that benefit comes with poor resource utilization due to the significant redundancy incurred in the resources. In this demonstration, we present Thrifty, a Parallel-Database-as-a-Service operated using the "shared-process" approach. Compared with Redshift, each tenant in Thrifty does not occupy an exclusive amount of resource but share the database processes together, leading to better resource utilization. To avoid contention among tenants, Thrifty uses a proper cluster design, a tenant placement scheme, and a query routing mechanism to achieve soft isolation. In the demonstration, an attendee will be invited to register with Thrifty as a tenant to rent a parallel database instance. Then the attendee will be allowed to view the dashboard of a Thrifty's administrator. Next, the attendee will be invited to control (e.g., increase) the workload of the tenant so as to see how Thrifty carries out online re- consolidation and elastic scaling. Petrie Wong, Zhian He, Ziqiang Feng, Wenjian Xu, Eric Lo 0001 |
SIGMOD Conference | 3 |
| 2013 | Cell selection in two-tier femtocell networks with open/closed access using evolutionary gameabstractCell selection is an important issue in femtocell networks, which can balance the utilization of the whole network. In this paper, we investigate cell selection problem in a two-tier femtocell network that contains a micro base station (MBS) and several femtocells with different access methods and coverage areas. We propose the evolutionary game model to describe the dynamics of the cell selection process and consider the evolutionary equilibrium as the solution. In order to achieve the evolutionary equilibrium, we introduce the reinforcement learning algorithm that can help distributed individual users make selection decisions independently. With their own knowledge of the past, the users can learn to achieve the evolutionary equilibrium without complete knowledge of other users. Finally, the performance of the evolutionary game and reinforcement learning algorithm is analyzed, and simulation results show the convergence and effectiveness of the proposed algorithm. Ziqiang Feng, Lingyang Song, Zhu Han 0001, Dusit Niyato, Xiaowu Zhao |
WCNC | 1 |