EDBT 2026 Demo / reviewers in the wild / expert
Yanqing Peng
dblp:16/10885
· DBLP profile ↗
11ranked-venue papers
2as first author
3since 2021 · last 2021
0000-0002-0805-7795ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Computer networks · 4Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Constrained Non-Affine Alignment of EmbeddingsabstractEmbeddings are one of the fundamental building blocks for data analysis tasks. Embeddings are already essential tools for large language models and image analysis, and their use is being extended to many other research domains. The generation of these distributed representations is often a data-and computation-expensive process; yet the holistic analysis and adjustment of them after they have been created is still a developing area. In this paper, we first propose a very general quantitatively measure for the presence of features in the embedding data based on if it can be learned. We then devise a method to remove or alleviate undesired features in the embedding while retaining the essential structure of the data. We use a Domain Adversarial Network (DAN) to generate a non-affine transformation, but we add constraints to ensure the essential structure of the embedding is preserved. Our empirical results demonstrate that the proposed algorithm significantly outperforms the state-of-art unsupervised algorithm on several data sets, including novel applications from the industry. Yan Zheng 0001, Yanqing Peng, Chin-Chia Michael Yeh, Zhongfang Zhuang, Mahashweta Das, Mangesh Bendre, Feifei Li 0001, Wei Zhang 0189, Jeff M. Phillips |
ICDM | 3 |
| 2021 | At-the-time and Back-in-time Persistent SketchesabstractIn the era of big data, more and more applications require the information of historical data to support rich analytics, learning, and mining operations. In these cases, it is highly desirable to retrieve information of previous versions of data. Traditionally, multi-version databases can be used to store all historical values of the data in order to support historical queries. However, storing all the historical data can be impractical due to its large space consumption. In this paper, we propose the concept of at-the-time persistent (ATTP) and back-in-time persistent (BITP) sketches, which are sketches that approximately answer queries on previous versions of data with small space. We then provide several implementations of ATTP/BITP sketches which are shown to be more efficient compared to existing state-of-the-art solutions in our empirical studies. Benwei Shi, Zhuoyue Zhao 0001, Yanqing Peng, Feifei Li 0001, Jeff M. Phillips |
SIGMOD Conference | 3 |
| 2021 | VeriDB: An SGX-based Verifiable DatabaseabstractThe emergence of trusted hardwares (such as Intel SGX) provides a new avenue towards verifiable database. Such trust hardwares act as an additional trust anchor, allowing great simplification and, in turn, performance improvement in the design of verifiable databases. In this paper, we introduce the design and implementation of VeriDB, an SGX-based verifiable database that supports relational tables, multiple access methods and general SQL queries. Built on top of write-read consistent memory, VeriDB provides verifiable page-structured storage, where results of storage operations can be efficiently verified with low, constant overhead. VeriDB further provides verifiable query execution that supports general SQL queries. Through a series of evaluation using practical workload, we demonstrate that VeriDB incurs low overhead for achieving verifiability: an overhead of 1-2 microseconds for read/write operations, and a 9% - 39% overhead for representative analytical workloads. Wenchao Zhou, Yifan Cai 0001, Yanqing Peng, Sheng Wang 0011, Feifei Li 0001 |
SIGMOD Conference | 3 |
| 2020 | Two-Level Data Compression using Machine Learning in Time Series DatabaseabstractThe explosion of time series advances the development of time series databases. To reduce storage overhead in these systems, data compression is widely adopted. Most existing compression algorithms utilize the overall characteristics of the entire time series to achieve high compression ratio, but ignore local contexts around individual points. In this way, they are effective for certain data patterns, and may suffer inherent pattern changes in real-world time series. It is therefore strongly desired to have a compression method that can always achieve high compression ratio in the existence of pattern diversity. In this paper, we propose a two-level compression model that selects a proper compression scheme for each individual point, so that diverse patterns can be captured at a fine granularity. Based on this model, we design and implement AMMMO framework, where a set of control parameters is defined to distill and categorize data patterns. At the top level, we evaluate each sub-sequence to fill in these parameters, generating a set of compression scheme candidates (i.e., major mode selection). At the bottom level, we choose the best scheme from these candidates for each data point respectively (i.e., sub-mode selection). To effectively handle diverse data patterns, we introduce a reinforcement learning based approach to learn parameter values automatically. Our experimental evaluation shows that our approach improves compression ratio by up to 120% (with an average of 50%), compared to other time-series compression methods. Xinyang Yu, Yanqing Peng, Feifei Li 0001, Sheng Wang 0011, Huijun Mai |
ICDE | 2 |
| 2020 | FalconDB: Blockchain-based Collaborative DatabaseabstractNowadays an emerging class of applications are based oncollaboration over a shared database among different entities. However, the existing solutions on shared database may require trust on others, have high hardware demand that is unaffordable for individual users, or have relatively low performance. In other words, there is a trilemma among security, compatibility and efficiency. In this paper, we present FalconDB, which enables different parties with limited hardware resources to efficiently and securely collaborate on a database. FalconDB adopts database servers with verification interfaces accessible to clients and stores the digests for query/update authentications on a blockchain. Using blockchain as a consensus platform and a distributed ledger, FalconDB is able to work without any trust on each other. Meanwhile, FalconDB requires only minimal storage cost on each client, and provides anywhere-available, real-time and concurrent access to the database. As a result, FalconDB over-comes the disadvantages of previous solutions, and enables individual users to participate in the collaboration with high efficiency, low storage cost and blockchain-level security guarantees. Yanqing Peng, Min Du 0003, Feifei Li 0001, Raymond Cheng 0001, Dawn Song |
SIGMOD Conference | 1 |
| 2020 | ARETE: On Designing Joint Online Pricing and Reward Sharing Mechanisms for Mobile Data MarketsabstractAlthough data has become an important kind of commercial goods, there are few appropriate online platforms to facilitate the trading of mobile crowd-sensed data so far. In this paper, we present the first architecture of mobile crowd-sensed data market, and conduct an in-depth study of the design problem of online data pricing and reward sharing. To build a practical mobile crowd-sensed data market, we have to consider four major design challenges: data uncertainty, economic-robustness (arbitrage-freeness in particular), profit maximization, and fair reward sharing. By jointly considering the design challenges, we propose an online query-bAsed cRowd-sensEd daTa pricing mEchanism, namely ARETE-PR, to determine the trading price of crowd-sensed data. Our theoretical analysis shows that ARETE-PR guarantees both arbitrage-freeness and a constant competitive ratio in terms of profit maximization. Based on some fairness criterions, we further design a reward sharing scheme, namely ARETE-SH, which is closely coupled with ARETE-PR, to incentivize data providers to contribute data. We have evaluated ARETE on a real-world sensory data set collected by Intel Berkeley lab. Evaluation results show that ARETE-PR outperforms the state-of-the-art pricing mechanisms, and achieves around 90 percent of the optimal revenue. ARETE-SH distributes the reward among data providers in a fair way. Zhenzhe Zheng 0001, Yanqing Peng, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2019 | Bursty Event Detection Throughout HistoriesabstractThe widespread use of social media and the active trend of moving towards more web-and mobile-based reporting for traditional media outlets have created an avalanche of information streams. These information streams bring in first-hand reporting on live events to massive crowds in real time as they are happening. It is important to study the phenomenon of burst in this context so that end-users can quickly identify important events that are emerging and developing in their early stages. In this paper, we investigate the problem of bursty event detection where we define burst as the acceleration over the incoming rate of an event mentioning. Existing works focus on the detection of current trending events, but it is important to be able to go back in time and explore bursty events throughout the history, while without the needs of storing and traversing the entire information stream from the past. We present a succinct probabilistic data structure and its associated query strategy to find bursty events at any time instance for the entire history. Extensive empirical results on real event streams have demonstrated the effectiveness of our approach. Debjyoti Paul, Yanqing Peng, Feifei Li 0001 |
ICDE | 2 |
| 2018 | Persistent Bloom Filter: Membership Testing for the Entire HistoryabstractMembership testing is the problem of testing whether an element is in a set of elements. Performing the test exactly is expensive space-wise, requiring the storage of all elements in a set. In many applications, an approximate testing that can be done quickly using small space is often desired. Bloom filter (BF) was designed and has witnessed great success across numerous application domains. But there is no compact structure that supports set membership testing for temporal queries, e.g., has person A visited a web server between 9:30am and 9:40am? And has the same person visited the web server again between 9:45am and 9:50am? It is possible to support such "temporal membership testing" using a BF, but we will show that this is fairly expensive. To that end, this paper designs persistent bloom filter (PBF), a novel data structure for temporal membership testing with compact space. Yanqing Peng, Jinwei Guo, Feifei Li 0001, Weining Qian, Aoying Zhou |
SIGMOD Conference | 1 |
| 2017 | An Online Pricing Mechanism for Mobile Crowdsensing Data MarketsabstractAlthough data has become an important kind of commercial goods, there are few appropriate online platforms to facilitate the trading of mobile crowd-sensed data so far. In this paper, we present the first architecture of mobile crowd-sensed data market, and conduct an in-depth study of the design problem of online data pricing. To build a practical mobile crowd-sensed data market, we have to consider three major design challenges: data uncertainty, economic-robustness (arbitrage-freeness in particular), revenue maximization. By jointly considering the design challenges, we propose a novel online query-bAsed cRowd-sensEd daTa pricing mEchanism, namely ARETE, to determine the trading price of crowd-sensed data. Our theoretical analysis shows that ARETE guarantees both arbitrage-freeness and a constant competitive ratio in terms of revenue maximization. We have evaluated ARETE on a real-world sensory data set collected by Intel Berkeley lab. Evaluation results show that ARETE outperforms the state-of-the-art pricing mechanisms, and achieves around 90% of the optimal revenue. Zhenzhe Zheng 0001, Yanqing Peng, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen |
MobiHoc | 2 |
| 2017 | Trading Data in the Crowd: Profit-Driven Data Acquisition for Mobile CrowdsensingabstractAs a significant business paradigm, data trading has attracted increasing attention. However, the study of data acquisition in data markets is still in its infancy. Mobile crowdsensing has been recognized as an efficient and scalable way to acquire large-scale data. Designing a practical data acquisition scheme for crowd-sensed data markets has to consider three major challenges: crowd-sensed data trading format determination, profit maximization with polynomial computational complexity, and payment minimization in strategic environments. In this paper, we jointly consider these design challenges, and propose VENUS, which is the first profit-driVEN data acqUiSition framework for crowd-sensed data markets. Specifically, VENUS consists of two complementary mechanisms: VENUS-PRO for profit maximization and VENUS-PAY for payment minimization. Given the expected payment for each of the data acquisition points, VENUS-PRO greedily selects the most “cost-efficient” data acquisition points to achieve a sub-optimal profit. To determine the minimum payment for each data acquisition point, we further design VENUS-PAY, which is a data procurement auction in Bayesian setting. Our theoretical analysis shows that VENUS-PAY can achieve both strategy-proofness and optimal expected payment. We evaluate VENUS on a public sensory data set, collected by Intel Research, Berkeley Laboratory. Our evaluation results show that VENUS-PRO approaches the optimal profit, and VENUS-PAY outperforms the canonical second-price reverse auction, in terms of total payment. Zhenzhe Zheng 0001, Yanqing Peng, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen |
IEEE J. Sel. Areas Commun. | 2 |
| 2016 | ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable HardwareabstractHighly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size. Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005 |
SIGCOMM | 4 |