Andreas Pfadler

dblp:86/8030 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-8018-8028ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2Computer networks · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Estimation of Doubly-Dispersive Channels in Linearly Precoded Multicarrier Systems Using Smoothness Regularization
abstract
In this paper, we propose a novel channel estimation scheme for pulse-shaped multicarrier systems using smoothness regularization for ultra-reliable low-latency communication (URLLC). It can be applied to any multicarrier system with or without linear precoding to estimate challenging doubly-dispersive channels. A recently proposed modulation scheme using orthogonal precoding is orthogonal time-frequency and space modulation (OTFS). In OTFS, pilot and data symbols are placed in delay-Doppler (DD) domain and are jointly precoded to the time-frequency (TF) domain. On the one hand, such orthogonal precoding increases the achievable channel estimation accuracy and enables high TF diversity at the receiver. On the other hand, it introduces leakage effects which requires extensive leakage suppression when the piloting is jointly precoded with the data. To avoid this, we propose to precode the data symbols only, place pilot symbols without precoding into the TF domain, and estimate the channel coefficients by interpolating smooth functions from the pilot samples. Furthermore, we present a piloting scheme enabling a smooth control of the number and position of the pilot symbols. Our numerical results suggest that the proposed scheme provides accurate channel estimation with reduced signaling overhead compared to standard estimators using Wiener filtering in the discrete DD domain.
Andreas Pfadler, Tom Szollmann, Peter Jung 0001, Slawomir Stanczak
IEEE Trans. Wirel. Commun.1
2023 Lero: A Learning-to-Rank Query Optimizer
abstract
A recent line of works apply machine learning techniques to assist or rebuild cost-based query optimizers in DBMS. While exhibiting superiority in some benchmarks, their deficiencies, e.g., unstable performance, high training cost, and slow model updating, stem from the inherent hardness of predicting the cost or latency of execution plans using machine learning models. In this paper, we introduce a learning-to-rank query optimizer, called Lero, which builds on top of a native query optimizer and continuously learns to improve the optimization performance. The key observation is that the relative order or rank of plans, rather than the exact cost or latency, is sufficient for query optimization. Lero employs a pairwise approach to train a classifier to compare any two plans and tell which one is better. Such a binary classification task is much easier than the regression task to predict the cost or latency, in terms of model efficiency and accuracy. Rather than building a learned optimizer from scratch, Lero is designed to leverage decades of wisdom of databases and improve the native query optimizer. With its non-intrusive design, Lero can be implemented on top of any existing DBMS with minimal integration efforts. We implement Lero and demonstrate its outstanding performance using PostgreSQL. In our experiments, Lero achieves near optimal performance on several benchmarks. It reduces the plan execution time of the native optimizer in PostgreSQL by up to 70% and other learned query optimizers by up to 37%. Meanwhile, Lero continuously learns and automatically adapts to query workloads and changes in data.
Wei Chen 0133, Bolin Ding, Xingguang Chen, Andreas Pfadler, Ziniu Wu, Jingren Zhou 0001
Proc. VLDB Endow.5
2022 Learned Query Optimizer: At the Forefront of AI-Driven Databases
Ziniu Wu, Chengliang Chai, Andreas Pfadler, Bolin Ding, Guoliang Li 0001, Jingren Zhou 0001
EDBT4
2021 Efficient and Scalable Structure Learning for Bayesian Networks: Algorithms and Applications
abstract
Structure Learning for Bayesian network (BN) is an important problem with extensive research. It plays central roles in a wide variety of applications in Alibaba Group. However, existing structure learning algorithms suffer from considerable limitations in real-world applications due to their low efficiency and poor scalability. To resolve this, we propose a new structure learning algorithm LEAST, which comprehensively fulfills our business requirements as it attains high accuracy, efficiency and scalability at the same time. The core idea of LEAST is to formulate the structure learning into a continuous constrained optimization problem, with a novel differentiable constraint function measuring the acyclicity of the resulting graph. Unlike with existing work, our constraint function is built on the spectral radius of the graph and could be evaluated in near linear time w.r.t. the graph node size. Based on it, LEAST can be efficiently implemented with low storage overhead. According to our benchmark evaluation, LEAST runs 1-2 orders of magnitude faster than state-of-the-art method with comparable accuracy, and it is able to scale on BNs with up to hundreds of thousands of variables. In our production environment, LEAST is deployed and serves for more than 20 applications with thousands of executions per day. We describe a concrete scenario in a ticket booking service in Alibaba, where LEAST is applied to build a near real-time automatic anomaly detection and root error cause analysis system. We also show that LEAST unlocks the possibility of applying BN structure learning in new areas, such as large-scale gene expression data analysis and explainable recommendation system.
Andreas Pfadler, Ziniu Wu, Yuxing Han 0002, Xiaoke Yang, Zhenping Qian, Jingren Zhou 0001, Bin Cui 0001
ICDE2
2021 Spectrum Needs of Cooperative, Connected and Automated Mobility
abstract
The introduction of cooperative, connected and automated mobility (CCAM) services has a promising business potential. Connected vehicles, however, represent a new kind of resource consumer to mobile network operators. It needs to be assured that existing and future mobile communication networks are capable of supporting the CCAM services. In this analysis, we present an evaluation of the spectrum demands of connected vehicles by comparing their data rate requirements to the resources available in typical 4G and 5G deployments.
Sebastian Euler, Andreas Pfadler, Luis Fernández Ferreira, Hongxia Zhao
VTC Spring2
2021 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation
abstract
Cardinality estimation (CardEst) plays a significant role in generating high-quality query plans for a query optimizer in DBMS. In the last decade, an increasing number of advanced CardEst methods (especially ML-based) have been proposed with outstanding estimation accuracy and inference latency. However, there exists no study that systematically evaluates the quality of these methods and answer the fundamental problem: to what extent can these methods improve the performance of query optimizer in real-world settings, which is the ultimate goal of a CardEst method. In this paper, we comprehensively and systematically compare the effectiveness of CardEst methods in a real DBMS. We establish a new benchmark for CardEst, which contains a new complex real-world dataset STATS and a diverse query workload STATS-CEB. We integrate multiple most representative CardEst methods into an open-source DBMS PostgreSQL, and comprehensively evaluate their true effectiveness in improving query plan quality, and other important aspects affecting their applicability. We obtain a number of key findings under different data and query settings. Furthermore, we find that the widely used estimation accuracy metric (Q-Error) cannot distinguish the importance of different sub-plan queries during query optimization and thus cannot truly reflect the generated query plan quality. Therefore, we propose a new metric P-Error to evaluate the performance of CardEst methods, which overcomes the limitation of Q-Error and is able to reflect the overall end-to-end performance of CardEst methods. It could serve as a better optimization objective for future CardEst methods.
Yuxing Han 0002, Ziniu Wu, Peizhi Wu, Liang Wei Tan, Kai Zeng 0002, Gao Cong, Yanzhao Qin, Andreas Pfadler, Zhengping Qian, Jingren Zhou 0001, Jiangneng Li, Bin Cui 0001
Proc. VLDB Endow.10
2021 FLAT: Fast, Lightweight and Accurate Method for Cardinality Estimation
abstract
Query optimizers rely on accurate cardinality estimation (CardEst) to produce good execution plans. The core problem of CardEst is how to model the rich joint distribution of attributes in an accurate and compact manner. Despite decades of research, existing methods either over-simplify the models only using independent factorization which leads to inaccurate estimates, or over-complicate them by lossless conditional factorization without any independent assumption which results in slow probability computation. In this paper, we propose FLAT, a CardEst method that is simultaneously fast in probability computation, lightweight in model size and accurate in estimation quality. The key idea of FLAT is a novel unsupervised graphical model, called FSPN. It utilizes both independent and conditional factorization to adaptively model different levels of attributes correlations, and thus combines their advantages. FLAT supports efficient online probability computation in near linear time on the underlying FSPN model, provides effective offline model construction and enables incremental model updates. It can estimate cardinality for both single table queries and multi-table join queries. Extensive experimental study demonstrates the superiority of FLAT over existing CardEst methods: FLAT achieves 1--5 orders of magnitude better accuracy, 1--3 orders of magnitude faster probability computation speed and 1--2 orders of magnitude lower storage cost. We also integrate FLAT into Postgres to perform an end-to-end test. It improves the query execution time by 12.9% on the well-known IMDB benchmark workload, which is very close to the optimal result 14.2% using the true cardinality.
Ziniu Wu, Yuxing Han 0002, Kai Zeng 0002, Andreas Pfadler, Zhengping Qian, Jingren Zhou 0001, Bin Cui 0001
Proc. VLDB Endow.5
2020 Mobility Modes for Pulse-Shaped OTFS with Linear Equalizer
abstract
Orthogonal time frequency and space (OTFS) modulation is a pulse-shaped Gabor signaling scheme with additional time-frequency (TF) spreading using the symplectic finite Fourier transform (SFFT). With a sufficient amount of accurate channel information and sophisticated equalizers, it promises performance gains in terms of robustness for high mobility users. To fully exploit diversity in OTFS, the 2D-deconvolution implemented by a linear equalizer should approximately invert the doubly dispersive channel operation, which however is a twisted convolution. In theory, this is achieved in a first step by matching the TF grid and the Gabor synthesis and analysis pulses to the delay and Doppler spread of the channel. However, in practice, one always has to balance between supporting high granularity in delay-Doppler (DD) spread, and multi-user and network aspects. In this paper, we propose mobility modes with distinct grid and pulse matching for different doubly dispersive channels. To account for remaining self-interference, we tune the minimum mean square error (MMSE) linear equalizer without the need of estimating channel cross-talk coefficients. We evaluate our approach with the QuaDRiGa channel simulator and with OTFS transceiver architecture based on a polyphase implementation for orthogonalized Gaussian pulses. In addition, we compare OTFS to a IEEE 802.11p compliant design of cyclic prefix (CP) based orthogonal frequency-division multiplexing (OFDM). Our results indicate that with an appropriate mobility mode, the potential OTFS gains can be indeed achieved with linear equalizers to significantly outperform OFDM.
Andreas Pfadler, Peter Jung 0001, Slawomir Stanczak
GLOBECOM1
2020 Billion-scale Recommendation with Heterogeneous Side Information at Taobao
abstract
In recent years, embedding models based on skip-gram algorithm have been widely applied to real-world recommendation systems (RSs). When designing embedding-based methods for recommendation at Taobao, there are three main challenges: scalability, sparsity and cold start. The first problem is inherently caused by the extremely large numbers of users and items (in the order of billions), while the remaining two problems are caused by the fact that most items have only very few (or none at all) user interactions. To address these challenges, in this work, we present a flexible and highly scalable Side Information (SI) enhanced Skip-Gram (SISG) framework, which is deployed at Taobao. SISG overcomes the drawbacks of existing embedding-based models by modeling user metadata and capturing asymmetries of user behavior. Furthermore, as training SISG can be performed using any SGNS implementation, we present our production deployment of SISG on a custom-built word2vec engine, which allows us to compute item and SI embedding vectors for billion-scale sets of products in a join semantic space on a daily basis. Finally, using offline and online experiments we demonstrate the significant superiority of SISG over our previously deployed framework, EGES, and a well-tuned CF, as well as present evidence supporting our scalability claims.
Andreas Pfadler, Huan Zhao 0002, Jizhe Wang, Pipei Huang, Dik Lun Lee
ICDE1
2020 Learning Efficient Parameter Server Synchronization Policies for Distributed SGD
Andreas Pfadler, Zhengping Qian, Jingren Zhou 0001
ICLR3
2020 Predictive Quality of Service: Adaptation of Platoon Inter-Vehicle Distance to Packet Inter-Reception Time
abstract
Vehicle-to-everything (V2X) communication is seen as an enabler of high-density platooning as part of more environmentally friendly future transportation systems. Indeed, in high-density platooning, trucks are able to reduce their overall fuel consumption. Compared to platooning systems exclusively based on sensors, V2X enabled platooning systems can drive smaller inter-vehicle distances. They are then able to achieve this fuel consumption reduction thanks to the decreased air drag. It has been shown that the performance of the application is dependent on the performance of the communications system. The application therefore needs to be aware of the maximal tolerable communication degradation that keeps the platoon safe considering its driving parameters. In this article, we derive the relationship between the maximal tolerable packet losses, measured as the packet inter-reception time, and the intervehicle distance. We first study the relationship between these parameters through the analysis of simulation data. We then derive a functional link by fitting different statistical models. Finally, we apply the resulting models to packet inter-reception time measurements obtained in simulation of platoons supported by IEEE 802. 11p driving through varying surrounding traffic densities.
Andreas Pfadler, Guillaume Jomod, Ahmad El Assaad, Peter Jung 0001
VTC Spring1
2019 POG: Personalized Outfit Generation for Fashion Recommendation at Alibaba iFashion
abstract
Increasing demand for fashion recommendation raises a lot of challenges for online shopping platforms and fashion communities. In particular, there exist two requirements for fashion outfit recommendation: the Compatibility of the generated fashion outfits, and the Personalization in the recommendation process. In this paper, we demonstrate these two requirements can be satisfied via building a bridge between outfit generation and recommendation. Through large data analysis, we observe that people have similar tastes in individual items and outfits. Therefore, we propose a Personalized Outfit Generation (POG) model, which connects user preferences regarding individual items and outfits with Transformer architecture. Extensive offline and online experiments provide strong quantitative evidence that our method outperforms alternative methods regarding both compatibility and personalization metrics. Furthermore, we deploy POG on a platform named Dida in Alibaba to generate personalized outfits for the users of the online application iFashion.
Wen Chen 0026, Pipei Huang, Fei Sun 0001, Andreas Pfadler, Binqiang Zhao
KDD8