EDBT 2026 Demo / reviewers in the wild / expert
Thanh T. L. Tran
dblp:t/ThanhTran2
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 5 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Data stream processing · 43% Query processing and optimization · 23% Data models and query languages · 16% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 100% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
uncertain data stream |
0.3 | 2 | 2012 | CLARO: modeling and processing uncertain data streams · VLDB J. 2012 Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Data models and query languages
uncertain data |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Query processing and optimization › query execution
user-defined function execution |
0.2 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Data stream processing
continuous query processing |
0.1 | 1 | 2012 | CLARO: modeling and processing uncertain data streams · VLDB J. 2012 |
Machine learning › Probabilistic and Bayesian machine learning
distribution approximation |
0.1 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Query processing and optimization
aggregate query processing |
0.1 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Data integration and cleaning › data preprocessing
data cleaning |
0.1 | 1 | 2009 | Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009 |
Data integration and cleaning › data preprocessing › data cleaning
probabilistic data cleaning |
0.1 | 1 | 2009 | Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009 |
Internet of things and sensor networks › RFID systems › RFID data management
RFID data processing |
0.1 | 1 | 2009 | Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009 |
Data models and query languages
uncertain data management |
0.0 | 1 | 2013 | Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013 |
Data mining
anomaly detection |
0.0 | 1 | 2010 | PODS: a new model and processing algorithms for uncertain data streams · SIGMOD Conference 2010 |
Query processing and optimization
approximate query processing |
0.0 | 1 | 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010 |
Methods — techniques the papers use, named apart from their topics
online algorithm · 0.3gaussian process · 0.3randomized approximation · 0.2deterministic approximation · 0.2spatial indexing · 0.2particle filtering · 0.2belief compression · 0.2probabilistic modeling · 0.1statistical approximation · 0.1sampling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | PriPeARL: A Framework for Privacy-Preserving Analytics and Reporting at LinkedInabstractPreserving privacy of users is a key requirement of web-scale analytics and reporting applications, and has witnessed a renewed focus in light of recent data breaches and new regulations such as GDPR. We focus on the problem of computing robust, reliable analytics in a privacy-preserving manner, while satisfying product requirements. We present PriPeARL, a framework for privacy-preserving analytics and reporting, inspired by differential privacy. We describe the overall design and architecture, and the key modeling components, focusing on the unique challenges associated with privacy, coverage, utility, and consistency. We perform an experimental study in the context of ads analytics and reporting at LinkedIn, thereby demonstrating the tradeoffs between privacy and utility needs, and the applicability of privacy-preserving mechanisms to real-world data. We also highlight the lessons learned from the production deployment of our system at LinkedIn. Krishnaram Kenthapadi, Thanh T. L. Tran |
CIKM | 2 |
| 2013 | Supporting User-Defined Functions on Uncertain DataabstractUncertain data management has become crucial in many sensing and scientific applications. As user-defined functions (UDFs) become widely used in these applications, an important task is to capture result uncertainty for queries that evaluate UDFs on uncertain data. In this work, we provide a general framework for supporting UDFs on uncertain data. Specifically, we propose a learning approach based on Gaussian processes (GPs) to compute approximate output distributions of a UDF when evaluated on uncertain input, with guaranteed error bounds. We also devise an online algorithm to compute such output distributions, which employs a suite of optimizations to improve accuracy and performance. Our evaluation using both real-world and synthetic functions shows that our proposed GP approach can outperform the state-of-the-art sampling approach with up to two orders of magnitude improvement for a variety of UDFs. Thanh T. L. Tran, Yanlei Diao, Charles Sutton, Anna Liu |
Proc. VLDB Endow. | 1 |
| 2012 | Differentially private summaries for sparse dataabstractDifferential privacy is fast becoming the method of choice for releasing data under strong privacy guarantees. A standard mechanism is to add noise to the counts in contingency tables derived from the dataset. However, when the dataset is sparse in its underlying domain, this vastly increases the size of the published data, to the point of making the mechanism infeasible. Graham Cormode, Cecilia M. Procopiuc, Divesh Srivastava, Thanh T. L. Tran |
ICDT | 4 |
| 2012 | CLARO: modeling and processing uncertain data streams
Thanh T. L. Tran, Liping Peng, Yanlei Diao, Andrew McGregor 0001, Anna Liu |
VLDB J. | 1 |
| 2010 | PODS: a new model and processing algorithms for uncertain data streamsabstractUncertain data streams, where data is incomplete, imprecise, and even misleading, have been observed in many environments. Feeding such data streams to existing stream systems produces results of unknown quality, which is of paramount concern to monitoring applications. In this paper, we present the PODS system that supports stream processing for uncertain data naturally captured using continuous random variables. PODS employs a unique data model that is flexible and allows efficient computation. Built on this model, we develop evaluation techniques for complex relational operators, i.e., aggregates and joins, by exploring advanced statistical theory and approximation. Evaluation results show that our techniques can achieve high performance while satisfying accuracy requirements, and significantly outperform a state-of-the-art sampling method. A case study further shows that our techniques can enable a tornado detection system (for the first time) to produce detection results at stream speed and with much improved quality. Thanh T. L. Tran, Liping Peng, Boduo Li, Yanlei Diao, Anna Liu |
SIGMOD Conference | 1 |
| 2010 | Conditioning and Aggregating Uncertain Data Streams: Going Beyond ExpectationsabstractUncertain data streams are increasingly common in real-world deployments and monitoring applications require the evaluation of complex queries on such streams. In this paper, we consider complex queries involving conditioning (e.g., selections and group by's) and aggregation operations on uncertain data streams. To characterize the uncertainty of answers to these queries, one generally has to compute the full probability distribution of each operation used in the query. Computing distributions of aggregates given conditioned tuple distributions is a hard, unsolved problem. Our work employs a new evaluation framework that includes a general data model, approximation metrics, and approximate representations. Within this framework we design fast data-stream algorithms, both deterministic and randomized, for returning approximate distributions with bounded errors as answers to those complex queries. Our experimental results demonstrate the accuracy and efficiency of our approximation techniques and offer insights into the strengths and limitations of deterministic and randomized algorithms. Thanh T. L. Tran, Andrew McGregor 0001, Yanlei Diao, Liping Peng, Anna Liu |
Proc. VLDB Endow. | 1 |
| 2009 | Capturing Data Uncertainty in High-Volume Stream Processing
Yanlei Diao, Boduo Li, Anna Liu, Liping Peng, Charles Sutton, Thanh T. L. Tran, Michael Zink |
CIDR | 6 |
| 2009 | Probabilistic Inference over RFID Streams in Mobile EnvironmentsabstractRecent innovations in RFID technology are enabling large-scale cost-effective deployments in retail, healthcare, pharmaceuticals and supply chain management. The advent of mobile or handheld readers adds significant new challenges to RFID stream processing due to the inherent reader mobility, increased noise, and incomplete data. In this paper, we address the problem of translating noisy, incomplete raw streams from mobile RFID readers into clean, precise event streams with location information. Specifically we propose a probabilistic model to capture the mobility of the reader, object dynamics, and noisy readings. Our model can self-calibrate by automatically estimating key parameters from observed data. Based on this model, we employ a sampling-based technique called particle filtering to infer clean, precise information about object locations from raw streams from mobile RFID readers. Since inference based on standard particle filtering is neither scalable nor efficient in our settings, we propose three enhancements-particle factorization, spatial indexing, and belief compression-for scalable inference over large numbers of objects and high-volume streams. Our experiments show that our approach can offer 49% error reduction over a state-of-the-art data cleaning approach such as SMURF while also being scalable and efficient. Thanh T. L. Tran, Charles Sutton, Richard Cocci, Yanming Nie, Yanlei Diao, Prashant J. Shenoy |
ICDE | 1 |
| 2008 | Efficient Data Interpretation and Compression over RFID StreamsabstractDespite its promise, RFID technology presents numerous challenges, including incomplete data, lack of location and containment information, and very high volumes. In this work, we present a novel data interpretation and compression substrate over RFID streams to address these challenges in enterprise supply-chain environments. Our results show that our inference techniques provide good accuracy while retaining efficiency, and our compression algorithm yields significant reduction in data volume. Richard Cocci, Thanh T. L. Tran, Yanlei Diao, Prashant J. Shenoy |
ICDE | 2 |