Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Thanh T. L. Tran

dblp:t/ThanhTran2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 5 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Data stream processing · 43% Query processing and optimization · 23% Data models and query languages · 16%
Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 100%
Computer networks
1 paper
Internet of things and sensor networks · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
uncertain data stream
0.322012
CLARO: modeling and processing uncertain data streams · VLDB J. 2012
Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212013
Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013
Data models and query languages
uncertain data
0.212013
Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013
Query processing and optimization › query execution
user-defined function execution
0.212013
Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013
Data stream processing
continuous query processing
0.112012
CLARO: modeling and processing uncertain data streams · VLDB J. 2012
Machine learning › Probabilistic and Bayesian machine learning
distribution approximation
0.112010
Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010
Query processing and optimization
aggregate query processing
0.112010
Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010
Data integration and cleaning › data preprocessing
data cleaning
0.112009
Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009
Data integration and cleaning › data preprocessing › data cleaning
probabilistic data cleaning
0.112009
Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009
Internet of things and sensor networks › RFID systems › RFID data management
RFID data processing
0.112009
Probabilistic Inference over RFID Streams in Mobile Environments · ICDE 2009
Data models and query languages
uncertain data management
0.012013
Supporting User-Defined Functions on Uncertain Data · Proc. VLDB Endow. 2013
Data mining
anomaly detection
0.012010
PODS: a new model and processing algorithms for uncertain data streams · SIGMOD Conference 2010
Query processing and optimization
approximate query processing
0.012010
Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

online algorithm · 0.3gaussian process · 0.3randomized approximation · 0.2deterministic approximation · 0.2spatial indexing · 0.2particle filtering · 0.2belief compression · 0.2probabilistic modeling · 0.1statistical approximation · 0.1sampling · 0.1
YearPublicationVenuePosition
2018 PriPeARL: A Framework for Privacy-Preserving Analytics and Reporting at LinkedIn
abstract
Preserving privacy of users is a key requirement of web-scale analytics and reporting applications, and has witnessed a renewed focus in light of recent data breaches and new regulations such as GDPR. We focus on the problem of computing robust, reliable analytics in a privacy-preserving manner, while satisfying product requirements. We present PriPeARL, a framework for privacy-preserving analytics and reporting, inspired by differential privacy. We describe the overall design and architecture, and the key modeling components, focusing on the unique challenges associated with privacy, coverage, utility, and consistency. We perform an experimental study in the context of ads analytics and reporting at LinkedIn, thereby demonstrating the tradeoffs between privacy and utility needs, and the applicability of privacy-preserving mechanisms to real-world data. We also highlight the lessons learned from the production deployment of our system at LinkedIn.
Krishnaram Kenthapadi, Thanh T. L. Tran
CIKM2
2013 Supporting User-Defined Functions on Uncertain Data
abstract
Uncertain data management has become crucial in many sensing and scientific applications. As user-defined functions (UDFs) become widely used in these applications, an important task is to capture result uncertainty for queries that evaluate UDFs on uncertain data. In this work, we provide a general framework for supporting UDFs on uncertain data. Specifically, we propose a learning approach based on Gaussian processes (GPs) to compute approximate output distributions of a UDF when evaluated on uncertain input, with guaranteed error bounds. We also devise an online algorithm to compute such output distributions, which employs a suite of optimizations to improve accuracy and performance. Our evaluation using both real-world and synthetic functions shows that our proposed GP approach can outperform the state-of-the-art sampling approach with up to two orders of magnitude improvement for a variety of UDFs.
Thanh T. L. Tran, Yanlei Diao, Charles Sutton, Anna Liu
Proc. VLDB Endow.1
2012 Differentially private summaries for sparse data
abstract
Differential privacy is fast becoming the method of choice for releasing data under strong privacy guarantees. A standard mechanism is to add noise to the counts in contingency tables derived from the dataset. However, when the dataset is sparse in its underlying domain, this vastly increases the size of the published data, to the point of making the mechanism infeasible.
Graham Cormode, Cecilia M. Procopiuc, Divesh Srivastava, Thanh T. L. Tran
ICDT4
2012 CLARO: modeling and processing uncertain data streams
Thanh T. L. Tran, Liping Peng, Yanlei Diao, Andrew McGregor 0001, Anna Liu
VLDB J.1
2010 PODS: a new model and processing algorithms for uncertain data streams
abstract
Uncertain data streams, where data is incomplete, imprecise, and even misleading, have been observed in many environments. Feeding such data streams to existing stream systems produces results of unknown quality, which is of paramount concern to monitoring applications. In this paper, we present the PODS system that supports stream processing for uncertain data naturally captured using continuous random variables. PODS employs a unique data model that is flexible and allows efficient computation. Built on this model, we develop evaluation techniques for complex relational operators, i.e., aggregates and joins, by exploring advanced statistical theory and approximation. Evaluation results show that our techniques can achieve high performance while satisfying accuracy requirements, and significantly outperform a state-of-the-art sampling method. A case study further shows that our techniques can enable a tornado detection system (for the first time) to produce detection results at stream speed and with much improved quality.
Thanh T. L. Tran, Liping Peng, Boduo Li, Yanlei Diao, Anna Liu
SIGMOD Conference1
2010 Conditioning and Aggregating Uncertain Data Streams: Going Beyond Expectations
abstract
Uncertain data streams are increasingly common in real-world deployments and monitoring applications require the evaluation of complex queries on such streams. In this paper, we consider complex queries involving conditioning (e.g., selections and group by's) and aggregation operations on uncertain data streams. To characterize the uncertainty of answers to these queries, one generally has to compute the full probability distribution of each operation used in the query. Computing distributions of aggregates given conditioned tuple distributions is a hard, unsolved problem. Our work employs a new evaluation framework that includes a general data model, approximation metrics, and approximate representations. Within this framework we design fast data-stream algorithms, both deterministic and randomized, for returning approximate distributions with bounded errors as answers to those complex queries. Our experimental results demonstrate the accuracy and efficiency of our approximation techniques and offer insights into the strengths and limitations of deterministic and randomized algorithms.
Thanh T. L. Tran, Andrew McGregor 0001, Yanlei Diao, Liping Peng, Anna Liu
Proc. VLDB Endow.1
2009 Capturing Data Uncertainty in High-Volume Stream Processing
Yanlei Diao, Boduo Li, Anna Liu, Liping Peng, Charles Sutton, Thanh T. L. Tran, Michael Zink
CIDR6
2009 Probabilistic Inference over RFID Streams in Mobile Environments
abstract
Recent innovations in RFID technology are enabling large-scale cost-effective deployments in retail, healthcare, pharmaceuticals and supply chain management. The advent of mobile or handheld readers adds significant new challenges to RFID stream processing due to the inherent reader mobility, increased noise, and incomplete data. In this paper, we address the problem of translating noisy, incomplete raw streams from mobile RFID readers into clean, precise event streams with location information. Specifically we propose a probabilistic model to capture the mobility of the reader, object dynamics, and noisy readings. Our model can self-calibrate by automatically estimating key parameters from observed data. Based on this model, we employ a sampling-based technique called particle filtering to infer clean, precise information about object locations from raw streams from mobile RFID readers. Since inference based on standard particle filtering is neither scalable nor efficient in our settings, we propose three enhancements-particle factorization, spatial indexing, and belief compression-for scalable inference over large numbers of objects and high-volume streams. Our experiments show that our approach can offer 49% error reduction over a state-of-the-art data cleaning approach such as SMURF while also being scalable and efficient.
Thanh T. L. Tran, Charles Sutton, Richard Cocci, Yanming Nie, Yanlei Diao, Prashant J. Shenoy
ICDE1
2008 Efficient Data Interpretation and Compression over RFID Streams
abstract
Despite its promise, RFID technology presents numerous challenges, including incomplete data, lack of location and containment information, and very high volumes. In this work, we present a novel data interpretation and compression substrate over RFID streams to address these challenges in enterprise supply-chain environments. Our results show that our inference techniques provide good accuracy while retaining efficiency, and our compression algorithm yields significant reduction in data volume.
Richard Cocci, Thanh T. L. Tran, Yanlei Diao, Prashant J. Shenoy
ICDE2