Trong Duc Nguyen

dblp:179/8635 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
1since 2021 · last 2021
0009-0003-7957-1331ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 92% Information retrieval · 8%
Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 76% Program synthesis and code generation · 24%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.412020
Random Sampling for Group-By Queries · ICDE 2020
Query processing and optimization › approximate query processing
stratified sampling
0.412020
Random Sampling for Group-By Queries · ICDE 2020
Software maintenance and evolution › code search
API knowledge retrieval
0.312018
Complementing global and local contexts in representing API descriptions to improve API retrieval tasks · ESEC/SIGSOFT FSE 2018
Software maintenance and evolution
API mapping
0.312017
Exploring API embedding for API usages and applications · ICSE 2017
Program synthesis and code generation
code completion
0.312017
Exploring API embedding for API usages and applications · ICSE 2017
Software maintenance and evolution › software reengineering › software modernization › software migration
code migration
0.312017
Exploring API embedding for API usages and applications · ICSE 2017
Query processing and optimization
aggregate query processing
0.112020
Random Sampling for Group-By Queries · ICDE 2020
Query processing and optimization › aggregate query processing
group-by query
0.112020
Random Sampling for Group-By Queries · ICDE 2020
Information retrieval › text analysis
text representation
0.112018
Complementing global and local contexts in representing API descriptions to improve API retrieval tasks · ESEC/SIGSOFT FSE 2018

Methods — techniques the papers use, named apart from their topics

word2vec · 0.9vector representation · 0.7coefficient of variation optimization · 0.4neural network embedding · 0.3
YearPublicationVenuePosition
2021 Stratified random sampling from streaming and stored data
Trong Duc Nguyen, Ming-Hung Shih, Divesh Srivastava, Srikanta Tirthapura, Bojian Xu
Distributed Parallel Databases1
2020 Random Sampling for Group-By Queries
abstract
Random sampling has been widely used in approximate query processing on large databases, due to its potential to significantly reduce resource usage and response times, at the cost of a small approximation error. We consider random sampling for answering the ubiquitous class of group-by queries, which first group data according to one or more attributes, and then aggregate within each group after filtering through a predicate. The challenge with group-by queries is that a sampling method cannot focus on optimizing the quality of a single answer (e.g. the mean of selected data), but must simultaneously optimize the quality of a set of answers (one per group). We present CVOPT, a query- and data-driven sampling framework for a set of group-by queries. To evaluate the quality of a sample, CVOPT defines a metric based on the norm (e.g. ℓ2or ℓ∞) of the coefficients of variation (CVs) of different answers, and constructs a stratified sample that provably optimizes the metric. CVOPT can handle group-by queries on data where groups have vastly different statistical characteristics, such as frequencies, means, or variances. CVOPT jointly optimizes for multiple aggregations and multiple group-by clauses, and provides a way to prioritize specific groups or aggregates. It can be tuned to cases when partial information about a query workload is known, such as a data warehouse where queries are run periodically. Our experimental results show that CVOPT outperforms the current state-of-the-art on sample quality and estimation accuracy for group-by queries. On a set of queries on two real-world data sets, CVOPT yields relative errors that are 5× smaller than competing approaches, under the same space budget.
Trong Duc Nguyen, Ming-Hung Shih, Sai Sree Parvathaneni, Bojian Xu, Divesh Srivastava, Srikanta Tirthapura
ICDE1
2019 Stratified Random Sampling over Streaming and Stored Data
abstract
Stratified random sampling (SRS) is a widely used sampling technique for approximate query processing. We consider SRS on continuously arriving data streams, and make the following contributions. We present a lower bound that shows that any streaming algorithm for SRS must have (in the worst case) a variance that is Ω(r ) factor away from the optimal, where r is the number of strata. We present S-VOILA, a streaming algorithm for SRS that is locally variance-optimal. Results from experiments on real and synthetic data show that S-VOILA results in a variance that is typically close to an optimal offline algorithm, which was given the entire input beforehand. We also present a variance-optimal offline algorithm VOILA for stratified random sampling. VOILA is a strict generalization of the well-known Neyman allocation, which is optimal only under the assumption that each stratum is abundant, i.e. has a large number of data points to choose from. Experiments show that VOILA can have significantly smaller variance (1.4x to 50x) than Neyman allocation on real-world data.
Trong Duc Nguyen, Ming-Hung Shih, Divesh Srivastava, Srikanta Tirthapura, Bojian Xu
EDBT1
2018 Complementing global and local contexts in representing API descriptions to improve API retrieval tasks
abstract
When being trained on API documentation and tutorials, Word2vec produces vector representations to estimate the relevance between texts and API elements. However, existing Word2vec-based approaches to measure document similarities aggregate Word2vec vectors of individual words or APIs to build the representation of a document as if the words are independent. Thus, the semantics of API descriptions or code fragments are not well represented.
Thanh Van Nguyen, Ngoc M. Tran, Hung Phan, Trong Duc Nguyen, Linh H. Truong, Anh Tuan Nguyen 0001, Hoan Anh Nguyen, Tien N. Nguyen
ESEC/SIGSOFT FSE4
2018 A deep neural network language model with contexts for source code
abstract
Statistical language models (LMs) have been applied in several software engineering applications. However, they have issues in dealing with ambiguities in the names of program and API elements (classes and method calls). In this paper, inspired by the success of Deep Neural Network (DNN) in natural language processing, we present Dnn4C, a DNN language model that complements the local context of lexical code elements with both syntactic and type contexts. We designed a context-incorporating method to use with syntactic and type annotations for source code in order to learn to distinguish the lexical tokens in different syntactic and type contexts. Our empirical evaluation on code completion for real-world projects shows that Dnn4C relatively improves 11.6%, 16.3%, 27.1%, and 44.7% top-1 accuracy over the state-of-the-art language models for source code used with the same features: RNN LM, DNN LM, SLAMC, and n-gram LM, respectively. For another application, we showed that Dnn4C helps improve accuracy over n-gram LM in migrating source code from Java to C# with a machine translation model.
Anh Tuan Nguyen 0001, Trong Duc Nguyen, Hung Dang Phan, Tien N. Nguyen
SANER2
2017 Exploring API embedding for API usages and applications
abstract
Word2Vec is a class of neural network models that as being trainedfrom a large corpus of texts, they can produce for each unique word acorresponding vector in a continuous space in which linguisticcontexts of words can be observed. In this work, we study thecharacteristics of Word2Vec vectors, called API2VEC or API embeddings, for the API elements within the API sequences in source code. Ourempirical study shows that the close proximity of the API2VEC vectorsfor API elements reflects the similar usage contexts containing thesurrounding APIs of those API elements. Moreover, API2VEC can captureseveral similar semantic relations between API elements in API usagesvia vector offsets. We demonstrate the usefulness of API2VEC vectorsfor API elements in three applications. First, we build a tool thatmines the pairs of API elements that share the same usage relationsamong them. The other applications are in the code migrationdomain. We develop API2API, a tool to automatically learn the APImappings between Java and C# using a characteristic of the API2VECvectors for API elements in the two languages: semantic relationsamong API elements in their usages are observed in the two vectorspaces for the two languages as similar geometric arrangements amongtheir API2VEC vectors. Our empirical evaluation shows that API2APIrelatively improves 22.6% and 40.1% top-1 and top-5 accuracy over astate-of-the-art mining approach for API mappings. Finally, as anotherapplication in code migration, we are able to migrate equivalent APIusages from Java to C# with up to 90.6% recall and 87.2% precision.
Trong Duc Nguyen, Anh Tuan Nguyen 0001, Hung Dang Phan, Tien N. Nguyen
ICSE1