Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Thien Q. Tran

dblp:245/5958 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0001-8377-501XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 50% Language models and text generation · 14% Reinforcement learning · 14%
Databases, data mining, and information retrieval
3 papers
Data mining · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › pattern mining › utility mining
high utility pattern mining
1.122023
Statistically Significant Pattern Mining With Ordinal Utility · IEEE Trans. Knowl. Data Eng. 2023
Statistically Significant Pattern Mining with Ordinal Utility · KDD 2020
Data mining
pattern mining
1.122023
Statistically Significant Pattern Mining With Ordinal Utility · IEEE Trans. Knowl. Data Eng. 2023
Statistically Significant Pattern Mining with Ordinal Utility · KDD 2020
Data mining › pattern mining › interesting pattern mining
significant pattern mining
1.122023
Statistically Significant Pattern Mining With Ordinal Utility · IEEE Trans. Knowl. Data Eng. 2023
Statistically Significant Pattern Mining with Ordinal Utility · KDD 2020
Natural language and speech › Language models and text generation
alignment
0.812024
Stepwise Alignment for Constrained Language Model Policy Optimization · NeurIPS 2024
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.812024
Stepwise Alignment for Constrained Language Model Policy Optimization · NeurIPS 2024
Machine learning › Trustworthy machine learning › AI safety
harmlessness
0.812024
Stepwise Alignment for Constrained Language Model Policy Optimization · NeurIPS 2024
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.812024
Stepwise Alignment for Constrained Language Model Policy Optimization · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation
0.612022
Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation · AAAI 2022
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.612022
Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation · AAAI 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation · AAAI 2022
Machine learning › Generative modeling
variational autoencoder
0.612022
Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation · AAAI 2022
Medical and health informatics › public health › public health informatics
epidemic forecasting
0.412019
Seasonal-adjustment Based Feature Selection Method for Predicting Epidemic with Large-scale Search Engine Logs · KDD 2019
Data mining › dimensionality reduction
feature selection
0.412019
Seasonal-adjustment Based Feature Selection Method for Predicting Epidemic with Large-scale Search Engine Logs · KDD 2019
Data mining › time series analysis
time series decomposition
0.112019
Seasonal-adjustment Based Feature Selection Method for Predicting Epidemic with Large-scale Search Engine Logs · KDD 2019

Methods — techniques the papers use, named apart from their topics

familywise error rate control · 1.1time series decomposition · 0.8seasonal adjustment · 0.8direct preference optimization · 0.8constrained optimization · 0.8multiple testing · 0.7variational autoencoder · 0.6causal discovery · 0.6tarone-bonferroni · 0.4multiple hypothesis testing · 0.4
YearPublicationVenuePosition
2024 Stepwise Alignment for Constrained Language Model Policy Optimization
abstract
Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment as an optimization problem of the language model policy to maximize reward under a safety constraint, and then proposes an algorithm, Stepwise Alignment for Constrained Policy Optimization (SACPO). One key idea behind SACPO, supported by theory, is that the optimal policy incorporating reward and safety can be directly obtained from a reward-aligned policy. Building on this key idea, SACPO aligns LLMs step-wise with each metric while leveraging simple yet powerful alignment algorithms such as direct preference optimization (DPO). SACPO offers several advantages, including simplicity, stability, computational efficiency, and flexibility of algorithms and datasets. Under mild assumptions, our theoretical analysis provides the upper bounds on optimality and safety constraint violation. Our experimental results show that SACPO can fine-tune Alpaca-7B better than the state-of-the-art method in terms of both helpfulness and harmlessness.
Akifumi Wachi, Thien Q. Tran, Rei Sato, Takumi Tanabe, Youhei Akimoto
NeurIPS2
2023 Statistically Significant Pattern Mining With Ordinal Utility
abstract
Statistically significant pattern mining (SSPM), which evaluates each pattern via a hypothesis test, is an essential and challenging data mining task for knowledge discovery. We introduce a preference relation between patterns and aim to discover the most preferred patterns under the constraint of statistical significance, which has never been considered in existing SSPM problems. We propose an iterative multiple testing procedure that can alternately reject a hypothesis and safely ignore the less useful hypotheses than the rejected one. By filtering out patterns with low utility, we can avoid the significance budget consumption of rejecting useless (uninteresting) patterns and focus the significance budget on more useful patterns, leading to more useful discoveries. We show that the proposed method can control the familywise error rate (FWER) under certain assumptions, which can be satisfied by a realistic problem class in SSPM. We also show that the proposed method always discovers equally or more useful patterns than Tarone-Bonferroni and Subfamily-wise Multiple Testing (SMT). Finally, we conducted several experiments with both synthetic and real-world data to evaluate the performance of our method. The proposed method discovered many more useful patterns in the experiments with real-world datasets than the existing method for all five conducted tasks.
Thien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun Sakuma
IEEE Trans. Knowl. Data Eng.1
2022 Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation
abstract
We aim to explain a black-box classifier with the form: "data X is classified as class Y because X has A, B and does not have C" in which A, B, and C are high-level concepts. The challenge is that we have to discover in an unsupervised manner a set of concepts, i.e., A, B and C, that is useful for explaining the classifier. We first introduce a structural generative model that is suitable to express and discover such concepts. We then propose a learning process that simultaneously learns the data distribution and encourages certain concepts to have a large causal influence on the classifier output. Our method also allows easy integration of user's prior knowledge to induce high interpretability of concepts. Finally, using multiple datasets, we demonstrate that the proposed method can discover useful concepts for explanation in this form.
Thien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun Sakuma
AAAI1
2020 Statistically Significant Pattern Mining with Ordinal Utility
abstract
Statistically significant patterns mining (SSPM) is an essential and challenging data mining task in the field of knowledge discovery in databases (KDD), in which each pattern is evaluated via a hypothesis test. Our study aims to introduce a preference relation into patterns and to discover the most preferred patterns under the constraint of statistical significance, which has never been considered in existing SSPM problems. We propose an iterative multiple testing procedure that can alternately reject a hypothesis and safely ignore the hypotheses that are less useful than the rejected hypothesis. One advantage of filtering out patterns with low utility is that it avoids consumption of the significance budget by rejection of useless (that is, uninteresting) patterns. This allows the significance budget to be focused on useful patterns, leading to more useful discoveries. We show that the proposed method can control the familywise error rate (FWER) under certain assumptions, that can be satisfied by a realistic problem class in SSPM. We also show that the proposed method always discovers a set of patterns that is at least equally or more useful than those discovered using the standard Tarone-Bonferroni method SSPM. Finally, we conducted several experiments with both synthetic and real-world data to evaluate the performance of our method. As a result, in the experiments with real-world datasets, the proposed method discovered a larger number of more useful patterns than the existing method for all five conducted tasks.
Thien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun Sakuma
KDD1
2019 Seasonal-adjustment Based Feature Selection Method for Predicting Epidemic with Large-scale Search Engine Logs
abstract
Search engine logs have a great potential in tracking and predicting outbreaks of infectious disease. More precisely, one can use the search volume of some search terms to predict the infection rate of an infectious disease in nearly real-time. However, conducting accurate and stable prediction of outbreaks using search engine logs is a challenging task due to the following two-way instability characteristics of the search logs. First, the search volume of a search term may change irregularly in the short-term, for example, due to environmental factors such as the amount of media or news. Second, the search volume may also change in the long-term due to the demographic change of the search engine. That is to say, if a model is trained with such search logs with ignoring such characteristic, the resulting prediction would contain serious mispredictions when these changes occur. In this work, we proposed a novel feature selection method to overcome this instability problem. In particular, we employ a seasonal-adjustment method that decomposes each time series into three components: seasonal, trend and irregular component and build prediction models for each component individually. We also carefully design a feature selection method to select proper search terms to predict each component. We conducted comprehensive experiments on ten different kinds of infectious diseases. The experimental results show that the proposed method outperforms all comparative methods in prediction accuracy for seven of ten diseases, in both now-casting and forecasting setting. Also, the proposed method is more successful in selecting search terms that are semantically related to target diseases.
Thien Q. Tran, Jun Sakuma
KDD1