Truyen Tran 0001

dblp:55/2269 · also Tran The Truyen · DBLP profile ↗
← Back
117ranked-venue papers
18as first author
49since 2021 · last 2026
0000-0001-6531-8907ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 27 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 10 · 3 since 2021
YearPublicationVenuePosition
2026 Rethinking Deep Alignment Through the Lens of Incomplete Safety Learning
abstract
Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment through supervised fine-tuning and reinforcement learning from human feedback. These vulnerabilities manifest as differential safety behavior across token positions, with safety modifications concentrating in early positions while later positions show minimal distributional changes from base models. We provide a mechanistic analysis of safety alignment training dynamics, revealing that gradient concentration during autoregressive training creates signal decay across token positions. This leads to incomplete distributional learning where safety training fails to fully transform model preferences in later response regions. We introduce base-favored tokens as computational indicators of incomplete safety learning. Analysis reveals that while early positions undergo substantial distributional changes, later positions retain concerning base model preferences in safety-critical contexts, indicating systematic incomplete learning due to insufficient training signals. We develop a targeted completion method that addresses these undertrained regions through adaptive penalties and hybrid teacher distillation. Experimental evaluation across Llama and Qwen model families demonstrates remarkable improvements in adversarial robustness, with dramatic reductions in attack success rates across multiple attack types while fully preserving general capabilities.
Thong Bach, Dung Nguyen 0001, Thao Minh Le, Truyen Tran 0001
AAAI4
2026 Robust SDE Parameter Estimation Under Missing Time Information Setting
abstract
Recent advances in stochastic differential equations (SDEs) have enabled robust modeling of real-world dynamical processes across diverse domains, such as finance, health, and systems biology. However, parameter estimation for SDEs typically relies on accurately time-stamped observational data. When temporal ordering information is corrupted, missing, or deliberately hidden (e.g., for privacy), existing estimation methods often fail. In this paper, we investigate the conditions under which temporal order can be recovered and introduce a novel framework that simultaneously reconstructs temporal information and estimates SDE parameters. Our approach exploits asymmetries between forward and backward processes, deriving a score-matching criterion to infer the correct temporal order between pairs of observations. We then recover the total order via a sorting procedure and estimate SDE parameters from the reconstructed sequence using maximum likelihood. Finally, we conduct extensive experiments on synthetic and real-world datasets to demonstrate the effectiveness of our method, extending parameter estimation to settings with missing temporal order and broadening applicability in sensitive domains.
Long Van Tran, Truyen Tran 0001, Phuoc Nguyen
AAAI2
2026 Confident and Trustworthy Model for Fidgety Movement Classification
abstract
General movements (GMs) are part of the spontaneous movement repertoire and are present from early fetal life onwards up to age five months. GMs are connected to infants' neurological development and can be qualitatively assessed via the General Movement Assessment (GMA). In particular, between the age of three to five months, typically developing infants produce Fidgety Movements (FM) and their absence provides strong evidence for the presence of cerebral palsy (CP). To improve accessibility to the GMA, automated GMA solutions have been a key research area with proposed models becoming increasingly more accurate and interpretable. However, current models cannot gauge their ability to make decisions, which may lead to overconfident mistakes. To address this issue, we propose a Deep learning-based approach that not only classifies movements as fidgety or non-fidgety but also selectively abstains from classification when uncertain. Through two novel regularization losses, our model maintains a balanced coverage across the two movement types, which prevents bias toward an easy-to-classify subset of movements. We show that our proposed model learns to gauge its own confidence on movement classification, and our proposed regularization losses effectively ensure that the model maintains a similar confidence across movement types. We also show that the local movement abstentions have little impact on the video-level coverage and that relying on the most confident predictions improves the video-level performance.
Romero F. A. B. de Morais, Thao Minh Le, Truyen Tran 0001, Caroline Alexander, Natasha Amery, Catherine Morgan, Alicia J. Spittle, Vuong Le, Nadia Badawi, Alison Salt, Jane Valentine, Catherine Elliott, Elizabeth M. Hurrion, Paul A. Dawson, Svetha Venkatesh
IEEE J. Biomed. Health Informatics3
2025 Learning Structural Causal Models from Ordering: Identifiable Flow Models
abstract
In this study, we address causal inference when only observational data and a valid causal ordering from the causal graph are available. We introduce a set of flow models that can recover component-wise, invertible transformation of exogenous variables. Our flow-based methods offer flexible model design while maintaining causal consistency regardless of the number of discretization steps. We propose design improvements that enable simultaneous learning of all causal mechanisms and reduce abduction and prediction complexity to linear O(n) relative to the number of layers, independent of the number of causal variables. Empirically, we demonstrate that our method outperforms previous state-of-the-art approaches and delivers consistent performance across a wide range of structural causal models in answering observational, interventional, and counterfactual questions. Additionally, our method achieves a significant reduction in computational time compared to existing diffusion-based techniques, making it practical for large structural causal models.
Minh Khoa Le, Kien Do, Truyen Tran 0001
AAAI3
2025 Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
abstract
Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between entities. This paper introduces Progressive multi-granular Vision-Language alignments (PromViL), a novel framework to enhance LVLMs' ability in performing grounded compositional visual reasoning tasks. Our approach constructs a hierarchical structure of multi-modal alignments, ranging from simple to complex concepts. By progressively aligning textual descriptions with corresponding visual regions, our model learns to leverage contextual information from lower levels to inform higher-level reasoning. To facilitate this learning process, we introduce a data generation process that creates a novel dataset derived from Visual Genome, providing a wide range of nested compositional vision-language pairs. Experimental results demonstrate that our PromViL framework significantly outperforms baselines on various visual grounding and compositional question answering tasks.
Quang-Hung Le, Long Hoang Dang, Ngan Hoang Le, Truyen Tran 0001, Thao Minh Le
AAAI4
2025 Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
abstract
Vision-language alignment in video must address the complexity of language, evolving interacting entities, their action chains, and semantic gaps between language and vision. This work introduces Planner-Refiner, a framework to overcome these challenges. Planner-Refiner bridges the semantic gap by iteratively refining visual elements’ space-time representation, guided by language until semantic gaps are minimal. A Planner module schedules language guidance by decomposing complex linguistic prompts into short sentence chains. The Refiner processes each short sentence—a noun-phrase and verb-phrase pair—to direct visual tokens’ self-attention across space then time, achieving efficient single-step refinement. A recurrent system chains these steps, maintaining refined visual token representations. The final representation feeds into task-specific heads for alignment generation. We demonstrate Planner-Refiner’s effectiveness on two video-language alignment tasks: Referring Video Object Segmentation and Temporal Grounding with varying language complexity. We further introduce a new MeViS-X benchmark to assess models’ capability with long queries. Superior performance versus state-of-the-art methods on these benchmarks shows the approach’s potential, especially for complex prompts.
Tuyen Tran, Thao Minh Le, Quang-Hung Le, Truyen Tran 0001
ECAI4
2025 Score-Based Integrated Gradient for Root Cause Explanations of Outliers
abstract
Identifying the root causes of outliers is a fundamental problem in causal inference and anomaly detection. Traditional approaches based on heuristics or counterfactual reasoning often struggle under uncertainty and high-dimensional dependencies. We introduce SIREN, a novel and scalable method that attributes the root causes of outliers by estimating the score functions of the data likelihood. Attribution is computed via integrated gradients that accumulate score contributions along paths from the outlier toward the normal data distribution. Our method satisfies three of the four classic Shapley value axioms-dummy, efficiency, and linearity-as well as an asymmetry axiom derived from the underlying causal structure. Unlike prior work, SIREN operates directly on the score function, enabling tractable and uncertainty-aware root cause attribution in nonlinear, high-dimensional, and heteroscedastic causal models. Extensive experiments on synthetic random graphs and real-world cloud service and supply chain datasets show that SIREN outperforms state-of-the-art baselines in both attribution accuracy and computational efficiency.
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Svetha Venkatesh
ICDM2
2025 Rapid Selection and Ordering of In-Context Demonstrations via Prompt Embedding Clustering
abstract
While Large Language Models (LLMs) excel at in-context learning (ICL) using just a few demonstrations, their performances are sensitive to demonstration orders. The reasons behind this sensitivity remain poorly understood. In this paper, we investigate the prompt embedding space to bridge the gap between the order sensitivity of ICL with inner workings of decoder-only LLMs, uncovering the clustering property: prompts sharing the first and last demonstrations have closer embeddings, with first-demonstration clustering usually being stronger in practice. We explain this property through extensive theoretical analyses and empirical evidences. Our finding suggests that the positional encoding and the causal attention mask are key contributors to the clustering phenomenon. Leveraging this clustering insight, we introduce Cluster-based Search, a novel method that accelerates the selection and ordering of demonstrations in self-adaptive ICL settings. Our approach substantially decreases the time complexity from factorial to quadratic, saving 92% to nearly 100% execution time while maintaining comparable performance to exhaustive search.
Kha Pham, Hung Le 0002, Man Ngo, Truyen Tran 0001
ICLR4
2025 Navigating Social Dilemmas with LLM-based Agents via Consideration of Future Consequences
Dung Nguyen 0001, Hung Le 0002, Kien Do, Sunil Gupta 0001, Svetha Venkatesh, Truyen Tran 0001
AAMAS6
2025 Navigating Social Dilemmas with LLM-based Agents via Consideration of Future Consequences
abstract
Artificial agents with the aid of large language models (LLMs) are effective in various real-world scenarios but struggle to cooperate in social dilemmas. When making decisions under the strain of selecting between long-term consequences and short-term benefits in commonly shared resources, LLM-based agents often exploit the environment, leading to early depletion. Inspired by the concept of consideration of future consequences (CFC), which is well-known in social psychology, we propose a framework to enable the ability to consider future consequences for LLM-based agents, which results in a new kind of agent that we term the CFC-Agent. We enable the CFC-Agent to act toward different levels of consideration for future consequences. Our first set of experiments, where LLM is directly asked to make decisions, shows that agents considering future consequences exhibit sustainable behaviour and achieve high common rewards for the population. Extensive experiments in complex environments showed that the CFC-Agent can manage a sequence of calls to LLM for reasoning and engaging in communication to cooperate with others to resolve the common dilemma better. Finally, our analysis showed that considering future consequences not only affects the final decision but also improves the conversations between LLM-based agents toward a better resolution of social dilemmas.
Dung Nguyen 0001, Hung Le 0002, Kien Do, Sunil Gupta 0001, Svetha Venkatesh, Truyen Tran 0001
IJCAI6
2025 LaGR-SEQ: Language-guided reinforcement learning with sample-efficient querying
abstract
Abstract Large language models (LLMs) have recently demonstrated their impressive ability to provide context-aware responses via text. This ability could potentially be used to predict plausible solutions in sequential decision making tasks pertaining to pattern completion. For example, by observing a partial stack of cubes, LLMs can predict the correct sequence in which the remaining cubes should be stacked by extrapolating the observed patterns (e.g., cube sizes, colors or other attributes) in the partial stack. In this work, we introduce LaGR (language-guided reinforcement learning), which uses this predictive ability of LLMs to propose solutions to tasks that have been partially completed by a primary reinforcement learning (RL) agent, in order to subsequently guide the latter’s training. However, as RL training is generally not sample-efficient, deploying this approach would inherently imply that the LLM be repeatedly queried for solutions; a process that can be expensive and infeasible. To address this issue, we introduce SEQ (sample-efficient querying), where we simultaneously train a secondary RL agent to decide when the LLM should be queried for solutions. Specifically, we use the quality of the solutions emanating from the LLM as the reward to train this agent. We show that our proposed framework LaGR-SEQ enables more efficient primary RL training, while simultaneously minimizing the number of queries to the LLM. We demonstrate our approach on a series of tasks and highlight the advantages of our approach, along with its limitations and potential future research directions.
Thommen George Karimpanal, Buddhika Laknath Semage, Santu Rana, Hung Le 0002, Truyen Tran 0001, Sunil Gupta 0001, Svetha Venkatesh
Neural Comput. Appl.5
2025 Fine-Grained Fidgety Movement Classification Using Active Learning
abstract
Typically developing infants, between the corrected age of 9-20 weeks, produce fidgety movements. These movements can be identified with the General Movement Assessment, but their identification requires trained professionals to conduct the assessment from video recordings. Since trained professionals are expensive and their demand may be higher than their availability, computer vision-based solutions have been developed to assist practitioners. However, most solutions to date treat the problem as a direct mapping from video to infant status, without modeling fidgety movements throughout the video. To address that, we propose to directly model infants' short movements and classify them as fidgety or non-fidgety. In this way, we model the explanatory factor behind the infant's status and improve model interpretability. The issue with our proposal is that labels for an infant's short movements are not available, which precludes us to train such a model. We overcome this issue with active learning. Active learning is a framework that minimizes the amount of labeled data required to train a model, by only labeling examples that are considered "informative" to the model. The assumption is that a model trained on informative examples reaches a higher performance level than a model trained with randomly selected examples. We validate our framework by modeling the movements of infants' hips on two representative cohorts: typically developing and at-risk infants. Our results show that active learning is suitable to our problem and that it works adequately even when the models are trained with labels provided by a novice annotator.
Romero F. A. B. de Morais, Truyen Tran 0001, Caroline Alexander, Natasha Amery, Catherine Morgan, Alicia J. Spittle, Vuong Le, Nadia Badawi, Alison Salt, Jane Valentine, Catherine Elliott, Elizabeth M. Hurrion, Paul A. Dawson, Svetha Venkatesh
IEEE J. Biomed. Health Informatics2
2024 Root Cause Explanation of Outliers under Noisy Mechanisms
abstract
Identifying root causes of anomalies in causal processes is vital across disciplines. Once identified, one can isolate the root causes and implement necessary measures to restore the normal operation. Causal processes are often modelled as graphs with entities being nodes and their paths/interconnections as edge. Existing work only consider the contribution of nodes in the generative process, thus can not attribute the outlier score to the edges of the mechanism if the anomaly occurs in the connections. In this paper, we consider both individual edge and node of each mechanism when identifying the root causes. We introduce a noisy functional causal model to account for this purpose. Then, we employ Bayesian learning and inference methods to infer the noises of the nodes and edges. We then represent the functional form of a target outlier leaf as a function of the node and edge noises. Finally, we propose an efficient gradient-based attribution method to compute the anomaly attribution scores which scales linearly with the number of nodes and edges. Experiments on simulated datasets and two real-world scenario datasets show better anomaly attribution performance of the proposed method compared to the baselines. Our method scales to larger graphs with more nodes and edges.
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Thin Nguyen, Svetha Venkatesh
AAAI2
2024 Unified Compositional Query Machine with Multimodal Consistency for Video-based Human Activity Recognition
Tuyen Tran, Thao Minh Le, Duy Hung Tran, Truyen Tran 0001
BMVC4
2024 Revisiting the Dataset Bias Problem from a Statistical Perspective
abstract
In this paper, we study the “dataset bias” problem from a statistical standpoint, and identify the main cause of the problem as the strong correlation between a class attribute u and a non-class attribute b in the input x, represented by p(u|b) differing significantly from p(u). Since p(u|b) appears as part of the sampling distributions in the standard maximum log-likelihood (MLL) objective, a model trained on a biased dataset via MLL inherently incorporates such correlation into its parameters, leading to poor generalization to unbiased test data. From this observation, we propose to mitigate dataset bias via either weighting the objective of each sample n by 1 / p(un|bn) or sampling that sample with a weight proportional to 1 / p(un|bn). While both methods are statistically equivalent, the former proves more stable and effective in practice. Additionally, we establish a connection between our debiasing approach and causal reasoning, reinforcing our method’s theoretical foundation. However, when the bias label is unavailable, computing p(u|b) exactly is difficult. To overcome this challenge, we propose to approximate 1 / p(u|b) using a biased classifier trained with “bias amplification” losses. Extensive experiments on various biased datasets demonstrate the superiority of our method over existing debiasing techniques in most settings, validating our theoretical analysis.
Kien Do, Dung Nguyen 0001, Hung Le 0002, Thao Le 0003, Dang Nguyen 0002, Haripriya Harikumar, Truyen Tran 0001, Santu Rana, Svetha Venkatesh
ECAI7
2024 Diversifying Training Pool Predictability for Zero-shot Coordination: A Theory of Mind Approach
Dung Nguyen 0001, Hung Le 0002, Kien Do, Sunil Gupta 0001, Svetha Venkatesh, Truyen Tran 0001
IJCAI6
2024 Learning evolving relations for multivariate time series forecasting
abstract
Abstract Multivariate time series forecasting is essential in various fields, including healthcare and traffic management, but it is a challenging task due to the strong dynamics in both intra-channel relations (temporal patterns within individual variables) and inter-channel relations (the relationships between variables), which can evolve over time with abrupt changes. This paper proposes ERAN (Evolving Relational Attention Network), a framework for multivariate time series forecasting, that is capable to capture such dynamics of these relations. On the one hand, ERAN represents inter-channel relations with a graph which evolves over time, modeled using a recurrent neural network. On the other hand, ERAN represents the intra-channel relations using a temporal attentional convolution, which captures the local temporal dependencies adaptively with the input data. The elvoving graph structure and the temporal attentional convolution are intergrated in a unified model to capture both types of relations. The model is experimented on a large number of real-life datasets including traffic flows, energy consumption, and COVID-19 transmission data. The experimental results show a significant improvement over the state-of-the-art methods in multivariate time series forecasting particularly for non-stationary data.
Binh Nguyen-Thai, Vuong Le, Ngoc-Dung T. Tieu, Truyen Tran 0001, Svetha Venkatesh, Naeem Ramzan
Appl. Intell.4
2023 Memory-Augmented Theory of Mind Network
abstract
Social reasoning necessitates the capacity of theory of mind (ToM), the ability to contextualise and attribute mental states to others without having access to their internal cognitive structure. Recent machine learning approaches to ToM have demonstrated that we can train the observer to read the past and present behaviours of other agents and infer their beliefs (including false beliefs about things that no longer exist), goals, intentions and future actions. The challenges arise when the behavioural space is complex, demanding skilful space navigation for rapidly changing contexts for an extended period. We tackle the challenges by equipping the observer with novel neural memory mechanisms to encode, and hierarchical attention to selectively retrieve information about others. The memories allow rapid, selective querying of distal related past behaviours of others to deliberatively reason about their current mental state, beliefs and future behaviours. This results in ToMMY, a theory of mind model that learns to reason while making little assumptions about the underlying mental processes. We also construct a new suite of experiments to demonstrate that memories facilitate the learning process and achieve better theory of mind performance, especially for high-demand false-belief tasks that require inferring through multiple steps of changes.
Dung Nguyen 0001, Phuoc Nguyen, Hung Le 0002, Kien Do, Svetha Venkatesh, Truyen Tran 0001
AAAI6
2023 Persistent-Transient Duality: A Multi-mechanism Approach for Modeling Human-Object Interaction
abstract
Humans are highly adaptable, swiftly switching between different modes to progressively handle different tasks, situations and contexts. In Human-object interaction (HOI) activities, these modes can be attributed to two mechanisms: (1) the large-scale consistent plan for the whole activity and (2) the small-scale children interactive actions that start and end along the timeline. While neuroscience and cognitive science have confirmed this multi-mechanism nature of human behavior, machine modeling approaches for human motion are trailing behind. While attempting to use gradually morphing structures (e.g., graph attention networks) to model the dynamic HOI patterns, they miss the expeditious and discrete mode-switching nature of the human motion. To bridge that gap, this work proposes to model two concurrent mechanisms that jointly control human motion: the Persistent process that runs continually on the global scale, and the Transient sub-processes that operate intermittently on the local context of the human while interacting with objects. These two mechanisms form an interactive Persistent-Transient Duality that synergistically governs the activity sequences. We model this conceptual duality by a parent-child neural network of Persistent and Transient channels with a dedicated neural module for dynamic mechanism switching. The framework is trialed on HOI motion forecasting. On two rich datasets and a wide variety of settings, the model consistently delivers superior performances, proving its suitability for the challenge.
Vuong Le, Svetha Venkatesh, Truyen Tran 0001
ICCV4
2023 Improving Out-of-distribution Generalization with Indirection Representations
Kha Pham, Hung Le 0002, Man Ngo, Truyen Tran 0001
ICLR4
2023 Social Motivation for Modelling Other Agents under Partial Observability in Decentralised Training
abstract
Understanding other agents is a key challenge in constructing artificial social agents. Current works focus on centralised training, wherein agents are allowed to know all the information about others and the environmental state during training. In contrast, this work studies decentralised training, wherein agents must learn the model of other agents in order to cooperate with them under partially-observable conditions, even during training, i.e. learning agents are myopic. The intrinsic motivation for artificial agents is modelled on the concept of human social motivation that entices humans to meet and understand each other, especially when experiencing a utility loss. Our intrinsic motivation encourages agents to stay near each other to obtain better observations and construct a model of others. They do so when their model of other agents is poor, or the overall task performance is bad during the learning phase. This simple but effective method facilitates the processes of modelling others, resulting in an improvement of the performance in cooperative tasks significantly. Our experiments demonstrate that the socially-motivated agent can model others better and promote cooperation across different tasks.
Dung Nguyen 0001, Hung Le 0002, Kien Do, Svetha Venkatesh, Truyen Tran 0001
IJCAI5
2023 Guiding Visual Question Answering with Attention Priors
abstract
The current success of modern visual reasoning systems is arguably attributed to cross-modality attention mechanisms. However, in deliberative reasoning such as in VQA, attention is unconstrained at each step, and thus may serve as a statistical pooling mechanism rather than a semantic operation intended to select information relevant to inference. This is because at training time, attention is only guided by a very sparse signal (i.e. the answer label) at the end of the inference chain. This causes the cross-modality attention weights to deviate from the desired visual-language bindings. To rectify this deviation, we propose to guide the attention mechanism using explicit linguistic-visual grounding. This grounding is derived by connecting structured linguistic concepts in the query to their referents among the visual objects. Here we learn the grounding from the pairing of questions and images alone, without the need for answer annotation or external grounding supervision. This grounding guides the attention mechanism inside VQA models through a duality of mechanisms: pre-training attention weight calculation and directly guiding the weights at inference time on a case- by-case basis. The resultant algorithm is capable of probing attention-based reasoning models, injecting relevant associative knowledge, and regulating the core reasoning process. This scalable enhancement improves the performance of VQA models, fortifies their robustness to limited access to supervised data, and increases interpretability.
Thao Minh Le, Vuong Le, Sunil Gupta 0001, Svetha Venkatesh, Truyen Tran 0001
WACV5
2023 Balanced Q-learning: Combining the influence of optimistic and pessimistic targets
abstract
The optimistic nature of the Q−learning target leads to an overestimation bias, which is an inherent problem associated with standard Q−learning. Such a bias fails to account for the possibility of low returns, particularly in risky scenarios. However, the existence of biases, whether overestimation or underestimation, need not necessarily be undesirable. In this paper, we analytically examine the utility of biased learning, and show that specific types of biases may be preferable, depending on the scenario. Based on this finding, we design a novel reinforcement learning algorithm, Balanced Q-learning, in which the target is modified to be a convex combination of a pessimistic and an optimistic term, whose associated weights are determined online, analytically. Such a balanced target inherently promotes risk-averse behavior, which we examine through the lens of the agent's exploration. We prove the convergence of this algorithm in a tabular setting, and empirically demonstrate its consistently good learning performance in various environments.
Thommen George Karimpanal, Hung Le 0002, Majid Abdolshah, Santu Rana, Sunil Gupta 0001, Truyen Tran 0001, Svetha Venkatesh
Artif. Intell.6
2023 Explaining Black Box Drug Target Prediction Through Model Agnostic Counterfactual Samples
abstract
Many high-performance DTA deep learning models have been proposed, but they are mostly black-box and thus lack human interpretability. Explainable AI (XAI) can make DTA models more trustworthy, and allows to distill biological knowledge from the models. Counterfactual explanation is one popular approach to explaining the behaviour of a deep neural network, which works by systematically answering the question "How would the model output change if the inputs were changed in this way?". We propose a multi-agent reinforcement learning framework, Multi-Agent Counterfactual Drug-target binding Affinity (MACDA), to generate counterfactual explanations for the drug-protein complex. Our proposed framework provides human-interpretable counterfactual instances while optimizing both the input drug and target for counterfactual generation at the same time. We benchmark the proposed MACDA framework using the Davis and PDBBind dataset and find that our framework produces more parsimonious explanations with no loss in explanation validity, as measured by encoding similarity. We then present a case study involving ABL1 and Nilotinib to demonstrate how MACDA can explain the behaviour of a DTA model in the underlying substructure interaction between inputs in its prediction, revealing mechanisms that align with prior domain knowledge.
Tri Minh Nguyen 0005, Thomas P. Quinn, Thin Nguyen, Truyen Tran 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Robust and Interpretable General Movement Assessment Using Fidgety Movement Detection
abstract
Fidgety movements occur in infants between the age of 9 to 20 weeks post-term, and their absence are a strong indicator that an infant has cerebral palsy. Prechtl's General Movement Assessment method evaluates whether an infant has fidgety movements, but requires a trained expert to conduct it. Timely evaluation facilitates early interventions, and thus computer-based methods have been developed to aid domain experts. However, current solutions rely on complex models or high-dimensional representations of the data, which hinder their interpretability and generalization ability. To address that we propose [Formula: see text], a method that detects fidgety movements and uses them towards an assessment of the quality of an infant's general movements. [Formula: see text] is true to the domain expert process, more accurate, and highly interpretable due to its fine-grained scoring system. The main idea behind [Formula: see text] is to specify signal properties of fidgety movements that are measurable and quantifiable. In particular, we measure the movement direction variability of joints of interest, for movements of small amplitude in short video segments. [Formula: see text] also comprises a strategy to reduce those measurements to a single score that quantifies the quality of an infant's general movements; the strategy is a direct translation of the qualitative procedure domain experts use to assess infants. This brings [Formula: see text] closer to the process a domain expert applies to decide whether an infant produced enough fidgety movements. We evaluated [Formula: see text] on the largest clinical dataset reported, where it showed to be interpretable and more accurate than many methods published to date.
Romero F. A. B. de Morais, Vuong Le, Catherine Morgan, Alicia J. Spittle, Nadia Badawi, Jane Valentine, Elizabeth M. Hurrion, Paul A. Dawson, Truyen Tran 0001, Svetha Venkatesh
IEEE J. Biomed. Health Informatics9
2022 Towards Effective and Robust Neural Trojan Defenses via Input Filtering
Kien Do, Haripriya Harikumar, Hung Le 0002, Dung Nguyen 0001, Truyen Tran 0001, Santu Rana, Dang Nguyen 0002, Willy Susilo, Svetha Venkatesh
ECCV (5)5
2022 Video Dialog as Conversation About Objects Living in Space-Time
Thao Minh Le, Vuong Le, Tu Minh Phuong, Truyen Tran 0001
ECCV (39)5
2022 Generative Pseudo-Inverse Memory
Kha Pham, Hung Le 0002, Man Ngo, Truyen Tran 0001, Bao Ho, Svetha Venkatesh
ICLR4
2022 Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge Distillation
abstract
Data-free Knowledge Distillation (DFKD) has attracted attention recently thanks to its appealing capability of transferring knowledge from a teacher network to a student network without using training data. The main idea is to use a generator to synthesize data for training the student. As the generator gets updated, the distribution of synthetic data will change. Such distribution shift could be large if the generator and the student are trained adversarially, causing the student to forget the knowledge it acquired at the previous steps. To alleviate this problem, we propose a simple yet effective method called Momentum Adversarial Distillation (MAD) which maintains an exponential moving average (EMA) copy of the generator and uses synthetic samples from both the generator and the EMA generator to train the student. Since the EMA generator can be considered as an ensemble of the generator's old versions and often undergoes a smaller change in updates compared to the generator, training on its synthetic samples can help the student recall the past knowledge and prevent the student from adapting too quickly to the new updates of the generator. Our experiments on six benchmark datasets including big datasets like ImageNet and Places365 demonstrate the superior performance of MAD over competing methods for handling the large distribution shift problem. Our method also compares favorably to existing DFKD methods and even achieves state-of-the-art results in some cases.
Kien Do, Hung Le 0002, Dung Nguyen 0001, Dang Nguyen 0002, Haripriya Harikumar, Truyen Tran 0001, Santu Rana, Svetha Venkatesh
NeurIPS6
2022 Functional Indirection Neural Estimator for Better Out-of-distribution Generalization
abstract
The capacity to achieve out-of-distribution (OOD) generalization is a hallmark of human intelligence and yet remains out of reach for machines. This remarkable capability has been attributed to our abilities to make conceptual abstraction and analogy, and to a mechanism known as indirection, which binds two representations and uses one representation to refer to the other. Inspired by these mechanisms, we hypothesize that OOD generalization may be achieved by performing analogy-making and indirection in the functional space instead of the data space as in current methods. To realize this, we design FINE (Functional Indirection Neural Estimator), a neural framework that learns to compose functions that map data input to output on-the-fly. FINE consists of a backbone network and a trainable semantic memory of basis weight matrices. Upon seeing a new input-output data pair, FINE dynamically constructs the backbone weights by mixing the basis weights. The mixing coefficients are indirectly computed through querying a separate corresponding semantic memory using the data pair. We demonstrate empirically that FINE can strongly improve out-of-distribution generalization on IQ tasks that involve geometric transformations. In particular, we train FINE and competing models on IQ tasks using images from the MNIST, Omniglot and CIFAR100 datasets and test on tasks with unseen image classes from one or different datasets and unseen transformation rules. FINE not only achieves the best performance on all tasks but also is able to adapt to small-scale data scenarios.
Kha Pham, Hung Le 0002, Man Ngo, Truyen Tran 0001
NeurIPS4
2022 Mitigating cold-start problems in drug-target affinity prediction with interaction knowledge transferring
abstract
Predicting the drug-target interaction is crucial for drug discovery as well as drug repurposing. Machine learning is commonly used in drug-target affinity (DTA) problem. However, the machine learning model faces the cold-start problem where the model performance drops when predicting the interaction of a novel drug or target. Previous works try to solve the cold start problem by learning the drug or target representation using unsupervised learning. While the drug or target representation can be learned in an unsupervised manner, it still lacks the interaction information, which is critical in drug-target interaction. To incorporate the interaction information into the drug and protein interaction, we proposed using transfer learning from chemical-chemical interaction (CCI) and protein-protein interaction (PPI) task to drug-target interaction task. The representation learned by CCI and PPI tasks can be transferred smoothly to the DTA task due to the similar nature of the tasks. The result on the DTA datasets shows that our proposed method has advantages compared to other pre-training methods in the DTA task.
Tri Minh Nguyen 0005, Thin Nguyen, Truyen Tran 0001
Briefings Bioinform.3
2022 GEFA: Early Fusion Approach in Drug-Target Affinity Prediction
abstract
Predicting the interaction between a compound and a target is crucial for rapid drug repurposing. Deep learning has been successfully applied in drug-target affinity (DTA)problem. However, previous deep learning-based methods ignore modeling the direct interactions between drug and protein residues. This would lead to inaccurate learning of target representation which may change due to the drug binding effects. In addition, previous DTA methods learn protein representation solely based on a small number of protein sequences in DTA datasets while neglecting the use of proteins outside of the DTA datasets. We propose GEFA (Graph Early Fusion Affinity), a novel graph-in-graph neural network with attention mechanism to address the changes in target representation because of the binding effects. Specifically, a drug is modeled as a graph of atoms, which then serves as a node in a larger graph of residues-drug complex. The resulting model is an expressive deep nested graph neural network. We also use pre-trained protein representation powered by the recent effort of learning contextualized protein representation. The experiments are conducted under different settings to evaluate scenarios such as novel drugs or targets. The results demonstrate the effectiveness of the pre-trained protein embedding and the advantages our GEFA in modeling the nested graph for drug-target interaction.
Tri Minh Nguyen 0005, Thin Nguyen, Thao Minh Le, Truyen Tran 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Semi-Supervised Learning with Variational Bayesian Inference and Maximum Uncertainty Regularization
Kien Do, Truyen Tran 0001, Svetha Venkatesh
AAAI2
2021 Learning Asynchronous and Sparse Human-Object Interaction in Videos
abstract
Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structure of the activities, such as the progression of the sub-activities. Automatically recognizing such structure from raw video signal is a new capability that promises authentic modeling and successful recognition of human-object interactions. Toward this goal, we introduce Asynchronous-Sparse Interaction Graph Networks (ASSIGN), a recurrent graph network that is able to automatically detect the structure of interaction events associated with entities in a video scene. ASSIGN pioneers learning of autonomous behavior of video entities including their dynamic structure and their interaction with the coexisting neighbors. Entities’ lives in our model are asynchronous to those of others therefore more flexible in adapting to complex scenarios. Their interactions are sparse in time hence more faithful to the true underlying nature and more robust in inference and learning. ASSIGN is tested on human-object interaction recognition and shows superior performance in segmenting and labeling of human sub-activities and object affordances from raw videos. The native ability of ASSIGN in discovering temporal structure also eliminates the dependence on external segmentation that was previously mandatory for this task.
Romero F. A. B. de Morais, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
CVPR4
2021 Clustering by Maximizing Mutual Information Across Views
abstract
We propose a novel framework for image clustering that incorporates joint representation learning and clustering. Our method consists of two heads that share the same backbone network - a "representation learning" head and a "clustering" head. The "representation learning" head captures fine-grained patterns of objects at the instance level which serve as clues for the "clustering" head to extract coarse-grain information that separates objects into clusters. The whole model is trained in an end-to-end manner by minimizing the weighted sum of two sample-oriented contrastive losses applied to the outputs of the two heads. To ensure that the contrastive loss corresponding to the "clustering" head is optimal, we introduce a novel critic function called "log-of-dot-product". Extensive experimental results demonstrate that our method significantly outperforms state-of-the-art single-stage clustering methods across a variety of image datasets, improving over the best baseline by about 5-7% in accuracy on CIFAR10/20, STL10, and ImageNet-Dogs. Further, the "two-stage" variant of our method also achieves better results than baselines on three challenging ImageNet subsets.
Kien Do, Truyen Tran 0001, Svetha Venkatesh
ICCV2
2021 DeepProcess: Supporting Business Process Execution Using a MANN-Based Recommender System
Muhammad Asjad Khan, Hung Le 0002, Kien Do, Truyen Tran 0001, Aditya Ghose, Khanh Hoa Dam, Renuka Sindhgatta
ICSOC4
2021 Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering
abstract
Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visual and linguistic domains in space-time. High-level reasoning demands lifting from associative visual pattern recognition to symbol like manipulation over objects, their behavior and interactions. Toward reaching this goal we propose an object-oriented reasoning approach in that video is abstracted as a dynamic stream of interacting objects. At each stage of the video event flow, these objects interact with each other, and their interactions are reasoned about with respect to the query and under the overall context of a video. This mechanism is materialized into a family of general-purpose neural units and their multi-level architecture called Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR) networks. This neural model maintains the objects' consistent lifelines in the form of a hierarchically nested spatio-temporal graph. Within this graph, the dynamic interactive object-oriented representations are built up along the video sequence, hierarchically abstracted in a bottom-up manner, and converge toward the key information for the correct answer. The method is evaluated on multiple major Video QA datasets and establishes new state-of-the-arts in these tasks. Analysis into the model's behavior indicates that object-oriented reasoning is a reliable, interpretable and efficient approach to Video QA.
Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran 0001
IJCAI4
2021 Object-Centric Representation Learning for Video Question Answering
abstract
Video question answering (Video QA) presents a powerful testbed for human-like intelligent behaviors. The task demands new capabilities to integrate video processing, language understanding, binding abstract linguistic concepts to concrete visual artifacts, and deliberative reasoning over spacetime. Neural networks offer a promising approach to reach this potential through learning from examples rather than handcrafting features and rules. However, neural networks are predominantly feature-based - they map data to unstructured vectorial representation and thus can fall into the trap of exploiting shortcuts through surface statistics instead of true systematic reasoning seen in symbolic systems. To tackle this issue, we advocate for object-centric representation as a basis for constructing spatio-temporal structures from videos, essentially bridging the semantic gap between low-level pattern recognition and high-level symbolic algebra. To this end, we propose a new query-guided representation framework to turn a video into an evolving relational graph of objects, whose features and interactions are dynamically and conditionally inferred. The object lives are then summarized into résumés, lending naturally for deliberative relational reasoning that produces an answer to the query. The framework is evaluated on major Video QA datasets, demonstrating clear benefits of the object-centric approach to video reasoning.
Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran 0001
IJCNN4
2021 From Deep Learning to Deep Reasoning
abstract
The rise of big data and big compute has brought modern neural networks to many walks of digital life, thanks to the relative ease of constructing large models that scale to the real world. Current successes of Transformers and self-supervised pretraining on massive data have led some to believe that deep neural networks will be able to do almost everything once we have sufficient data and computational resources. However, neural networks are fast to exploit surface statistics but fail miserably to generalize to novel combinations. This is because they are not designed for deliberate reasoning -- the capacity to deliberately deduce new knowledge out of the contextualized data. This tutorial reviews recent developments to extend the capacity of neural networks to "learning-to-reason'' from data, where the task is to determine if the data entails a conclusion. This capacity opens up new ways to generate insights from data through arbitrary compositional querying without the need of predefining a narrow set of tasks. The tutorial consists of four parts. The first part covers the learning-to-reason framework, and explains how neural networks can serve as a strong backbone for reasoning through its natural operations such as binding, attention & dynamic computational graphs. The second part goes into more detail on how neural networks perform reasoning over unstructured and structured data, and across modalities. The third part reviews neural memories and their role in reasoning. The last part discusses generalization to novel combinations, under less supervision and with more knowledge.
Truyen Tran 0001, Vuong Le, Hung Le 0002, Thao Minh Le
KDD1
2021 Model-Based Episodic Memory Induces Dynamic Hybrid Controls
abstract
Episodic control enables sample efficiency in reinforcement learning by recalling past experiences from an episodic memory. We propose a new model-based episodic memory of trajectories addressing current limitations of episodic control. Our memory estimates trajectory values, guiding the agent towards good policies. Built upon the memory, we construct a complementary learning model via a dynamic hybrid control unifying model-based, episodic and habitual learning into a single architecture. Experiments demonstrate that our model allows significantly faster and better learning than other strong reinforcement learning agents across a variety of environments including stochastic and non-Markovian settings.
Hung Le 0002, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 0001, Svetha Venkatesh
NeurIPS4
2021 Knowledge Distillation with Distribution Mismatch
Dang Nguyen 0002, Sunil Gupta 0001, Trong Nguyen, Santu Rana, Phuoc Nguyen, Truyen Tran 0001, Ky Le, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (2)6
2021 Variational Hyper-encoding Networks
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Hieu-Chi Dam, Svetha Venkatesh
ECML/PKDD (2)2
2021 Fast Conditional Network Compression Using Bayesian HyperNetworks
Phuoc Nguyen, Truyen Tran 0001, Ky Le, Sunil Gupta 0001, Santu Rana, Dang Nguyen 0002, Trong Nguyen, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (3)2
2021 Goal-driven Long-Term Trajectory Prediction
abstract
The prediction of humans' short-term trajectories has advanced significantly with the use of powerful sequential modeling and rich environment feature extraction. However, long-term prediction is still a major challenge for the current methods as the errors could accumulate along the way. Indeed, consistent and stable prediction far to the end of a trajectory inherently requires deeper analysis into the overall structure of that trajectory, which is related to the pedestrian's intention on the destination of the journey. In this work, we propose to model a hypothetical process that determines pedestrians' goals and the impact of such process on long-term future trajectories. We design Goal-driven Trajectory Prediction model - a dual-channel neural network that realizes such intuition. The two channels of the network take their dedicated roles and collaborate to generate future trajectories. Different than conventional goal-conditioned, planning-based methods, the model architecture is designed to generalize the patterns and work across different scenes with arbitrary geometrical and semantic structures. The model is shown to outperform the state-of-the-art in various settings, especially in large prediction horizons. This result is another evidence for the effectiveness of adaptive structured representation of visual and geometrical features in human behavior analysis.
Vuong Le, Truyen Tran 0001
WACV3
2021 Automatically recommending components for issue reports using deep learning
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Chaiyong Ragkhitwetsagul, Aditya Ghose
Empir. Softw. Eng.3
2021 Hierarchical Conditional Relation Networks for Multimodal Video Question Answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
Int. J. Comput. Vis.4
2021 PAN: Personalized Annotation-Based Networks for the Prediction of Breast Cancer Relapse
abstract
The classification of clinical samples based on gene expression data is an important part of precision medicine. In this manuscript, we show how transforming gene expression data into a set of personalized (sample-specific) networks can allow us to harness existing graph-based methods to improve classifier performance. Existing approaches to personalized gene networks have the limitation that they depend on other samples in the data and must get re-computed whenever a new sample is introduced. Here, we propose a novel method, called Personalized Annotation-based Networks (PAN), that avoids this limitation by using curated annotation databases to transform gene expression data into a graph. Unlike competing methods, PANs are calculated for each sample independent of the population, making it a more efficient way to obtain single-sample networks. Using three breast cancer datasets as a case study, we show that PAN classifiers not only predict cancer relapse better than gene features alone, but also outperform PPI (protein-protein interactions) and population-level graph-based classifiers. This work demonstrates the practical advantages of graph-based classification for high-dimensional genomic data, while offering a new approach to making sample-specific networks. Supplementary information: PAN and the baselines are implemented in Python. Source code and data are available at https://github.com/thinng/PAN.
Thin Nguyen, Samuel C. Lee, Thomas P. Quinn, Buu Minh Thanh Truong, Truyen Tran 0001, Svetha Venkatesh, Thuc Duy Le
IEEE ACM Trans. Comput. Biol. Bioinform.6
2021 A Spatio-Temporal Attention-Based Model for Infant Movement Assessment From Videos
abstract
The absence or abnormality of fidgety movements of joints or limbs is strongly indicative of cerebral palsy in infants. Developing computer-based methods for assessing infant movements in videos is pivotal for improved cerebral palsy screening. Most existing methods use appearance-based features and are thus sensitive to strong but irrelevant signals caused by background clutter or a moving camera. Moreover, these features are computed over the whole frame, thus they measure gross whole body movements rather than specific joint/limb motion. Addressing these challenges, we develop and validate a new method for fidgety movement assessment from consumer-grade videos using human poses extracted from short clips. Human poses capture only relevant motion profiles of joints and limbs and are thus free from irrelevant appearance artifacts. The dynamics and coordination between joints are modeled using spatio-temporal graph convolutional networks. Frames and body parts that contain discriminative information about fidgety movements are selected through a spatio-temporal attention mechanism. We validate the proposed model on the cerebral palsy screening task using a real-life consumer-grade video dataset collected at an Australian hospital through the Cerebral Palsy Alliance, Australia. Our experiments show that the proposed method achieves the ROC-AUC score of 81.87%, significantly outperforming existing competing methods with better interpretability.
Binh Nguyen-Thai, Vuong Le, Catherine Morgan, Nadia Badawi, Truyen Tran 0001, Svetha Venkatesh
IEEE J. Biomed. Health Informatics5
2021 Automatic Feature Learning for Predicting Vulnerable Software Components
abstract
Code flaws or vulnerabilities are prevalent in software systems and can potentially cause a variety of problems including deadlock, hacking, information loss and system failure. A variety of approaches have been developed to try and detect the most likely locations of such code vulnerabilities in large code bases. Most of them rely on manually designing code features (e.g., complexity metrics or frequencies of code tokens) that represent the characteristics of the potentially problematic code to locate. However, all suffer from challenges in sufficiently capturing both semantic and syntactic representation of source code, an important capability for building accurate prediction models. In this paper, we describe a new approach, built upon the powerful deep learning Long Short Term Memory model, to automatically learn both semantic and syntactic features of code. Our evaluation on 18 Android applications and the Firefox application demonstrates that the prediction power obtained from our learned features is better than what is achieved by state of the art vulnerability prediction models, for both within-project prediction and cross-project prediction.
Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Shien Wee Ng, John C. Grundy, Aditya Ghose
IEEE Trans. Software Eng.2
2020 Theory of Mind with Guilt Aversion Facilitates Cooperative Reinforcement Learning
abstract
Guilt aversion induces experience of a utility loss in people if they believe they have disappointed others, and this promotes cooperative behaviour in human. In psychological game theory, guilt aversion necessitates modelling of agents that have theory about what other agents think, also known as Theory of Mind (ToM). We aim to build a new kind of affective reinforcement learning agents, called Theory of Mind Agents with Guilt Aversion (ToMAGA), which are equipped with an ability to think about the wellbeing of others instead of just self-interest. To validate the agent design, we use a general-sum game known as Stag Hunt as a test bed. As standard reinforcement learning agents could learn suboptimal policies in social dilemmas like Stag Hunt, we propose to use belief-based guilt aversion as a reward shaping mechanism. We show that our belief-based guilt averse agents can efficiently learn cooperative behaviours in Stag Hunt Games.
Dung Nguyen 0001, Svetha Venkatesh, Phuoc Nguyen, Truyen Tran 0001
ACML4
2020 Learning to Abstract and Predict Human Actions
Romero F. A. B. de Morais, Vuong Le, Truyen Tran 0001, Svetha Venkatesh
BMVC3
2020 Hierarchical Conditional Relation Networks for Video Question Answering
abstract
Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts. We introduce a general-purpose reusable neural unit called Conditional Relation Network (CRN) that serves as a building block to construct more sophisticated structures for representation and reasoning over video. CRN takes as input an array of tensorial objects and a conditioning feature, and computes an array of encoded output objects. Model building becomes a simple exercise of replication, rearrangement and stacking of these reusable units for diverse modalities and contextual information. This design thus supports high-order relational and multi-step reasoning. The resulting architecture for VideoQA is a CRN hierarchy whose branches represent sub-videos or clips, all sharing the same question as the contextual condition. Our evaluations on well-known datasets achieved new SoTA results, demonstrating the impact of building a general-purpose reasoning unit on complex domains such as VideoQA.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
CVPR4
2020 Theory and Evaluation Metrics for Learning Disentangled Representations
Kien Do, Truyen Tran 0001
ICLR2
2020 Neural Stored-program Memory
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh
ICLR2
2020 Self-Attentive Associative Memory
abstract
Heretofore, neural networks with external memory are restricted to single memory with lossy representations of memory interactions. A rich representation of relationships between memory pieces urges a high-order and segregated relational memory. In this paper, we propose to separate the storage of individual experiences (item memory) and their occurring relationships (relational memory). The idea is implemented through a novel Self-attentive Associative Memory (SAM) operator. Found upon outer product, SAM forms a set of associative memories that represent the hypothetical high-order relationships between arbitrary pairs of memory elements, through which a relational memory is constructed from an item memory. The two memories are wired into a single sequential model capable of both memorization and relational reasoning. We achieve competitive results with our proposed two-memory model in a diversity of machine learning tasks, from challenging synthetic problems to practical testbeds such as geometry, graph, reinforcement learning, and question answering.
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh
ICML2
2020 Dynamic Language Binding in Relational Visual Reasoning
abstract
We present Language-binding Object Graph Network, the first neural reasoning method with dynamic relational structures across both visual and textual domains with applications in visual question answering. Relaxing the common assumption made by current models that the object predicates pre-exist and stay static, passive to the reasoning process, we propose that these dynamic predicates expand across the domain borders to include pair-wise visual-linguistic object binding. In our method, these contextualized object links are actively found within each recurrent reasoning step without relying on external predicative priors. These dynamic structures reflect the conditional dual-domain object dependency given the evolving context of the reasoning through co-attention. Such discovered dynamic graphs facilitate multi-step knowledge combination and refinements that iteratively deduce the compact representation of the final answer. The effectiveness of this model is demonstrated on image question answering demonstrating favorable performance on major VQA datasets. Our method outperforms other methods in sophisticated question-answering tasks wherein multiple object relations are involved. The graph structure effectively assists the progress of training, and therefore the network learns efficiently compared to other reasoning models.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
IJCAI4
2020 Learning Transferable Domain Priors for Safe Exploration in Reinforcement Learning
abstract
Prior access to domain knowledge could significantly improve the performance of a reinforcement learning agent. In particular, it could help agents avoid potentially catastrophic exploratory actions, which would otherwise have to be experienced during learning. In this work, we identify consistently undesirable actions in a set of previously learned tasks, and use pseudo-rewards associated with them to learn a prior policy. In addition to enabling safer exploratory behaviors in subsequent tasks in the domain, we show that these priors are transferable to similar environments, and can be learned off-policy and in parallel with the learning of other tasks in the domain. We compare our approach to established, state-of-the-art algorithms in both discrete as well as continuous environments, and demonstrate that it exhibits a safer exploratory behavior while learning to perform arbitrary tasks in the domain. We also present a theoretical analysis to support these results, and briefly discuss the implications and some alternative formulations of this approach, which could also be useful in certain scenarios.
Thommen George Karimpanal, Santu Rana, Sunil Gupta 0001, Truyen Tran 0001, Svetha Venkatesh
IJCNN4
2020 Neural Reasoning, Fast and Slow, for Video Question Answering
abstract
What does it take to design a machine that learns to answer natural questions about a video? A Video QA system must simultaneously understand language, represent visual content over space-time, and iteratively transform these representations in response to lingual content in the query, and finally arriving at a sensible answer. While recent advances in lingual and visual question answering have enabled sophisticated representations and neural reasoning mechanisms, major challenges in Video QA remain on dynamic grounding of concepts, relations and actions to support the reasoning process. Inspired by the dual-process account of human reasoning, we design a dual process neural architecture, which is composed of a question-guided video processing module (System 1, fast and reactive) followed by a generic reasoning module (System 2, slow and deliberative). System 1 is a hierarchical model that encodes visual patterns about objects, actions and relations in space-time given the textual cues from the question. The encoded representation is a set of high-level visual features, which are then passed to System 2. Here multi-step inference follows to iteratively chain visual elements as instructed by the textual elements. The system is evaluated on the SVQA (synthetic) and TGIF-QA datasets (real), demonstrating competitive results, with a large margin in the case of multi-step reasoning.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
IJCNN4
2020 Catastrophic forgetting and mode collapse in GANs
abstract
In this paper, we show that Generative Adversarial Networks (GANs) suffer from catastrophic forgetting even when they are trained to approximate a single target distribution. We show that GAN training is a continual learning problem in which the sequence of changing model distributions is the sequence of tasks to the discriminator. The level of mismatch between tasks in the sequence determines the level of forgetting. Catastrophic forgetting is interrelated to mode collapse and can make the training of GANs non-convergent. We investigate the landscape of the discriminator's output in different variants of GANs and find that when a GAN converges to a good equilibrium, real training datapoints are wide local maxima of the discriminator. We empirically show the relationship between the sharpness of local maxima and mode collapse and generalization in GANs. We show how catastrophic forgetting prevents the discriminator from making real datapoints local maxima, and thus causes non-convergence. Finally, we study methods for preventing catastrophic forgetting in GANs.
Hoang Thanh-Tung, Truyen Tran 0001
IJCNN2
2019 Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos
abstract
Appearance features have been widely used in video anomaly detection even though they contain complex entangled factors. We propose a new method to model the normal patterns of human movements in surveillance video for anomaly detection using dynamic skeleton features. We decompose the skeletal movements into two sub-components: global body movement and local body posture. We model the dynamics and interaction of the coupled features in our novel Message-Passing Encoder-Decoder Recurrent Network. We observed that the decoupled features collaboratively interact in our spatio-temporal model to accurately identify human-related irregular events from surveillance video sequences. Compared to traditional appearance-based models, our method achieves superior outlier detection performance. Our model also offers “open-box” examination and decision explanation made possible by the semantically understandable features and a network architecture supporting interpretability.
Romero F. A. B. de Morais, Vuong Le, Truyen Tran 0001, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh
CVPR3
2019 Learning to Remember More with Less Memorization
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh
ICLR2
2019 Improving Generalization and Stability of Generative Adversarial Networks
Hoang Thanh-Tung, Truyen Tran 0001, Svetha Venkatesh
ICLR (Poster)2
2019 Graph Transformation Policy Network for Chemical Reaction Prediction
abstract
We address a fundamental problem in chemistry known as chemical reaction product prediction. Our main insight is that the input reactant and reagent molecules can be jointly represented as a graph, and the process of generating product molecules from reactant molecules can be formulated as a sequence of graph transformations. To this end, we propose Graph Transformation Policy Network (GTPN) - a novel generic method that combines the strengths of graph neural networks and reinforcement learning to learn reactions directly from data with minimal chemical knowledge. Compared to previous methods, GTPN has some appealing properties such as: end-to-end learning, and making no assumption about the length or the order of graph transformations. In order to guide model search through the complex discrete space of sets of bond changes effectively, we extend the standard policy gradient loss by adding useful constraints. Evaluation results show that GTPN improves the top-1 accuracy over the current state-of-the-art method by about 3% on the large USPTO dataset.
Kien Do, Truyen Tran 0001, Svetha Venkatesh
KDD2
2019 Lessons learned from using a deep tree-based model for software defect prediction in practice
abstract
Defects are common in software systems and cause many problems for software users. Different methods have been developed to make early prediction about the most likely defective modules in large codebases. Most focus on designing features (e.g. complexity metrics) that correlate with potentially defective code. Those approaches however do not sufficiently capture the syntax and multiple levels of semantics of source code, a potentially important capability for building accurate prediction models. In this paper, we report on our experience of deploying a new deep learning tree-based defect prediction model in practice. This model is built upon the tree-structured Long Short Term Memory network which directly matches with the Abstract Syntax Tree representation of source code. We discuss a number of lessons learned from developing the model and evaluating it on two datasets, one from open source projects contributed by our industry partner Samsung and the other from the public PROMISE repository.
Khanh Hoa Dam, Trang Pham, Shien Wee Ng, Truyen Tran 0001, John C. Grundy, Aditya Ghose, Taeksu Kim, Chul-Joo Kim
MSR4
2019 Incomplete Conditional Density Estimation for Fast Materials Discovery
abstract
Designing new physical products and processes requires enormous experimentation. The scientific simulators play a fundamental role for such design tasks. To design a new product with certain target characteristics, a search is performed in the design space by trying out a large number of design combinations through simulators before reaching to the target characteristics. However, searching for the target design using simulators is generally expensive and becomes prohibitive when the target is either revised or only partially specified. To address this problem, we use a machine learning model to predict the design in single step using the target product specifications as input. We overcome two technical challenges: the first caused due to one-to-many mapping when learning the inverse problem and the second caused due to a user specifying the target specifications only partially. We unify a conditional variational auto-encoder model (to address the partial target specification) with mixture density networks (to address the one-to-many mapping) and train an end-to-end model to predict the optimum design.
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Matthew Barnett, Svetha Venkatesh
SDM2
2019 Attentional multilabel learning over graphs: a message passing approach
Kien Do, Truyen Tran 0001, Thin Nguyen, Svetha Venkatesh
Mach. Learn.2
2019 A Deep Learning Model for Estimating Story Points
abstract
Although there has been substantial research in software analytics for effort estimation in traditional software projects, little work has been done for estimation in agile projects, especially estimating the effort required for completing user stories or issues. Story points are the most common unit of measure used for estimating the effort involved in completing a user story or resolving an issue. In this paper, we propose a prediction model for estimating story points based on a novel combination of two powerful deep learning architectures: long short-term memory and recurrent highway network. Our prediction system is end-to-end trainable from raw input data to prediction outcomes without any manual feature engineering. We offer a comprehensive dataset for story points-based estimation that contains 23,313 issues from 16 open source projects. An empirical evaluation demonstrates that our approach consistently outperforms three common baselines (Random Guessing, Mean, and Median methods) and six alternatives (e.g., using Doc2Vec and Random Forests) in Mean Absolute Error, Median Absolute Error, and the Standardized Accuracy.
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Aditya Ghose, Tim Menzies
IEEE Trans. Software Eng.3
2018 Knowledge Graph Embedding with Multiple Relation Projections
abstract
Knowledge graphs contain rich relational structures of the world, and thus complement data-driven knowledge discovery from heterogeneous data. Relational inference between distant entities in large-scale knowledge graphs demands fast relation-specific algebraic manipulations. One of the most effective methods is to embed symbolic relations and entities into continuous spaces, where relations are approximately linear translation between projected images of entities in the relation space. However, state-of-art relation projection methods such as TransR, TransD or TransSparse do not model the correlation between relations, and thus are not scalable to complex knowledge graphs with thousands of relations, both in term of computational demand and statistical robustness. To this end we introduce TransF, a novel translation-based method which mitigates the burden of relation projection by explicitly modeling the basis subspaces of projection matrices. As a result, TransF is far more light weight than the existing projection methods, and is robust when facing a high number of relations. Experimental results on canonical link prediction and triples classification tasks show that our proposed model outperforms competing rivals by a large margin and achieves state-of-the-art performance. Especially, TransF improves by 9% (5%) on the head/tail entity prediction task with N-to-l (l-to-N) over the best performing translation-based method.
Kien Do, Truyen Tran 0001, Svetha Venkatesh
ICPR2
2018 Graph Memory Networks for Molecular Activity Prediction
abstract
Molecular activity prediction is critical in drug design. Machine learning techniques such as kernel methods and random forests have been successful for this task. These models require fixed-size feature vectors as input while the molecules are variable in size and structure. As a result, fixed-size fingerprint representation is poor in handling substructures for large molecules. Here we approach the problem through deep neural networks as they are flexible in modeling structured data such as grids, sequences and graphs. We train multiple BioAssays using a multi-task learning framework, which combines information from multiple sources to improve the performance of prediction, especially on small datasets. We propose Graph Memory Network (GraphMem), a memory-augmented neural network to model the graph structure in molecules. GraphMem consists of a recurrent controller coupled with an external memory whose cells dynamically interact and change through a multi-hop reasoning process. Applied to the molecules, the dynamic interactions enable an iterative refinement of the representation of molecular graphs with multiple bond types. GraphMem is capable of jointly training on multiple datasets by using a specific-task query fed to the controller as an input. We demonstrate the effectiveness of the proposed model for separately and jointly training on more than 100K measurements, spanning across 9 BioAssay activity tests.
Trang Pham, Truyen Tran 0001, Svetha Venkatesh
ICPR2
2018 Resset: A Recurrent Model for Sequence of Sets with Applications to Electronic Medical Records
abstract
Modern healthcare is ripe for disruption by AI. A game changer would be automatic understanding the latent processes from electronic medical records, which are being collected for billions of people worldwide. However, these healthcare processes are complicated by the interaction between at least three dynamic components: the illness which involves multiple diseases, the care which involves multiple treatments, and the recording practice which is biased and erroneous. Existing methods are inadequate in capturing the dynamic structure of care. We propose Resset, an end-to-end recurrent model that reads medical record and predicts future risk. The model adopts the algebraic view in that discrete medical objects are embedded into continuous vectors lying in the same space. We formulate the problem as modeling sequences of sets, a novel setting that have rarely, if not, been addressed. Within Resset, the bag of diseases recorded at each clinic visit is modeled as function of sets. The same hold for the bag of treatments. The interaction between the disease bag and the treatment bag at a visit is modeled in several, one of which as residual of diseases minus the treatments. Finally, the health trajectory, which is a sequence of visits, is modeled using a recurrent neural network. We report results on over a hundred thousand hospital visits by patients suffered from two costly chronic diseases - diabetes and mental health. Resset shows promises in multiple predictive tasks such as readmission prediction, treatments recommendation and diseases progression.
Phuoc Nguyen, Truyen Tran 0001, Svetha Venkatesh
IJCNN2
2018 Dual Memory Neural Computer for Asynchronous Two-view Sequential Learning
abstract
One of the core tasks in multi-view learning is to capture relations among views. For sequential data, the relations not only span across views, but also extend throughout the view length to form long-term intra-view and inter-view interactions. In this paper, we present a new memory augmented neural network that aims to model these complex interactions between two asynchronous sequential views. Our model uses two encoders for reading from and writing to two external memories for encoding input views. The intra-view interactions and the long-term dependencies are captured by the use of memories during this encoding process. There are two modes of memory accessing in our system: late-fusion and early-fusion, corresponding to late and early inter-view interactions. In the late-fusion mode, the two memories are separated, containing only view-specific contents. In the early-fusion mode, the two memories share the same addressing space, allowing cross-memory accessing. In both cases, the knowledge from the memories will be combined by a decoder to make predictions over the output space. The resulting dual memory neural computer is demonstrated on a comprehensive set of experiments, including a synthetic task of summing two sequences and the tasks of drug prescription and disease progression in healthcare. The results demonstrate competitive performance over both traditional algorithms and deep learning methods designed for multi-view problems.
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh
KDD2
2018 Variational Memory Encoder-Decoder
abstract
Introducing variability while maintaining coherence is a core task in learning to generate utterances in conversation. Standard neural encoder-decoder models and their extensions using conditional variational autoencoder often result in either trivial or digressive responses. To overcome this, we explore a novel approach that injects variability into neural encoder-decoder via the use of external memory as a mixture model, namely Variational Memory Encoder-Decoder (VMED). By associating each memory read with a mode in the latent mixture distribution at each timestep, our model can capture the variability observed in sequential data such as natural conversations. We empirically compare the proposed model against other recent approaches on various conversational datasets. The results show that VMED consistently achieves significant improvement over others in both metric-based and qualitative evaluations.
Hung Le 0002, Truyen Tran 0001, Thin Nguyen, Svetha Venkatesh
NeurIPS2
2018 Dual Control Memory Augmented Neural Networks for Treatment Recommendations
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh
PAKDD (3)2
2018 Energy-based anomaly detection for mixed data
Kien Do, Truyen Tran 0001, Svetha Venkatesh
Knowl. Inf. Syst.2
2018 Predicting Delivery Capability in Iterative Software Development
abstract
Iterative software development has become widely practiced in industry. Since modern software projects require fast, incremental delivery for every iteration of software development, it is essential to monitor the execution of an iteration, and foresee a capability to deliver quality products as the iteration progresses. This paper presents a novel, data-driven approach to providing automated support for project managers and other decision makers in predicting delivery capability for an ongoing iteration. Our approach leverages a history of project iterations and associated issues, and in particular, we extract characteristics of previous iterations and their issues in the form of features. In addition, our approach characterizes an iteration using a novel combination of techniques including feature aggregation statistics, automatic feature learning using the Bag-of-Words approach, and graph-based complexity measures. An extensive evaluation of the technique on five large open source projects demonstrates that our predictive models outperform three common baseline methods in Normalized Mean Absolute Error and are highly accurate in predicting the outcome of an ongoing iteration.
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose, John C. Grundy
IEEE Trans. Software Eng.3
2017 Column Networks for Collective Classification
abstract
Relational learning deals with data that are characterized by relational structures. An important task is collective classification, which is to jointly classify networked objects. While it holds a great promise to produce a better accuracy than non-collective classifiers, collective classification is computationally challenging and has not leveraged on the recent breakthroughs of deep learning. We present Column Network (CLN), a novel deep learning model for collective classification in multi-relational domains. CLN has many desirable theoretical properties: (i) it encodes multi-relations between any two instances; (ii) it is deep and compact, allowing complex functions to be approximated at the network level with a small set of free parameters; (iii) local and relational features are learned simultaneously; (iv) long-range, higher-order dependencies between instances are supported naturally; and (v) crucially, learning and inference are efficient with linear complexity in the size of the network and the number of relations. We evaluate CLN on multiple real-world applications: (a) delay prediction in software projects, (b) PubMed Diabetes publication classification and (c) film genre classification. In all of these applications, CLN demonstrates a higher accuracy than state-of-the-art rivals.
Trang Pham, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
AAAI2
2017 Hierarchical semi-Markov conditional random fields for deep recursive sequential data
Truyen Tran 0001, Dinh Q. Phung, Hung Hai Bui, Svetha Venkatesh
Artif. Intell.1
2017 Predicting the delay of issues with due dates in software projects
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose
Empir. Softw. Eng.3
2017 Predicting healthcare trajectories from medical records: A deep learning approach
Trang Pham, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
J. Biomed. Informatics2
2017 Preference Relation-based Markov Random Fields for Recommender Systems
Shaowu Liu, Gang Li 0009, Truyen Tran 0001, Yuan Jiang 0001
Mach. Learn.3
2017 Erratum to: Preference Relation-based Markov Random Fields for Recommender Systems
Shaowu Liu, Gang Li 0009, Truyen Tran 0001, Yuan Jiang 0001
Mach. Learn.3
2017 Deepr: A Convolutional Net for Medical Records
abstract
Feature engineering remains a major bottleneck when creating predictive systems from electronic medical records. At present, an important missing element is detecting predictive regular clinical motifs from irregular episodic records. We present Deepr (short for Deep record), a new end-to-end deep learning system that learns to extract features from medical records and predicts future risk automatically. Deepr transforms a record into a sequence of discrete elements separated by coded time gaps and hospital transfers. On top of the sequence is a convolutional neural net that detects and combines predictive local clinical motifs to stratify the risk. Deepr permits transparent inspection and visualization of its inner working. We validate Deepr on hospital data to predict unplanned readmission after discharge. Deepr achieves superior accuracy compared to traditional techniques, detects meaningful clinical motifs, and uncovers the underlying structure of the disease and intervention space.
Phuoc Nguyen, Truyen Tran 0001, Nilmini Wickramasinghe, Svetha Venkatesh
IEEE J. Biomed. Health Informatics2
2016 Outlier Detection on Mixed-Type Data: An Energy-Based Approach
Kien Do, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ADMA2
2016 Stabilizing Linear Prediction Models Using Autoencoder
Shivapratap Gopakumar, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ADMA2
2016 Faster training of very deep networks via p-norm gates
abstract
A major contributing factor to the recent advances in deep neural networks is structural units that let sensory information and gradients to propagate easily. Gating is one such structure that acts as a flow control. Gates are employed in many recent state-of-the-art recurrent models such as LSTM and GRU, and feedforward models such as Residual Nets and Highway Networks. This enables learning in very deep networks with hundred layers and helps achieve record-breaking results in vision (e.g., ImageNet with Residual Nets) and NLP (e.g., machine translation with GRU). However, there is limited work in analysing the role of gating in the learning process. In this paper, we propose a flexible p-norm gating scheme, which allows user-controllable flow and as a consequence, improve the learning speed. This scheme subsumes other existing gating schemes, including those in GRU, Highway Networks and Residual Nets as special cases. Experiments on large sequence and vector datasets demonstrate that the proposed gating scheme helps improve the learning speed significantly without extra overhead.
Trang Pham, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ICPR2
2016 DeepCare: A Deep Dynamic Memory Model for Predictive Medicine
Trang Pham, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (2)2
2016 DeepSoft: a vision for a deep model of software
abstract
Although software analytics has experienced rapid growth as a research area, it has not yet reached its full potential for wide industrial adoption. Most of the existing work in software analytics still relies heavily on costly manual feature engineering processes, and they mainly address the traditional classification problems, as opposed to predicting future events. We present a vision for DeepSoft, an end-to-end generic framework for modeling software and its development process to predict future risks and recommend interventions. DeepSoft, partly inspired by human memory, is built upon the powerful deep learning-based Long Short Term Memory architecture that is capable of learning long-term temporal dependencies that occur in software evolution. Such deep learned patterns of software can be used to address a range of challenging problems such as code and task recommendation and prediction. DeepSoft provides a new approach for research into modeling of source code, risk prediction and mitigation, developer modeling, and automatically generating code patches from bug reports.
Khanh Hoa Dam, Truyen Tran 0001, John C. Grundy, Aditya Ghose
SIGSOFT FSE2
2016 Graph-induced restricted Boltzmann machines for document modeling
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
Inf. Sci.2
2016 Collaborative filtering via sparse Markov random fields
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
Inf. Sci.1
2016 Modelling human preferences for ranking and collaborative filtering: a probabilistic ordered partition approach
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
Knowl. Inf. Syst.1
2015 Tensor-Variate Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machines (RBMs) are an important class of latent variable models for representing vector data. An under-explored area is multimode data, where each data point is a matrix or a tensor. Standard RBMs applying to such data would require vectorizing matrices and tensors, thus resulting in unnecessarily high dimensionality and at the same time, destroying the inherent higher-order interaction structures. This paper introduces Tensor-variate Restricted Boltzmann Machines (TvRBMs) which generalize RBMs to capture the multiplicative interaction between data modes and the latent variables. TvRBMs are highly compact in that the number of free parameters grows only linear with the number of modes. We demonstrate the capacity of TvRBMs on three real-world applications: handwritten digit classification, face recognition and EEG-based alcoholic diagnosis. The learnt features of the model are more discriminative than the rivals, resulting in better classification performance.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
AAAI2
2015 Preference Relation-based Markov Random Fields for Recommender Systems
Shaowu Liu, Gang Li 0009, Truyen Tran 0001, Yuan Jiang 0001
ACML3
2015 Predicting Delays in Software Projects Using Networked Classification (T)
abstract
Software projects have a high risk of cost and schedule overruns, which has been a source of concern for the software engineering community for a long time. One of the challenges in software project management is to make reliable prediction of delays in the context of constant and rapid changes inherent in software projects. This paper presents a novel approach to providing automated support for project managers and other decision makers in predicting whether a subset of software tasks (among the hundreds to thousands of ongoing tasks) in a software project have a risk of being delayed. Our approach makes use of not only features specific to individual software tasks (i.e. local data) -- as done in previous work -- but also their relationships (i.e. networked data). In addition, using collective classification, our approach can simultaneously predict the degree of delay for a group of related tasks. Our evaluation results show a significant improvement over traditional approaches which perform classification on each task independently: achieving 46% -- 97% precision (49% improved), 46% -- 97% recall (28% improved), 56% -- 75% F-measure (39% improved), and 78% -- 95% Area Under the ROC Curve (16% improved).
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose
ASE3
2015 Characterization and Prediction of Issue-Related Risks in Software Projects
abstract
Identifying risks relevant to a software project and planning measures to deal with them are critical to the success of the project. Current practices in risk assessment mostly rely on high-level, generic guidance or the subjective judgements of experts. In this paper, we propose a novel approach to risk assessment using historical data associated with a software project. Specifically, our approach identifies patterns of past events that caused project delays, and uses this knowledge to identify risks in the current state of the project. A set of risk factors characterizing “risky” software tasks (in the form of issues) were extracted from five open source projects: Apache, Duraspace, JBoss, Moodle, and Spring. In addition, we performed feature selection using a sparse logistic regression model to select risk factors with good discriminative power. Based on these risk factors, we built predictive models to predict if an issue will cause a project delay. Our predictive models are able to predict both the risk impact (i.e. the extend of the delay) and the likelihood of a risk occurring. The evaluation results demonstrate the effectiveness of our predictive models, achieving on average 48%-81% precision, 23%-90% recall, 29%-71% F-measure, and 70%-92% Area Under the ROC Curve. Our predictive models also have low error rates: 0.39-0.75 for Macro-averaged Mean Cost-Error and 0.7-1.2 for Macro-averaged Mean Absolute Error.
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose
MSR3
2015 Stabilizing Sparse Cox Model Using Statistic and Semantic Structures in Electronic Medical Records
Shivapratap Gopakumar, Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (2)3
2015 Learning vector representation of medical objects via EMR-driven nonnegative restricted Boltzmann machines (eNRBM)
Truyen Tran 0001, Tu Dinh Nguyen, Dinh Q. Phung, Svetha Venkatesh
J. Biomed. Informatics1
2015 Stabilized sparse ordinal regression for medical risk stratification
Truyen Tran 0001, Dinh Q. Phung, Wei Luo 0001, Svetha Venkatesh
Knowl. Inf. Syst.1
2015 Stabilizing High-Dimensional Prediction Models Using Feature Graphs
abstract
We investigate feature stability in the context of clinical prognosis derived from high-dimensional electronic medical records. To reduce variance in the selected features that are predictive, we introduce Laplacian-based regularization into a regression model. The Laplacian is derived on a feature graph that captures both the temporal and hierarchic relations between hospital events, diseases, and interventions. Using a cohort of patients with heart failure, we demonstrate better feature stability and goodness-of-fit through feature graph stabilization.
Shivapratap Gopakumar, Truyen Tran 0001, Tu Dinh Nguyen, Dinh Q. Phung, Svetha Venkatesh
IEEE J. Biomed. Health Informatics2
2014 Ordinal Random Fields for Recommender Systems
Shaowu Liu, Truyen Tran 0001, Gang Li 0009
ACML2
2014 iPoll: Automatic Polling Using Online Search
Thin Nguyen, Dinh Q. Phung, Wei Luo 0001, Truyen Tran 0001, Svetha Venkatesh
WISE (1)4
2014 A framework for feature extraction from hospital medical data with applications in risk prediction
abstract
BACKGROUND: Feature engineering is a time consuming component of predictive modeling. We propose a versatile platform to automatically extract features for risk prediction, based on a pre-defined and extensible entity schema. The extraction is independent of disease type or risk prediction task. We contrast auto-extracted features to baselines generated from the Elixhauser comorbidities. RESULTS: Hospital medical records was transformed to event sequences, to which filters were applied to extract feature sets capturing diversity in temporal scales and data types. The features were evaluated on a readmission prediction task, comparing with baseline feature sets generated from the Elixhauser comorbidities. The prediction model was through logistic regression with elastic net regularization. Predictions horizons of 1, 2, 3, 6, 12 months were considered for four diverse diseases: diabetes, COPD, mental disorders and pneumonia, with derivation and validation cohorts defined on non-overlapping data-collection periods. For unplanned readmissions, auto-extracted feature set using socio-demographic information and medical records, outperformed baselines derived from the socio-demographic information and Elixhauser comorbidities, over 20 settings (5 prediction horizons over 4 diseases). In particular over 30-day prediction, the AUCs are: COPD-baseline: 0.60 (95% CI: 0.57, 0.63), auto-extracted: 0.67 (0.64, 0.70); diabetes-baseline: 0.60 (0.58, 0.63), auto-extracted: 0.67 (0.64, 0.69); mental disorders-baseline: 0.57 (0.54, 0.60), auto-extracted: 0.69 (0.64,0.70); pneumonia-baseline: 0.61 (0.59, 0.63), auto-extracted: 0.70 (0.67, 0.72). CONCLUSIONS: The advantages of auto-extracted standard features from complex medical records, in a disease and task agnostic manner were demonstrated. Auto-extracted features have good predictive power over multiple time horizons. Such feature sets have potential to form the foundation of complex automated analytic tasks.
Truyen Tran 0001, Wei Luo 0001, Dinh Q. Phung, Sunil Gupta 0001, Santu Rana, Richard Kennedy, Ann Larkins, Svetha Venkatesh
BMC Bioinform.1
2013 Learning Parts-based Representations with Nonnegative Restricted Boltzmann Machine
abstract
The success of any machine learning system depends critically on effective representations of data. In many cases, especially those in vision, it is desirable that a representation scheme uncovers the parts-based, additive nature of the data. Of current representation learning schemes, restricted Boltzmann machines (RBMs) have proved to be highly effective in unsupervised settings. However, when it comes to parts-based discovery, RBMs do not usually produce satisfactory results. We enhance such capacity of RBMs by introducing nonnegativity into the model weights, resulting in a variant called \emphnonnegative restricted Boltzmann machine (NRBM). The NRBM produces not only controllable decomposition of data into interpretable parts but also offers a way to estimate the intrinsic nonlinear dimensionality of data. We demonstrate the capacity of our model on well-known datasets of handwritten digits, faces and documents. The decomposition quality on images is comparable with or better than what produced by the nonnegative matrix factorisation (NMF), and the thematic features uncovered from text are qualitatively interpretable in a similar manner to that of the latent Dirichlet allocation (LDA). However, the learnt features, when used for classification, are more discriminative than those discovered by both NMF and LDA and comparable with those by RBM.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ACML2
2013 Learning sparse latent representation and distance metric for image retrieval
abstract
The performance of image retrieval depends critically on the semantic representation and the distance function used to estimate the similarity of two images. A good representation should integrate multiple visual and textual (e.g., tag) features and offer a step closer to the true semantics of interest (e.g., concepts). As the distance function operates on the representation, they are interdependent, and thus should be addressed at the same time. We propose a probabilistic solution to learn both the representation from multiple feature types and modalities and the distance metric from data. The learning is regularised so that the learned representation and information-theoretic metric will (i) preserve the regularities of the visual/textual spaces, (ii) enhance structured sparsity, (iii) encourage small intra-concept distances, and (iv) keep inter-concept images separated. We demonstrate the capacity of our method on the NUS-WIDE data. For the well-studied 13 animal subset, our method outperforms state-of-the-art rivals. On the subset of single-concept images, we gain 79:5% improvement over the standard nearest neighbours approach on the MAP score, and 45.7% on the NDCG.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ICME2
2013 Thurstonian Boltzmann Machines: Learning from Multiple Inequalities
abstract
We introduce Thurstonian Boltzmann Machines (TBM), a unified architecture that can naturally incorporate a wide range of data inputs at the same time. Our motivation rests in the Thurstonian view that many discrete data types can be considered as being generated from a subset of underlying latent continuous variables, and in the observation that each realisation of a discrete type imposes certain inequalities on those variables. Thus learning and inference in TBM reduce to making sense of a set of inequalities. Our proposed TBM naturally supports the following types: Gaussian, intervals, censored, binary, categorical, muticategorical, ordinal, (in)-complete rank with and without ties. We demonstrate the versatility and capacity of the proposed model on three applications of very different natures; namely handwritten digit recognition, collaborative filtering and complex social survey analysis.
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ICML (2)1
2013 An integrated framework for suicide risk prediction
abstract
Suicide is a major concern in society. Despite of great attention paid by the community with very substantive medico-legal implications, there has been no satisfying method that can reliably predict the future attempted or completed suicide. We present an integrated machine learning framework to tackle this challenge. Our proposed framework consists of a novel feature extraction scheme, an embedded feature selection process, a set of risk classifiers and finally, a risk calibration procedure. For temporal feature extraction, we cast the patient's clinical history into a temporal image to which a bank of one-side filters are applied. The responses are then partly transformed into mid-level features and then selected in l1-norm framework under the extreme value theory. A set of probabilistic ordinal risk classifiers are then applied to compute the risk probabilities and further re-rank the features. Finally, the predicted risks are calibrated. Together with our Australian partner, we perform comprehensive study on data collected for the mental health cohort, and the experiments validate that our proposed framework outperforms risk assessment instruments by medical practitioners.
Truyen Tran 0001, Dinh Q. Phung, Wei Luo 0001, Richard Harvey 0002, Michael Berk, Svetha Venkatesh
KDD1
2013 Latent Patient Profile Modelling and Applications with Mixed-Variate Restricted Boltzmann Machine
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (1)2
2012 A Sequential Decision Approach to Ordinal Preferences in Recommender Systems
abstract
We propose a novel sequential decision approach to modeling ordinal ratings in collaborative filtering problems. The rating process is assumed to start from the lowest level, evaluates against the latent utility at the corresponding level and moves up until a suitable ordinal level is found. Crucial to this generative process is the underlying utility random variables that govern the generation of ratings and their modelling choices. To this end, we make a novel use of the generalised extreme value distributions, which is found to be particularly suitable for our modeling tasks and at the same time, facilitate our inference and learning procedure. The proposed approach is flexible to incorporate features from both the user and the item. We evaluate the proposed framework on three well-known datasets: MovieLens, Dating Agency and Netflix. In all cases, it is demonstrated that the proposed work is competitive against state-of-the-art collaborative filtering methods.
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
AAAI1
2012 Embedded Restricted Boltzmann Machines for fusion of mixed data types and applications in social measurements analysis
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
FUSION1
2012 Learning Boltzmann Distance Metric for Face Recognition
abstract
We introduce a new method for face recognition using a versatile probabilistic model known as Restricted Boltzmann Machine (RBM). In particular, we propose to regularise the standard data likelihood learning with an information-theoretic distance metric defined on intra-personal images. This results in an effective face representation which captures the regularities in the face space and minimises the intra-personal variations. In addition, our method allows easy incorporation of multiple feature sets with controllable level of sparsity. Our experiments on a high variation dataset show that the proposed method is competitive against other metric learning rivals. We also investigated the RBM method under a variety of settings, including fusing facial parts and utilising localised feature detectors under varying resolutions. In particular, the accuracy is boosted from 71.8% with the standard whole-face pixels to 99.2% with combination of facial parts, localised feature extractors and appropriate resolutions.
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ICME1
2011 Probabilistic Models over Ordered Partitions with Applications in Document Ranking and Collaborative Filtering
abstract
Ranking is an important task for handling a large amount of content. Ideally, training data for supervised ranking would include a complete rank of documents (or other objects such as images or videos) for a particular query. However, this is only possible for small sets of documents. In practice, one often resorts to document rating, in that a subset of documents is assigned with a small number indicating the degree of relevance. This poses a general problem of modelling and learning rank data with ties. In this paper, we propose a probabilistic generative model, that models the process as permutations over partitions. This results in super-exponential combinatorial state space with unknown numbers of partitions and unknown ordering among them. We approach the problem from the discrete choice theory, where subsets are chosen in a stagewise manner, reducing the state space per each stage significantly. Further, we show that with suitable parameterisation, we can still learn the models in linear time. We evaluate the proposed models on two application areas: (i) document ranking with the data from the recently held Yahoo! challenge, and (ii) collaborative filtering with movie data. The results demonstrate that the models are competitive against well-known rivals.
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
SDM1
2010 Nonnegative shared subspace learning and its application to social media retrieval
abstract
Although tagging has become increasingly popular in online image and video sharing systems, tags are known to be noisy, ambiguous, incomplete and subjective. These factors can seriously affect the precision of a social tag-based web retrieval system. Therefore improving the precision performance of these social tag-based web retrieval systems has become an increasingly important research topic. To this end, we propose a shared subspace learning framework to leverage a secondary source to improve retrieval performance from a primary dataset. This is achieved by learning a shared subspace between the two sources under a joint Nonnegative Matrix Factorization in which the level of subspace sharing can be explicitly controlled. We derive an efficient algorithm for learning the factorization, analyze its complexity, and provide proof of convergence. We validate the framework on image and video retrieval tasks in which tags from the LabelMe dataset are used to improve image retrieval performance from a Flickr dataset and video retrieval performance from a YouTube dataset. This has implications for how to exploit and transfer knowledge from readily available auxiliary tagging resources to improve another social web retrieval system. Our shared subspace learning framework is applicable to a range of problems where one needs to exploit the strengths existing among multiple and heterogeneous datasets.
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh
KDD4
2010 Classification and Pattern Discovery of Mood in Weblogs
Thin Nguyen, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh
PAKDD (2)4
2009 Ordinal Boltzmann Machines for Collaborative Filtering
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
UAI1
2008 Hierarchical Semi-Markov Conditional Random Fields for Recursive Sequential Data
abstract
Inspired by the hierarchical hidden Markov models (HHMM), we present the hierarchical semi-Markov conditional random field (HSCRF), a generalisation of embedded undirected Markov chains to model complex hierarchical, nested Markov processes. It is parameterised in a discriminative framework and has polynomial time algorithms for learning and inference. Importantly, we develop efficient algorithms for learning and constrained inference in a partially-supervised setting, which is important issue in practice where labels can only be obtained sparsely. We demonstrate the HSCRF in two applications: (i) recognising human activities of daily living (ADLs) from indoor surveillance cameras, and (ii) noun-phrase chunking. We show that the HSCRF is capable of learning rich hierarchical models with reasonable accuracy in both fully and partially observed data cases.
Truyen Tran 0001, Dinh Q. Phung, Hung Hai Bui, Svetha Venkatesh
NIPS1
2008 Learning Discriminative Sequence Models from Partially Labelled Data for Activity Recognition
Truyen Tran 0001, Hung Hai Bui, Dinh Q. Phung, Svetha Venkatesh
PRICAI1
2008 Constrained Sequence Classification for Lexical Disambiguation
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PRICAI1
2006 AdaBoost.MRF: Boosted Markov Random Forests and Application to Multilevel Activity Recognition
abstract
Activity recognition is an important issue in building intelligent monitoring systems. We address the recognition of multilevel activities in this paper via a conditional Markov random field (MRF), known as the dynamic conditional random field (DCRF). Parameter estimation in general MRFs using maximum likelihood is known to be computationally challenging (except for extreme cases), and thus we propose an efficient boosting-based algorithm AdaBoost.MRF for this task. Distinct from most existing work, our algorithm can handle hidden variables (missing labels) and is particularly attractive for smarthouse domains where reliable labels are often sparsely observed. Furthermore, our method works exclusively on trees and thus is guaranteed to converge. We apply the AdaBoost.MRF algorithmto a home video surveillance application and demonstrate its efficacy.
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh, Hung Hai Bui
CVPR (2)1