Hung Le 0003

dblp:45/466-3 · DBLP profile ↗
← Back
17ranked-venue papers
14as first author
10since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 14 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
abstract
Jierui Li, Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jierui Li, Hung Le 0003, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese, Doyen Sahoo
NAACL (Long Papers)2
2024 CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
abstract
Large Language Models (LLMs) have already become quite proficient at solving simpler programming tasks like those in HumanEval or MBPP benchmarks. However, solving more complex and competitive programming tasks is still quite challenging for these models - possibly due to their tendency to generate solutions as monolithic code blocks instead of decomposing them into logical sub-tasks and sub-modules. On the other hand, experienced programmers instinctively write modularized code with abstraction for solving complex tasks, often reusing previously developed modules. To address this gap, we propose CodeChain, a novel framework for inference that elicits modularized code generation through a chain of self-revisions, each being guided by some representative sub-modules generated in previous iterations. Concretely, CodeChain first instructs the LLM to generate modularized codes through chain-of-thought prompting. Then it applies a chain of self-revisions by iterating the two steps: 1) extracting and clustering the generated sub-modules and selecting the cluster representatives as the more generic and re-usable implementations, and 2) augmenting the original chain-of-thought prompt with these selected module-implementations and instructing the LLM to re-generate new modularized solutions. We find that by naturally encouraging the LLM to reuse the previously developed and verified sub-modules, CodeChain can significantly boost both modularity as well as correctness of the generated solutions, achieving relative pass@1 improvements of 35\% on APPS and 76\% on CodeContests. It is shown to be effective on both OpenAI LLMs as well as open-sourced LLMs like WizardCoder. We also conduct comprehensive ablation studies with different methods of prompting, number of clusters, model sizes, program qualities, etc., to provide useful insights that underpin CodeChain's success.
Hung Le 0003, Hailin Chen, Amrita Saha, Akash Gokul, Doyen Sahoo, Shafiq R. Joty
ICLR1
2024 INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
abstract
Large language models (LLMs) for code are typically trained to align with natural language instructions to closely follow their intentions and requirements. However, in many practical scenarios, it becomes increasingly challenging for these models to navigate the intricate boundary between helpfulness and safety, especially against highly complex yet potentially malicious instructions. In this work, we introduce INDICT: a new framework that empowers LLMs with Internal Dialogues of Critiques for both safety and helpfulness guidance. The internal dialogue is a dual cooperative system between a safety-driven critic and a helpfulness-driven critic. Each critic provides analysis against the given task and corresponding generated response, equipped with external knowledge queried through relevant code snippets and tools like web search and code interpreter. We engage the dual critic system in both code generation stage as well as code execution stage, providing preemptive and post-hoc guidance respectively to LLMs. We evaluated INDICT on 8 diverse tasks across 8 programming languages from 5 benchmarks, using LLMs from 7B to 70B parameters. We observed that our approach can provide an advanced level of critiques of both safety and helpfulness analysis, significantly improving the quality of output codes (+10% absolute improvements in all models).
Hung Le 0003, Doyen Sahoo, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese
NeurIPS1
2023 CodeT5+: Open Code Large Language Models for Code Understanding and Generation
abstract
Large language models (LLMs) pretrained on vast source code have achieved prominent progress in code intelligence. However, existing code LLMs have two main limitations. First, they often adopt a specific architecture (encoder-only or decoder-only) or rely on a unified encoder-decoder network for different downstream tasks, lacking the flexibility to operate in the optimal architecture for a specific task. Secondly, they often employ a limited set of pretraining objectives which might not be relevant to some tasks and hence result in substantial performance degrade. To address these limitations, we propose “CodeT5+”, a family of encoder-decoder LLMs for code in which component modules can be flexibly combined to suit a wide range of code tasks. Such flexibility is enabled by our proposed mixture of pretraining objectives, which cover span denoising, contrastive learning, text-code matching, and causal LM pretraining tasks, on both unimodal and bimodal multilingual code corpora. Furthermore, we propose to initialize CodeT5+ with frozen off-the-shelf LLMs without training from scratch to efficiently scale up our models, and explore instruction-tuning to align with natural language instructions. We extensively evaluate CodeT5+ on over 20 code-related benchmarks in different settings, including zero-shot, finetuning, and instruction-tuning. We observe state-of-the-art (SoTA) performance on various code-related tasks, and our instruction-tuned CodeT5+ 16B achieves new SoTA results of 35.0% pass@1 and 54.5% pass@10 on the HumanEval code generation task against other open code LLMs, even surpassing the OpenAI code-cushman-001 model.
Yue Wang 0034, Hung Le 0003, Akhilesh Gotmare, Nghi D. Q. Bui, Junnan Li 0001, Steven C. H. Hoi
EMNLP2
2023 C3: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues
abstract
Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context.Most existing approaches employ deep learning models and have achieved remarkable performance, given the relatively small datasets available.However, the results are partially accomplished by exploiting biases in the datasets rather than developing multimodal reasoning, resulting in limited generalization.In this paper, we propose a novel approach of Compositional Counterfactual Contrastive Learning (C 3 ) to develop contrastive training between factual and counterfactual samples in videogrounded dialogues.Specifically, we design factual/counterfactual samples based on the temporal steps in videos and tokens in dialogues and propose contrastive loss functions that exploit object-level or action-level variance.Different from prior approaches, we focus on contrastive hidden state representations among compositional output tokens to optimize the representation space in a generation setting.We achieved promising performance gains on the Audio-Visual Scene-Aware Dialogues (AVSD) benchmark and showed the benefits of our approach in grounding video and dialogue context.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
SIGDIAL1
2022 VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems
abstract
Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images.However, very limited work on NMN has been studied in the video-grounded dialogue tasks.These tasks extend the complexity of traditional visual tasks with the additional visual temporal variance and language cross-turn dependencies.Motivated by recent NMN approaches on image-grounded tasks, we introduce Videogrounded Neural Module Network (VGNMN) to model the information retrieval process in video-grounded language tasks as a pipeline of neural modules.VGNMN first decomposes all language components in dialogues to explicitly resolve any entity references and detect corresponding action-based inputs from the question.The detected entities and actions are used as parameters to instantiate neural module networks and extract visual cues from the video.Our experiments show that VGNMN can achieve promising performance on a challenging video-grounded dialogue benchmark as well as a video QA benchmark.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
NAACL-HLT1
2022 Multimodal Dialogue State Tracking
abstract
Designed for tracking user goals in dialogues, a dialogue state tracker is an essential component in a dialogue system.However, the research of dialogue state tracking has largely been limited to unimodality, in which slots and slot values are limited by knowledge domains (e.g.restaurant domain with slots of restaurant name and price range) and are defined by specific database schema.In this paper, we propose to extend the definition of dialogue state tracking to multimodality.Specifically, we introduce a novel dialogue state tracking task to track the information of visual objects that are mentioned in video-grounded dialogues.Each new dialogue utterance may introduce a new video segment, new visual objects, or new object attributes and a state tracker is required to update these information slots accordingly.We created a new synthetic benchmark and designed a novel baseline, Video-Dialogue Transformer Network (VDTN), for this task.VDTN combines both object-level features and segment-level features and learns contextual dependencies between videos and dialogues to generate multimodal dialogue states.We optimized VDTN for a state generation task as well as a self-supervised video understanding task which recovers video segment or object representations.Finally, we trained VDTN to use the decoded states in a response prediction task.Together with comprehensive ablation and qualitative analysis, we discovered interesting insights towards building more capable multimodal dialogue systems.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
NAACL-HLT1
2022 CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
abstract
Program synthesis or code generation aims to generate a program that satisfies a problem specification. Recent approaches using large-scale pretrained language models (LMs) have shown promising results, yet they have some critical limitations. In particular, they often follow a standard supervised fine-tuning procedure to train a code generation model from natural language problem descriptions and ground-truth programs only. Such paradigm largely ignores some important but potentially useful signals in the problem specification such as unit tests, which thus results in poor performance when solving complex unseen coding tasks. We propose “CodeRL” to address the limitations, a new framework for program synthesis tasks through pretrained LMs and deep reinforcement learning (RL). Specifically, during training, we treat the code-generating LM as an actor network, and introduce a critic network that is trained to predict the functional correctness of generated programs and provide dense feedback signals to the actor. During inference, we introduce a new generation procedure with a critical sampling strategy that allows a model to automatically regenerate programs based on feedback from example unit tests and critic scores. For the model backbones, we extended the encoder-decoder architecture of CodeT5 with enhanced learning objectives, larger model sizes, and better pretraining data. Our method not only achieves new SOTA results on the challenging APPS benchmark, but also shows strong zero-shot transfer capability with new SOTA results on the simpler MBPP benchmark.
Hung Le 0003, Yue Wang 0034, Akhilesh Gotmare, Silvio Savarese, Steven C. H. Hoi
NeurIPS1
2021 DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue
abstract
Hung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hung Le 0003, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur
ACL/IJCNLP (1)1
2021 Learning Reasoning Paths over Semantic Graphs for Video-grounded Dialogues
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
ICLR1
2020 Video-Grounded Dialogues with Pretrained Generation Language Models
abstract
Pre-trained language models have shown remarkable success in improving various downstream NLP tasks due to their ability to capture dependencies in textual data and generate natural responses.In this paper, we leverage the power of pre-trained language models for improving video-grounded dialogue, which is very challenging and involves complex features of different dynamics: (1) Video features which can extend across both spatial and temporal dimensions; and (2) Dialogue features which involve semantic dependencies over multiple dialogue turns.We propose a framework by extending GPT-2 models to tackle these challenges by formulating videogrounded dialogue tasks as a sequence-tosequence task, combining both visual and textual representation into a structured sequence, and fine-tuning a large pre-trained GPT-2 network.Our framework allows fine-tuning language models to capture dependencies across multiple modalities over different levels of information: spatio-temporal level in video and token-sentence level in dialogue context.We achieve promising improvement on the Audio-Visual Scene-Aware Dialogues (AVSD) benchmark from DSTC7, which supports a potential direction in this line of research.* This work was mostly done when Hung Le was an intern at Salesforce Research Asia, Singapore.
Hung Le 0003, Steven C. H. Hoi
ACL1
2020 BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues
abstract
Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances which query different segments and/or different objects in videos over multiple dialogue turns.However, existing approaches to video-grounded dialogues often focus on superficial temporal-level visual cues, but neglect more fine-grained spatial signals from videos.To address this drawback, we propose Bi-directional Spatio-Temporal Learning (BiST), a vision-language neural framework for high-resolution queries in videos based on textual cues.Specifically, our approach not only exploits both spatial and temporal-level information, but also learns dynamic information diffusion between the two feature spaces through spatial-to-temporal and temporal-tospatial reasoning.The bidirectional strategy aims to tackle the evolving semantics of user queries in the dialogue setting.The retrieved visual cues are used as contextual information to construct relevant responses to the users.Our empirical results and comprehensive qualitative analysis show that BiST achieves competitive performance and generates reasonable responses on a large-scale AVSD benchmark.We also adapt our BiST models to the Video QA setting, and substantially outperform prior approaches on the TGIF-QA benchmark.* This work was mostly
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
EMNLP (1)1
2020 UniConv: A Unified Conversational Neural Architecture for Multi-domain Task-oriented Dialogues
abstract
Building an end-to-end conversational agent for multi-domain task-oriented dialogues has been an open challenge for two main reasons.First, tracking dialogue states of multiple domains is non-trivial as the dialogue agent must obtain complete states from all relevant domains, some of which might have shared slots among domains as well as unique slots specifically for one domain only.Second, the dialogue agent must also process various types of information across domains, including dialogue context, dialogue states, and database, to generate natural responses to users.Unlike the existing approaches that are often designed to train each module separately, we propose "UniConv" -a novel unified neural architecture for end-to-end conversational systems in multi-domain task-oriented dialogues, which is designed to jointly train (i) a Bi-level State Tracker which tracks dialogue states by learning signals at both slot and domain level independently, and (ii) a Joint Dialogue Act and Response Generator which incorporates information from various input components and models dialogue acts and target responses simultaneously.We conduct comprehensive experiments in dialogue state tracking, contextto-text, and end-to-end settings on the Multi-WOZ2.1 benchmark, achieving superior performance over competitive baselines.
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
EMNLP (1)1
2020 Non-Autoregressive Dialog State Tracking
Hung Le 0003, Richard Socher, Steven C. H. Hoi
ICLR1
2020 Hierarchical multimodal attention for end-to-end audio-visual scene-aware dialogue response generation
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
Comput. Speech Lang.1
2019 Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems
abstract
Developing Video-Grounded Dialogue Systems (VGDS), where a dialogue is conducted based on visual and audio aspects of a given video, is significantly more challenging than traditional image or text-grounded dialogue systems because (1) feature space of videos span across multiple picture frames, making it difficult to obtain semantic information; and(2) a dialogue agent must perceive and process information from different modalities (audio, video, caption, etc.) to obtain a comprehensive understanding.Most existing work is based on RNNs and sequence-to-sequence architectures, which are not very effective for capturing complex long-term dependencies (like in videos).To overcome this, we propose Multimodal Transformer Networks (MTN) to encode videos and incorporate information from different modalities.We also propose queryaware attention through an auto-encoder to extract query-aware features from non-text modalities.We develop a training procedure to simulate token-level decoding to improve the quality of generated responses during inference.We get state of the art performance on Dialogue System Technology Challenge 7 (DSTC7).Our model also generalizes to another multimodal visual-grounded dialogue task, and obtains promising performance.
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
ACL (1)1
2019 FoodAI: Food Image Recognition via Deep Learning for Smart Food Logging
abstract
An important aspect of health monitoring is effective logging of food consumption. This can help management of diet-related diseases like obesity, diabetes, and even cardiovascular diseases. Moreover, food logging can help fitness enthusiasts, and people who wanting to achieve a target weight. However, food-logging is cumbersome, and requires not only taking additional effort to note down the food item consumed regularly, but also sufficient knowledge of the food item consumed (which is difficult due to the availability of a wide variety of cuisines). With increasing reliance on smart devices, we exploit the convenience offered through the use of smart phones and propose a smart-food logging system: FoodAI, which offers state-of-the-art deep-learning based image recognition capabilities. FoodAI has been developed in Singapore and is particularly focused on food items commonly consumed in Singapore. FoodAI models were trained on a corpus of 400,000 food images from 756 different classes.
Doyen Sahoo, Hao Wang 0094, Shu Ke, Xiongwei Wu, Hung Le 0003, Palakorn Achananuparp, Ee-Peng Lim, Steven C. H. Hoi
KDD5