Tenghao Huang

dblp:79/11059 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 56% Efficient and distributed learning · 12% Planning, search and constraint satisfaction · 8%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 100%

Topics — the 9 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
web agents
1.012026
GTA: Generating Long-horizon Tasks for Web Agents at Scale · ACL (1) 2026
Natural language and speech › Language models and text generation › LLM agents
agent memory
0.912025
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory · ACL (1) 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists · KDD (2) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists · KDD (2) 2025
Natural language and speech › Language models and text generation › text generation
scientific hypothesis generation
0.912025
FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists · KDD (2) 2025
Natural language and speech › Language models and text generation › text generation
story generation
0.812024
Are Large Language Models Capable of Generating Human-Level Narratives? · EMNLP 2024
Requirements engineering and software design › model-driven engineering
collaborative modeling
0.712023
Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models · ICML 2023
Natural language and speech › Language models and text generation
in-context learning
0.612022
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.612022
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.6in-context learning · 1.7model merging · 1.3task generation · 1.0language model agent · 1.0reflection · 0.9large language model · 0.9red teaming · 0.8large language model prompting · 0.8human evaluation · 0.8version control · 0.7
YearPublicationVenuePosition
2026 GTA: Generating Long-horizon Tasks for Web Agents at Scale
abstract
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen 0001, Jonathan May, Chien-Sheng Wu
ACL (1)1
2025 R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory
abstract
Tenghao Huang, Kinjal Basu, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, Muhao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tenghao Huang, Kinjal Basu 0002, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, Muhao Chen 0001
ACL (1)1
2025 NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational Interviews
abstract
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho, Tenghao Huang, Weiyan Shi, Jonathan May. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Cho, Tenghao Huang, Weiyan Shi 0001, Jonathan May
ACL (1)5
2025 FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists
abstract
Flavor development in the food industry is increasingly challenged by the need for rapid innovation and precise flavor profile creation. Traditional flavor research methods typically rely on iterative, subjective testing, which lacks the efficiency and scalability required for modern demands. This paper presents three contributions to address these challenges. Firstly, we define a new problem domain for scientific agents in flavor science, conceptualized as the generation of hypotheses for flavor profile sourcing and understanding. By leveraging their capacity to identify relevant evidence and reason within large context spaces, language model-backed agents can perform the labor-intensive tasks of flavor sourcing and understanding with enhanced efficiency and precision. To facilitate research in this area, we introduce the FoodPuzzle dataset, a challenging benchmark consisting of 978 food items and 1,766 flavor molecule profiles. We propose a novel Scientific Agent approach, integrating in-context learning and retrieval augmented techniques to generate grounded hypotheses in the domain of food science. Experimental results indicate that our model significantly surpasses traditional methods in flavor profile prediction tasks, demonstrating its potential to transform flavor development practices.
Tenghao Huang, John Sweeney, Jiatong Shi, Emily Steliotes, Matthew Lange, Jonathan May, Muhao Chen 0001
KDD (2)1
2024 Are Large Language Models Capable of Generating Human-Level Narratives?
abstract
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nanyun Peng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen 0001, Jonathan May, Nanyun Peng 0001
EMNLP2
2024 Red Teaming Language Models for Processing Contradictory Dialogues
abstract
Most language models currently available are prone to self-contradiction during dialogues.To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in a conversation.This task is inspired by research on context faithfulness and dialogue comprehension, which have demonstrated that the detection and understanding of contradictions often necessitate detailed explanations.We develop a dataset comprising contradictory dialogues, in which one side of the conversation contradicts itself.Each dialogue is accompanied by an explanatory label that highlights the location and details of the contradiction.With this dataset, we present a Red Teaming framework for contradictory dialogue processing.The framework detects and attempts to explain the dialogue, then modifies the existing contradictory content using the explanation.Our experiments demonstrate that the framework improves the ability to detect contradictory dialogues and provides valid explanations.Additionally, it showcases distinct capabilities for modifying such dialogues.Our study highlights the importance of the logical inconsistency problem in conversational AI 1 Prompts Instructions Explanation
Xiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao Chen 0001
EMNLP3
2023 Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models
abstract
Currently, most machine learning models are trained by centralized teams and are rarely updated. In contrast, open-source software development involves the iterative development of a shared artifact through distributed collaboration using a version control system. In the interest of enabling collaborative and continual improvement of machine learning models (Raffel, 2023), we introduce Git-Theta, a version control system for machine learning models. Git-Theta is an extension to Git, the most widely used version control software, that allows fine-grained tracking of changes to model parameters alongside code and other artifacts. Unlike existing version control systems that treat a model checkpoint as a blob of data, Git-Theta leverages the structure of checkpoints to support communication-efficient updates, automatic model merges, and meaningful reporting about the difference between two versions of a model. In addition, Git-Theta includes a plug-in system that enables users to easily add support for new functionality. In this paper, we introduce Git-Theta’s design and features and include an example use-case of Git-Theta where a pre-trained model is continually adapted and modified. We publicly release Git-Theta in hopes of kickstarting a new era of collaborative model development. https://github.com/r-three/git-theta/
Nikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas, Monty Evans, Vishal Baskaran, Tenghao Huang, Haokun Liu, Colin Raffel
ICML7
2022 Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
abstract
Few-shot in-context learning (ICL) enables pre-trained language models to perform a previously-unseen task without any gradient-based training by feeding a small number of training examples as part of the input. ICL incurs substantial computational, memory, and storage costs because it involves processing all of the training examples every time a prediction is made. Parameter-efficient fine-tuning (PEFT) (e.g. adapter modules, prompt tuning, sparse update methods, etc.) offers an alternative paradigm where a small set of parameters are trained to enable a model to perform the new task. In this paper, we rigorously compare few-shot ICL and PEFT and demonstrate that the latter offers better accuracy as well as dramatically lower computational costs. Along the way, we introduce a new PEFT method called (IA)^3 that scales activations by learned vectors, attaining stronger performance while only introducing a relatively tiny amount of new parameters. We also propose a simple recipe based on the T0 model called T-Few that can be applied to new tasks without task-specific tuning or modifications. We validate the effectiveness of T-Few on completely unseen tasks by applying it to the RAFT benchmark, attaining super-human performance for the first time and outperforming the state-of-the-art by 6% absolute. All of the code used in our experiments will be publicly available.
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, Colin Raffel
NeurIPS5