Hongyin Luo

dblp:147/4317 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2025 Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
abstract
Knowledge Graphs (KGs) can serve as reliable knowledge sources for question answering (QA) due to their structured representation of knowledge.Existing research on the utilization of KG for large language models (LLMs) prevalently relies on subgraph retriever or iterative prompting, overlooking the potential synergy of LLMs' step-wise reasoning capabilities and KGs' structural nature.In this paper, we present DoG (Decoding on Graphs), a novel framework that facilitates a deep synergy between LLMs and KGs.We first define a concept, well-formed chain, which consists of a sequence of interrelated fact triplets on the KGs, starting from question entities and leading to answers.We argue that this concept can serve as a principle for making faithful and sound reasoning for KGQA.To enable LLMs to generate well-formed chains, we propose graph-aware constrained decoding, in which a constraint derived from the topology of the KG regulates the decoding process of the LLMs.This constrained decoding method ensures the generation of well-formed chains while making full use of the step-wise reasoning capabilities of LLMs.Based on the above, DOG, a trainingfree approach, is able to provide faithful and sound reasoning trajectories grounded on the KGs.Experiments across various KGQA tasks with different background KGs demonstrate that DOG achieves superior and robust performance.DOG also shows general applicability with various open-source LLMs 1 .* Equal contribution. 1 The code is available here.
Kun Li 0003, Tianhua Zhang, Xixin Wu, Hongyin Luo, James R. Glass, Helen M. Meng
ACL (1)4
2025 RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement Learning
abstract
Robust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems.However, current LLM-based evaluation frameworks predominantly rely on directly prompting resource-intensive models with complex multi-stage prompts, underutilizing models' reasoning capabilities and introducing significant computational cost.In this paper, we present RAG-Zeval (RAG-Zero Evaluator), a novel end-to-end framework that formulates faithfulness and correctness evaluation of RAG systems as a rule-guided reasoning task.Our approach trains evaluators with reinforcement learning, facilitating compact models to generate comprehensive and sound assessments with detailed explanation in onepass.We introduce a ranking-based outcome reward mechanism, using preference judgments rather than absolute scores, to address the challenge of obtaining precise pointwise reward signals.To this end, we synthesize the ranking references by generating quality-controlled responses with zero human annotation.Experiments demonstrate RAG-Zeval's superior performance, achieving the strongest correlation with human judgments and outperforming baselines that rely on LLMs with 10 -100× more parameters.Our approach also exhibits superior interpretability in response evaluation 1 .
Kun Li 0003, Tianhua Zhang, Hongyin Luo, Xixin Wu, James R. Glass, Helen M. Meng
EMNLP4
2025 Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
abstract
We present Self-MoE, an approach that transforms a monolithic LLM into a compositional, modular system of self-specialized experts, named MiXSE (MiXture of Self-specialized Experts). Our approach leverages self-specialization, which constructs expert modules using self-generated synthetic data, each equipping a shared base LLM with distinct domain-specific capabilities, activated via self-optimized routing. This allows for dynamic and capability-specific handling of various target tasks, enhancing overall capabilities, without extensive human-labeled data and added parameters. Our empirical results reveal that specializing LLMs may exhibit potential trade-offs in performances on non-specialized tasks. On the other hand, our Self-MoE demonstrates substantial improvements (6.5%p on average) over the base LLM across diverse benchmarks such as knowledge, reasoning, math, and coding. It also consistently outperforms other methods, including instance merging and weight merging, while offering better flexibility and interpretability by design with semantic experts and routing. Our findings highlight the critical role of modularity, the applicability of Self-MoE to multiple base LLMs, and the potential of self-improvement in achieving efficient, scalable, and adaptable systems.
Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang 0041, Jacob A. Hansen, James R. Glass, David D. Cox, Rameswar Panda, Rogério Feris, Alan Ritter
ICLR3
2025 Quantifying Generalization Complexity for Large Language Models
abstract
While large language models (LLMs) have shown exceptional capabilities in understanding complex queries and performing sophisticated tasks, their generalization abilities are often deeply entangled with memorization, necessitating more precise evaluation. To address this challenge, we introduce Scylla, a dynamic evaluation framework that quantitatively measures the generalization abilities of LLMs. Scylla disentangles generalization from memorization via assessing model performance on both in-distribution (ID) and out-of-distribution (OOD) data through 20 tasks across 5 levels of complexity. Through extensive experiments, we uncover a non-monotonic relationship between task complexity and the performance gap between ID and OOD data, which we term the generalization valley. Specifically, this phenomenon reveals a critical threshold---referred to as critical complexity---where reliance on non-generalizable behavior peaks, indicating the upper bound of LLMs' generalization capabilities. As model size increases, the critical complexity shifts toward higher levels of task complexity, suggesting that larger models can handle more complex reasoning tasks before over-relying on memorization. Leveraging Scylla and the concept of critical complexity, we benchmark 28 LLMs including both open-sourced models such as LLaMA and Qwen families, and closed-sourced models like Claude and GPT, providing a more robust evaluation and establishing a clearer understanding of LLMs' generalization capabilities.
Zhenting Qi, Hongyin Luo, Xuliang Huang, Zhuokai Zhao, Yibo Jiang, Xiangjun Fan, Himabindu Lakkaraju, James R. Glass
ICLR2
2025 THREAD: Thinking Deeper with Recursive Spawning
abstract
Philip Schroeder, Nathaniel W. Morgan, Hongyin Luo, James R. Glass. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Philip Schroeder, Nathaniel Morgan, Hongyin Luo, James R. Glass
NAACL (Long Papers)3
2025 ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
abstract
Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over long frame sequences from a continuous stream of visual input at each moment of a task attempt. To address this limitation, we propose ROVER (Reasoning Over VidEo Recursively), a framework that enables the model to recursively decompose long-horizon video trajectories into segments corresponding to shorter subtasks within the trajectory. In doing so, ROVER facilitates more focused and accurate reasoning over temporally localized frame sequences without losing global context. We evaluate ROVER, implemented using an in-context learning approach, on diverse OpenX Embodiment videos and on a new dataset derived from RoboCasa that consists of 543 videos showing both expert and perturbed non-expert trajectories across 27 manipulation tasks. ROVER outperforms strong baselines across three video reasoning tasks: task progress estimation, frame-level natural language reasoning, and video question answering. We observe that, by reducing the number of frames the model reasons over at each timestep, ROVER mitigates model hallucinations, especially during unexpected or non-optimal moments of a trajectory. In addition, by enabling the implementation of a subtask-specific sliding context window, ROVER's time complexity scales linearly with video length, an asymptotic improvement over baselines.
Philip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo, James R. Glass
NeurIPS4
2024 Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
abstract
Query rewriting is a crucial technique for passage retrieval in open-domain conversational question answering (CQA).It decontexualizes conversational queries into self-contained questions suitable for off-the-shelf retrievers.Existing methods attempt to incorporate retriever's preference during the training of rewriting models.However, these approaches typically rely on extensive annotations such as in-domain rewrites and/or relevant passage labels, limiting the models' generalization and adaptation capabilities.In this paper, we introduce AdaQR (Adaptive Query Rewriting), a framework for training query rewriting models with limited rewrite annotations from seed datasets and completely no passage label.Our approach begins by fine-tuning compact large language models using only ~10% of rewrite annotations from the seed dataset training split.The models are then utilized to self-sample rewrite candidates for each query instance, further eliminating the expense for human labeling or larger language model prompting often adopted in curating preference data.A novel approach is then proposed to assess retriever's preference for these candidates with the probability of answers conditioned on the conversational query by marginalizing the Top-K passages.This serves as the reward for optimizing the rewriter further using Direct Preference Optimization (DPO), a process free of rewrite and retrieval annotations.Experimental results on four open-domain CQA datasets demonstrate that AdaQR not only enhances the in-domain capabilities of the rewriter with limited annotation requirement, but also adapts effectively to out-of-domain datasets.
Tianhua Zhang, Kun Li 0003, Hongyin Luo, Xixin Wu, James R. Glass, Helen M. Meng
EMNLP3
2024 Listen, Think, and Understand
abstract
The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this area since the development of AudioSet, most existing models are designed to map audio inputs to pre-defined, discrete sound label sets. In contrast, humans possess the ability to not only classify sounds into general categories, but also to listen to the finer details of the sounds, explain the reason for the predictions, think about what the sound infers, and understand the scene and what action needs to be taken, if any. Such capabilities beyond perception are not yet present in existing audio models. On the other hand, modern large language models (LLMs) exhibit emerging reasoning ability but they lack audio perception capabilities. Therefore, we ask the question: can we build a model that has both audio perception and reasoning ability? In this paper, we propose a new audio foundation model, called LTU (Listen, Think, and Understand). To train LTU, we created a new OpenAQA-5M dataset consisting of 1.9 million closed-ended and 3.7 million open-ended, diverse (audio, question, answer) tuples, and have used an autoregressive training framework with a perception-to-understanding curriculum. LTU demonstrates strong performance and generalization ability on conventional audio tasks such as classification and captioning. More importantly, it exhibits emerging audio reasoning and comprehension abilities that are absent in existing audio models. To the best of our knowledge, LTU is the first multimodal large language model that focuses on general audio (rather than just speech) understanding.
Yuan Gong 0001, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James R. Glass
ICLR2
2024 DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
abstract
Despite their impressive capabilities, large language models (LLMs) are prone to hallucinations, i.e., generating content that deviates from facts seen during pretraining. We propose a simple decoding strategy for reducing hallucinations with pretrained LLMs that does not require conditioning on retrieved external knowledge nor additional fine-tuning. Our approach obtains the next-token distribution by contrasting the differences in logits obtained from projecting the later layers versus earlier layers to the vocabulary space, exploiting the fact that factual knowledge in an LLMs has generally been shown to be localized to particular transformer layers. We find that this **D**ecoding by C**o**ntrasting **La**yers (DoLa) approach is able to better surface factual knowledge and reduce the generation of incorrect facts. DoLa consistently improves the truthfulness across multiple choices tasks and open-ended generation tasks, for example improving the performance of LLaMA family models on TruthfulQA by 12-17% absolute points, demonstrating its potential in making LLMs reliably generate truthful facts.
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, James R. Glass
ICLR3
2024 HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
abstract
While large vision-language models (LVLMs) have demonstrated impressive capabilities in interpreting multi-modal contexts, they invariably suffer from object hallucinations (OH). We introduce HALC, a novel decoding algorithm designed to mitigate OH in LVLMs. HALC leverages distinct fine-grained optimal visual information in vision-language tasks and operates on both local and global contexts simultaneously. Specifically, HALC integrates a robust auto-focal grounding mechanism (locally) to correct hallucinated tokens on the fly, and a specialized beam search algorithm (globally) to significantly reduce OH while preserving text generation quality. Additionally, HALC can be integrated into any LVLMs as a plug-and-play module without extra training. Extensive experimental studies demonstrate HALC’s effectiveness in reducing OH, outperforming state-of-the-arts across four benchmarks. Code is released at https://github.com/BillChan226/HALC.
Zhaorun Chen, Zhuokai Zhao, Hongyin Luo, Huaxiu Yao, Bo Li 0026
ICML3
2023 Entailment as Robust Self-Learner
abstract
Entailment has been recognized as an important metric for evaluating natural language understanding (NLU) models, and recent studies have found that entailment pretraining benefits weakly supervised fine-tuning.In this work, we design a prompting strategy that formulates a number of different NLU tasks as contextual entailment.This approach improves the zero-shot adaptation of pretrained entailment models.Secondly, we notice that self-training entailment-based models with unlabeled data can significantly improve the adaptation performance on downstream tasks.To achieve more stable improvement, we propose the Simple Pseudo-Label Editing (SimPLE) algorithm for better pseudo-labeling quality in self-training.We also found that both pretrained entailmentbased models and the self-trained models are robust against adversarial evaluation data.Experiments on binary and multi-class classification tasks show that SimPLE leads to more robust self-training results, indicating that the self-trained entailment models are more efficient and trustworthy than large language models on language understanding tasks.
Jiaxin Ge, Hongyin Luo, James R. Glass
ACL (1)2
2023 Joint Audio and Speech Understanding
abstract
Humans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profound comprehension of the relationship between them, constitute fundamental cognitive capabilities. For the first time, we build a machine learning model, called LTU-AS, that has a conceptually similar universal audio perception and advanced reasoning ability. Specifically, by integrating Whisper [1] as a perception module and LLaMA [2] as a reasoning module, LTU-AS can simultaneously recognize and jointly understand spoken text, speech paralinguistics, and non-speech audio events - almost everything perceivable from audio signals.
Yuan Gong 0001, Alexander H. Liu, Hongyin Luo, Leonid Karlinsky, James R. Glass
ASRU3
2023 Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning
abstract
Due to their similarity-based learning objectives, pretrained sentence encoders often internalize stereotypical assumptions that reflect the social biases that exist within their training corpora.In this paper, we describe several kinds of stereotypes concerning different communities that are present in popular sentence representation models, including pretrained next sentence prediction and contrastive sentence representation models.We compare such models to textual entailment models that learn language logic for a variety of downstream language understanding tasks.By comparing strong pretrained models based on text similarity with textual entailment learning, we conclude that the explicit logic learning with textual entailment can significantly reduce bias and improve the recognition of social communities, without an explicit de-biasing process.The code, model, and data associated with this work are publicly available at https: //github.com/luohongyin/ESP.git.
Hongyin Luo, James R. Glass
EACL1
2022 DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings
abstract
Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang 0001, Shiyu Chang, Marin Soljacic, Shang-Wen Li 0001, Scott Yih, James R. Glass
NAACL-HLT3
2022 Cooperative Self-training of Machine Reading Comprehension
abstract
Hongyin Luo, Shang-Wen Li, Mingye Gao, Seunghak Yu, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Hongyin Luo, Shang-Wen Li 0001, Mingye Gao, Seunghak Yu, James R. Glass
NAACL-HLT1
2021 Joint Retrieval-Extraction Training for Evidence-Aware Dialog Response Selection
Hongyin Luo, James R. Glass, Garima Lalwani, Shang-Wen Li 0001
Interspeech1
2020 Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption
abstract
Spoken dialog systems have seen applications in many domains, including medical for automatic conversational diagnosis.State-of-the-art dialog managers are usually driven by deep reinforcement learning models, such as deep Q networks (DQNs), which learn by interacting with a simulator to explore the entire action space since real conversations are limited.However, the DQN-based automatic diagnosis models do not achieve satisfying performances when adapted to new, unseen diseases with only a few training samples.In this work, we propose the Prototypical Q Networks (ProtoQN) as the dialog manager for the automatic diagnosis systems.The model calculates prototype embeddings with real conversations between doctors and patients, learning from them and simulator-augmented dialogs more efficiently.We create both supervised and few-shot learning tasks with the Muzhi corpus.Experiments showed that the ProtoQN significantly outperformed the baseline DQN model in both supervised and few-shot learning scenarios, and achieves state-of-the-art few-shot learning performances.
Hongyin Luo, Shang-Wen Li 0001, James R. Glass
INTERSPEECH1
2019 Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
abstract
Common language models typically predict the next word given the context.In this work, we propose a method that improves language modeling by learning to align the given context and the following phrase.The model does not require any linguistic annotation of phrase segmentation.Instead, we define syntactic heights and phrase segmentation rules, enabling the model to automatically induce phrases, recognize their task-specific heads, and generate phrase embeddings in an unsupervised learning manner.Our method can easily be applied to language models with different network architectures since an independent module is used for phrase induction and context-phrase alignment, and no change is required in the underlying language modeling network.Experiments have shown that our model outperformed several strong baseline models on different data sets.We achieved a new state-of-the-art performance of 17.4 perplexity on the Wikitext-103 dataset.Additionally, visualizing the outputs of the phrase induction module showed that our model is able to learn approximate phrase-level structural knowledge without any annotation.
Hongyin Luo, Yonatan Belinkov, James R. Glass
ACL (1)1
2019 Integrating Video Retrieval and Moment Detection in a Unified Corpus for Video Question Answering
Hongyin Luo, Mitra Mohtarami, James R. Glass, Karthik Krishnamurthy, Brigitte Richardson
INTERSPEECH1
2018 Learning Word Representations with Cross-Sentence Dependencyfor End-to-End Co-reference Resolution
abstract
In this work, we present a word embedding model that learns cross-sentence dependency for improving end-to-end co-reference resolution (E2E-CR).While the traditional E2E-CR model generates word representations by running long short-term memory (LSTM) recurrent neural networks on each sentence of an input article or conversation separately, we propose linear sentence linking and attentional sentence linking models to learn crosssentence dependency.Both sentence linking strategies enable the LSTMs to make use of valuable information from context sentences while calculating the representation of the current input word.With this approach, the LSTMs learn word embeddings considering knowledge not only from the current sentence but also from the entire input document.Experiments show that learning cross-sentence dependency enriches information contained by the word representations, and improves the performance of the co-reference resolution model compared with our baseline.
Hongyin Luo, James R. Glass
EMNLP1
2018 Race-Condition-Aware and Hardware-Oriented Task Partitioning and Scheduling Using Entropy Maximization
abstract
In a multithreaded execution environment, race condition leads to computational errors and system hazards. Up to date, a series of task scheduling strategies have been presented in the literature to reduce the risk of race condition. Because of the increasing complexity of this problem in multi-core systems, existing task scheduling approaches are not very efficient. To deal with this challenge, in this work, we develop a race-condition-aware and hardware-oriented task partitioning and scheduling algorithm using entropy maximization model. We model uncertainty as a probabilistic occurrence within a time interval, and hence the characteristics of event ordering are analyzed through an uncertainty matrix. Next, a metric is developed to measure the uncertainty of task execution in various execution environments. Finally, a maximum entropy model is generated to ensure the lowest probability of race condition during task execution. The smallest one among maximum entropy values is chosen and used in our proposed task scheduling algorithm. Experimental results show that the proposed task scheduling strategy based on our maximum entropy model outperforms existing state-of-the-art approaches. For example, in a 128-core computing system, the task execution time, CPU utilization ratio, and throughput of our proposed task scheduling is improved by$15.3\sim 36.4$percent,$8.2\sim 17.6$percent, and$20.7\sim 41.4$percent, respectively. Moreover, our proposed scheduling algorithm exhibits low computational complexity and good adaptivity to diverse execution environments.
Sizhao Li, Yuanzhi Zhang 0004, Hongyin Luo, Chao Lu 0005, Donghui Guo
IEEE Trans. Parallel Distributed Syst.3
2016 DrMAD: Distilling Reverse-Mode Automatic Differentiation for Optimizing Hyperparameters of Deep Neural Networks
Jie Fu 0001, Hongyin Luo, Jiashi Feng, Kian Hsiang Low, Tat-Seng Chua
IJCAI2
2015 Online Learning of Interpretable Word Embeddings
abstract
Word embeddings encode semantic meanings of words into low-dimension word vectors.In most word embeddings, one cannot interpret the meanings of specific dimensions of those word vectors.Nonnegative matrix factorization (NMF) has been proposed to learn interpretable word embeddings via non-negative constraints.However, NMF methods suffer from scale and memory issue because they have to maintain a global matrix for learning.To alleviate this challenge, we propose online learning of interpretable word embeddings from streaming text data.Experiments show that our model consistently outperforms the state-of-the-art word embedding methods in both representation ability and interpretability.The source code of this paper can be obtained from http: //github.com/skTim/OIWE.
Hongyin Luo, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001
EMNLP1
2014 Hybrid circuit-switched network for on-chip communication in large-scale chip-multiprocessors
Hongyin Luo, Shaojun Wei, Deming Chen, Donghui Guo
J. Parallel Distributed Comput.1