EDBT 2026 Demo / reviewers in the wild / expert
Xinting Huang
dblp:240/7147
· DBLP profile ↗
18ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0001-6827-7426ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context CompressionabstractChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu 0001, Zhicheng Dou |
ACL (1) | 5 |
| 2025 | LoGU: Long-form Generation with Uncertainty ExpressionsabstractRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Nigel Collier, Dong Yu 0001, Deqing Yang |
ACL (1) | 4 |
| 2025 | Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsabstractLarge language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms. Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001 |
ACL (1) | 3 |
| 2025 | UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationabstractLarge Language Models (LLMs) are prone to hallucination, particularly in long-form generations.A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge.However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation.To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long-and short-form question answering (QA).UNCLE covers five domains and includes more than 1,000 entities, each with paired short-and long-form QA items.Our dataset is the first to directly link short-and long-form QA through aligned questions and gold-standard answers.Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty.We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation.We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains.Further analysis of alignment gaps between short-and long-form uncertainty expression highlights promising directions for future research using UNCLE. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Dong Yu 0001, Nigel Collier, Deqing Yang |
EMNLP | 4 |
| 2025 | A Formal Framework for Understanding Length Generalization in TransformersabstractA major challenge for transformers is generalizing to sequences longer than those observed during training. While previous works have empirically shown that transformers can either succeed or fail at length generalization depending on the task, theoretical understanding of this phenomenon remains limited. In this work, we introduce a rigorous theoretical framework to analyze length generalization in causal transformers with learnable absolute positional encodings. In particular, we characterize those functions that are identifiable in the limit from sufficiently long inputs with absolute positional encodings under an idealized inference scheme using a norm-based regularizer. This enables us to prove the possibility of length generalization for a rich family of problems. We experimentally validate the theory as a predictor of success and failure of length generalization across a range of algorithmic and formal language tasks. Our theory not only explains a broad set of empirical observations but also opens the way to provably predicting length generalization capabilities in transformers. Xinting Huang, Andy Yang, Satwik Bhattamishra, Yash Raj Sarrof, Andreas Krebs, Hattie Zhou, Preetum Nakkiran, Michael Hahn 0001 |
ICLR | 1 |
| 2025 | Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention TransformersabstractChain-of-thought reasoning and scratchpads have emerged as critical tools for enhancing the computational capabilities of transformers. While theoretical results show that polynomial-length scratchpads can extend transformers’ expressivity from $TC^0$ to $PTIME$, their required length remains poorly understood. Empirical evidence even suggests that transformers need scratchpads even for many problems in $TC^0$, such as Parity or Multiplication, challenging optimistic bounds derived from circuit complexity. In this work, we initiate the study of systematic lower bounds for the number of CoT steps across different algorithmic problems, in the hard-attention regime. We study a variety of algorithmic problems, and provide bounds that are tight up to logarithmic factors. Overall, these results contribute to emerging understanding of the power and limitations of chain-of-thought reasoning. Alireza Amiri Bavandpour, Xinting Huang, Mark Rofin, Michael Hahn 0001 |
ICML | 2 |
| 2024 | SEGO: Sequential Subgoal Optimization for Mathematical Problem-SolvingabstractLarge Language Models (LLMs) have driven substantial progress in artificial intelligence in recent years, exhibiting impressive capabilities across a wide range of tasks, including mathematical problem-solving.Inspired by the success of subgoal-based methods, we propose a novel framework called SEquential subGoal Optimization (SEGO) to enhance LLMs' ability to solve mathematical problems.By establishing a connection between the subgoal breakdown process and the probability of solving problems, SEGO aims to identify better subgoals with theoretical guarantees.Addressing the challenge of identifying suitable subgoals in a large solution space, our framework generates problem-specific subgoals and adjusts them according to carefully designed criteria.Incorporating these optimized subgoals into the policy model training leads to significant improvements in problem-solving performance.We validate SEGO's efficacy through experiments on two benchmarks, GSM8K and MATH, where our approach outperforms existing methods, highlighting the potential of SEGO in AI-driven mathematical problemsolving. * This work Xueliang Zhao, Xinting Huang, Wei Bi, Lingpeng Kong |
ACL (1) | 2 |
| 2024 | Knowledge Verification to Nip Hallucination in the BudabstractWhile large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.In this paper, we demonstrate the feasibility of mitigating hallucinations by verifying and minimizing the inconsistency between external knowledge present in the alignment data and the intrinsic knowledge embedded within foundation LLMs.Specifically, we propose a novel approach called Knowledge Consistent Alignment (KCA), which employs a well-aligned LLM to automatically formulate assessments based on external knowledge to evaluate the knowledge boundaries of foundation LLMs.To address knowledge inconsistencies in the alignment data, KCA implements several specific strategies to deal with these data instances.We demonstrate the superior efficacy of KCA in reducing hallucinations across six benchmarks, utilizing foundation LLMs of varying backbones and scales.This confirms the effectiveness of mitigating hallucinations by reducing knowledge inconsistency.Our code, model weights, and data are openly accessible at https://github.com/fanqiwan/KCA.* Part of the work was done during his internship at Tencent AI Lab. Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
EMNLP | 2 |
| 2024 | Knowledge Fusion of Large Language ModelsabstractWhile training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures—Llama-2, MPT, and OpenLLaMA—across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at \url{https://github.com/fanqiwan/FuseLLM}. Fanqi Wan, Xinting Huang, Deng Cai 0002, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
ICLR | 2 |
| 2024 | See or Guess: Counterfactually Regularized Image CaptioningabstractImage captioning, which generates natural language descriptions of images, is a crucial task in vision-language research. Previous models have typically addressed this task by aligning the generative capabilities of machines with humans through statistical fitting existing datasets. While effective for normal images, they may struggle to accurately describe those where certain parts of the image are obscured or edited, unlike humans who excel in such cases. These weaknesses, including hallucinations and limited interpretability, often hinder performance in scenarios with shifted association patterns. In this paper, we present a generic image captioning framework that employs causal inference to make existing models more capable of interventional tasks, and counterfactually explainable. Our approach includes two variants leveraging either total effect or natural direct effect. Integrating them into the training process enables models to handle counterfactual scenarios, increasing their generalizability. Extensive experiments on various datasets show that our method effectively reduces hallucinations and improves the model's faithfulness to images, demonstrating high portability across both small-scale and large-scale image-to-text models. The code is available at https://github.com/Aman-4-Real/See-or-Guess. Qian Cao 0001, Xu Chen 0017, Ruihua Song, Xiting Wang, Xinting Huang, Yuchen Ren 0005 |
ACM Multimedia | 5 |
| 2024 | DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction WrappingabstractYongrui Chen, Haiyun Jiang, Xinting Huang, Shuming Shi, Guilin Qi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yongrui Chen 0002, Haiyun Jiang, Xinting Huang, Shuming Shi 0001, Guilin Qi |
NAACL-HLT | 3 |
| 2024 | InversionView: A General-Purpose Method for Reading Information from Neural ActivationsabstractThe inner workings of neural networks can be better understood if we can fully decipher the information encoded in neural activations. In this paper, we argue that this information is embodied by the subset of inputs that give rise to similar activations. We propose InversionView, which allows us to practically inspect this subset by sampling from a trained decoder model conditioned on activations. This helps uncover the information content of activation vectors, and facilitates understanding of the algorithms implemented by transformer models. We present four case studies where we investigate models ranging from small transformers to GPT-2. In these studies, we show that InversionView can reveal clear information contained in activations, including basic information about tokens appearing in the context, as well as more complex information, such as the count of certain tokens, their relative positions, and abstract knowledge about the subject. We also provide causally verified circuits to confirm the decoded information. Xinting Huang, Madhur Panwar, Navin Goyal, Michael Hahn 0001 |
NeurIPS | 1 |
| 2023 | Pre-training Multi-party Dialogue Models with Latent Discourse InferenceabstractMulti-party dialogues are more difficult for models to understand than one-to-one twoparty dialogues, since they involve multiple interlocutors, resulting in interweaving reply-to relations and information flows.To step over these obstacles, an effective way is to pre-train a model that understands the discourse structure of multi-party dialogues, namely, to whom each utterance is replying.However, due to the lack of explicitly annotated discourse labels in multi-party dialogue corpora, previous works fail to scale up the pre-training process by putting aside the unlabeled multi-party conversational data for nothing.To fully utilize the unlabeled data, we propose to treat the discourse structures as latent variables, then jointly infer them and pre-train the discourse-aware model by unsupervised latent variable inference methods.Experiments on multiple downstream tasks show that our pre-trained model outperforms strong baselines by large margins and achieves state-of-the-art (SOTA) results, justifying the effectiveness of our method.The official implementation of this paper is available at https://github.com/EricLee8/MPD_EMVI. Yiyang Li 0002, Xinting Huang, Wei Bi, Hai Zhao 0001 |
ACL (1) | 2 |
| 2023 | Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationabstractInstruction-tuning can be substantially optimized through enhanced diversity, resulting in models capable of handling a broader spectrum of tasks.However, existing data employed for such tuning often exhibit an inadequate coverage of individual domains, limiting the scope for nuanced comprehension and interactions within these areas.To address this deficiency, we propose EXPLORE-INSTRUCT, a novel approach to enhance the data coverage to be used in domain-specific instruction-tuning through active exploration via Large Language Models (LLMs).Built upon representative domain use cases, EXPLORE-INSTRUCT explores a multitude of variations or possibilities by implementing a search algorithm to obtain diversified and domain-focused instruction-tuning data.Our data-centric analysis validates the effectiveness of this proposed approach in improving domain-specific instruction coverage.Moreover, our model's performance demonstrates considerable advancements over multiple baselines, including those utilizing domainspecific data enhancement.Our findings offer a promising opportunity to improve instruction coverage, especially in domain-specific contexts, thereby advancing the development of adaptable language models.Our code, model weights, and data are public at https:// github.com/fanqiwan/Explore-Instruct. Fanqi Wan, Xinting Huang, Tao Yang 0033, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
EMNLP | 2 |
| 2020 | MALA: Cross-Domain Dialogue Generation with Action LearningabstractResponse generation for task-oriented dialogues involves two basic components: dialogue planning and surface realization. These two components, however, have a discrepancy in their objectives, i.e., task completion and language quality. To deal with such discrepancy, conditioned response generation has been introduced where the generation process is factorized into action decision and language generation via explicit action representations. To obtain action representations, recent studies learn latent actions in an unsupervised manner based on the utterance lexical similarity. Such an action learning approach is prone to diversities of language surfaces, which may impinge task completion and language quality. To address this issue, we propose multi-stage adaptive latent action learning (MALA) that learns semantic latent actions by distinguishing the effects of utterances on dialogue progress. We model the utterance effect using the transition of dialogue states caused by the utterance and develop a semantic similarity measurement that estimates whether utterances have similar effects. For learning semantic actions on domains without dialogue states, MALA extends the semantic similarity measurement across domains progressively, i.e., from aligning shared actions to learning domain-specific actions. Experiments using multi-domain datasets, SMD and MultiWOZ, show that our proposed model achieves consistent improvements over the baselines models in terms of both task completion and language quality. Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003 |
AAAI | 1 |
| 2020 | Semi-Supervised Dialogue Policy Learning via Stochastic Reward EstimationabstractDialogue policy optimization often obtains feedback until task completion in taskoriented dialogue systems.This is insufficient for training intermediate dialogue turns since supervision signals (or rewards) are only provided at the end of dialogues.To address this issue, reward learning has been introduced to learn from state-action pairs of an optimal policy to provide turn-by-turn rewards.This approach requires complete state-action annotations of human-to-human dialogues (i.e., expert demonstrations), which is labor intensive.To overcome this limitation, we propose a novel reward learning approach for semisupervised policy learning.The proposed approach learns a dynamics model as the reward function which models dialogue progress (i.e., state-action sequences) based on expert demonstrations, either with or without annotations.The dynamics model computes rewards by predicting whether the dialogue progress is consistent with expert demonstrations.We further propose to learn action embeddings for a better generalization of the reward function.The proposed approach outperforms competitive policy learning baselines on MultiWOZ, a benchmark multi-domain dataset. Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003 |
ACL | 1 |
| 2019 | CARL: Aggregated Search with Context-Aware Module Embedding LearningabstractAggregated search aims to construct search result pages (SERPs) from blue-links and heterogeneous modules (such as news, images, and videos). Existing studies have largely ignored the correlations between blue-links and heterogeneous modules when selecting the heterogeneous modules to be presented. We observe that the top ranked blue-links, which we refer to as the context, can provide important information about query intent and helps identify the relevant heterogeneous modules. For example, informative terms like "streamed" and "recorded" in the context imply that a video module may better satisfy the query. To model and utilize the context information for aggregated search, we propose a model with context attention and representation learning (CARL). Our model applies a recurrent neural network with attention mechanism to encode the context, and incorporates the encoded context information into module embeddings. The context-aware module embeddings together with the ranking policy are jointly optimized under the Markov decision process (MDP) formulation. To achieve a more effective joint learning, we further propose an optimization function with self-supervision loss to provide auxiliary supervision signals. Experimental results based on two public datasets demonstrate the superiority of CARL over multiple baseline approaches, and confirm the effectiveness of the proposed optimization function in boosting the joint learning process. Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003, Hai-Tao Zheng 0002 |
IJCNN | 1 |
| 2019 | Enhancing intraday stock price manipulation detection by leveraging recurrent neural networks with ensemble learning
Qili Wang, Wei Xu 0008, Xinting Huang |
Neurocomputing | 3 |