VLDB 2026 Research / reviewers in the wild / expert
Xiang Kong
dblp:128/8032
· DBLP profile ↗
20ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte CarloabstractAugmenting the multi-step reasoning abilities of Large Language Models (LLMs) has been a persistent challenge. Recently, verification has shown promise in improving solution consistency by evaluating generated outputs. However, current verification approaches suffer from sampling inefficiencies, requiring a large number of samples to achieve satisfactory performance. Additionally, training an effective verifier often depends on extensive process supervision, which is costly to acquire. In this paper, we address these limitations by introducing a novel verification method based on Twisted Sequential Monte Carlo (TSMC). TSMC sequentially refines its sampling effort to focus exploration on promising candidates, resulting in more efficient generation of high-quality solutions. We apply TSMC to LLMs by estimating the expected future rewards at partial solutions. This approach results in a more straightforward training target that eliminates the need for step-wise human annotations. We empirically demonstrate the advantages of our method across multiple math benchmarks, and also validate our theoretical analysis of both our approach and existing verification methods. Shengyu Feng, Xiang Kong, Aonan Zhang, Ruoming Pang, Yiming Yang 0002 |
ICLR | 2 |
| 2025 | TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated WeightsabstractDirect Preference Optimization (DPO) has been widely adopted for preference alignment of Large Language Models (LLMs) due to its simplicity and effectiveness.
However, DPO is derived as a bandit problem in which the whole response is treated as a single arm, ignoring the importance differences between tokens, which may affect optimization efficiency and make it difficult to achieve optimal results.
In this work, we propose that the optimal data for DPO has equal expected rewards for each token in winning and losing responses, as there is no difference in token importance.
However, since the optimal dataset is unavailable in practice, we propose using the original dataset for importance sampling to achieve unbiased optimization.
Accordingly, we propose a token-level importance sampling DPO objective named TIS-DPO that assigns importance weights to each token based on its reward.
Inspired by previous works, we estimate the token importance weights using the difference in prediction probabilities from a pair of contrastive LLMs. We explore three methods to construct these contrastive LLMs: (1) guiding the original LLM with contrastive prompts, (2) training two separate LLMs using winning and losing responses, and (3) performing forward and reverse DPO training with winning and losing responses.
Experiments show that TIS-DPO significantly outperforms various baseline methods on harmlessness and helpfulness alignment and summarization tasks. We also visualize the estimated weights, demonstrating their ability to identify key token positions. Aiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun, Xiang Kong, Xiaoming Simon Wang, Jiulong Shan, Albin Madappally Jose, Xiaojiang Liu, Lijie Wen 0001, Philip S. Yu |
ICLR | 5 |
| 2025 | Checklists Are Better Than Reward Models For Aligning Language ModelsabstractLanguage models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this —typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item—using both AI judges and specialized verifier programs—then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods on top of a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks — RLCF is the only method to help on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. We show that RLCF can also be used off-policy to improve Llama 3.1 8B Instruct and OLMo 2 7B Instruct. These results establish rubrics as a key tool for improving language models' support of queries that express a multitude of needs. We release our our dataset of rubrics (WildChecklists), models, and code to the public. Vijay Viswanathan 0002, Yanchao Sun, Xiang Kong, Graham Neubig, Sherry Tongshuang Wu |
NeurIPS | 3 |
| 2024 | Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationabstractAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong, Xiaoming Wang, Jiulong Shan, Meng Cao, Lijie Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Aiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong, Jiulong Shan, Lijie Wen 0001 |
ACL (1) | 4 |
| 2024 | MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang |
ECCV (29) | 17 |
| 2023 | Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, Luke Zettlemoyer |
ICLR | 3 |
| 2022 | BLT: Bidirectional Layout Transformer for Controllable Layout Generation
Xiang Kong, Lu Jiang 0004, Huiwen Chang, Han Zhang 0010, Yuan Hao, Haifeng Gong, Irfan A. Essa |
ECCV (17) | 1 |
| 2021 | Multilingual Neural Machine Translation with Deep Encoder and Multiple Shallow DecodersabstractXiang Kong, Adithya Renduchintala, James Cross, Yuqing Tang, Jiatao Gu, Xian Li. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Xiang Kong, Adithya Renduchintala, James Cross 0003, Jiatao Gu, Xian Li 0003 |
EACL | 1 |
| 2021 | Decoupling Global and Local Representations via Invertible Generative Flows
Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy |
ICLR | 2 |
| 2021 | Luna: Linear Unified Nested AttentionabstractThe quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Luna, a linear unified nested attention mechanism that approximates softmax attention with two nested linear attention functions, yielding only linear (as opposed to quadratic) time and space complexity. Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length. Then, the packed sequence is unpacked using the second attention function. As compared to a more traditional attention mechanism, Luna introduces an additional sequence with a fixed length as input and an additional corresponding output, which allows Luna to perform attention operation linearly, while also storing adequate contextual information. We perform extensive evaluations on three benchmarks of sequence modeling tasks: long-context sequence modelling, neural machine translation and masked language modeling for large-scale pretraining. Competitive or even better experimental results demonstrate both the effectiveness and efficiency of Luna compared to a variety of strong baseline methods including the full-rank attention and other efficient sparse and dense attention methods. Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma 0001, Luke Zettlemoyer |
NeurIPS | 2 |
| 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From ExaminationsabstractWe introduce SCDE, a dataset to evaluate the performance of computational models through sentence prediction.SCDE is a humancreated sentence cloze dataset, collected from public school English examinations.Our task requires a model to fill up multiple blanks in a passage from a shared candidate set with distractors designed by English teachers.Experimental results demonstrate that this task requires the use of non-local, discourse-level context beyond the immediate sentence neighborhood.The blanks require joint solving and significantly impair each other's context.Furthermore, through ablations, we show that the distractors are of high quality and make the task more challenging.Our experiments show that there is a significant performance gap between advanced models (72%) and humans (87%), encouraging future models to bridge this gap. 1 2 Passage: A student's life is never easy.And it is even more difficult if you will have to complete your study in a foreign land. 1The following are some basic things you need to do before even seizing that passport and boarding on the plane.Knowing the country.You shouldn't bother researching the country's hottest tourist spots or historical places.You won't go there as a tourist, but as a student.2 In addition, read about their laws.You surely don't want to face legal problems, especially if you're away from home.3 Don't expect that you can graduate abroad without knowing even the basics of the language.Before leaving your home country, take online lessons to at least master some of their words and sentences.This will be useful in living and studying there.Doing this will also prepare you in communicating with those who can't speak English.Preparing for other needs.Check the conversion of your money to their local currency.4. The Internet of your intended school will be very helpful in findings an apartment and helping you understand local currency.Remember, you're not only carrying your own reputation but your country's reputation as well.If you act foolishly, people there might think that all of your countrymen are foolish as well. 5Candidates: A. Studying their language.B. That would surely be a very bad start for your study abroad program.C. Going with their trends will keep it from being too obvious that you're a foreigner.D. Set up your bank account so you can use it there, get an insurance, and find an apartment.E. It'll be helpful to read the most important points in their history and to read up on their culture.F. A lot of preparations are needed so you can be sure to go back home with a diploma and a bright future waiting for you.G. Packing your clothes.Answers with Reasoning Type: 1→F (Summary) , 2→E (Inference) , 3→A (Paraphrase) , 4→D (WordMatch), 5→B (Inference) (C and G are distractors) Discussion: Blank 3 is the easiest to solve, since "Studying their language" is a near-paraphrase of "Knowing even the basics of the language".Blank 2 needs to be reasoned out by Inference -specifically E can be inferred from the previous sentence.Note however that C is also a possible inference from the previous sentence -it is only after reading the entire context, which seems to be about learning various aspects of a country, that E seems to fit better.Blank 1 needs Summary → it requires understanding several later sentences and abstracting out that they all refer to lots of preparations.Finally, Blank 5 can be mapped to B by inferring that people thinking all your countrymen are foolish is bad, while Blank 4 is a easy WordMatch on apartment to D. The other distractor G, although topically related to preparation for going abroad, does not directly fit into any of the blank contexts Xiang Kong, Varun Gangal, Eduard H. Hovy |
ACL | 1 |
| 2020 | A Two-Step Approach for Implicit Event Argument DetectionabstractIn this work, we explore the implicit event argument detection task, which studies event arguments beyond sentence boundaries.The addition of cross-sentence argument candidates imposes great challenges for modeling.To reduce the number of candidates, we adopt a two-step approach, decomposing the problem into two sub-problems: argument head-word detection and head-to-span expansion.Evaluated on the recent RAMS dataset (Ebner et al., 2020), our model achieves overall better performance than a strong sequence labeling baseline.We further provide detailed error analysis, presenting where the model mainly makes errors and indicating directions for future improvements.It remains a challenge to detect implicit arguments, calling for more future work of document-level modeling for this task. Zhisong Zhang, Xiang Kong, Zhengzhong Liu 0001, Xuezhe Ma, Eduard H. Hovy |
ACL | 2 |
| 2020 | Incorporating a Local Translation Mechanism into Non-autoregressive TranslationabstractIn this work, we introduce a novel local autoregressive translation (LAT) mechanism into non-autoregressive translation (NAT) models so as to capture local dependencies among target outputs.Specifically, for each target decoding position, instead of only one token, we predict a short sequence of tokens in an autoregressive way.We further design an efficient merging algorithm to align and merge the output pieces into one final output sequence.We integrate LAT into the conditional masked language model (CMLM; Ghazvininejad et al., 2019) and similarly adopt iterative decoding.Empirical results on five translation tasks show that compared with CMLM, our method achieves comparable or better performance with fewer decoding iterations, bringing a 2.5x speedup.Further analysis indicates that our method reduces repeated translations and performs better at longer sentences.The code for our model is available at https://github. com/shawnkx/NAT-with-Local-AT. Xiang Kong, Zhisong Zhang, Eduard H. Hovy |
EMNLP (1) | 1 |
| 2020 | Deep Transformers with Latent DepthabstractThe Transformer model has achieved state-of-the-art performance in many sequence modeling tasks. However, how to leverage model capacity with large or variable depths is still an open challenge. We present a probabilistic framework to automatically learn which layer(s) to use by learning the posterior distributions of layer selection. As an extension of this framework, we propose a novel method to train one shared Transformer network for multilingual machine translation with different layer selection posteriors for each language pair. The proposed method alleviates the vanishing gradient issue and enables stable training of deep Transformers (e.g. 100 layers). We evaluate on WMT English-German machine translation and masked language modeling tasks, where our method outperforms existing approaches for training deeper Transformers. Experiments on multilingual machine translation demonstrate that this approach can effectively leverage increased model capacity and bring universal improvement for both many-to-one and one-to-many translation with diverse language pairs. Xian Li 0003, Asa Cooper Stickland, Xiang Kong |
NeurIPS | 4 |
| 2019 | Neural Machine Translation with Adequacy-Oriented LearningabstractAlthough Neural Machine Translation (NMT) models have advanced state-of-the-art performance in machine translation, they face problems like the inadequate translation. We attribute this to that the standard Maximum Likelihood Estimation (MLE) cannot judge the real translation quality due to its several limitations. In this work, we propose an adequacyoriented learning mechanism for NMT by casting translation as a stochastic policy in Reinforcement Learning (RL), where the reward is estimated by explicitly measuring translation adequacy. Benefiting from the sequence-level training of RL strategy and a more accurate reward designed specifically for translation, our model outperforms multiple strong baselines, including (1) standard and coverage-augmented attention models with MLE-based training, and (2) advanced reinforcement and adversarial training strategies with rewards based on both word-level BLEU and character-level CHRF3. Quantitative and qualitative analyses on different language pairs and NMT architectures demonstrate the effectiveness and universality of the proposed approach. Xiang Kong, Zhaopeng Tu, Shuming Shi 0001, Eduard H. Hovy, Tong Zhang 0001 |
AAAI | 1 |
| 2019 | Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language GenerationabstractMixture of Softmaxes (MoS) has been shown to be effective at addressing the expressiveness limitation of Softmax-based models. Despite the known advantage, MoS is practically sealed by its large consumption of memory and computational time due to the need of computing multiple Softmaxes. In this work, we set out to unleash the power of MoS in practical applications by investigating improved word coding schemes, which could effectively reduce the vocabulary size and hence relieve the memory and computation burden. We show both BPE and our proposed Hybrid-LightRNN lead to improved encoding mechanisms that can halve the time and memory consumption of MoS without performance losses. With MoS, we achieve an improvement of 1.5 BLEU scores on IWSLT 2014 German-to-English corpus and an improvement of 0.76 CIDEr score on image captioning. Moreover, on the larger WMT 2014 machine translation dataset, our MoSboosted Transformer yields 29.6 BLEU score for English-toGerman and 42.1 BLEU score for English-to-French, outperforming the single-Softmax Transformer by 0.9 and 0.4 BLEU scores respectively and achieving the state-of-the-art result on WMT 2014 English-to-German task. Xiang Kong, Qizhe Xie, Zihang Dai, Eduard H. Hovy |
AAAI | 1 |
| 2019 | Generalized Data Augmentation for Low-Resource TranslationabstractTranslation to or from low-resource languages (LRLs) poses challenges for machine translation in terms of both adequacy and fluency.Data augmentation utilizing large amounts of monolingual data is regarded as an effective way to alleviate these problems.In this paper, we propose a general framework for data augmentation in low-resource machine translation that not only uses target-side monolingual data, but also pivots through a related highresource language (HRL).Specifically, we experiment with a two-step pivoting method to convert high-resource data to the LRL, making use of available resources to better approximate the true data distribution of the LRL.First, we inject LRL words into HRL sentences through an induced bilingual dictionary.Second, we further edit these modified sentences using a modified unsupervised machine translation framework.Extensive experiments on four low-resource datasets show that under extreme low-resource settings, our data augmentation techniques improve translation quality by up to 1.5 to 8 BLEU points compared to supervised back-translation baselines.1 Mengzhou Xia, Xiang Kong, Antonios Anastasopoulos, Graham Neubig |
ACL (1) | 2 |
| 2019 | MaCow: Masked Convolutional Generative FlowabstractFlow-based generative models, conceptually attractive due to tractability of both the exact log-likelihood computation and latent-variable inference, and efficiency of both training and sampling, has led to a number of impressive empirical successes and spawned many advanced variants and theoretical investigations. Despite their computational efficiency, the density estimation performance of flow-based generative models significantly falls behind those of state-of-the-art autoregressive models. In this work, we introduce masked convolutional generative flow (MaCow), a simple yet effective architecture of generative flow using masked convolution. By restricting the local connectivity in a small kernel, MaCow enjoys the properties of fast and stable training, and efficient sampling, while achieving significant improvements over Glow for density estimation on standard image benchmarks, considerably narrowing the gap to autoregressive models. Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy |
NeurIPS | 2 |
| 2017 | Evaluating automatic speech recognition systems in comparison with human perception results using distinctive feature measuresabstractThis paper describes methods for evaluating automatic speech recognition (ASR) systems in comparison with human perception results, using measures derived from linguistic distinctive features. Error patterns in terms of manner, place and voicing are presented, along with an examination of confusion matrices via a distinctive-feature-distance metric. These evaluation methods contrast with conventional performance criteria that focus on the phone or word level, and are intended to provide a more detailed profile of ASR system performance, as well as a means for direct comparison with human perception results at the sub-phonemic level. Xiang Kong, Jeung-Yoon Choi, Stefanie Shattuck-Hufnagel |
ICASSP | 1 |
| 2003 | Dynamic routing with endpoint admission control for V61P networksabstractIn this paper, we propose a dynamic routing with endpoint measurement based admission control (EMBAC) for VoIP networks with mesh topology. The basic idea of this admission control is to use probe flows to measure two routes, which are direct route and one of the two-link routes between each VoIP call's originating-terminating node pair. Based on the probe flow's QoS evaluation, the node decides a route for each call. We compare performance of direct routing, basic dynamic routing and advanced dynamic routing by using simulation. It is shown that an advanced dynamic routing with EMBAC works better for minimizing call blocking probability and packet loss rate than direct routing. Xiang Kong, Kenichi Mase |
ICC | 1 |