VLDB 2026 Research / reviewers in the wild / expert
Kaiqiang Song
dblp:222/2920
· DBLP profile ↗
23ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0001-8203-9723ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VC4VG: Optimizing Video Captions for Text-to-Video GenerationabstractRecent advances in text-to-video (T2V) generation highlight the critical role of highquality video-text pairs in training models capable of producing coherent and instructionaligned videos.However, strategies for optimizing video captions specifically for T2V training remain underexplored.In this paper, we introduce VC4VG (Video Captioning for Video Generation), a comprehensive caption optimization framework tailored to the needs of T2V models.We begin by analyzing caption content from a T2V perspective, decomposing the essential elements required for video reconstruction into multiple dimensions, and proposing a principled caption design methodology.To support evaluation, we construct VC4VG-Bench, a new benchmark featuring fine-grained, multi-dimensional, and necessity-graded metrics aligned with T2Vspecific requirements.Extensive T2V finetuning experiments demonstrate a strong correlation between improved caption quality and video generation performance, validating the effectiveness of our approach.We release all benchmark tools and code 1 to support further research. Yang Du 0011, Zhuoran Lin, Kaiqiang Song, Zhicheng Zheng, Tiezheng Ge, Bo Zheng 0007, Qin Jin |
EMNLP | 3 |
| 2025 | Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyabstractLarge Language Models (LLMs) are susceptible to security and safety threats, such as prompt injection, prompt extraction, and harmful requests.
One major cause of these vulnerabilities is the lack of an instruction hierarchy.
Modern LLM architectures treat all inputs equally, failing to distinguish between and prioritize various types of instructions, such as system messages, user prompts, and data.
As a result, lower-priority user prompts may override more critical system instructions, including safety protocols.
Existing approaches to achieving instruction hierarchy, such as delimiters and instruction-based training, do not address this issue at the architectural level.
We introduce the $\textbf{I}$nstructional $\textbf{S}$egment $\textbf{E}$mbedding (ISE) technique, inspired by BERT, to modern large language models, which embeds instruction priority information directly into the model.
This approach enables models to explicitly differentiate and prioritize various instruction types, significantly improving safety against malicious prompts that attempt to override priority rules.
Our experiments on the Structured Query and Instruction Hierarchy benchmarks demonstrate an average robust accuracy increase of up to 15.75\% and 18.68\%, respectively.
Furthermore, we observe an improvement in the instruction-following capability of up to 4.1\% on AlpacaEval.
Overall, our approach offers a promising direction for enhancing the safety and effectiveness of LLM architectures. Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Reddy Indurthi, Chong Xiang 0001, Prateek Mittal, Wenxuan Zhou 0005 |
ICLR | 3 |
| 2024 | SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMsabstractYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004 |
ACL (1) | 2 |
| 2024 | When Reasoning Meets Information Aggregation: A Case Study with Sports NarrativesabstractReasoning is most powerful when an LLM accurately aggregates relevant information.We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives.To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions.We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives.By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information.Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns.Open-source models like Llama-3 further suffer from significant score hallucinations.Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work. https://github.com/YebowenHu/SportsGenAnalyze the team-player affiliations and play-by-play descriptions below to determine the total points scored by each team (player).Please explain your reasoning step by step and provide the final results in JSON format.Start with: {New York Knicks: 0, Denver Nuggets: 0} ({Andrea Bargnani: 0, Timofey Mozgov: 0, ... Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Wenlin Yao, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004 |
EMNLP | 2 |
| 2024 | WPO: Enhancing RLHF with Weighted Preference OptimizationabstractWenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wenxuan Zhou 0005, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu |
EMNLP | 6 |
| 2024 | Polarity Calibration for Opinion SummarizationabstractYuanyuan Lei, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Ruihong Huang, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanyuan Lei 0001, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Ruihong Huang, Dong Yu 0001 |
NAACL-HLT | 2 |
| 2024 | MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction TuningabstractFuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fuxiao Liu, Xiaoyang Wang 0001, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu 0001 |
NAACL-HLT | 5 |
| 2023 | Generating User-Engaging News HeadlinesabstractPengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang, Xiaoyang Wang, Hong Yu, Fei Liu, Dong Yu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Pengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang 0010, Xiaoyang Wang 0001, Hong Yu 0001, Fei Liu 0004, Dong Yu 0001 |
ACL (1) | 2 |
| 2023 | How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation MethodabstractWenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
EACL | 5 |
| 2023 | DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4abstractHuman preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values.Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics.Despite their significance, however, there has been limited research probing these pairwise or kwise comparisons.The collective impact and relative importance of factors such as output length, informativeness, fluency, and factual consistency are still not well understood.It is also unclear if there are other hidden factors influencing human judgments.In this paper, we conduct an in-depth examination of a collection of pairwise human judgments released by Ope-nAI.Utilizing the Bradley-Terry-Luce (BTL) model, we reveal the inherent preferences embedded in these human judgments.We find that the most favored factors vary across tasks and genres, whereas the least favored factors tend to be consistent, e.g., outputs are too brief, contain excessive off-focus content or hallucinated facts.Our findings have implications on the construction of balanced datasets in human preference evaluations, which is a crucial step in shaping the behaviors of future LLMs. Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Fei Liu 0004 |
EMNLP | 2 |
| 2023 | Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsabstractTraditional sentence embedding models encode sentences into vector representations to capture useful properties such as the semantic similarity between sentences.However, in addition to similarity, sentence semantics can also be interpreted via compositional operations such as sentence fusion or difference.It is unclear whether the compositional semantics of sentences can be directly reflected as compositional operations in the embedding space.To more effectively bridge the continuous embedding and discrete text spaces, we explore the plausibility of incorporating various compositional properties into the sentence embedding space that allows us to interpret embedding transformations as compositional sentence operations.We propose INTERSENT, an end-toend framework for learning interpretable sentence embeddings that supports compositional sentence operations in the embedding space.Our method optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddings.Experimental results demonstrate that our method significantly improves the interpretability of sentence embeddings on four textual generation tasks over existing approaches while maintaining strong performance on traditional semantic similarity tasks. 1 . James Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 0009, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 3 |
| 2022 | Towards Abstractive Grounded Summarization of Podcast TranscriptsabstractPodcasts have shown a recent rise in popularity.Summarization of podcasts is of practical benefit to both content providers and consumers.It helps people quickly decide whether they will listen to a podcast and/or reduces the cognitive load of content providers to write summaries.Nevertheless, podcast summarization faces significant challenges including factual inconsistencies of summaries with respect to the inputs.The problem is exacerbated by speech disfluencies and recognition errors in transcripts of spoken language.In this paper, we explore a novel abstractive summarization method to alleviate these issues.Our approach learns to produce an abstractive summary while grounding summary segments in specific regions of the transcript to allow for full inspection of summary details.We conduct a series of analyses of the proposed approach on a large podcast dataset and show that the approach can achieve promising results.Grounded summaries bring clear benefits in locating the summary and transcript segments that contain inconsistent information, and hence improve summarization quality in terms of automatic and human evaluation. Kaiqiang Song, Chen Li 0003, Xiaoyang Wang 0001, Dong Yu 0001, Fei Liu 0004 |
ACL (1) | 1 |
| 2022 | Toward Unifying Text Segmentation and Long Document SummarizationabstractText segmentation is important for signaling a document's structure.Without segmenting a long document into topically coherent sections, it is difficult for readers to comprehend the text, let alone find important information.The problem is only exacerbated by a lack of segmentation in transcripts of audio/video recordings.In this paper, we explore the role that section segmentation plays in extractive summarization of written and spoken documents.Our approach learns robust sentence representations by performing summarization and segmentation simultaneously, which is further enhanced by an optimization-based regularizer to promote selection of diverse summary sentences.We conduct experiments on multiple datasets ranging from scientific articles to spoken transcripts to evaluate the model's performance.Our findings suggest that the model can not only achieve state-of-the-art performance on publicly available benchmarks, but demonstrate better crossgenre transferability when equipped with text segmentation.We perform a series of analyses to quantify the impact of section segmentation on summarizing written and spoken documents of substantial length and complexity. Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang 0001, Fei Liu 0004, Dong Yu 0001 |
EMNLP | 2 |
| 2022 | Salience Allocation as Guidance for Abstractive SummarizationabstractFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 2 |
| 2022 | Meta-learning without data via Wasserstein distributionally-robust model fusionabstractExisting meta-learning works assume that each task has available training and testing data. However, there are many available pre-trained models without accessing their training data in practice. We often need a single model to solve different tasks simultaneously as this is much more convenient to deploy the models. Our work aims to meta-learn a model initialization from these pre-trained models without using corresponding training data. We name this challenging problem setting as Data-Free Learning To Learn (DFL2L). We propose a distributionally robust optimization (DRO) framework to learn a black-box model to fuse and compress all the pre-trained models into a single network to address this problem. To encourage good generalization to the unseen new tasks, the proposed DRO framework diversifies the learned task embedding associated with each pre-trained model to cover the diversity in the underlying training task distributions. A model initialization is sampled from the black-box network during meta-testing as the meta learned initialization. Extensive experiments on offline and online DFL2L settings and several real image datasets demonstrate the effectiveness of the proposed methods. Zhenyi Wang 0001, Xiaoyang Wang 0001, Li Shen 0008, Qiuling Suo, Kaiqiang Song, Dong Yu 0001, Yan Shen 0002, Mingchen Gao |
UAI | 5 |
| 2021 | CATE: Computation-aware Neural Architecture Encoding with TransformersabstractRecent works (White et al., 2020a; Yan et al., 2020) demonstrate the importance of architecture encodings in Neural Architecture Search (NAS). These encodings encode either structure or computation information of the neural architectures. Compared to structure-aware encodings, computation-aware encodings map architectures with similar accuracies to the same region, which improves the downstream architecture search performance (Zhang et al., 2019; White et al., 2020a). In this work, we introduce a Computation-Aware Transformer-based Encoding method called CATE. Different from existing computation-aware encodings based on fixed transformation (e.g. path encoding), CATE employs a pairwise pre-training scheme to learn computation-aware encodings using Transformers with cross-attention. Such learned encodings contain dense and contextualized computation information of neural architectures. We compare CATE with eleven encodings under three major encoding-dependent NAS subroutines in both small and large search spaces. Our experiments show that CATE is beneficial to the downstream search, especially in the large search space. Moreover, the outside search space experiment demonstrates its superior generalization ability beyond the search space on which it was trained. Our code is available at: https://github.com/MSU-MLSys-Lab/CATE. Shen Yan 0008, Kaiqiang Song, Fei Liu 0004, Mi Zhang 0002 |
ICML | 2 |
| 2021 | A New Approach to Overgenerating and Scoring Abstractive SummariesabstractWe propose a new approach to generate multiple variants of the target summary with diverse content and varying lengths, then score and select admissible ones according to users' needs.Abstractive summarizers trained on single reference summaries may struggle to produce outputs that achieve multiple desirable properties, i.e., capturing the most important information, being faithful to the original, grammatical and fluent.In this paper, we propose a two-staged strategy to generate a diverse set of candidate summaries from the source text in stage one, then score and select admissible ones in stage two.Importantly, our generator gives a precise control over the length of the summary, which is especially well-suited when space is limited.Our selectors are designed to predict the optimal summary length and put special emphasis on faithfulness to the original text.Both stages can be effectively trained, optimized and evaluated.Our experiments on benchmark summarization datasets suggest that this paradigm can achieve state-of-the-art performance. Kaiqiang Song, Zhe Feng 0003, Fei Liu 0004 |
NAACL-HLT | 1 |
| 2020 | Joint Parsing and Generation for Abstractive SummarizationabstractSentences produced by abstractive summarization systems can be ungrammatical and fail to preserve the original meanings, despite being locally fluent. In this paper we propose to remedy this problem by jointly generating a sentence and its syntactic dependency parse while performing abstraction. If generating a word can introduce an erroneous relation to the summary, the behavior must be discouraged. The proposed method thus holds promise for producing grammatical sentences and encouraging the summary to stay true-to-original. Our contributions of this work are twofold. First, we present a novel neural architecture for abstractive summarization that combines a sequential decoder with a tree-based decoder in a synchronized manner to generate a summary sentence and its syntactic parse. Secondly, we describe a novel human evaluation protocol to assess if, and to what extent, a summary remains true to its original meanings. We evaluate our method on a number of summarization datasets and demonstrate competitive results against strong baselines. Kaiqiang Song, Logan Lebanoff, Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Chen Li 0003, Dong Yu 0001, Fei Liu 0004 |
AAAI | 1 |
| 2020 | Controlling the Amount of Verbatim Copying in Abstractive SummarizationabstractAn abstract must not change the meaning of the original text. A single most effective way to achieve that is to increase the amount of copying while still allowing for text abstraction. Human editors can usually exercise control over copying, resulting in summaries that are more extractive than abstractive, or vice versa. However, it remains poorly understood whether modern neural abstractive summarizers can provide the same flexibility, i.e., learning from single reference summaries to generate multiple summary hypotheses with varying degrees of copying. In this paper, we present a neural summarization model that, by learning from single human abstracts, can produce a broad spectrum of summaries ranging from purely extractive to highly generative ones. We frame the task of summarization as language modeling and exploit alternative mechanisms to generate summary hypotheses. Our method allows for control over copying during both training and decoding stages of a neural summarization model. Through extensive experiments we illustrate the significance of our proposed method on controlling the amount of verbatim copying and achieve competitive results over strong baselines. Our analysis further reveals interesting and unobvious facts. Kaiqiang Song, Zhe Feng 0003, Fei Liu 0004 |
AAAI | 1 |
| 2020 | Better Highlighting: Creating Sub-Sentence Summary HighlightsabstractAmongst the best means to summarize is highlighting. In this paper, we aim to generate summary highlights to be overlaid on the original documents to make it easier for readers to sift through a large amount of text. The method allows summaries to be understood in context to prevent a summarizer from distorting the original meaning, of which abstractive summarizers usually fall short. In particular, we present a new method to produce self-contained highlights that are understandable on their own to avoid confusion. Our method combines determinantal point processes and deep contextualized representations to identify an optimal set of sub-sentence segments that are both important and non-redundant to form summary highlights. To demonstrate the flexibility and modeling power of our method, we conduct extensive experiments on summarization datasets. Our analysis provides evidence that highlighting is a promising avenue of research towards future summarization. Sangwoo Cho, Kaiqiang Song, Chen Li 0003, Dong Yu 0001, Hassan Foroosh, Fei Liu 0004 |
EMNLP (1) | 2 |
| 2019 | Scoring Sentence Singletons and Pairs for Abstractive SummarizationabstractWhen writing a summary, humans tend to choose content from one or two sentences and merge them into a single summary sentence.However, the mechanisms behind the selection of one or multiple source sentences remain poorly understood.Sentence fusion assumes multi-sentence input; yet sentence selection methods only work with single sentences and not combinations of them.There is thus a crucial gap between sentence selection and fusion to support summarizing by both compressing single sentences and fusing pairs.This paper attempts to bridge the gap by ranking sentence singletons and pairs together in a unified space.Our proposed framework attempts to model human methodology by selecting either a single sentence or a pair of sentences, then compressing or fusing the sentence(s) to produce a summary sentence.We conduct extensive experiments on both single-and multidocument summarization datasets and report findings on sentence selection and abstraction. Logan Lebanoff, Kaiqiang Song, Franck Dernoncourt, Doo Soon Kim, Seokhwan Kim, Walter Chang, Fei Liu 0004 |
ACL (1) | 2 |
| 2018 | Structure-Infused Copy Mechanisms for Abstractive SummarizationabstractSeq2seq learning has produced promising results on summarization. However, in many cases, system summaries still struggle to keep the meaning of the original intact. They may miss out important words or relations that play critical roles in the syntactic structure of source sentences. In this paper, we present structure-infused copy mechanisms to facilitate copying important words and relations from the source sentence to summary sentence. The approach naturally combines source dependency structure with the copy mechanism of an abstractive sentence summarizer. Experimental results demonstrate the effectiveness of incorporating source-side syntactic information in the system, and our proposed approach compares favorably to state-of-the-art methods. Kaiqiang Song, Fei Liu 0004 |
COLING | 1 |
| 2018 | Adapting the Neural Encoder-Decoder Framework from Single to Multi-Document SummarizationabstractGenerating a text abstract from a set of documents remains a challenging task.The neural encoder-decoder framework has recently been exploited to summarize single documents, but its success can in part be attributed to the availability of large parallel data automatically acquired from the Web.In contrast, parallel data for multi-document summarization are scarce and costly to obtain.There is a pressing need to adapt an encoder-decoder model trained on single-document summarization data to work with multiple-document input.In this paper, we present an initial investigation into a novel adaptation method.It exploits the maximal marginal relevance method to select representative sentences from multi-document input, and leverages an abstractive encoder-decoder model to fuse disparate sentences to an abstractive summary.The adaptation method is robust and itself requires no training data.Our system compares favorably to state-of-the-art extractive and abstractive approaches judged by automatic metrics and human assessors. Logan Lebanoff, Kaiqiang Song, Fei Liu 0004 |
EMNLP | 2 |