Xinnian Liang

dblp:275/9970 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
19since 2021 · last 2025
0000-0002-4744-6179ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
abstract
Chain-of-thought (CoT) has emerged as a powerful technique to elicit reasoning in large language models and improve a variety of downstream tasks. CoT mainly demonstrates excellent performance in English, but its usage in low-resource languages is constrained due to poor language generalization. To bridge the gap among different languages, we propose a cross-lingual instruction fine-tuning framework (xCoT) to transfer knowledge from high-resource languages to low-resource languages. Specifically, the multilingual instruction training data (xCoT-Instruct) is created to encourage the semantic alignment of multiple languages. We introduce cross-lingual in-context few-shot learning (xICL) to accelerate multilingual agreement in instruction tuning, where some fragments of source languages in examples are randomly substituted by their counterpart translations of target languages. During multilingual instruction tuning, we adopt the randomly online CoT strategy to enhance the multilingual reasoning ability of the large language model by first translating the query to another language and then answering in English. To further facilitate the language transfer, we leverage the high-resource CoT to supervise the training of low-resource languages with cross-lingual distillation. Experimental results demonstrate the superior performance of xCoT in reducing the gap among different languages, highlighting its potential to reduce the cross-lingual gap.
Linzheng Chai, Jian Yang 0030, Tao Sun 0016, Hongcheng Guo, Xinnian Liang, Jiaqi Bai 0001, Tongliang Li, Qiyao Peng 0001, Zhoujun Li 0001
AAAI7
2025 MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL
abstract
Recent LLM-based Text-to-SQL methods usually suffer from significant performance degradation on “huge” databases and complex user questions that require multi-step reasoning. Moreover, most existing methods neglect the crucial significance of LLMs utilizing external tools and model collaboration. To address these challenges, we introduce MAC-SQL, a novel LLM-based multi-agent collaborative framework. Our framework comprises a core decomposer agent for Text-to-SQL generation with few-shot chain-of-thought reasoning, accompanied by two auxiliary agents that utilize external tools or models to acquire smaller sub-databases and refine erroneous SQL queries. The decomposer agent collaborates with auxiliary agents, which are activated as needed and can be expanded to accommodate new features or tools for effective Text-to-SQL parsing. In our framework, We initially leverage GPT-4 as the strong backbone LLM for all agent tasks to determine the upper bound of our framework. We then fine-tune an open-sourced instruction-followed model, SQL-Llama, by leveraging Code Llama 7B, to accomplish all tasks as GPT-4 does. Experiments show that SQL-Llama achieves a comparable execution accuracy of 43.94, compared to the baseline accuracy of 46.35 for vanilla GPT-4. At the time of writing, MAC-SQL+GPT-4 achieves an execution accuracy of 59.59 when evaluated on the BIRD benchmark, establishing a new state-of-the-art (SOTA) on its holdout test set.
Changyu Ren, Jian Yang 0030, Xinnian Liang, Jiaqi Bai 0001, Linzheng Chai, Qian-Wen Zhang, Xing Sun 0001, Zhoujun Li 0001
COLING4
2025 SCM: Enhancing Large Language Model with Self-Controlled Memory Framework
Xinnian Liang, Jian Yang 0003, Hui Huang 0021, Zhenhe Wu, Shuangzhi Wu, Zejun Ma 0001, Zhoujun Li 0001
DASFAA (6)2
2025 One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
abstract
Leveraging mathematical Large Language Models (LLMs) for proof generation is a fundamental topic in LLMs research. We argue that the ability of current LLMs to prove statements largely depends on whether they have encountered the relevant proof process during training. This reliance limits their deeper understanding of mathematical theorems and related concepts. Inspired by the pedagogical method of "proof by counterexamples" commonly used in human mathematics education, our work aims to enhance LLMs’ ability to conduct mathematical reasoning and proof through counterexamples. Specifically, we manually create a high-quality, university-level mathematical benchmark, COUNTERMATH, which requires LLMs to prove mathematical statements by providing counterexamples, thereby assessing their grasp of mathematical concepts. Additionally, we develop a data engineering framework to automatically obtain training data for further model improvement. Extensive experiments and detailed analyses demonstrate that COUNTERMATH is challenging, indicating that LLMs, such as OpenAI o1, have insufficient counterexample-driven proof capabilities. Moreover, our exploration into model training reveals that strengthening LLMs’ counterexample-driven conceptual reasoning abilities is crucial for improving their overall mathematical capabilities. We believe that our work offers new perspectives on the community of mathematical LLMs.
Jiayi Kuang, Haojing Huang 0001, Zhikun Xu, Xinnian Liang, Wenlian Lu, Yangning Li, Xiaoyu Tan, Chao Qu, Ying Shen 0001, Hai-Tao Zheng 0002, Philip S. Yu
ICML5
2025 Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
abstract
Large Language Models (LLMs) have demonstrated outstanding performance in mathematical reasoning capabilities. However, we argue that current large-scale reasoning models primarily rely on scaling up training datasets with diverse mathematical problems and long thinking chains, which raises questions about whether LLMs genuinely acquire mathematical concepts and reasoning principles or merely remember the training data. In contrast, humans tend to break down complex problems into multiple fundamental atomic capabilities. Inspired by this, we propose a new paradigm for evaluating mathematical atomic capabilities. Our work categorizes atomic abilities into two dimensions: (1) field-specific abilities across four major mathematical fields, algebra, geometry, analysis, and topology, and (2) logical abilities at different levels, including conceptual understanding, forward multi-step reasoning with formal math language, and counterexample-driven backward reasoning. We propose corresponding training and evaluation datasets for each atomic capability unit, and conduct extensive experiments about how different atomic capabilities influence others, to explore the strategies to elicit the required specific atomic capability. Evaluation and experimental results on advanced models show many interesting discoveries and inspirations about the different performances of models on various atomic capabilities and the interactions between atomic capabilities. Our findings highlight the importance of decoupling mathematical intelligence into atomic components, providing new insights into model cognition and guiding the development of training strategies toward a more efficient, transferable, and cognitively grounded paradigm of "atomic thinking".
Jiayi Kuang, Haojing Huang 0001, Xinnian Liang, Zhikun Xu, Yangning Li, Xiaoyu Tan, Chao Qu, Meishan Zhang, Ying Shen 0001, Philip S. Yu
NeurIPS4
2024 m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt
abstract
Multilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario.
Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001
LREC/COLING7
2024 Align vision-language semantics by multi-task learning for multi-modal summarization
Chenhao Cui, Xinnian Liang, Shuangzhi Wu, Zhoujun Li 0001
Neural Comput. Appl.2
2023 GripRank: Bridging the Gap between Retrieval and Generation via the Generative Knowledge Improved Passage Ranking
abstract
Retrieval-enhanced text generation has shown remarkable progress on knowledge-intensive language tasks, such as open-domain question answering and knowledge-enhanced dialogue generation, by leveraging passages retrieved from a large passage corpus for delivering a proper answer given the input query. However, the retrieved passages are not ideal for guiding answer generation because of the discrepancy between retrieval and generation, i.e., the candidate passages are all treated equally during the retrieval procedure without considering their potential to generate a proper answer. This discrepancy makes a passage retriever deliver a sub-optimal collection of candidate passages to generate the answer. In this paper, we propose the GeneRative Knowledge Improved Passage Ranking (GripRank) approach, addressing the above challenge by distilling knowledge from a generative passage estimator (GPE) to a passage ranker, where the GPE is a generative language model used to measure how likely the candidate passages can generate the proper answer. We realize the distillation procedure by teaching the passage ranker learning to rank the passages ordered by the GPE. Furthermore, we improve the distillation quality by devising a curriculum knowledge distillation mechanism, which allows the knowledge provided by the GPE can be progressively distilled to the ranker through an easy-to-hard curriculum, enabling the passage ranker to correctly recognize the provenance of the answer from many plausible candidates. We conduct extensive experiments on four datasets across three knowledge-intensive language tasks. Experimental results show advantages over the state-of-the-art methods for both passage ranking and answer generation on the KILT benchmark.
Jiaqi Bai 0001, Hongcheng Guo, Jian Yang 0030, Xinnian Liang, Zhoujun Li 0001
CIKM5
2023 Enhancing Dialogue Summarization with Topic-Aware Global- and Local- Level Centrality
abstract
Dialogue summarization aims to condense a given dialogue into a simple and focused summary text.Typically, both the roles' viewpoints and conversational topics change in the dialogue stream.Thus how to effectively handle the shifting topics and select the most salient utterance becomes one of the major challenges of this task.In this paper, we propose a novel topic-aware Global-Local Centrality (GLC) model to help select the salient context from all sub-topics.The centralities are constructed at both the global and local levels.The global one aims to identify vital sub-topics in the dialogue and the local one aims to select the most important context in each sub-topic.Specifically, the GLC collects sub-topic based on the utterance representations.And each utterance is aligned with one sub-topic.Based on the sub-topics, the GLC calculates globaland local-level centralities.Finally, we combine the two to guide the model to capture both salient context and sub-topics when generating summaries.Experimental results show that our model outperforms strong baselines on three public dialogue summarization datasets: CSDS, MC, and SAMSUM.Further analysis demonstrates that our GLC can exactly identify vital contents from sub-topics. 1
Xinnian Liang, Shuangzhi Wu, Chenhao Cui, Jiaqi Bai 0001, Chao Bian 0006, Zhoujun Li 0001
EACL1
2023 Learning Unified Video-Language Representations via Joint Modeling and Contrastive Learning for Natural Language Video Localization
abstract
Natural language video localization (NLVL) aims to locate the matching span relevant to a given query sentence from an untrimmed video. This task requires not only understanding video and text but also aligning the semantics between video and language. Existing methods obtain vision-language representations via separate encoders, cross-modal interactions are not fine-grained enough, and the semantics are not fully aligned. In this paper, we address the vision-language alignment via joint modeling and contrastive learning. We propose a unified Video-Language Representation Network (UniNet), employing a transformer encoder to learn vision-language representations aligned. Simultaneously taking video and text as input, the encoder jointly learns the representations of both and captures the inter-relations between video and text. Then the representations are used by the predictor to locate the grounding video span. Besides, we train our model with contrastive learning to enhance vision-language representations in the training stage. Experiments on three benchmark datasets show that UniNet outperforms the baseline methods and adopting unified representation and contrastive learning can improve vision-language semantic alignment.
Chenhao Cui, Xinnian Liang, Shuangzhi Wu, Zhoujun Li 0001
IJCNN2
2023 Towards Making the Most of LLM for Translation Quality Estimation
Hui Huang 0021, Shuangzhi Wu, Xinnian Liang, Yanrui Shi, Peihao Wu, Muyun Yang, Tiejun Zhao
NLPCC (1)3
2023 KnowPrefix-Tuning: A Two-Stage Prefix-Tuning Framework for Knowledge-Grounded Dialogue Generation
Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Xinnian Liang, Hongcheng Guo, Zhoujun Li 0001
ECML/PKDD (2)5
2023 Unsupervised abstractive summarization via sentence rewriting
Zhihao Zhang 0004, Xinnian Liang, Yuan Zuo, Zhoujun Li 0001
Comput. Speech Lang.2
2023 Improving unsupervised keyphrase extraction by modeling hierarchical multi-granularity features
abstract
Existing unsupervised keyphrase extraction methods typically emphasize the importance of the candidate keyphrase itself, ignoring other important factors such as the influence of uninformative sentences. We hypothesize that the salient sentences of a document are particularly important as they are most likely to contain keyphrases, especially for long documents. To our knowledge, our work is the first attempt to exploit sentence salience for unsupervised keyphrase extraction by modeling hierarchical multi-granularity features. Specifically, we propose a novel position-aware graph-based unsupervised keyphrase extraction model, which includes two model variants. The pipeline model first extracts salient sentences from the document, followed by keyphrase extraction from the extracted salient sentences. In contrast to the pipeline model which models multi-granularity features in a two-stage paradigm, the joint model accounts for both sentence and phrase representations of the source document simultaneously via hierarchical graphs. Concretely, the sentence nodes are introduced as an inductive bias, injecting sentence-level information for determining the importance of candidate keyphrases. We compare our model against strong baselines on three benchmark datasets including Inspec, DUC 2001, and SemEval 2010. Experimental results show that the simple pipeline-based approach achieves promising results, indicating that keyphrase extraction task benefits from the salient sentence extraction task. The joint model, which mitigates the potential accumulated error of the pipeline model, gives the best performance and achieves new state-of-the-art results while generalizing better on data from different domains and with different lengths. In particular, for the SemEval 2010 dataset consisting of long documents, our joint model outperforms the strongest baseline UKERank by 3.48%, 3.69% and 4.84% in terms of [email protected], [email protected] and [email protected], respectively. We also conduct qualitative experiments to validate the effectiveness of our model components.
Zhihao Zhang 0004, Xinnian Liang, Yuan Zuo, Chenghua Lin 0002
Inf. Process. Manag.2
2022 An Efficient Coarse-to-Fine Facet-Aware Unsupervised Summarization Framework Based on Semantic Blocks
abstract
Unsupervised summarization methods have achieved remarkable results by incorporating representations from pre-trained language models. However, existing methods fail to consider efficiency and effectiveness at the same time when the input document is extremely long. To tackle this problem, in this paper, we proposed an efficient Coarse-to-Fine Facet-Aware Ranking (C2F-FAR) framework for unsupervised long document summarization, which is based on the semantic block. The semantic block refers to continuous sentences in the document that describe the same facet. Specifically, we address this problem by converting the one-step ranking method into the hierarchical multi-granularity two-stage ranking. In the coarse-level stage, we proposed a new segment algorithm to split the document into facet-aware semantic blocks and then filter insignificant blocks. In the fine-level stage, we select salient sentences in each block and then extract the final summary from selected sentences. We evaluate our framework on four long document summarization datasets: Gov-Report, BillSum, arXiv, and PubMed. Our C2F-FAR can achieve new state-of-the-art unsupervised summarization results on Gov-Report and BillSum. In addition, our method speeds up 4-28 times more than previous methods.
Xinnian Liang, Shuangzhi Wu, Jiali Zeng, Yufan Jiang, Mu Li 0001, Zhoujun Li 0001
COLING1
2022 Modeling Multi-Granularity Hierarchical Features for Relation Extraction
abstract
Relation extraction is a key task in Natural Language Processing (NLP), which aims to extract relations between entity pairs from given texts.Recently, relation extraction (RE) has achieved remarkable progress with the development of deep neural networks.Most existing research focuses on constructing explicit structured features using external knowledge such as knowledge graph and dependency tree.In this paper, we propose a novel method to extract multi-granularity features based solely on the original input sentences.We show that effective structured features can be attained even without external knowledge.Three kinds of features based on the input sentences are fully exploited, which are in entity mention level, segment level, and sentence level.All the three are jointly and hierarchically modeled.We evaluate our method on three public benchmarks: SemEval 2010 Task 8, Tacred, and Tacred Revisited.To verify the effectiveness, we apply our method to different encoders such as LSTM and BERT.Experimental results show that our method significantly outperforms existing state-of-the-art models that even use external knowledge.Extensive analyses demonstrate that the performance of our model is contributed by the capture of multi-granularity features and the model of their hierarchical structure.
Xinnian Liang, Shuangzhi Wu, Mu Li 0001, Zhoujun Li 0001
NAACL-HLT1
2022 Improving Unsupervised Extractive Summarization by Jointly Modeling Facet and Redundancy
abstract
Unsupervised extractive summarization aims to extract salient sentences from documents without labeled corpus. Existing methods are mostly graph-based by computing sentence centrality. These methods have two main problems: facet bias and redundant problems. Facet bias problem leads summarization models tend to select sentences within the same facet, which often leads to the ignoring of other vital facets, especially on long-document and multi-documents. First, to address the facet bias problem, we proposed a novel Facet-Aware centrality-based Ranking model (FAR). We let the model pay more attention to different facets by introducing a sentence-document weight. The weight is added to the sentence centrality score. FAR can alleviate redundancy to some extent. Then, to further reduce redundancy, we proposed a novel Redundancy- and Facet-Aware Ranking model (RFAR) which jointly models facet and redundancy by incorporating Determinantal Point Process (DPP) into the previous proposed FAR. We evaluate our FAR and RFAR on a wide range of summarization tasks that include 8 representative benchmark datasets. Experimental results show that FAR and RFAR consistently outperforms strong baselines, especially in long- and multi-document scenarios, and even perform comparably to some supervised models. Besides, we find that our methods can alleviate the position bias problem.
Xinnian Liang, Shuangzhi Wu, Mu Li 0001, Zhoujun Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Learning to Copy Coherent Knowledge for Response Generation
abstract
Knowledge-driven dialog has shown remarkable performance to alleviate the problem of generating uninformative responses in the dialog system. However, incorporating knowledge coherently and accurately into response generation is still far from being solved. Previous works dropped into the paradigm of non-goal-oriented knowledge-driven dialog, they are prone to ignore the effect of dialog goal, which has potential impacts on knowledge exploitation and response generation. To address this problem, this paper proposes a Goal-Oriented Knowledge Copy network, GOKC. Specifically, a goal-oriented knowledge discernment mechanism is designed to help the model discern the knowledge facts that are highly correlated to the dialog goal and the dialog context. Besides, a context manager is devised to copy facts not only from the discerned knowledge but also from the dialog goal and the dialog context, which allows the model to accurately restate the facts in the generated response. The empirical studies are conducted on two benchmarks of goal-oriented knowledge-driven dialog generation. The results show that our model can significantly outperform several state-of-the-art models in terms of both automatic evaluation and human judgments.
Jiaqi Bai 0001, Ze Yang 0001, Xinnian Liang, Wei Wang 0301, Zhoujun Li 0001
AAAI3
2021 Unsupervised Keyphrase Extraction by Jointly Modeling Local and Global Context
abstract
Embedding based methods are widely used for unsupervised keyphrase extraction (UKE) tasks.Generally, these methods simply calculate similarities between phrase embeddings and document embedding, which is insufficient to capture different context for a more effective UKE model.In this paper, we propose a novel method for UKE, where local and global contexts are jointly modeled.From a global view, we calculate the similarity between a certain phrase and the whole document in the vector space as transitional embedding based models do.In terms of the local view, we first build a graph structure based on the document where phrases are regarded as vertices and the edges are similarities between vertices.Then, we proposed a new centrality computation method to capture local salient information based on the graph structure.Finally, we further combine the modeling of global and local context for ranking.We evaluate our models on three public benchmarks (Inspec, DUC 2001, SemEval 2010) and compare with existing state-of-the-art models.The results show that our model outperforms most models while generalizing better on input documents with different domains and length.Additional ablation study shows that both the local and global information is crucial for unsupervised keyphrase extraction tasks.
Xinnian Liang, Shuangzhi Wu, Mu Li 0001, Zhoujun Li 0001
EMNLP (1)1