Tian Lan 0003

dblp:31/83-3 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-5200-1537ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 WikiREVIEW: A Multi-Perspective Review Framework for Automatic Wiki-Style Article Generation
abstract
As a knowledge-intensive and challenging task, automatic generation of long-form wiki-style articles has garnered increasing attention from researchers due to its ability to efficiently integrate, organize and present vast amounts of both structured and unstructured knowledge. To the best of our knowledge, most of the existing mainstream state-of-the-art methods for automatic wiki-style article generation typically follow a "one-shot generation" paradigm: given a topic, (1) first generating a structured outline, (2) then independently and in parallel generating the content of each outline chapter in a one-shot using the chapter title and references. However, the core limitation of the paradigm lies in its disregards inter-chapter correlation and lacks post-generation revision and refinement, resulting in content redundancy, weak relevance and logical inconsistency. To address these issues, we propose WikiREVIEW, a novel multi-perspective review framework for automatic wiki-style article generation. Specifically, our proposed method introduces multi-perspective experts to review the content of each outline chapter at both chapter and paragraph levels following the initial generation, offering evaluation feedback and continuously refining the numerous deficiencies in the initial long-form article, ultimately achieving high-quality wiki-style article generation. Extensive experimental results on the public English dataset FreshWiki and our own constructed high-quality Chinese dataset ChineseWiki, demonstrate that our proposed WikiREVIEW significantly outperforms existing state-of-the-art automatic wiki-style article generation methods across all automatic evaluation metrics and human evaluation.
Guo-Biao Zhang, Zhijing Wu 0001, Tian Lan 0003, Ding-Yuan Liu, Yu-Shi Zhu, Xianling Mao
AAAI3
2026 Your Reasoning Model Knows What Counts: Self-Guided Chain-of-Thought Pruning for Efficient Reasoning
abstract
Chain-of-Thought (CoT) reasoning is crucial for the performance of Large Reasoning Models (LRMs) but is often hindered by redundant and distracting segments, which incur excessive inference costs and degrade robustness.Existing approaches try to solve this problem by enforcing brevity through external supervision, such as length-based penalties or heuristic truncation.However, these approaches often degrade performance because they disregard the model's intrinsic reasoning dependency and thus fail to distinguish between essential and redundant CoT segments.To address this problem, we propose SGP-CoT, a novel Self-Guided Pruning framework that leverages the model's intrinsic likelihood landscape to identify segments that are extraneous to its specific reasoning pattern.Specifically, SGP-CoT treats the reasoning trajectory as a sequence of semantic units and assesses the necessity of each one via internal likelihood signals, measuring its contribution to the answer and local coherence.Based on this, it selectively removes non-essential segments and then forms high-quality pruning-based preference pairs, enabling the model to learn focused reasoning via self-optimization.Extensive experiments across diverse benchmarks demonstrate that the proposed SGP-CoT significantly reduces output length while maintaining or improving accuracy.These results validate that LRMs intrinsically possess the capability to discern reasoning utility, positioning SGP-CoT as a robust pathway toward scalable inference.
Zi-Ao Ma, Xianling Mao, Tian Lan 0003, Zhijing Wu 0001
ACL (1)3
2026 PUPPET: Neural-Symbolic Standardized Patients for Mental Health
abstract
Chen Xu, Yu ji, Zhenyu Lv, Yang Yi, Yizhe Yang, Luyao Ji, Chaoyi Chen, Xianyang Wang, Tian Lan, Zhihua Wang, Juan Wang, Xunde Dong, Fuze Tian, Qunxi Dong, Bin Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenyu Lv, Yizhe Yang, Luyao Ji, Chaoyi Chen, Xianyang Wang, Tian Lan 0003, Xunde Dong, Fuze Tian, Qunxi Dong, Bin Hu 0001
ACL (1)9
2026 MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
abstract
Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, Jingyun Liao, YiMing Cheng, Xuefeng Chen, Xian-Ling Mao, Yousheng Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan 0003, Ziqin Zhou, Jingyun Liao, YiMing Cheng, Xianling Mao, Yousheng Feng
ACL (1)7
2026 Bridging the gap between data distribution and model: Dynamic data distribution optimization for improving critique capabilities of large language models
abstract
Critique ability, defined as the capacity to identify and rectify flaws in text generation, is crucial for the applications of Large Language Models (LLMs). As a meta-cognitive capability, enhancing the critique ability of LLMs poses significant challenges. Recent studies have proposed improving this ability through fine-tuning on critique datasets. However, the static data distribution of existing datasets often leads to a mismatch between the training data and the diverse optimization needs of target models, thereby hindering their effectiveness. To address this issue, we introduce a novel Dynamic Iterative Data Distribution Optimization Method (DIDD) that dynamically adjusts training data distributions to align with the specific optimization requirements of target models. Specifically, DIDD detects the vulnerable data distribution of target optimization models by conducting the meta-critique on synthesized test set. The detected vulnerable data distribution are then leveraged to construct the training dataset that aligns with target model more closely, improving the effectiveness of the training dataset. Extensive experimental results across four benchmarks demonstrate that our proposed DIDD effectively alleviates the mismatch between the training dataset and target optimization models.
Tian Lan 0003, Zhenyu Lv, Qunxi Dong, Jieshuo Zhang, Heyan Huang, Minqiang Yang, Bin Hu 0001
Expert Syst. Appl.2
2025 SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
abstract
Automatic evaluation for Open Domain Event Detection (ODED) is a highly challenging task, because ODED is characterized by a vast diversity of un-constrained output labels from various domains. Nearly all existing evaluation methods for ODED usually first construct evaluation benchmarks with limited labels and domain coverage, and then evaluate ODED methods using metrics based on token-level label matching rules. However, this kind of evaluation framework faces two issues: (1) The limited evaluation benchmarks lack representatives of the real world, making it difficult to accurately reflect the performance of various ODED methods in real-world scenarios; (2) Evaluation metrics based on token-level matching rules fail to capture semantic similarity between predictions and golden labels. To address these two problems above, we propose a scalable and reliable Semantic-level Evaluation framework for Open domain Event detection (SEOE) by constructing a more representative evaluation benchmark and introducing a semantic evaluation metric. Specifically, our proposed framework first constructs a scalable evaluation benchmark that currently includes 564 event types covering 7 major domains, with a cost-effective supplementary annotation strategy to ensure the benchmark’s representativeness. The strategy also allows for the supplement of new event types and domains in the future. Then, the proposed SEOE leverages large language models (LLMs) as automatic evaluation agents to compute a semantic F1-score, incorporating fine-grained definitions of semantically similar labels to enhance the reliability of the evaluation. Extensive experiments validate the representatives of the benchmark and the reliability of the semantic evaluation metric. Existing ODED methods are thoroughly evaluated, and the error patterns of predictions are analyzed, revealing several insightful findings.
Yi-Fan Lu, Xianling Mao, Tian Lan 0003, Yu-Shi Zhu, Heyan Huang
ACL (1)3
2025 Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
abstract
Driven by the remarkable progress in diffusion models, text-to-image generation has achieved substantial advancements, underscoring the urgent need for robust automatic quality assessment.This task is inherently complex, requiring evaluations that range from object presence and attribute correctness to relational consistency and visual fidelity.Consequently, current state-of-the-art MLLM-based approaches often rely on powerful commercial models such as GPT-4o, which offer superior reasoning and instruction-following capabilities but are not universally accessible.In contrast, while opensource MLLMs demonstrate promising skills in vision and language understanding, they underperform in comprehensive image quality assessment.To address these challenges, we propose a task decomposition evaluation framework based on GPT-4o to automatically construct a specialized training dataset, breaking down the multifaceted evaluation process into simpler sub-tasks and thus reducing learning complexity.Building on this dataset, we design novel training strategies to distill GPT-4o's evaluation capabilities into a 7B open-source MLLM, MiniCPM-V-2.6,enabling it to better follow instructions across diverse assessment criteria.Furthermore, to reliably and comprehensively assess prior works and our proposed model, we manually annotate a meta-evaluation benchmark that includes chain-of-thought explanations alongside quality scores for generated images.Experimental results demonstrate that our distilled open-source MLLM significantly outperforms the current state-of-the-art GPT-4o-base baseline, VIEScore, with over 4.6% improvement in Spearman and Kendall correlations with human judgments.
Rongcheng Tu, Zi-Ao Ma, Tian Lan 0003, Yuehao Zhao, Heyan Huang, Xianling Mao
ACL (1)3
2025 Automatic Evaluating Scientific Reviews Through Meta-reviewer's Lens: A Reliable Benchmark for Peer Review Generation
Shu-Hang Liu, Tian Lan 0003, Yun-He Zhang, Heyan Huang, Zhijing Wu 0001, Xianling Mao
NLPCC (2)2
2025 Prospective Layout-Guided Multi-Modal Online Hashing
abstract
In real-world scenarios, the data usually appears in a streaming fashion. To achieve remarkable retrieval performance in such scenarios, online multi-modal hashing has drawn great research attention due to its high retrieval speed and low storage cost. However, existing online multi-modal hashing methods still fail to achieve satisfactory retrieval performance in the scenarios where the new streaming datapoints all belong to the new classes. Therefore, to further improve the retrieval performance in these scenarios, we propose a novel Prospective Layout-Guided Multi-modal Online Hashing, termed PLG-MOH. Specifically, PLG-MOH first establishes the layout of the Hamming space by generating a series of hashing centers to split the space. Each hashing center will be gradually assigned to a new appearing class, and these assigned centers correspond one-to-one with the classes. Moreover, we propose a novel prospective layout-guided loss, which leverages all the hashing centers, including those not yet assigned to the classes, to supervise the training of hashing model. As the unassigned hashing centers will be designated to the new classes emerging in the future, it signifies that during each round of training, PLG-MOH has already considered the forthcoming data from new classes in the future rounds. Consequently, PLG-MOH can effectively adapt its hashing functions to address the new arriving samples and learn semantic similarity-preserved hash codes for them, meanwhile it can effectively retain the information learned from the old data. Extensive experiments on two public datasets demonstrate that the proposed PLG-MOH achieves better retrieval performance than state-of-the-art baselines on online scenarios.
Rongcheng Tu, Xianling Mao, Jin-Yu Liu, Zi-Ao Ma, Tian Lan 0003, Heyan Huang
IEEE Trans. Image Process.5
2025 Decider: A Dual-System Rule-Controllable Decoding Framework for Language Generation
abstract
Constrained decoding approaches aim to control the meaning or style of text generated by a Pre-trained Language Model (PLM) for various task-specific objectives at inference time. However, these methods often guide plausible continuations by greedily and explicitly selecting targets, which, while fulfilling the task requirements, may overlook the natural patterns of human language generation. In this work, we propose a novel decoding framework,Decider, which enables us to program high-level rules on how we might effectively complete tasks to control a PLM. Differing from previous works, our framework transforms the encouragement of concrete target words into the encouragement of all words that satisfy the high-level rules. Specifically,Decideris a dual system in which a PLM is equipped and controlled by a First-Order Logic (FOL) reasoner to express and evaluate the rules, along with a decision function that merges the outputs from both systems to guide the generation. Experiments on CommonGen and PersonaChat demonstrate thatDecidercan effectively follow given rules to guide a PLM in achieving generation tasks in a more human-like manner.
Tian Lan 0003, Changlong Yu, Wei Wang 0138, Qunxi Dong, Kun Qian 0003, Piji Li, Wei Bi, Bin Hu 0001
IEEE Trans. Knowl. Data Eng.2
2024 A Hierarchical Context Augmentation Method to Improve Retrieval-Augmented LLMs on Scientific Papers
abstract
Scientific papers of a large scale on the Internet encompass a wealth of data and knowledge, attracting the attention of numerous researchers. To fully utilize these knowledge, Retrieval-Augmented Large Language Models (LLMs) usually leverage large-scale scientific corpus to train and then retrieve relevant passages from external memory to improve generation, which have demonstrated outstanding performance. However, existing methods can only capture one-dimension fragmented textual information without incorporating hierarchical structural knowledge, eg. the deduction relationship of abstract and main body, which makes it difficult to grasp the central thought of papers. To tackle this problem, we propose a hierarchical context augmentation method, which helps Retrieval-Augmented LLMs to autoregressively learn the structure knowledge of scientific papers. Specifically, we utilize the document tree to represent the hierarchical relationship of a paper and enhance the structure information of scientific context from three aspects: scale, format and global information. First, we think each top-bottom path of document tree is a logical independent context, which can be used to largely increase the scale of extracted structural corpus. Second, we propose a novel label-based format to represent the structure of context in textual sequences, unified between training and inference. Third, we introduce the global information of retrieved passages to further enhance the structure of context. Extensive experiments on three scientific tasks show that the proposed method significantly improves the performance of Retrieval-Augmented LLMs on all tasks. Besides, our method achieves start-of-art performance in Question Answer task and outperforms ChatGPT. Moreover, it also brings considerate gains with irrelevant retrieval passages, illustrating its effectiveness on practical application scenarios.
Tian-Yi Che, Xianling Mao, Tian Lan 0003, Heyan Huang
KDD3
2024 CriticEval: Evaluating Large-scale Language Model as Critic
abstract
Critique ability, i.e., the capability of Large Language Models (LLMs) to identify and rectify flaws in responses, is crucial for their applications in self-improvement and scalable oversight. While numerous studies have been proposed to evaluate critique ability of LLMs, their comprehensiveness and reliability are still limited. To overcome this problem, we introduce CriticEval, a novel benchmark designed to comprehensively and reliably evaluate critique ability of LLMs. Specifically, to ensure the comprehensiveness, CriticEval evaluates critique ability from four dimensions across nine diverse task scenarios. It evaluates both scalar-valued and textual critiques, targeting responses of varying quality. To ensure the reliability, a large number of critiques are annotated to serve as references, enabling GPT-4 to evaluate textual critiques reliably. Extensive evaluations of open-source and closed-source LLMs first validate the reliability of evaluation in CriticEval. Then, experimental results demonstrate the promising potential of open-source LLMs, the effectiveness of critique datasets and several intriguing relationships between the critique ability and some critical factors, including task types, response qualities and critique dimensions.
Tian Lan 0003, Heyan Huang, Dahua Lin, Kai Chen 0026, Xianling Mao
NeurIPS1
2024 Exploring Dense Retrieval for Dialogue Response Selection
abstract
Recent progress in deep learning has continuously improved the accuracy of dialogue response selection. However, in real-world scenarios, the high computation cost forces existing dialogue response selection models to rank only a small number of candidates, recalled by a coarse-grained model, precluding many high-quality candidates. To overcome this problem, we present a novel and efficient response selection model and a set of tailor-designed learning strategies to train it effectively. The proposed model consists of a dense retrieval module and an interaction layer, which could directly select the proper response from a large corpus. We conduct re-rank and full-rank evaluations on widely used benchmarks to evaluate our proposed model. Extensive experimental results demonstrate that our proposed model notably outperforms the state-of-the-art baselines on both re-rank and full-rank evaluations. Moreover, human evaluation results show that the response quality could be improved further by enlarging the candidate pool with nonparallel corpora. In addition, we also release high-quality benchmarks that are carefully annotated for more accurate dialogue response selection evaluation. All source codes, datasets, model parameters, and other related resources have been publicly available. 1
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Yixuan Su, Heyan Huang, Xianling Mao
ACM Trans. Inf. Syst.1
2024 Towards Efficient Coarse-grained Dialogue Response Selection
abstract
Coarse-grained response selection is a fundamental and essential subsystem for the widely used retrieval-based chatbots, aiming to recall a coarse-grained candidate set from a large-scale dataset. The dense retrieval technique has recently been proven very effective in building such a subsystem. However, dialogue dense retrieval models face two problems in real scenarios: (1) the multi-turn dialogue history is re-computed in each turn, leading to inefficient inference; (2) the index storage of the offline index is enormous, significantly increasing the deployment cost. To address these problems, we propose an efficient coarse-grained response selection subsystem consisting of two novel methods. Specifically, to address the first problem, we propose the H ierarchical D ense R etrieval. It caches rich multi-vector representations of the dialogue history and only encodes the latest user’s utterance, leading to better inference efficiency. Then, to address the second problem, we design the D eep S emantic H ashing to reduce the index storage while effectively saving its recall accuracy notably. Extensive experimental results prove the advantages of the two proposed methods over previous works. Specifically, with the limited performance loss, our proposed coarse-grained response selection model achieves over 5x FLOPs speedup and over 192x storage compression ratio. Moreover, our source codes have been publicly released. 1
Tian Lan 0003, Xianling Mao, Wei Wei 0002, Xiaoyan Gao 0001, Heyan Huang
ACM Trans. Inf. Syst.1
2023 Copy is All You Need
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Heyan Huang, Xianling Mao
ICLR1
2023 Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective
abstract
There are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning.
Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su
NeurIPS2
2023 LASH: Large-Scale Academic Deep Semantic Hashing
abstract
With the explosively increasing of academic papers, efficient academic document retrieval is becoming an essential requirement for large-scale information retrieval systems. Inspired by the success of deep semantic hashing in normal document retrieval, deep semantic hashing is a promising approach for academic document retrieval by mapping academic documents into efficient hash codes. However, for academic document retrieval, the existing deep semantic hashing methods suffer from following two problems: (1) they cannot differentiate the importance of different field labels; (2) they cannot plenty utilize the structure information in paper citations. To address these problems, we propose a novel Large-scale Academic deep Semantic Hashing, called LASH. Specifically, LASH first treats paper citations as a citation network, and then employs a multi-input variational deep autoencoder to directly encode both structure information of the citation network and semantic information of academic documents into unified hash codes. Moreover, a weighted percentage similarity is designed to measure the importance of different field labels, which is a linear combination of Jaccard and Cosine similarity. Supervised by the similarity, the learned unified hash codes can further preserve the importance of different field labels. Extensive experiments show LASH significantly outperforms state-of-the-art baselines over proposed three real-world large-scale academic datasets.
Jia-Nan Guo, Xianling Mao, Tian Lan 0003, Rongxin Tu, Wei Wei 0002, Heyan Huang
IEEE Trans. Knowl. Data Eng.3
2022 Cross-Lingual Phrase Retrieval
abstract
Heqi Zheng, Xiao Zhang, Zewen Chi, Heyan Huang, Yan Tan, Tian Lan, Wei Wei, Xian-Ling Mao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Heqi Zheng, Xiao Zhang 0036, Zewen Chi, Heyan Huang, Tian Lan 0003, Wei Wei 0002, Xianling Mao
ACL (1)6
2022 A Contrastive Framework for Neural Text Generation
abstract
Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease the probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method---contrastive search---to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach outperforms state-of-the-art text generation methods as evaluated by both human and automatic metrics.
Yixuan Su, Tian Lan 0003, Yan Wang 0060, Dani Yogatama, Lingpeng Kong, Nigel Collier
NeurIPS2
2022 Food recommendation with graph convolutional network
Xiaoyan Gao 0001, Fuli Feng, Heyan Huang, Xianling Mao, Tian Lan 0003, Zewen Chi
Inf. Sci.5
2020 PONE: A Novel Automatic Evaluation Metric for Open-domain Generative Dialogue Systems
abstract
Open-domain generative dialogue systems have attracted considerable attention over the past few years. Currently, how to automatically evaluate them is still a big challenge. As far as we know, there are three kinds of automatic evaluations for open-domain generative dialogue systems: (1) Word-overlap-based metrics; (2) Embedding-based metrics; (3) Learning-based metrics. Due to the lack of systematic comparison, it is not clear which kind of metrics is more effective. In this article, we first measure systematically all kinds of metrics to check which kind is best. Extensive experiments demonstrate that learning-based metrics are the most effective evaluation metrics for open-domain generative dialogue systems. Moreover, we observe that nearly all learning-based metrics depend on the negative sampling mechanism, which obtains extremely imbalanced and low-quality samples to train a score model. To address this issue, we propose a novel learning-based metric that significantly improves the correlation with human judgments by using augmented PO sitive samples and valuable NE gative samples, called PONE. Extensive experiments demonstrate that PONE significantly outperforms the state-of-the-art learning-based evaluation method. Besides, we have publicly released the codes of our proposed metric and state-of-the-art baselines. 1
Tian Lan 0003, Xianling Mao, Wei Wei 0002, Xiaoyan Gao 0001, Heyan Huang
ACM Trans. Inf. Syst.1