EDBT 2026 Demo / reviewers in the wild / expert
Feng Jiang 0007
dblp:75/1693-7
· DBLP profile ↗
37ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0002-3465-311XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 10 first-author · 27 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical GenerationabstractTheme detection is a fundamental task in user-centric dialogue systems, aiming to identify the latent topic of each utterance without relying on predefined schemas. Unlike intent induction, which operates within fixed label spaces, theme detection requires cross-dialogue consistency and alignment with personalized user preferences, posing significant challenges. Existing methods often struggle with sparse, short utterances for accurate topic representation and fail to capture user-level thematic preferences across dialogues. To address these challenges, we propose CATCH (Controllable Theme Detection with Contextualized Clustering and Hierarchical Generation), a unified framework that integrates three core components: (1) context-aware topic representation, which enriches utterance-level semantics using surrounding topic segments; (2) preference-guided topic clustering, which jointly models semantic proximity and personalized feedback to align themes across dialogue; and (3) a hierarchical theme generation mechanism designed to suppress noise and produce robust, coherent topic labels. Experiments on a multi-domain customer dialogue benchmark (DSTC-12) demonstrate the effectiveness of CATCH with 8B LLM in both theme clustering and topic generation quality. Rui Ke, Shenghao Yang 0001, Kuang Wang, Feng Jiang 0007, Haizhou Li 0001 |
AAAI | 5 |
| 2026 | Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive InquirersabstractXin Chen, Feng Jiang, Yiqian Zhang, Hardy Chen, Shuo Yan, Wenya Xie, Min Yang, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xin Chen 0032, Feng Jiang 0007, Hardy Chen, Wenya Xie, Min Yang 0007, Shujian Huang |
ACL (1) | 2 |
| 2026 | S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsabstractFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue, Fan Bu, Yuhao Du, Xiangying Chen, Benyou Wang, Haizhou Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Feng Jiang 0007, Liumeng Xue, Xiangying Chen, Benyou Wang, Haizhou Li 0001 |
ACL (1) | 1 |
| 2026 | GATHER: Convergence-Centric Hyper-Entity Retrieval for Zero-Shot Cell-Type AnnotationabstractZero-shot single-cell cell-type annotation aims to determine a cell's type from a given set of expressed genes without any training. Existing knowledge-graph-based RAG approaches retrieve evidence by expanding from source entities and relying on iterative LLM reasoning. However, in this setting each query contains tens to hundreds of genes, where no single gene is decisive and the label emerges only from their collective co-occurrence. Such hyper-entity queries fundamentally challenge local, entity-wise exploration strategies, which reason from individual genes, leading to poor scalability and substantial LLM cost. We propose GATHER (Graph-Aware Traversal with Hyper-Entity Retrieval), a convergence-centric retriever tailored to hyper-entity queries. It performs global multi-source graph traversal and identifies topological convergence points, nodes jointly reachable from many input genes. These convergence nodes act as high-information hyper-entities that capture entity synergy. By incorporating node- and path-importance scoring, GATHER selects informative evidence entirely without LLM involvement during retrieval. Instantiated on a self-constructed cell-centric biological knowledge graph (VCKG), GATHER outperforms strong KG-RAG baselines (ToG, ToG-2, RoG, PoG) on two datasets (Immune and Lung), achieving the highest exact-match accuracy (27.45% and 59.64%) with only a single LLM call per sample, compared to 2 to 61 calls for KG-RAG baselines. Our results demonstrate that convergence nodes compress multi-entity signals into compact, high-information evidence that conveys more per item than multi-hop paths, providing an efficient global alternative to local entity-wise reasoning. Zhonghui Zhang, Feng Jiang 0007, Shaowei Qin, Min Yang 0007 |
SIGIR | 2 |
| 2025 | Aligning Language Models Using Follow-up Likelihood as Reward SignalabstractIn natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions, the machine can leverage the user's follow-up utterances as feedback signals to assess whether it has appropriately addressed the user's request. Therefore, we propose using the likelihood of follow-up utterances as rewards to differentiate preferred responses from less favored ones, without relying on human or commercial LLM-based preference annotations. Our proposed reward mechanism, ``Follow-up Likelihood as Reward" (FLR), matches the performance of strong reward models trained on large-scale human or GPT-4 annotated data on 8 pairwise-preference and 4 rating-based benchmarks. Building upon the FLR mechanism, we propose to automatically mine preference data from the online generations of a base policy model. The preference data are subsequently used to boost the helpfulness of the base model through direct alignment from preference (DAP) methods, such as direct preference optimization (DPO). Lastly, we demonstrate that fine-tuning the language model that provides follow-up likelihood with natural language feedback significantly enhances FLR's performance on reward modeling benchmarks and effectiveness in aligning the base policy model's helpfulness. Chen Zhang 0055, Dading Chong, Feng Jiang 0007, Chengguang Tang, Anningzhe Gao, Guohua Tang, Haizhou Li 0001 |
AAAI | 3 |
| 2025 | Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesabstractUser simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and user-level diversity, often hindered by role confusion and dependence on predefined profiles of well-known figures. In contrast, direct simulation focuses solely on text, neglecting implicit user traits like personality and conversation-level consistency. To address these issues, we introduce the User Simulator with Implicit Profiles (USP), a framework that infers implicit user profiles from human-machine interactions to simulate personalized and realistic dialogues. We first develop an LLM-driven extractor with a comprehensive profile schema, then refine the simulation using conditional supervised fine-tuning and reinforcement learning with cycle consistency, optimizing at both the utterance and conversation levels. Finally, a diverse profile sampler captures the distribution of real-world user profiles. Experimental results show that USP outperforms strong baselines in terms of authenticity and diversity while maintaining comparable consistency. Additionally, using USP to evaluate LLM on dynamic multi-turn aligns well with mainstream benchmarks, demonstrating its effectiveness in real-world applications. Kuang Wang, Xianfei Li, Shenghao Yang 0001, Li Zhou 0010, Feng Jiang 0007, Haizhou Li 0001 |
ACL (1) | 5 |
| 2025 | Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language ModelsabstractZiche Liu, Rui Ke, Yajiao Liu, Feng Jiang, Haizhou Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ziche Liu, Rui Ke, Yajiao Liu, Feng Jiang 0007, Haizhou Li 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementabstractThe rapid development of large language models (LLMs), like ChatGPT, has resulted in the widespread presence of LLM-generated content on social media platforms, raising concerns about misinformation, data biases, and privacy violations, which can undermine trust in online discourse. While detecting LLM-generated content is crucial for mitigating these risks, current methods often focus on binary classification, failing to address the complexities of real-world scenarios like human-LLM collaboration. To move beyond binary classification and address these challenges, we propose a new paradigm for detecting LLM-generated content. This approach introduces two novel tasks: LLM Role Recognition (LLM-RR), a multi-class classification task that identifies specific roles of an LLM in content generation, and LLM Involvement Measurement (LLM-IM), a regression task that quantifies the extent of LLM involvement in content creation. To support these tasks, we propose LLMDetect, a benchmark designed to evaluate detectors' performance on these new tasks. LLMDetect includes the Hybrid News Detection Corpus (HNDC) for training detectors, as well as DetectEval, a comprehensive evaluation suite that considers five distinct cross-context variations and two multi-intensity variations within the same LLM role. This allows for a thorough assessment of detectors' generalization and robustness across diverse contexts. Our empirical validation of 10 baseline detection methods demonstrates that fine-tuned Pre-trained Language Model (PLM)-based models consistently outperform others on both tasks, while advanced LLMs face challenges in accurately detecting their own generated content. Our experimental results and analysis offer insights for developing more effective detection models for LLM-generated content. This research enhances the understanding of LLM-generated content and establishes a foundation for more nuanced detection methodologies. Li Zhou 0010, Feng Jiang 0007, Benyou Wang, Haizhou Li 0001 |
WWW | 3 |
| 2024 | PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User SimulatorabstractThe unparalleled performance of closedsourced ChatGPT has sparked efforts towards its democratization, with notable strides made by leveraging real user and ChatGPT dialogues, as evidenced by Vicuna.However, due to challenges in gathering dialogues involving human participation, current endeavors like Baize and UltraChat rely on ChatGPT conducting roleplay to simulate humans based on instructions, resulting in overdependence on seeds, diminished human-likeness, limited topic diversity, and an absence of genuine multi-round conversational dynamics.To address the above issues, we propose a paradigm to simulate human behavior better and explore the benefits of incorporating more human-like questions in multiturn conversations.Specifically, we directly target human questions extracted from genuine human-machine conversations as a learning goal and provide a novel user simulator called 'Socratic'.The experimental results show our response model, 'PlatoLM', achieves SoTA performance among LLaMA-based 7B models in MT-Bench.Our findings further demonstrate that our method introduces highly human-like questioning patterns and rich topic structures, which can teach the response model better than previous works in multi-round conversations. Chuyi Kong, Yaxin Fan, Feng Jiang 0007, Benyou Wang |
ACL (1) | 4 |
| 2024 | Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical StudyabstractLarge language models, like ChatGPT, have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored, where it requires higher level capabilities of understanding and reasoning. In this paper, we aim to systematically inspect ChatGPT’s performance in two discourse analysis tasks: topic segmentation and discourse parsing, focusing on its deep semantic understanding of linear and hierarchical discourse structures underlying dialogue. To instruct ChatGPT to complete these tasks, we initially craft a prompt template consisting of the task description, output format, and structured input. Then, we conduct experiments on four popular topic segmentation datasets and two discourse parsing datasets. The experimental results showcase that ChatGPT demonstrates proficiency in identifying topic structures in general-domain conversations yet struggles considerably in specific-domain conversations. We also found that ChatGPT hardly understands rhetorical structures that are more complex than topic structures. Our deeper investigation indicates that ChatGPT can give more reasonable topic structures than human annotations but only linearly parses the hierarchical rhetorical structures. In addition, we delve into the impact of in-context learning (e.g., chain-of-thought) on ChatGPT and conduct the ablation study on various prompt components, which can provide a research foundation for future work. The code is available at https://github.com/yxfanSuda/GPTforDDA. Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Haizhou Li 0001 |
LREC/COLING | 2 |
| 2024 | Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and BenchmarkabstractTopic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and understand the overall context of the document from a higher level, benefitting many downstream tasks such as summarization, discourse parsing, and information retrieval. However, the lack of large-scale, high-quality Chinese paragraph-level topic structure corpora restrained relative research and applications. To fill this gap, we build the Chinese paragraph-level topic representation, corpus, and benchmark in this paper. Firstly, we propose a hierarchical paragraph-level topic structure representation with three layers to guide the corpus construction. Then, we employ a two-stage man-machine collaborative annotation method to construct the largest Chinese Paragraph-level Topic Structure corpus (CPTS), achieving high quality. We also build several strong baselines, including ChatGPT, to validate the computability of CPTS on two fundamental tasks (topic segmentation and outline generation) and preliminarily verified its usefulness for the downstream task (discourse parsing). Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Haizhou Li 0001 |
LREC/COLING | 1 |
| 2024 | Humans or LLMs as the Judge? A Study on Judgement BiasabstractAdopting human and large language models (LLM) as judges (a.k.a human-and LLM-as-ajudge) for evaluating the performance of LLMs has recently gained attention.Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation results.In this paper, we propose a novel framework that is free from referencing groundtruth annotations for investigating Misinformation Oversight Bias, Gender Bias, Authority Bias and Beauty Bias on LLM and human judges.We curate a dataset referring to the revised Bloom's Taxonomy and conduct thousands of evaluations.Results show that human and LLM judges are vulnerable to perturbations to various degrees, and that even the cutting-edge judges possess considerable biases.We further exploit these biases to conduct attacks on LLM judges.We hope that our work can notify the community of the bias and vulnerability of human-and LLMas-a-judge, as well as the urgency of developing robust evaluation systems 1 .Warning: we provide illustrative attack protocols to reveal the vulnerabilities of LLM judges, aiming to develop more robust ones. Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang 0007, Benyou Wang |
EMNLP | 4 |
| 2024 | Tailored Domain-Specific Summaries: A Two-Stage Method Combining Extractive and Abstractive Summarization Models
Feng Jiang 0007, Lingyi Yang, Haizhou Li 0001 |
ICONIP (9) | 1 |
| 2024 | CMB: A Comprehensive Medical Benchmark in ChineseabstractXidong Wang, Guiming Chen, Song Dingjie, Zhang Zhiyi, Zhihong Chen, Qingying Xiao, Junying Chen, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, Haizhou Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xidong Wang, Guiming Chen, Dingjie Song, Zhiyi Zhang 0007, Qingying Xiao, Feng Jiang 0007, Benyou Wang, Haizhou Li 0001 |
NAACL-HLT | 8 |
| 2023 | Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee RecognitionabstractDialogue discourse parsing aims to reflect the relation-based structure of dialogue by establishing discourse links according to discourse relations.To alleviate data sparsity, previous studies have adopted multitasking approaches to jointly learn dialogue discourse parsing with related tasks (e.g., reading comprehension) that require additional human annotation, thus limiting their generality.In this paper, we propose a multitasking framework that integrates dialogue discourse parsing with its neighboring task addressee recognition.Addressee recognition reveals the reply-to structure that partially overlaps with the relation-based structure, which can be exploited to facilitate relationbased structure learning.To this end, we first proposed a reinforcement learning agent to identify training examples from addressee recognition that are most helpful for dialog discourse parsing.Then, a task-aware structure transformer is designed to capture the shared and private dialogue structure of different tasks, thereby further promoting dialogue discourse parsing.Experimental results on both the Molweni and STAC datasets show that our proposed method can outperform the SOTA baselines.The code will be available at https://github.com/yxfanSuda/RLTST. Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Fang Kong 0001, Qiaoming Zhu |
EMNLP | 2 |
| 2023 | Topic Shift Detection in Chinese Dialogues: Corpus and Benchmark
Jiangyi Lin, Yaxin Fan, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
ICDAR (3) | 3 |
| 2023 | A Unified Document-Level Chinese Discourse Parser on Different Granularity Levels
Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICDAR (1) | 2 |
| 2023 | Recognizing Functional Pragmatics of Chinese Discourses on Data Augmentation and Dependency Graph
Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICIC (4) | 2 |
| 2023 | GrammarGPT: Exploring Open-Source LLMs for Native Chinese Grammatical Error Correction with Supervised Fine-Tuning
Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Haizhou Li 0001 |
NLPCC (3) | 2 |
| 2023 | Chinese Macro Discourse Parsing on Generative Fusion and Distant Supervision
Longwang He, Feng Jiang 0007, Xiaoyi Bao, Yaxin Fan, Peifeng Li 0001, Xiaomin Chu |
PRICAI (2) | 2 |
| 2022 | Automated Chinese Essay Scoring from Multiple TraitsabstractAutomatic Essay Scoring (AES) is the task of using the computer to evaluate the quality of essays automatically. Current research on AES focuses on scoring the overall quality or single trait of prompt-specific essays. However, the users not only expect to obtain the overall score but also the instant feedback from different traits to help their writing in the real world. Therefore, we first annotate a mutli-trait dataset ACEA including 1220 argumentative essays from four traits, i.e., essay organization, topic, logic, and language. And then we design a hierarchical multi-task trait scorer HMTS to evaluate the quality of writing by modeling these four traits. Moreover, we propose an inter-sequence attention mechanism to enhance information interaction between different tasks and design the trait-specific features for various tasks in AES. The experimental results on ACEA show that our HMTS can effectively score essays from multiple traits, outperforming several strong models. Yaqiong He, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
COLING | 2 |
| 2022 | Adversarial Fine-Grained Fact Graph for Factuality-Oriented Abstractive Summarization
Zhiguang Gao, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
NLPCC (1) | 2 |
| 2022 | Employing Internal and External Knowledge to Factuality-Oriented Abstractive Summarization
Zhiguang Gao, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
NLPCC (1) | 2 |
| 2022 | Bidirectional Macro-level Discourse Parser Based on Oracle Selection
Longwang He, Feng Jiang 0007, Xiaoyi Bao, Yaxin Fan, Peifeng Li 0001, Xiaomin Chu |
PRICAI (2) | 2 |
| 2021 | Hierarchical Macro Discourse Parsing Based on Topic SegmentationabstractHierarchically constructing micro (i.e., intra-sentence or inter-sentence) discourse structure trees using explicit boundaries (e.g., sentence and paragraph boundaries) has been proved to be an effective strategy. However, it is difficult to apply this strategy to document-level macro (i.e., inter-paragraph) discourse parsing, the more challenging task, due to the lack of explicit boundaries at the higher level. To alleviate this issue, we introduce a topic segmentation mechanism to detect implicit topic boundaries and then help the document-level macro discourse parser to construct better discourse trees hierarchically. In particular, our parser first splits a document into several sections using the topic boundaries that the topic segmentation detects. Then it builds a smaller and more accurate discourse sub-tree in each section and sequentially forms a whole tree for a document. The experimental results on both Chinese MCDTB and English RST-DT show that our proposed method outperforms the state-of-the-art baselines significantly. Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Fang Kong 0001 |
AAAI | 1 |
| 2021 | Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and GenerationabstractImplicit discourse relation recognition (IDRR)is a critical task in discourse analysis.Previous studies only regard it as a classification task and lack an in-depth understanding of the semantics of different relations.Therefore, we first view IDRR as a generation task and further propose a method joint modeling of the classification and generation.Specifically, we propose a joint model, CG-T5, to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously.Furthermore, we design three target sentence forms, including the question form, for the generation model to incorporate prior knowledge.To address the issue that large discourse units are hardly embedded into the target sentence, we also propose a target sentence construction mechanism that automatically extracts core sentences from those large discourse units.Experimental results both on Chinese MCDTB and English PDTB datasets show that our model CG-T5 achieves the best performance against several state-of-the-art systems. Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
EMNLP (1) | 1 |
| 2021 | More Than One-Hot: Chinese Macro Discourse Relation Recognition on Joint Relation Embedding
Junhao Zhou, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICONIP (5) | 2 |
| 2021 | Macro Discourse Relation Recogniztion Based on Micro Discourse Structure and Self-Interactive Attention NetworkabstractMacro discourse relation recognition is an important task of macro discourse analysis. The existing models ignore the micro discourse structure within paragraphs and could not accurately grasp the paragraph semantics. In addition, the traditional pre-trained models only used the representation of the [CLS] token for macro relation classification, lacking more detailed semantic interaction between discourse units. To solve the above issues, we proposed a macro discourse relation recognition model based on Micro Discourse Structure and Self-Interactive Network (MDSSIN) that mines the semantic representation of important parts within a paragraph and enhances semantic interaction between paragraphs. Specially, we first automatically build the micro-structure discourse tree for each paragraph and get the core clause in each paragraph according to nuclearity. Then, we use the self-interactive attention mechanism to capture more detailed semantic interaction between discourse units and between the core clauses of discourse units. Experimental results on Chinese MCDTB show that our method achieves the SOTA performance. Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Qiaoming Zhu |
IJCNN | 2 |
| 2021 | Recognizing Chinese Discourse Relations Based on Multi-Perspective and Hierarchical ModelingabstractRecognizing discourse relations is crucial for understanding semantic and logical connections between two discourse units in the text. The most of existing work only considers the single-level and single-perspective discourse relationship between discourse units, which will lead to inconsistencies between recognizing classes and sub-classes of the discourse relations and cannot distinguish very similar relations in semantics. Therefore, in this paper, we propose a Multi-perspective and Hierarchical Model (MHM) that can model hierarchical relationships and capture the implicit connections between discourse units from multi-perspective (i.e., rhetorical, co-referential, and temporal) to strengthen the distinction between different discourse relations. Specifically, we first build a hierarchical classification module with contrastive learning to mine the finer semantic in two-level discourse relations and recognize relations consistently. Then, we introduce recognizing co-referential and temporal relations as auxiliary tasks and build a mapping from rhetorical relations to them, modeling multi-perspective relationships between discourse units. The experimental results on MCDTB and CDTB show that our model achieves the best performance. Feng Jiang 0007, Peifeng Li 0001, Qiaoming Zhu |
IJCNN | 1 |
| 2021 | Chinese Macro Discourse Parsing on Dependency Graph Convolutional Network
Yaxin Fan, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
NLPCC (1) | 2 |
| 2020 | Chinese Paragraph-level Discourse Parsing with Global Backward and Local Reverse ReadingabstractDiscourse structure tree construction is the fundamental task of discourse parsing and most previous work focused on English.Due to the cultural and linguistic differences, existing successful methods on English discourse parsing cannot be transformed into Chinese directly, especially in paragraph level suffering from longer discourse units and fewer explicit connectives.To alleviate the above issues, we propose two reading modes, i.e., the global backward reading and the local reverse reading, to construct Chinese paragraph level discourse trees.The former processes discourse units from the end to the beginning in a document to utilize the left-branching bias of discourse structure in Chinese, while the latter reverses the position of paragraphs in a discourse unit to enhance the differentiation of coherence between adjacent discourse units.The experimental results on Chinese MCDTB demonstrate that our model outperforms all strong baselines. Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Fang Kong 0001, Qiaoming Zhu |
COLING | 1 |
| 2020 | Macro Discourse Relation Recognition via Discourse Argument Pair Graph
Zhenhua Sun, Feng Jiang 0007, Peifeng Li 0001, Qiaoming Zhu |
NLPCC (2) | 2 |
| 2019 | Joint Modeling of Recognizing Macro Chinese Discourse Nuclearity and Relation Based on Structure and Topic Gated Semantic Network
Feng Jiang 0007, Peifeng Li 0001, Qiaoming Zhu |
NLPCC (2) | 1 |
| 2018 | Joint Modeling of Structure Identification and Nuclearity Recognition in Macro Chinese Discourse TreebankabstractDiscourse parsing is a challenging task and plays a critical role in discourse analysis. This paper focus on the macro level discourse structure analysis, which has been less studied in the previous researches. We explore a macro discourse structure presentation schema to present the macro level discourse structure, and propose a corresponding corpus, named Macro Chinese Discourse Treebank. On these bases, we concentrate on two tasks of macro discourse structure analysis, including structure identification and nuclearity recognition. In order to reduce the error transmission between the associated tasks, we adopt a joint model of the two tasks, and an Integer Linear Programming approach is proposed to achieve global optimization with various kinds of constraints. Xiaomin Chu, Feng Jiang 0007, Guodong Zhou 0001, Qiaoming Zhu |
COLING | 2 |
| 2018 | MCDTB: A Macro-level Chinese Discourse TreeBankabstractIn view of the differences between the annotations of micro and macro discourse rela-tionships, this paper describes the relevant experiments on the construction of the Macro Chinese Discourse Treebank (MCDTB), a higher-level Chinese discourse corpus. Fol-lowing RST (Rhetorical Structure Theory), we annotate the macro discourse information, including discourse structure, nuclearity and relationship, and the additional discourse information, including topic sentences, lead and abstract, to make the macro discourse annotation more objective and accurate. Finally, we annotated 720 articles with a Kappa value greater than 0.6. Preliminary experiments on this corpus verify the computability of MCDTB. Feng Jiang 0007, Sheng Xu 0006, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Guodong Zhou 0001 |
COLING | 1 |
| 2018 | Building a Macro Chinese Discourse Treebank
Xiaomin Chu, Feng Jiang 0007, Sheng Xu 0006, Qiaoming Zhu |
LREC | 2 |
| 2018 | Recognizing Macro Chinese Discourse Structure on Label Degeneracy Combination Model
Feng Jiang 0007, Peifeng Li 0001, Xiaomin Chu, Qiaoming Zhu, Guodong Zhou 0001 |
NLPCC (2) | 1 |