VLDB 2026 Research / reviewers in the wild / expert
Xiaomin Chu
dblp:178/7275
· DBLP profile ↗
25ranked-venue papers
2as first author
19since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and BenchmarkabstractTopic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and understand the overall context of the document from a higher level, benefitting many downstream tasks such as summarization, discourse parsing, and information retrieval. However, the lack of large-scale, high-quality Chinese paragraph-level topic structure corpora restrained relative research and applications. To fill this gap, we build the Chinese paragraph-level topic representation, corpus, and benchmark in this paper. Firstly, we propose a hierarchical paragraph-level topic structure representation with three layers to guide the corpus construction. Then, we employ a two-stage man-machine collaborative annotation method to construct the largest Chinese Paragraph-level Topic Structure corpus (CPTS), achieving high quality. We also build several strong baselines, including ChatGPT, to validate the computability of CPTS on two fundamental tasks (topic segmentation and outline generation) and preliminarily verified its usefulness for the downstream task (discourse parsing). Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Haizhou Li 0001 |
LREC/COLING | 3 |
| 2023 | Topic Shift Detection in Chinese Dialogues: Corpus and Benchmark
Jiangyi Lin, Yaxin Fan, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
ICDAR (3) | 4 |
| 2023 | A Unified Document-Level Chinese Discourse Parser on Different Granularity Levels
Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICDAR (1) | 4 |
| 2023 | Employing Beautiful Sentence Evaluation to Automatic Chinese Essay Scoring
Yaqiong He, Xiaomin Chu, Peifeng Li 0001 |
ICIC (4) | 2 |
| 2023 | Multi-granularity Prompts for Topic Shift Detection in Dialogue
Jiangyi Lin, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICIC (4) | 3 |
| 2023 | Recognizing Functional Pragmatics in Chinese Discourses on Enhancing Paragraph Representation and Deep Differential Amplifier
Yaxin Fan, Peifeng Li 0001, Xiaomin Chu, Qiaoming Zhu |
ICIC (4) | 4 |
| 2023 | Recognizing Functional Pragmatics of Chinese Discourses on Data Augmentation and Dependency Graph
Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICIC (4) | 3 |
| 2023 | Discourse Parsing on Multi-Granularity InteractionabstractDiscourse parsing aims to construct a discourse structure tree to reflect the internal structure of a document. Most existing work only considers parsing documents from the paragraph-level or sentence-level granularity, ignoring the inter-action of different levels of granularity. Therefore, we propose a Multi-Granularity Interaction Method (MGIM) that facilitates discourse parsing through bidirectional information interaction at multi-granularity. We first introduce the structural information at the sentence level to boost paragraph-level parsing, and then use the functional pragmatics information at the paragraph level to guide sentence-level parsing. Moreover, we introduce an auxiliary task, discourse functional pragmatics recognition, to improve sentence-level parsing, which can guide sentence-level discourse tree construction from a macro perspective. Meanwhile, since the research field still lacks data for studying unified Chi-nese discourse parsing, we construct a Unified Chinese Discourse TreeBank UCDTB. Experimental results on both the Chinese UCDTB and the English RST-DT demonstrate the effectiveness of our proposed method. Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
IJCNN | 3 |
| 2023 | Chinese Macro Discourse Parsing on Generative Fusion and Distant Supervision
Longwang He, Feng Jiang 0007, Xiaoyi Bao, Yaxin Fan, Peifeng Li 0001, Xiaomin Chu |
PRICAI (2) | 6 |
| 2023 | Online meta-learning for POI recommendation
Yao Lv, Chong Tai, Wanjun Cheng, Jedi S. Shang, Jianfeng Qu, Xiaomin Chu, Ruoqian Zhang |
GeoInformatica | 7 |
| 2023 | How Deepbics Quantifies Intensities of Transcription Factor-DNA Binding and Facilitates Prediction of Single Nucleotide Variant Pathogenicity With a Deep Learning Model Trained On ChIP-Seq Data SetsabstractThe binding of DNA sequences to cell type-specific transcription factors is essential for regulating gene expression in all organisms. Many variants occurring in these binding regions play crucial roles in human disease by disrupting the cis-regulation of gene expression. We first implemented a sequence-based deep learning model called deepBICS to quantify the intensity of transcription factors-DNA binding. The experimental results not only showed the superiority of deepBICS on ChIP-seq data sets but also suggested deepBICS as a language model could help the classification of disease-related and neutral variants. We then built a language model-based method called deepBICS4SNV to predict the pathogenicity of single nucleotide variants. The good performance of deepBICS4SNV on 2 tests related to Mendelian disorders and viral diseases shows the sequence contextual information derived from language models can improve prediction accuracy and generalization capability. Lijun Quan, Xiaomin Chu, Xiaoyu Sun 0006, Tingfang Wu, Qiang Lyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Automated Chinese Essay Scoring from Multiple TraitsabstractAutomatic Essay Scoring (AES) is the task of using the computer to evaluate the quality of essays automatically. Current research on AES focuses on scoring the overall quality or single trait of prompt-specific essays. However, the users not only expect to obtain the overall score but also the instant feedback from different traits to help their writing in the real world. Therefore, we first annotate a mutli-trait dataset ACEA including 1220 argumentative essays from four traits, i.e., essay organization, topic, logic, and language. And then we design a hierarchical multi-task trait scorer HMTS to evaluate the quality of writing by modeling these four traits. Moreover, we propose an inter-sequence attention mechanism to enhance information interaction between different tasks and design the trait-specific features for various tasks in AES. The experimental results on ACEA show that our HMTS can effectively score essays from multiple traits, outperforming several strong models. Yaqiong He, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
COLING | 3 |
| 2022 | Adversarial Fine-Grained Fact Graph for Factuality-Oriented Abstractive Summarization
Zhiguang Gao, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
NLPCC (1) | 3 |
| 2022 | Employing Internal and External Knowledge to Factuality-Oriented Abstractive Summarization
Zhiguang Gao, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001 |
NLPCC (1) | 3 |
| 2022 | Bidirectional Macro-level Discourse Parser Based on Oracle Selection
Longwang He, Feng Jiang 0007, Xiaoyi Bao, Yaxin Fan, Peifeng Li 0001, Xiaomin Chu |
PRICAI (2) | 7 |
| 2021 | Hierarchical Macro Discourse Parsing Based on Topic SegmentationabstractHierarchically constructing micro (i.e., intra-sentence or inter-sentence) discourse structure trees using explicit boundaries (e.g., sentence and paragraph boundaries) has been proved to be an effective strategy. However, it is difficult to apply this strategy to document-level macro (i.e., inter-paragraph) discourse parsing, the more challenging task, due to the lack of explicit boundaries at the higher level. To alleviate this issue, we introduce a topic segmentation mechanism to detect implicit topic boundaries and then help the document-level macro discourse parser to construct better discourse trees hierarchically. In particular, our parser first splits a document into several sections using the topic boundaries that the topic segmentation detects. Then it builds a smaller and more accurate discourse sub-tree in each section and sequentially forms a whole tree for a document. The experimental results on both Chinese MCDTB and English RST-DT show that our proposed method outperforms the state-of-the-art baselines significantly. Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Fang Kong 0001 |
AAAI | 3 |
| 2021 | Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and GenerationabstractImplicit discourse relation recognition (IDRR)is a critical task in discourse analysis.Previous studies only regard it as a classification task and lack an in-depth understanding of the semantics of different relations.Therefore, we first view IDRR as a generation task and further propose a method joint modeling of the classification and generation.Specifically, we propose a joint model, CG-T5, to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously.Furthermore, we design three target sentence forms, including the question form, for the generation model to incorporate prior knowledge.To address the issue that large discourse units are hardly embedded into the target sentence, we also propose a target sentence construction mechanism that automatically extracts core sentences from those large discourse units.Experimental results both on Chinese MCDTB and English PDTB datasets show that our model CG-T5 achieves the best performance against several state-of-the-art systems. Feng Jiang 0007, Yaxin Fan, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
EMNLP (1) | 3 |
| 2021 | More Than One-Hot: Chinese Macro Discourse Relation Recognition on Joint Relation Embedding
Junhao Zhou, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
ICONIP (5) | 3 |
| 2021 | Chinese Macro Discourse Parsing on Dependency Graph Convolutional Network
Yaxin Fan, Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
NLPCC (1) | 3 |
| 2020 | Chinese Paragraph-level Discourse Parsing with Global Backward and Local Reverse ReadingabstractDiscourse structure tree construction is the fundamental task of discourse parsing and most previous work focused on English.Due to the cultural and linguistic differences, existing successful methods on English discourse parsing cannot be transformed into Chinese directly, especially in paragraph level suffering from longer discourse units and fewer explicit connectives.To alleviate the above issues, we propose two reading modes, i.e., the global backward reading and the local reverse reading, to construct Chinese paragraph level discourse trees.The former processes discourse units from the end to the beginning in a document to utilize the left-branching bias of discourse structure in Chinese, while the latter reverses the position of paragraphs in a discourse unit to enhance the differentiation of coherence between adjacent discourse units.The experimental results on Chinese MCDTB demonstrate that our model outperforms all strong baselines. Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Fang Kong 0001, Qiaoming Zhu |
COLING | 2 |
| 2019 | Constructing Chinese Macro Discourse Tree via Multiple Views and Word Pair Similarity
Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu |
NLPCC (1) | 2 |
| 2018 | Joint Modeling of Structure Identification and Nuclearity Recognition in Macro Chinese Discourse TreebankabstractDiscourse parsing is a challenging task and plays a critical role in discourse analysis. This paper focus on the macro level discourse structure analysis, which has been less studied in the previous researches. We explore a macro discourse structure presentation schema to present the macro level discourse structure, and propose a corresponding corpus, named Macro Chinese Discourse Treebank. On these bases, we concentrate on two tasks of macro discourse structure analysis, including structure identification and nuclearity recognition. In order to reduce the error transmission between the associated tasks, we adopt a joint model of the two tasks, and an Integer Linear Programming approach is proposed to achieve global optimization with various kinds of constraints. Xiaomin Chu, Feng Jiang 0007, Guodong Zhou 0001, Qiaoming Zhu |
COLING | 1 |
| 2018 | MCDTB: A Macro-level Chinese Discourse TreeBankabstractIn view of the differences between the annotations of micro and macro discourse rela-tionships, this paper describes the relevant experiments on the construction of the Macro Chinese Discourse Treebank (MCDTB), a higher-level Chinese discourse corpus. Fol-lowing RST (Rhetorical Structure Theory), we annotate the macro discourse information, including discourse structure, nuclearity and relationship, and the additional discourse information, including topic sentences, lead and abstract, to make the macro discourse annotation more objective and accurate. Finally, we annotated 720 articles with a Kappa value greater than 0.6. Preliminary experiments on this corpus verify the computability of MCDTB. Feng Jiang 0007, Sheng Xu 0006, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Guodong Zhou 0001 |
COLING | 3 |
| 2018 | Building a Macro Chinese Discourse Treebank
Xiaomin Chu, Feng Jiang 0007, Sheng Xu 0006, Qiaoming Zhu |
LREC | 1 |
| 2018 | Recognizing Macro Chinese Discourse Structure on Label Degeneracy Combination Model
Feng Jiang 0007, Peifeng Li 0001, Xiaomin Chu, Qiaoming Zhu, Guodong Zhou 0001 |
NLPCC (2) | 3 |