EDBT 2026 Demo / reviewers in the wild / expert
Jianzhu Bao
dblp:278/8133
· DBLP profile ↗
20ranked-venue papers
7as first author
19since 2021 · last 2026
0009-0004-9818-8765ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart UnderstandingabstractJianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu, Sarthak Ketanbhai Modi, Zi Pong Lim, Yon Shin Teo, Wenya Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu, Sarthak Ketanbhai Modi, Zi Pong Lim, Yon Shin Teo, Wenya Wang 0001 |
ACL (1) | 1 |
| 2026 | ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language ModelsabstractArgument generation is a fundamental NLP task that aims to automatically produce persuasive arguments.Effective human argumentation is inherently complex and multifaceted, integrating argumentative strategies, appropriate styles, and adaptation to target audiences, etc.However, existing studies focus on limited control signals such as topic, stance, or key aspects, failing to capture this complexity.As LLMs advance, the lack of benchmarks evaluating multifaceted argumentative control becomes a critical bottleneck.To address this, we introduce ArgGenBench, a novel benchmark containing complex instructions that integrate multi-dimensional control, including topic, stance, length, style, strategy, audience, and key points.Extensive evaluation across 15 LLMs reveals significant limitations: even the best-performing model achieves only 42.7% win rate against human-verified references.These results highlight the challenge of controlled argument generation and establish ArgGenBench as a rigorous testbed for developing more capable systems. Bojun Jin, Jianzhu Bao, Yice Zhang, Ruifeng Xu 0001 |
ACL (1) | 2 |
| 2026 | Deep-Reporter: Deep Research for Grounded Multimodal Long-Form GenerationabstractFangda Ye, Kuicai Dong, Xie Zhifei, Yuxin Hu, Yihang Yin, Shurui Huang, Shikai Dong, Chen Zhang, Jianzhu Bao, Shuicheng Yan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fangda Ye, Kuicai Dong, Zhifei Xie, Yihang Yin, Shurui Huang, Shikai Dong, Chen Zhang 0003, Jianzhu Bao, Shuicheng Yan |
ACL (1) | 9 |
| 2025 | A Multi-persona Framework for Argument Quality AssessmentabstractBojun Jin, Jianzhu Bao, Yufang Hou, Yang Sun, Yice Zhang, Huajie Wang, Bin Liang, Ruifeng Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bojun Jin, Jianzhu Bao, Yice Zhang, Huajie Wang, Bin Liang 0004, Ruifeng Xu 0001 |
ACL (1) | 2 |
| 2025 | Learning First-Order Logic Rules for Argumentation MiningabstractYang Sun, Guanrong Chen, Hamid Alinejad-Rokny, Jianzhu Bao, Yuqi Huang, Bin Liang, Kam-Fai Wong, Min Yang, Ruifeng Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guanrong Chen, Hamid Alinejad-Rokny, Jianzhu Bao, Bin Liang 0004, Kam-Fai Wong, Min Yang 0007, Ruifeng Xu 0001 |
ACL (1) | 4 |
| 2025 | Exploring Quality and Diversity in Synthetic Data Generation for Argument MiningabstractThe advancement of Argument Mining (AM) is hindered by a critical bottleneck: the scarcity of structure-annotated datasets, which are expensive to create manually.Inspired by recent successes in synthetic data generation across various NLP tasks, this paper explores methodologies for LLMs to generate synthetic data for AM.We investigate two complementary synthesis perspectives: a quality-oriented synthesis approach, which employs structure-aware paraphrasing to preserve annotation quality, and a diversity-oriented synthesis approach, which generates novel argumentative texts with diverse topics and argument structures.Experiments on three datasets show that augmenting original training data with our synthetic data, particularly when combining both quality-and diversity-oriented instances, significantly enhances the performance of existing AM models, both in full-data and low-resource settings.Moreover, the positive correlation between synthetic data volume and model performance highlights the scalability of our methods. Jianzhu Bao, Wenya Wang 0001, Yice Zhang, Bojun Jin, Ruifeng Xu 0001 |
EMNLP | 1 |
| 2025 | Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis ModelsabstractRecent efforts leverage knowledge distillation techniques to develop lightweight and practical sentiment analysis models.These methods are grounded in human-written instructions and large-scale user texts.Despite the promising results, two key challenges remain: (1) manually written instructions are limited in diversity and quantity, making them insufficient to ensure comprehensive coverage of distilled knowledge; (2) large-scale user texts incur high computational cost, hindering the practicality of these methods.To this end, we introduce COMPEFFDIST, a comprehensive and efficient distillation framework for sentiment analysis.Our framework consists of two key modules: attribute-based automatic instruction construction and difficulty-based data filtering, which correspondingly tackle the aforementioned challenges.Applying our method across multiple model series (Llama-3, Qwen-3, and Gemma-3), we enable 3B student models to match the performance of 20x larger teacher models on most tasks.In addition, our approach greatly outperforms baseline methods in data efficiency, attaining the same performance level with only 10% of the data.All codes are available at https://github.com/ HITSZ-HLT/COMPEFFDIST. Guangyu Xie, Yice Zhang, Jianzhu Bao, Qianlong Wang 0001, Ruifeng Xu 0001 |
EMNLP | 3 |
| 2025 | Targeted Distillation for Sentiment AnalysisabstractThis paper explores targeted distillation methods for sentiment analysis 1 , aiming to build compact and practical models that preserve strong and generalizable sentiment analysis capabilities.To this end, we conceptually decouple the distillation target into knowledge and alignment and accordingly propose a two-stage distillation framework.Moreover, we introduce SENTIBENCH, a comprehensive and systematic sentiment analysis benchmark that covers a diverse set of tasks across 12 datasets.We evaluate a wide range of models on this benchmark.Experimental results show that our approach substantially enhances the performance of compact models across diverse sentiment analysis tasks, and the resulting models demonstrate strong generalization to unseen tasks, showcasing robust competitiveness against existing small-scale models.2 1. collect a large and diverse set of user texts. construct large-scale distillation corpus.4. optimize student model using two-stage approach.Distillation Yice Zhang, Guangyu Xie, Jingjie Lin, Jianzhu Bao, Qianlong Wang 0001, Ruifeng Xu 0001 |
EMNLP | 4 |
| 2024 | PITA: Prompting Task Interaction for Argumentation MiningabstractYang Sun, Muyi Wang, Jianzhu Bao, Bin Liang, Xiaoyan Zhao, Caihua Yang, Min Yang, Ruifeng Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Muyi Wang, Jianzhu Bao, Bin Liang 0004, Xiaoyan Zhao 0005, Caihua Yang, Min Yang 0007, Ruifeng Xu 0001 |
ACL (1) | 3 |
| 2024 | BC-Prover: Backward Chaining Prover for Formal Theorem ProvingabstractDespite the remarkable progress made by large language models in mathematical reasoning, interactive theorem proving in formal logic still remains a prominent challenge.Previous methods resort to neural models for proofstep generation and search.However, they suffer from exploring possible proofsteps empirically in a large search space.Besides, they directly use a less rigorous informal proof for proofstep generation, neglecting the incomplete reasoning within.In this paper, we propose BC-Prover, a backward chaining framework guided by pseudo steps.Specifically, BC-Prover prioritizes pseudo steps to proofstep generation.The pseudo steps boost the proof construction in two aspects: (1) Backward Chaining that decomposes the proof into sub-goals for goaloriented exploration.(2) Step Planning that makes a fine-grained planning to bridge the gap between informal and formal proofs.Experiments on the miniF2F benchmark show significant performance gains by our framework over the state-of-the-art approaches.Our framework is also compatible with existing provers and further improves their performance with the backward chaining technique. Jihai Zhang 0001, Jianzhu Bao, Fangquan Lin, Cheng Yang 0008, Bing Qin 0001, Ruifeng Xu 0001, Wotao Yin |
EMNLP | 3 |
| 2024 | Enhancing Argumentative Relation Classification by Multi-Granularity Retrieval and Heterogeneous Graph ReasoningabstractArgumentative relation classification (ARC) aims to identify the relation between arguments. Previous methods that employ structured knowledge graphs to tackle the ARC task have achieved promising results. However, the prerequisite for structured knowledge to function is that the knowledge includes the topics of arguments. In practice, the topics of arguments are constantly emerging, making it impractical to construct a structured knowledge graph that contains all potential topics in advance. To address this issue, we investigate ARC from a novel perspective by utilizing unstructured knowledge to enhance the learning of ARC, where useful information for topics and arguments could be flexibly obtained from unstructured knowledge. Specifically, to retrieve diverse and comprehensive knowledge for topics and arguments, we first propose a multi-granularity retrieval method tailored for ARC, which acquires unstructured knowledge by dense retrieval at three levels of granularity: the concept level, the concept relation level, and the argument level. Further, we introduce a Knowledge-aware Heterogeneous Graph Reasoner (KHGR), which enables better utilization of retrieved knowledge to facilitate ARC. Extensive experiments on three publicly available datasets verify the superiority of our model compared with several state-of-the-art baselines. Further analysis shows that our method yields more significant benefits in low-resource scenarios. Caihua Yang, Jianzhu Bao, Bin Liang 0004, Ruifeng Xu 0001 |
ICASSP | 2 |
| 2024 | Leveraging temporal dependency for cross-subject-MI BCIs by contrastive learning and self-attention
Yi Ding 0012, Jianzhu Bao, Ke Qin, Chengxuan Tong, Jing Jin 0001, Cuntai Guan |
Neural Networks | 3 |
| 2023 | A Synthetic Data Generation Framework for Grounded DialoguesabstractTraining grounded response generation models often requires a large collection of grounded dialogues.However, it is costly to build such dialogues.In this paper, we present a synthetic data generation framework (SynDG) for grounded dialogues.The generation process utilizes large pre-trained language models and freely available knowledge data (e.g., Wikipedia pages, persona profiles, etc.).The key idea of designing SynDG is to consider dialogue flow and coherence in the generation process.Specifically, given knowledge data, we first heuristically determine a dialogue flow, which is a series of knowledge pieces.Then, we employ T5 to incrementally turn the dialogue flow into a dialogue.To ensure coherence of both the dialogue flow and the synthetic dialogue, we design a two-level filtering strategy, at the flow-level and the utterance-level respectively.Experiments on two public benchmarks show that the synthetic grounded dialogue data produced by our framework is able to significantly boost model performance in both full training data and low-resource scenarios. Jianzhu Bao, Rui Wang 0092, Yasheng Wang, Aixin Sun, Fei Mi, Ruifeng Xu 0001 |
ACL (1) | 1 |
| 2023 | Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue ModelsabstractRui Wang, Jianzhu Bao, Fei Mi, Yi Chen, Hongru Wang, Yasheng Wang, Yitong Li, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Rui Wang 0092, Jianzhu Bao, Fei Mi, Yi Chen 0007, Hongru Wang 0003, Yasheng Wang, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu 0001 |
ACL (1) | 2 |
| 2022 | A Generative Model for End-to-End Argument Mining with Reconstructed Positional Encoding and Constrained Pointer MechanismabstractArgument mining (AM) is a challenging task as it requires recognizing the complex argumentation structures involving multiple subtasks.To handle all subtasks of AM in an end-to-end fashion, previous works generally transform AM into a dependency parsing task.However, such methods largely require complex pre-and post-processing to realize the task transformation.In this paper, we investigate the endto-end AM task from a novel perspective by proposing a generative framework, in which the expected outputs of AM are framed as a simple target sequence.Then, we employ a pretrained sequence-to-sequence language model with a constrained pointer mechanism (CPM) to model the clues for all the subtasks of AM in the light of the target sequence.Furthermore, we devise a reconstructed positional encoding (RPE) to alleviate the order biases induced by the autoregressive generation paradigm.Experimental results show that our proposed framework achieves new state-of-the-art performance on two AM benchmarks. 1 Jianzhu Bao, Bin Liang 0004, Jiachen Du, Bing Qin 0001, Min Yang 0007, Ruifeng Xu 0001 |
EMNLP | 1 |
| 2022 | AEG: Argumentative Essay Generation via A Dual-Decoder Model with Content PlanningabstractArgument generation is an important but challenging task in computational argumentation.Existing studies have mainly focused on generating individual short arguments, while research on generating long and coherent argumentative essays is still under-explored.In this paper, we propose a new task, Argumentative Essay Generation (AEG).Given a writing prompt, the goal of AEG is to automatically generate an argumentative essay with strong persuasiveness.We construct a large-scale dataset, ArgEssay, for this new task and establish a strong model based on a dual-decoder Transformer architecture.Our proposed model contains two decoders, a planning decoder (PD) and a writing decoder (WD), where PD is used to generate a sequence for essay content planning and WD incorporates the planning information to write an essay.Further, we pre-train this model on a large news dataset to enhance the plan-and-write paradigm.Automatic and human evaluation results show that our model can generate more coherent and persuasive essays with higher diversity and less repetition compared to several baselines.1 Jianzhu Bao, Yasheng Wang, Fei Mi, Ruifeng Xu 0001 |
EMNLP | 1 |
| 2021 | A Neural Transition-based Model for Argumentation MiningabstractJianzhu Bao, Chuang Fan, Jipeng Wu, Yixue Dang, Jiachen Du, Ruifeng Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jianzhu Bao, Chuang Fan, Jipeng Wu, Yixue Dang, Jiachen Du, Ruifeng Xu 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Argument Pair Extraction with Mutual Guidance and Inter-sentence Relation GraphabstractArgument pair extraction (APE) aims to extract interactive argument pairs from two passages of a discussion.Previous work studied this task in the context of peer review and rebuttal, and decomposed it into a sequence labeling task and a sentence relation classification task.However, despite the promising performance, such an approach obtains the argument pairs implicitly by the two decomposed tasks, lacking explicitly modeling of the argument-level interactions between argument pairs.In this paper, we tackle the APE task by a mutual guidance framework, which could utilize the information of an argument in one passage to guide the identification of arguments that can form pairs with it in another passage.In this manner, two passages can mutually guide each other in the process of APE.Furthermore, we propose an inter-sentence relation graph to effectively model the interrelations between two sentences and thus facilitates the extraction of argument pairs.Our proposed method can better represent the holistic argument-level semantics and thus explicitly capture the complex correlations between argument pairs.Experimental results show that our approach significantly outperforms the current state-of-the-art model. Jianzhu Bao, Bin Liang 0004, Yice Zhang, Min Yang 0007, Ruifeng Xu 0001 |
EMNLP (1) | 1 |
| 2021 | A Hierarchical Sequence Labeling Model for Argument Pair Extraction
Qinglin Zhu, Jianzhu Bao, Jipeng Wu, Caihua Yang, Rui Wang 0092, Ruifeng Xu 0001 |
NLPCC (2) | 3 |
| 2020 | Emotion-Cause Pair Extraction as Sequence Labeling Based on A Novel Tagging SchemeabstractThe task of emotion-cause pair extraction deals with finding all emotions and the corresponding causes in unannotated emotion texts.Most recent studies are based on the likelihood of Cartesian product among all clause candidates, resulting in a high computational cost.Targeting this issue, we regard the task as a sequence labeling problem and propose a novel tagging scheme with coding the distance between linked components into the tags, so that emotions and the corresponding causes can be extracted simultaneously.Accordingly, an end-to-end model is presented to process the input texts from left to right, always with linear time complexity, leading to a speed up.Experimental results show that our proposed model achieves the best performance, outperforming the state-of-the-art method by 2.26% (p < 0.001) in F 1 measure. Chaofa Yuan, Chuang Fan, Jianzhu Bao, Ruifeng Xu 0001 |
EMNLP (1) | 3 |