Dianbo Sui

dblp:254/8270 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0002-5200-2265ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 21 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Editing the Moving World: Model Editing for Video LLMs
abstract
Qian Zhang, Xinye Li, Xiaokai Wu, Junhao Xu, Zhanyue Qin, Qingbin Liu, Junxian Cai, Xi Chen, Bolin Zhang, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinye Li 0001, Xiaokai Wu, Zhanyue Qin, Qingbin Liu, Junxian Cai, Xi Chen 0003, Zhiying Tu, Dianbo Sui
ACL (1)13
2025 VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing. Unlike traditional task-specific models, Video Large Language Models (video LLMs) can handle multiple tasks concurrently in a zero-shot manner. Consequently, exploring the application of video LLMs for VTG tasks has become a burgeoning research area. However, despite considerable advancements in video content understanding, video LLMs often struggle to accurately pinpoint timestamps within videos, limiting their effectiveness in VTG tasks. To address this, we introduce VTG-LLM, a model designed to enhance video LLMs' timestamp localization abilities. Our approach includes: (1) effectively integrating timestamp knowledge into visual tokens; (2) incorporating absolute-time tokens to manage timestamp knowledge without concept shifts; and (3) introducing a lightweight, high-performance, slot-based token compression technique designed to accommodate the demands of a large number of frames to be sampled for VTG tasks. Additionally, we present VTG-IT-120K, a collection of publicly available VTG datasets that we have re-annotated to improve upon low-quality annotations. Our comprehensive experiments demonstrate the superior performance of VTG-LLM in comparison to other video LLM methods across a variety of VTG tasks.
Yongxin Guo 0001, Dingxin Cheng, Xiaoying Tang 0002, Dianbo Sui, Qingbin Liu, Xi Chen 0003, Kevin Zhao
AAAI6
2025 HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object Tracking
abstract
Referring Multi-Object Tracking (RMOT) aims to track multiple objects based on a provided language expression. Although prior studies have sought to accomplish this by integrating an textual module into the multi-object tracker, these methods combine text and image features in a basic way, neglecting the importance of text features. In this study, we propose a Hierarchical Fine-grained text-image Fusion tracker, named HFF-Tracker, which can perform fine-grained fusion of pixel-level visual features and text features across various semantic levels. Specifically, we have devised a Hierarchical Multi-Modal Fusion (HMMF) module to merge text and image features at an early stage in a hierarchical and detailed manner. The Text-Guided Decoder (TGD) is designed to provide the query with prior semantic information during the decoding process. Additionally, we have crafted a Text-Guided Prediction Head (TGPH) that utilizes text information to enhance the performance of the prediction head. Furthermore, we have implemented an adaptive Look-Back training strategy to maximize the utilization of valuable labeled data. Extensive experiments on the Refer-KITTI dataset and the Refer-KITTI-V2 dataset demonstrate that our proposed HFF-Tracker outperforms other state-of-the-art methods with remarkable margins.
Zeyong Zhao, Yanchao Hao, Qingbin Liu, Dianbo Sui, Shizhu He, Xi Chen 0003
AAAI6
2025 Advanced News Event Clustering via Topic Enhanced Modeling with Multi-Aspect Contrastive Learning
abstract
News event clustering, a crucial task for discovering and comprehending real-world information, aims to aggregate news articles into fine-grained clusters based on specific key events. As the presence of topic-unrelated documents within clusters and redundant information within individual documents, it is challenging to learn a discriminative document representation. To address this issue, we introduce a novel method, TECL (Topic Enhanced modeling with Contrastive Learning), that leverages topic enhanced modeling with multi-aspect contrastive learning for news event clustering. The topic-enhanced modeling employs neural topic models to incorporate global and local semantics into document representations, while contrastive learning refines these representations from both inter-document and intra-document perspectives.Experiments conducted in both unsupervised and supervised scenarios indicate that the proposed method significantly improves performance, demonstrating its effectiveness.
Dianbo Sui
CIKM3
2025 A Framework for Effective Invocation Methods of Various LLM Services
abstract
Large Language Models (LLMs) have shown impressive abilities in solving various natural language processing tasks and are now widely offered as services. LLM services enable users to accomplish tasks without requiring specialized knowledge, simply by paying service providers. However, numerous providers offer various LLM services with variations in pricing, latency, and performance. These factors are also affected by different invocation methods, such as the choice of context and the use of cache, which lead to unpredictable and uncontrollable service cost and quality. Consequently, utilizing various LLM services invocation methods to construct an effective (cost-saving, low-latency and high-performance) invocation strategy that best meets task demands becomes a pressing challenge. This paper provides a comprehensive overview of methods help LLM services to be invoked efficiently. Technically, we define the problem of constructing an effective LLM services invocation strategy, and based on this, propose a unified LLM service invocation framework. The framework classifies existing methods into four categories: input abstraction, semantic cache, solution design, and output enhancement, which can be used separately or jointly during the invocation life cycle. We discuss the methods in each category and compare them to provide valuable guidance for researchers. Finally, we emphasize the open challenges in this domain and shed light on future research.
Dianbo Sui, Jiabao Kang, Zhidong Qiao, Zhiying Tu
COLING2
2025 SHARP: Steering Hallucination in LVLMs via Representation Engineering
abstract
Junfei Wu, Yue Ding, Guofan Liu, Tianze Xia, Ziyue Huang, Dianbo Sui, Qiang Liu, Shu Wu, Liang Wang, Tieniu Tan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junfei Wu, Yue Ding 0009, Guofan Liu, Tianze Xia, Dianbo Sui, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan
EMNLP6
2025 M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language Model
abstract
Multimodal knowledge editing is an important method for modifying outdated or incorrect knowledge in Multimodal Large Language Models (MLLMs).However, existing datasets for multimodal knowledge editing lack multi-granularity knowledge.In this paper, we present a more realistic dataset called M2Edit, which includes three distinct types of knowledge: entity, relation, and action.Additionally, existing knowledge editing methods for MLLMs lack the ability to handle multigranularity knowledge and generalize to multimodal data.To address these limitations, we propose the multimodal knowledge editing method MLE.This approach identifies key knowledge layers within different components and collaboratively edits the various components of MLLMs.As a result, we observe significant improvements in visual generality performance, ranging from 4.8% to 10.8%, and achieve the best overall performance on knowledge data of different granularities.
Yubo Chen 0001, Qingbin Liu, Dianbo Sui, Xi Chen 0003, Kang Liu 0001, Jun Zhao 0001
EMNLP5
2025 Maximizing Intermediate Checkpoint Value in LLM Pretraining with Bayesian Optimization
abstract
The rapid proliferation of large language models (LLMs), such as GPT-4 and Gemini, underscores the intense demand for resources during their training processes, posing significant challenges due to substantial computational and environmental costs. In this paper, we introduce a novel checkpoint merging strategy aimed at making efficient use of intermediate checkpoints during LLM pretraining. This method utilizes intermediate checkpoints with shared training trajectories, and is rooted in an extensive search space exploration for the best merging weight via Bayesian optimization. Through various experiments, we demonstrate that: (1) Our proposed methodology exhibits the capacity to augment pretraining, presenting an opportunity akin to obtaining substantial benefits at minimal cost; (2) Our proposed methodology, despite requiring a given held-out dataset, still demonstrates robust generalization capabilities across diverse domains, a pivotal aspect in pretraining.
Deyuan Liu, Zecheng Wang, Bingning Wang, Weipeng Chen, Chunshan Li, Zhiying Tu, Dianbo Sui
ICML8
2025 A Weighted Preference Optimization Service Recommendation Method Based on Knowledge Graph and Large Language Model
abstract
Knowledge graph (KG)-based service recommendation methods address issues such as data sparsity and cold start in real-world service recommendations by integrating external knowledge as auxiliary information. Recently, large language models (LLMs) have gained significant attention due to their powerful comprehension and reasoning capabilities. LLM-based recommendation systems also demonstrate advantages in interpretability and few-shot service reasoning. However, the integration of LLMs and KGs into existing service recommendation methods presents two major challenges: (1) the difficulty of aligning service recommendation tasks with language modeling tasks, and (2) the lack of interpretable quantification of the relationship between knowledge and personalized preferences. To address these challenges, this paper proposes WPKL (Weighted Preference Optimization based on KG and LLM). WPKL leverages external knowledge to assist LLMs in modeling user preferences and employs a hybrid graph neural network (GNN) framework to enhance preference representation. Additionally, a weighted preference optimization (WPO) approach is proposed to fine-tune the LLM, enabling interpretable quantification of user preferences and personalized knowledge. Extensive experimental results demonstrate that WPKL achieves high-quality service recommendations.
Hongliang Sun 0001, Zhiying Tu, Dianbo Sui, Yongchao Xing, Kai Zhang 0067, Bohai Zhao, Xiaofei Xu 0001
ICWS4
2025 Unlocking Hidden Capabilities: A Self-Improving Workflow for Chatbots to Utilize Unintegrated Services
abstract
Chatbots have advanced from basic conversational agents to versatile tools by integrating external services. However, traditional chatbots are constrained by predefined service boundaries, limiting their ability to handle complex tasks with unintegrated services. While most research focuses on improving service discovery and invocation through data-intensive pretraining, only 13.29% of services are well-documented, hindering practical deployment. This paper proposes a self-improving workflow for chatbots, using a “wide in, strict out” self-supervised learning approach to acquire domain knowledge efficiently and generate high-quality service documents. Compatible with existing methods, it eliminates the need for dataset collection or pre-training. Experiments demonstrate that our workflow significantly improves the pass and success rate of chatbots in utilizing unintegrated services, offering a powerful solution for real-world applications where service integration is limited.
Yongchao Xing, Bohai Zhao, Dianbo Sui, Zhiying Tu
ICWS5
2025 VPO: Reasoning Preferences Optimization Based on V-Usable Information
Zecheng Wang, Chunshan Li, Bingning Wang, Dianbo Sui
NeurIPS7
2025 Ask and Retrieve Knowledge: Towards Proactive Asking with Imperfect Information in Medical Multi-turn Dialogues
abstract
Large language models (LLMs) cannot effectively collaborate with humans who provide imperfect information at the initial stage of the dialogue, unless they learn to proactively ask questions.
Yangqin Jiang, Dianbo Sui, Zhiying Tu
SIGIR4
2025 A Federated Social Recommendation Approach with Enhanced Hypergraph Neural Network
abstract
In recent years, the development of online social network platforms has led to increased research efforts in social recommendation systems. Unlike traditional recommendation systems, social recommendation systems utilize both user-item interactions and user-user social relations to recommend relevant items, taking into account social homophily and social influence. Graph neural network (GNN)-based social recommendation methods have been proposed to model these item interactions and social relations effectively. However, existing GNN-based methods rely on centralized training, which raises privacy concerns and faces challenges in data collection due to regulations and privacy restrictions. Federated learning has emerged as a privacy-preserving alternative. Combining federated learning with GNN-based methods for social recommendation can leverage their respective advantages, but it also introduces new challenges: (1) existing federated recommendation systems often lack the capability to process heterogeneous data, such as user-item interactions and social relations; (2) due to the sparsity of data distributed across different clients, capturing the higher-order relationship information among users becomes challenging and is often overlooked by most federated recommendation systems. To overcome these challenges, we propose a federated social recommendation approach with enhanced hypergraph neural network (HGNN). We introduce HGNN to learn user and item embeddings in federated recommendation systems, leveraging the hypergraph structure to address the heterogeneity of data. Based on carefully crafted triangular motifs, we merge user and item nodes to construct hypergraphs on local clients, capturing specific triangular relations. Multiple HGNN channels are used to encode different categories of high-order relations, and an attention mechanism is applied to aggregate the embedded information from these channels. Our experiments on real-world social recommendation datasets demonstrate the effectiveness of the proposed approach. Extensive experiment results on three publicly available datasets validate the effectiveness of the proposed method.
Hongliang Sun 0001, Zhiying Tu, Dianbo Sui, Xiaofei Xu 0001
ACM Trans. Intell. Syst. Technol.3
2024 Analyzing Chain-of-thought Prompting in Black-Box Large Language Models via Estimated V-information
abstract
Chain-of-Thought (CoT) prompting combined with large language models (LLM) has shown great potential in improving performance on challenging reasoning tasks. While understanding why CoT prompting is effective is crucial for the application and improvement of CoT prompting, few studies have addressed this issue. Besides, almost no prior work has conducted theoretical analysis on CoT prompting in the context of black-box models. In this paper, we approach the analysis of CoT prompting in black-box LLMs from an information-theoretic perspective. Specifically, we propose a new metric, EPVI (Estimated Pointwise V-Information), which extends the concept of pointwise V-information to black-box models, quantifying the label-relevant new information introduced by CoT prompting beyond the pre-existing information in the input. Based on this, we conduct a series of experiments at both the task and instance levels to analyze CoT prompting, demonstrating that the effectiveness of CoT prompting can be attributed to its capacity to influence the difficulty of model inference by augmenting or reducing the model-usable information. Furthermore, we show that selecting high-quality demonstrations of CoT reasoning based on EPVI can improve the downstream performance of reasoning tasks.
Zecheng Wang, Chunshan Li, Zhao Yang 0004, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Dianbo Sui
LREC/COLING8
2024 Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
abstract
Deyuan Liu, Zhanyue Qin, Hairu Wang, Zhao Yang, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Bo Li, Xi Chen, Cunhang Fan, Zhao Lv, Dianhui Chu, Zhiying Tu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Deyuan Liu, Zhanyue Qin, Hairu Wang 0002, Zhao Yang 0004, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Cunhang Fan, Zhao Lv, Zhiying Tu, Dianbo Sui
EMNLP15
2024 UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
abstract
Zhanyue Qin, Haochuan Wang, Deyuan Liu, Ziyang Song, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhanyue Qin, Deyuan Liu, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei 0001, Zhiying Tu, Dianbo Sui
EMNLP12
2024 A Federated Graph to Embedding Approach for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) tasks have been developed to address the inherent incompleteness of KGs. Recently, knowledge graph embedding (KGE) methods have gained popularity for embedding entities and relations, proving effective in KGC. However, privacy concerns make it challenging to collect privacy KG data from different institutions in the actual application. Federated learning has emerged as a solution for training models with decentralized data, eliminating the need for collecting private data. However, existing federated KGE methods overlook the implicit graph structural information of entities and relations, resulting in fragmented and incomplete representations within federated clients. Moreover, these methods often struggle with capturing multiple relational representations. To address these challenges, we propose a Federated Graph to Embedding (FedGE) approach based on encoder-decoder to capture interactions among entities and relations. Extensive experiments on two common KG datasets demonstrate the superiority of our method. The code is available at https://github.com/s460305450/FedGE.git.
Hongliang Sun 0001, Xiaofeng Bi, Dianbo Sui, Zhiying Tu
ICASSP3
2024 Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
Dianbo Sui, Hongliang Sun 0001, Zhiying Tu
ICSOC (1)2
2024 Learning Dynamic Knowledge Graph Embedding in Evolving Service Ecosystems via Meta-Learning
abstract
In the context of dynamic service ecosystems, the inability of conventional knowledge graph embedding (KGE) methods to efficiently update incremental knowledge poses a significant challenge for the effectiveness of intelligent web applications. To address the continuous updating challenges of service knowledge, this paper introduces MetaHG, a meta-learning strategy for KGE. Unlike existing meta-learning KGE studies that focus solely on local entity information, MetaHG incorporates both local and potential global structural information from current snapshot’s seen knowledge graphs (KGs) to mitigate issues such as spatial deformation and enhance the representation of unseen entities. Our approach initializes entity embeddings using ‘in’ and ‘out’ relationship matrices and refines them through a hybrid graph neural network (GNN) framework, which includes a GNN layer for local information and a hypergraph neural network (HGNN) layer for potential global information. The meta-learning strategy embedded in MetaHG effectively transfers meta-knowledge for the accurate representation of emerging entities. Extensive experiments are conducted on a self-collected clothing industry service dataset and two publicly available open-source KG datasets. By comparing with several baselines, experiment results demonstrate the superior performance of MetaHG in generating high-quality embeddings for emerging entities and dynamically updating service knowledge.
Hongliang Sun 0001, Jinlan Liu 0001, Dianbo Sui, Zhiying Tu, Xiaofei Xu 0001
ICWS4
2024 RealWeb: A Benchmark for Universal Instruction Following in Realistic Web Services Navigation
abstract
Traditional methods of interacting with web pages, such as clicking and scrolling, greatly hinder users, especially those with disabilities and the elderly from conveniently accessing web services. By following user instructions, automatic web service navigation agents accomplish complex tasks on the websites, which is a natural interactions with web services. To study this task, previous works constructed simple web pages within simulated environments, but the realistic websites are far more intricate in true environments. Moreover, existing methods for this task rely on manually collecting human demonstrations on the given websites, which is time-consuming and labor-intensive, and reduce the generalization ability of service agents to unseen websites. Thus, we construct the first Chinese multimodal benchmark for web services navigation under the realistic settings: across domains and without human demonstrations. Our benchmark comprises a dataset (RealWeb) and a baseline method (WeServe). RealWeb consists of 40 real-world websites, 110 pages, and 11,739 language instructions. To detect and understand the feasible operations of pages in the visual mode, the screenshot of each page is annotated with 5 critical areas in RealWeb. WeServeis a multi-modal framework for web services navigation that combines visual and textual information, enabling universal navigation on any web pages with a success rate of 68.61%. Our dataset and codes are available at https://gitee.com/plabrolin/real-web.
Shiyun Xiong, Dianbo Sui, Yunzhe Xu, Zhiying Tu
ICWS3
2024 MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker Extraction
abstract
Speaker extraction aims to selectively extract the target speaker from the multi-talker environment under the guidance of auxiliary reference. Recent studies have shown that the attended speaker's information can be decoded by the auditory attention decoding from the listener's brain activity. However, how to more effectively utilize the common information about the target speaker contained in both electroencephalography (EEG) and speech is still an unresolved problem. In this paper, we propose a multi-scale fusion network (MSFNet) for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to make full use of the speech information, the mixed speech is encoded with multiple time scales so that the multi-scale embeddings are acquired. In addition, to effectively extract the non-Euclidean data of EEG, the graph convolutional networks are used as the EEG encoder. Finally, these multi-scale embeddings are separately fused with the EEG features. To facilitate research related to auditory attention decoding and further validate the effectiveness of the proposed method, we also construct the AVED dataset, a new EEG-Audio dataset. Experimental results on both the public Cocktail Party dataset and the newly proposed AVED dataset in this paper show that our MSFNet model significantly outperforms the state-of-the-art method in certain objective evaluation metrics.
Cunhang Fan, Wang Xiang, Jianhua Tao 0001, Jiangyan Yi, Dianbo Sui, Zhao Lv
ACM Multimedia8
2024 Can We Debias Multimodal Large Language Models via Model Editing?
abstract
Multimodal large language models (MLLM) have been observed to exhibit biases originating from their training datasets. Unlike unimodal LLMs, biases in MLLMs may stem from interactions between multiple modalities, which increases the complexity of multimodal debiasing. Conventional approaches like fine-tuning to alleviate biases in models are costly and data-hungry. Model editing methods, which focus on post-hoc modifications of model knowledge, have recently demonstrated significant potential across diverse applications. These methods can effectively and precisely adjust the behavior of models in specific knowledge domains, while minimizing the impact on the overall performance of the model. However, there is currently no comprehensive study to drive the application of model editing methods in debiasing MLLM and to analyze its pros and cons. To facilitate research in this field, we define the debiasing problem of MLLM as an editing problem and propose a novel set of evaluation metrics for MLLM debias editing. Through various experiments, we demonstrate that: (1) Existing model editing methods can effectively alleviate biases in MLLM and can generalize well to semantically equivalent image-text pairs. However, most methods tend to adversely affect the stability of the MLLM. (2) Compared to editing the visual modality of the MLLM, editing the textual modality yields better results in addressing MLLM biases. (3) Model editing based debiasing method can achieve generalization across different types of biases.
Zecheng Wang, Xinye Li 0001, Zhanyue Qin, Chunshan Li, Zhiying Tu, Dianbo Sui
ACM Multimedia7
2024 A Privacy-Preserving Framework for Medical Chatbot Based on LLM with Retrieval Augmented Generation
Chunshan Li, Zecheng Wang, Dianbo Sui, Jianen Yan
NLPCC (3)4
2024 Woodpecker: hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu 0001, Hao Wang 0076, Dianbo Sui, Yunhang Shen, Ke Li 0015, Xing Sun 0001, Enhong Chen
Sci. China Inf. Sci.6
2024 Explanation Guided Knowledge Distillation for Pre-trained Language Model Compression
abstract
Knowledge distillation is widely used in pre-trained language model compression, which can transfer knowledge from a cumbersome model to a lightweight one. Though knowledge distillation based model compression has achieved promising performance, we observe that explanations between the teacher model and the student model are not consistent. We argue that the student model should study not only the predictions of the teacher model but also the internal reasoning process. To this end, we propose Explanation Guided Knowledge Distillation (EGKD) in this article, which utilizes explanations to represent the thinking process and improve knowledge distillation. To obtain explanations in our distillation framework, we select three typical explanation methods rooted in different mechanisms, namely gradient-based , perturbation-based , and feature selection methods. Then, to improve computational efficiency, we propose different optimization strategies to utilize the explanations obtained by these three different explanation methods, which could provide the student model with better learning guidance. Experimental results on GLUE demonstrate that leveraging explanations can improve the performance of the student model. Moreover, our EGKD could also be applied to model compression with different architectures.
Zhao Yang 0004, Yuanzhe Zhang, Dianbo Sui, Yiming Ju, Jun Zhao 0001, Kang Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2024 Joint Entity and Relation Extraction With Set Prediction Networks
abstract
Joint entity and relation extraction is an important task in natural language processing, which aims to extract all relational triples mentioned in a given sentence. In essence, the relational triples mentioned in a sentence are in the form of a set, which has no intrinsic order between elements and exhibits the permutation invariant feature. However, previous seq2seq-based models require sorting the set of relational triples into a sequence beforehand with some heuristic global rules, which destroys the natural set structure. In order to break this bottleneck, we treat joint entity and relation extraction as a direct set prediction problem, so that the extraction model is not burdened with predicting the order of multiple triples. To solve this set prediction problem, we propose networks featured by transformers with non-autoregressive parallel decoding. In contrast to autoregressive approaches that generate triples one by one in a specific order, the proposed networks are able to directly output the final set of relational triples in one shot. Furthermore, we also design a set-based loss that forces unique predictions through bipartite matching. Compared with cross-entropy loss that highly penalizes small shifts in triple order, the proposed bipartite matching loss is invariant to any permutation of predictions; thus, it can provide the proposed networks with a more accurate training signal by ignoring triple order and focusing on relation types and entities. Various experiments on two benchmark datasets demonstrate that our proposed model significantly outperforms the current state-of-the-art (SoTA) models. Training code and trained models are now publicly available at https://github.com/DianboWork/SPN4RE.
Dianbo Sui, Xiangrong Zeng, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Learning with Partial Annotations for Event Detection
abstract
Event detection (ED) seeks to discover and classify event instances in plain texts.Previous methods for ED typically adopt supervised learning, requiring fully labeled and high-quality training data.However, in a realworld application, we may not obtain clean training data but only partially labeled one, which could substantially impede the learning process.In this work, we conduct a seminal study for learning with partial annotations for ED.We propose a new trigger localization formulation using contrastive learning to distinguish ground-truth triggers from contexts, showing a decent robustness for addressing partial annotation noise.Impressively, in an extreme scenario where more than 90% of events are unlabeled, our approach achieves an F1 score of over 60%.In addition, we reannotate and make available two fully annotated subsets of ACE 2005 to serve as an unbiased benchmark for event detection.We hope our approach and data will inspire future studies on this vital yet understudied problem.
Jian Liu 0032, Dianbo Sui, Kang Liu 0001, Haoyan Liu 0001, Zhe Zhao 0006
ACL (1)2
2023 Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point Process
abstract
Although In-Context Learning has proven effective across a broad array of tasks, its efficiency is noticeably influenced by the selection of demonstrations.Existing methods tend to select different demonstrations for each test instance, which is time-consuming and poses limitations in practical scenarios.Therefore, this study aims to address the challenge of selecting a representative subset of in-context demonstrations that can effectively prompt different test instances in a specific task.We propose that this representative subset should be of high quality and diversity.Our empirical analyses confirm that demonstrations that meet these criteria can indeed bolster model performance.To satisfy these criteria, this paper further introduces a two-stage Determinantal Point Process (DPP) method designed to incorporate both quality and diversity in the process of demonstration selection, thereby obtaining representative in-context demonstrations.Through comprehensive experimentation, we have confirmed the efficacy of our proposed method, paving the way for more practical and effective In-Context Learning.
Zhao Yang 0004, Yuanzhe Zhang, Dianbo Sui, Cao Liu, Jun Zhao 0001, Kang Liu 0001
EMNLP3
2021 A Large-Scale Chinese Multimodal NER Dataset with Speech Clues
abstract
Dianbo Sui, Zhengkun Tian, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dianbo Sui, Zhengkun Tian, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
ACL/IJCNLP (1)1
2021 Document-level Event Extraction via Parallel Prediction Networks
abstract
Hang Yang, Dianbo Sui, Yubo Chen, Kang Liu, Jun Zhao, Taifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dianbo Sui, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Taifeng Wang
ACL/IJCNLP (1)2
2021 Set Generation Networks for End-to-End Knowledge Base Population
abstract
The task of knowledge base population (KBP) aims to discover facts about entities from texts and expand a knowledge base with these facts.Previous studies shape end-to-end KBP as a machine translation task, which is required to convert unordered fact into a sequence according to a pre-specified order.However, the facts stated in a sentence are unordered in essence.In this paper, we formulate end-to-end KBP as a direct set generation problem, avoiding considering the order of multiple facts.To solve the set generation problem, we propose networks featured by transformers with nonautoregressive parallel decoding.Unlike previous approaches that use an autoregressive decoder to generate facts one by one, the proposed networks can directly output the final set of facts in one shot.Furthermore, to train the networks, we also design a set-based loss that forces unique predictions via bipartite matching.Compared with cross-entropy loss that highly penalizes small shifts in fact order, the proposed bipartite matching loss is invariant to any permutation of predictions.Benefiting from getting rid of the burden of predicting the order of multiple facts, our proposed networks achieve state-of-the-art (SoTA) performance on two benchmark datasets.
Dianbo Sui, Chenhao Wang 0004, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Wei Bi
EMNLP (1)1
2021 Knowledge Guided Metric Learning for Few-Shot Text Classification
abstract
Dianbo Sui, Yubo Chen, Binjie Mao, Delai Qiu, Kang Liu, Jun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dianbo Sui, Yubo Chen 0001, Binjie Mao, Delai Qiu, Kang Liu 0001, Jun Zhao 0001
NAACL-HLT1
2020 Graph-Based Knowledge Integration for Question Answering over Dialogue
abstract
Question answering over dialogue, a specialized machine reading comprehension task, aims to comprehend a dialogue and to answer specific questions.Despite many advances, existing approaches for this task did not consider dialogue structure and background knowledge (e.g., relationships between speakers).In this paper, we introduce a new approach for the task, featured by its novelty in structuring dialogue and integrating background knowledge for reasoning.Specifically, different from previous "structure-less" approaches, our method organizes a dialogue as a "relational graph", using edges to represent relationships between entities.To encode this relational graph, we devise a relational graph convolutional network (R-GCN), which can traverse the graph's topological structure and effectively encode multi-relational knowledge for reasoning.The extensive experiments have justified the effectiveness of our approach over competitive baselines.Moreover, a deeper analysis shows that our model is better at tackling complex questions requiring relational reasoning and defending adversarial attacks with distracting sentences.
Jian Liu 0032, Dianbo Sui, Kang Liu 0001, Jun Zhao 0001
COLING2
2020 FedED: Federated Learning via Ensemble Distillation for Medical Relation Extraction
abstract
Unlike other domains, medical texts are inevitably accompanied by private information, so sharing or copying these texts is strictly restricted.However, training a medical relation extraction model requires collecting these privacy-sensitive texts and storing them on one machine, which comes in conflict with privacy protection.In this paper, we propose a privacypreserving medical relation extraction model based on federated learning, which enables training a central model with no single piece of private local data being shared or exchanged.Though federated learning has distinct advantages in privacy protection, it suffers from the communication bottleneck, which is mainly caused by the need to upload cumbersome local parameters.To overcome this bottleneck, we leverage a strategy based on knowledge distillation.Such a strategy uses the uploaded predictions of ensemble local models to train the central model without requiring uploading local parameters.Experiments on three publicly available medical relation extraction datasets demonstrate the effectiveness of our method.
Dianbo Sui, Yubo Chen 0001, Jun Zhao 0001, Yantao Jia, Yuantao Xie, Weijian Sun
EMNLP (1)1
2019 Leverage Lexical Knowledge for Chinese Named Entity Recognition via Collaborative Graph Network
abstract
Dianbo Sui, Yubo Chen, Kang Liu, Jun Zhao, Shengping Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dianbo Sui, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Shengping Liu
EMNLP/IJCNLP (1)1