EDBT 2026 Demo / reviewers in the wild / expert
Weijian Sun
dblp:157/9221
· DBLP profile ↗
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-5494-5171ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language ModelsabstractYing Nie, Binwei Yan, Tianyu Guo, Hao Liu, Haoyu Wang, Wei He, Binfan Zheng, Weihao Wang, Qiang Li, Weijian Sun, Yunhe Wang, Dacheng Tao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Binwei Yan, Tianyu Guo 0001, Wei He 0001, Binfan Zheng, Qiang Li 0024, Weijian Sun, Yunhe Wang 0001, Dacheng Tao |
NAACL (Long Papers) | 10 |
| 2023 | KBQA: Accelerate Fuzzy Path Query on Knowledge Graph
Qiheng You, Jincheng Lu, Shizheng Liu, Weijian Sun, Rongqian Zhao |
DEXA (1) | 5 |
| 2023 | Multi-Turn and Multi-Granularity Reader for Document-Level Event ExtractionabstractMost existing event extraction works mainly focus on extracting events from one sentence. However, in real-world applications, arguments of one event may scatter across sentences and multiple events may co-occur in one document. Thus, these scenarios require document-level event extraction (DEE), which aims to extract events and their arguments across sentences from a document. Previous works cast DEE as a two-step paradigm: sentence-level event extraction (SEE) to document-level event fusion. However, this paradigm lacks integrating document-level information for SEE and suffers from the inherent limitations of error propagation. In this article, we propose a multi-turn and multi-granularity reader for DEE that can extract events from the document directly without the stage of preliminary SEE. Specifically, we propose a new paradigm of DEE by formulating it as a machine reading comprehension task (i.e., the extraction of event arguments is transformed to identify the answer span from the document). Beyond the framework of machine reading comprehension, we introduce a multi-turn and multi-granularity reader to capture the dependencies between arguments explicitly and model long texts effectively. The empirical results demonstrate that our method achieves superior performance on the MUC-4 and the ChFinAnn datasets. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Zuyu Zhao, Weijian Sun |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2022 | Document-Level Relation Extraction via Pair-Aware and Entity-Enhanced Representation LearningabstractDocument-level relation extraction aims to recognize relations among multiple entity pairs from a whole piece of article. Recent methods achieve considerable performance but still suffer from two challenges: a) the relational entity pairs are sparse, b) the representation of entity pairs is insufficient. In this paper, we propose Pair-Aware and Entity-Enhanced(PAEE) model to solve the aforementioned two challenges. For the first challenge, we design a Pair-Aware Representation module to predict potential relational entity pairs, which constrains the relation extraction to the predicted entity pairs subset rather than all pairs; For the second, we introduce a Entity-Enhanced Representation module to assemble directional entity pairs and obtain a holistic understanding of the entire document. Experimental results show that our approach can obtain state-of-the-art performance on four benchmark datasets DocRED, DWIE, CDR and GDA. Xiusheng Huang, Yubo Chen 0001, Jun Zhao 0001, Kang Liu 0001, Weijian Sun, Zuyu Zhao |
COLING | 6 |
| 2022 | Enhancing Document-Level Relation Extraction by Entity Knowledge Injection
Xinyi Wang 0010, Zitao Wang, Weijian Sun, Wei Hu 0007 |
ISWC | 3 |
| 2022 | Automatic extraction of building geometries based on centroid clustering and contour analysis on oblique images taken by unmanned aerial vehiclesabstractThis paper introduces a method based on centroid clustering and contour analysis to extract area and height measurements on buildings from the 3D model generated by oblique images. The method comprises three steps: (1) extract the contour plane from the fused data of the digital surface model (DSM) and digital orthophoto map (DOM); (2) identify building contour clusters based on the number of centroids contained in each category determined by mean-shift centroid clustering; (3) remove the mis-identified contours in a given building contour cluster by a contour analysis and obtain the geometric information of the building using map algebra. The proposed approach was tested against four datasets. Compared with other results, the detection has effective completeness, correctness, quality, and higher geometric accuracy. The maximum average relative error of building height and area extraction is less than 8%. The method is fast for a large-scale collection of building attributes and improves the applicability of oblique photography in GIS. Weijian Sun |
Int. J. Geogr. Inf. Sci. | 3 |
| 2021 | Operation Diagnosis on Procedure Graph: The Task and DatasetabstractUsers usually consult the manufacturers or the internet when they encounter operation questions with an electronics product. In this paper, we explore to represent an operation question as a procedure graph and formulate the problem of operation diagnosis as two sub-tasks, namely error node detection, and correction, on top of the graph. We construct the first benchmark for this task and propose a transformer-based model to integrate external knowledge and context information to enhance the performance. Experimental results show the effectiveness of our proposed model. Ruipu Luo, Qin Chen 0001, Zhongyu Wei, Weijian Sun, Shuang Tang |
CIKM | 6 |
| 2020 | FedED: Federated Learning via Ensemble Distillation for Medical Relation ExtractionabstractUnlike other domains, medical texts are inevitably accompanied by private information, so sharing or copying these texts is strictly restricted.However, training a medical relation extraction model requires collecting these privacy-sensitive texts and storing them on one machine, which comes in conflict with privacy protection.In this paper, we propose a privacypreserving medical relation extraction model based on federated learning, which enables training a central model with no single piece of private local data being shared or exchanged.Though federated learning has distinct advantages in privacy protection, it suffers from the communication bottleneck, which is mainly caused by the need to upload cumbersome local parameters.To overcome this bottleneck, we leverage a strategy based on knowledge distillation.Such a strategy uses the uploaded predictions of ensemble local models to train the central model without requiring uploading local parameters.Experiments on three publicly available medical relation extraction datasets demonstrate the effectiveness of our method. Dianbo Sui, Yubo Chen 0001, Jun Zhao 0001, Yantao Jia, Yuantao Xie, Weijian Sun |
EMNLP (1) | 6 |
| 2020 | Global-to-Local Neural Networks for Document-Level Relation ExtractionabstractRelation extraction (RE) aims to identify the semantic relations between named entities in text.Recent years have witnessed it raised to the document level, which requires complex reasoning with entities and mentions throughout an entire document.In this paper, we propose a novel model to document-level RE, by encoding the document information in terms of entity global and local representations as well as context relation representations.Entity global representations model the semantic information of all entities in the document, entity local representations aggregate the contextual information of multiple mentions of specific entities, and context relation representations encode the topic information of other relations.Experimental results demonstrate that our model achieves superior performance on two public datasets for document-level RE.It is particularly effective in extracting relations between entities of long distance and having multiple mentions. Difeng Wang, Wei Hu 0007, Ermei Cao, Weijian Sun |
EMNLP (1) | 4 |
| 2020 | PathQG: Neural Question Generation from FactsabstractExisting research for question generation encodes the input text as a sequence of tokens without explicitly modeling fact information.These models tend to generate irrelevant and uninformative questions.In this paper, we explore to incorporate facts in the text for question generation in a comprehensive way.We present a novel task of question generation given a query path in the knowledge graph constructed from the input text.We divide the task into two steps, namely, query representation learning and query-based question generation.We formulate query representation learning as a sequence labeling problem for identifying the involved facts to form a query and employ an RNN-based generator for question generation.We first train the two modules jointly in an end-to-end fashion, and further enforce the interaction between these two modules in a variational framework.We construct the experimental datasets on top of SQuAD and results show that our model outperforms other state-of-the-art approaches, and the performance margin is larger when target questions are complex.Human evaluation also proves that our model is able to generate relevant and informative questions. 1 Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Zengfeng Huang, Weijian Sun, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 5 |
| 2016 | Characterizing User Mobility from the View of 4G Cellular NetworkabstractThe mobility models obtained from mobile data are expected to affect numbers of fields, including urban planning, road traffic engineering, human sociology, epidemiology of infectious diseases, or telecommunication networks. Current user mobility models are mainly extracted from Call Detail Records (CDR) data or WiFi traces. However, CDR data only captures user movements during telephone calls or short message service (SMS) and WiFi traces can not provide intuitive understanding to mobility behavior of cellular network users in large scale. In this paper, we take the first step to investigate if the characteristics of mobility derived from 4G cellular data network is different from the previous findings utilizing other data sources, especially the widely used, CDR based approach. Utilizing our Hadoop based mobile big data processing platform together with a systematic analysis framework of user mobility, we present a comprehensive characterization of the mobility models from the 6 Terabyte (TB) 4G data traffic. Through the comparison with CDR models and 3G models from the view of user occurrence patterns, user movement patterns and dominant locations of users, we find that the 4G data traffic can provide finer granularity of mobility and location information. Weijian Sun, Dandan Miao, Xiaowei Qin, Guo Wei 0001 |
MDM | 1 |