EDBT 2026 Demo / reviewers in the wild / expert
Xianyu Bao
dblp:123/7150
· DBLP profile ↗
13ranked-venue papers
0as first author
10since 2021 · last 2023
0000-0003-1590-083XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Anomaly Detection in Directed Dynamic Graphs via RDGCN and LSTAN
Mark Junjie Li, Zukang Gao, Xianyu Bao, Meiting Li |
ICANN (3) | 4 |
| 2023 | Topic Modeling for Short Texts via Adaptive P$\acute{o}$lya Urn Dirichlet Multinomial Mixture
Mark Junjie Li, Rui Wang 0136, Xianyu Bao, Jueying He, Lijuan He |
ICONIP (14) | 4 |
| 2023 | Semi-Supervised Event Extraction Incorporated With Topic Event FrameabstractSupervised Meta-event extraction suffers from two limitations: (1) The extracted meta-events only contain local semantic information and do not present the core content of the text; (2) model performance is easily degraded because of labeled samples with insufficient number and poor quality. To overcome these limitations, this study presents an approach called frame-incorporated semi-supervised topic event extraction (FISTEE), which aims to extract topic events containing global semantic information. Inspired by the frame-based knowledge representation, a topic event frame is developed to integrate multiple meta-events into a topic event. Combined with the tri-training algorithm, a strategy for selecting unlabeled samples is designed to expand the training sets, and labeling models based on conditional random field (CRF) are constructed to label meta-events. The experimental results show that the event extraction performance of FISTEE is better than supervised learning-based approaches. Furthermore, the extracted topic events can present the core content of the text. Gong-Qing Wu, Zhuochun Miao, Shengjie Hu, Yinghuan Wang, Zan Zhang 0002, Xianyu Bao |
J. Database Manag. | 6 |
| 2023 | Crowdsourcing Truth Inference via Reliability-Driven Multi-View Graph EmbeddingabstractCrowdsourcing truth inference aims to assign a correct answer to each task from candidate answers that are provided by crowdsourced workers. A common approach is to generate workers’ reliabilities to represent the quality of answers. Although crowdsourced triples can be converted into various crowdsourced relationships, the available related methods are not effective in capturing these relationships to alleviate the harm to inference that is caused by conflicting answers. In this research, we propose aReliability-drivenMulti-viewGraphEmbedding framework forTruthinference (TiReMGE), which explores multiple crowdsourced relationships by organically integrating worker reliabilities into a graph space that is constructed from crowdsourced triples. Specifically, to create an interactive environment, we propose a reliability-driven initialization criterion for initializing vectors of tasks and workers as interactive carriers of reliabilities. From the perspective of multiple crowdsourced relationships, a multi-view graph embedding framework is proposed for reliability information interaction on a task-worker graph, which encodes latent crowdsourced relationships into vectors of workers and tasks for reliability update and truth inference. A heritable reliability updating method based on the Lagrange multiplier method is proposed to obtain reliabilities that match the quality of workers for interaction by a novel constraint law. Our ultimate goal is to minimize the Euclidean distance between the encoded task vector and the answer that is provided by a worker with high reliability. Extensive experimental results on nine real-world datasets demonstrate that TiReMGE significantly outperforms the nine state-of-the-art baselines. Gong-Qing Wu, Xingrui Zhuo, Xianyu Bao, Xuegang Hu, Richang Hong, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Crowdsourcing Truth Inference Based on Label Confidence ClusteringabstractTruth inference can help solve some difficult problems of data integration in crowdsourcing. Crowdsourced workers are not experts and their labeling ability varies greatly; therefore, in practical applications, it is difficult to determine whether the labels collected from a crowdsourcing platform are correct. This article proposes a novel algorithm called truth inference based on label confidence clustering (TILCC) to improve the quality of integrated labels for the single-choice classification problem in crowdsourcing labeling tasks. We obtain the label confidence via worker reliability, which is calculated from multiple noise labels using a truth discovery method, and then we generate the clustering features and use the K-means algorithm to cluster all the tasks into K different clusters. Each cluster corresponds to a specific class, and the tasks in the cluster are assigned a label. Compared with the performances of six state-of-the-art methods, MV, ZenCrowd, PM, CATD, GLAD, and GTIC, on 12 randomly selected real-world datasets, the performance of our algorithm showed many advantages: no need to set complex parameters, faster running speed, and significantly higher accuracy. Gong-Qing Wu, Liangzhu Zhou, Jiazhu Xia, Lei Li 0002, Xianyu Bao, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2023 | TIRA: Truth Inference via Reliability Aggregation on Object-Source GraphabstractCrowdsourcing platforms collect massive dirty claims that are provided by sources for crowdsourced objects, which prompts truth inference to be proposed for crowdsourcing data denoising. Although current graph-based truth-inference methods achieve remarkable success by capturing complex crowdsourcing relationships, they typically suffer from two challenges: 1) They fail to obtain complete crowdsourcing relationships because of the structural limitations of crowdsourcing relationship graphs; 2) Their vector initialization methods for objects and sources are disturbed by claim noise, which limits them from obtaining correct object and source semantics. To cope with these challenges, we propose a novelTruth-Inference method viaReliabilityAggregation (TIRA) on an object-source graph. Specifically, we propose a hierarchical graph auto-encoder to adapt to a reasonable object-source graph, which enables TIRA to capture complete crowdsourcing relationships from multiple perspectives. To better guide TIRA, we design a vector initialization method based on source reliabilities to map the denoised claims to a representation space of objects and sources. Finally, TIRA aggregates the reliability information on an object-source graph to generate object embeddings for truth inference. We conducted extensive experiments on 12 real-world datasets. The experimental results demonstrate that our method significantly outperforms 12 state-of-the-art baselines in terms of the$accuracy$and$weighted\_{F}1$. Gong-Qing Wu, Xingrui Zhuo, Liangzhu Zhou, Xianyu Bao, Richang Hong, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | DuSAG: An Anomaly Detection Method in Dynamic Graph Based on Dual Self-attention
Weiqin Lin, Xianyu Bao, Mark Junjie Li, Zukang Gao |
ICANN (1) | 2 |
| 2021 | CmaGraph: A TriBlocks Anomaly Detection Method in Dynamic Graph Using Evolutionary Community Representation Learning
Weiqin Lin, Xianyu Bao, Mark Junjie Li |
ICANN (1) | 2 |
| 2021 | BTGAN: Training GAN with Balanced Triplet Loss and Two-Branch ArchitectureabstractTripletGAN is a variant of Generative Adversarial Network (GAN) by replacing the classification loss of discriminator with a triplet loss. Although TripletGAN delivers better mode coverage than vanilla GAN thanks to the characteristics of adversarial triplet loss that maximizes the embedding distance between generated samples, its adversarial training method suffers from the drawback that some generated images tend to deviate from the real sample distribution and noisy images are produced as we increase the number of iterations of training. In this paper, we propose an adversarially balanced triplet loss with four dynamic coefficients to achieve a trade-off between the quality and the diversity of generated samples. We also design a novel network architecture to provide GANs with an auto-encoding ability. Extensive experiments demonstrate the effectiveness of our proposed methods in terms of alleviating the problem in TripletGAN and the superiority in terms of reconstruction over some methods that directly train generator and encoder such as O-GAN. Simin Yu, Kuntian Zhang, Chuan Xiao 0001, Xianyu Bao, Joshua Zhexue Huang, Mark Junjie Li |
IJCNN | 4 |
| 2021 | Attentive interaction-driven entity resolution over multi-source web information
Ying He 0008, Gong-Qing Wu, Desheng Cai, Shengjie Hu, Xianyu Bao, Xuegang Hu |
Neurocomputing | 5 |
| 2019 | Chinese Temporal Expression Recognition Combining Rules with a Statistical Model
Mengmeng Huang, Jiazhu Xia, Xianyu Bao, Gong-Qing Wu |
ICIC (3) | 3 |
| 2016 | Combining user-based and global lexicon features for sentiment analysis in twitterabstractGenerally speaking, sentiment lexicons employed in the majority of current sentiment analysis systems are trained globally from public data stream source or other large independent corpus. However, sentiments are rather subjective and personal states of mind that the individuality and diversity of characteristics, particular writing habit and idiolect could play a crucial role in the judgment of sentiment expressed by a specific user. In this paper, we present a novel feature construction method to combine user-based and global lexicon features in sentiment analysis for short social media text. After the creation of user-based sentiment lexicons from user-timeline corpus, a rule-based fusing approach is adopted subsequently to generate user-based lexicon features in combination with general lexicon features. Experiments show that user-based features may capture potential user preferences hence adjusting the bias caused by representing an individual's sentiment with an averaged lexicon score, and our proposed method yield better results in comparison with some of the state-of-the-art sentiment analysis systems in twitter. Yujiu Yang 0001, Xianyu Bao, Biqing Huang |
IJCNN | 3 |
| 2015 | A semantic approach for text clustering using WordNet and lexical chainsabstractTraditional clustering algorithms do not consider the semantic relationships among words so that cannot accurately represent the meaning of documents. To overcome this problem, introducing semantic information from ontology such as WordNet has been widely used to improve the quality of text clustering. However, there still exist several challenges, such as synonym and polysemy, high dimensionality, extracting core semantics from texts, and assigning appropriate description for the generated clusters. In this paper, we report our attempt towards integrating WordNet with lexical chains to alleviate these problems. The proposed approach exploits ontology hierarchical structure and relations to provide a more accurate assessment of the similarity between terms for word sense disambiguation. Furthermore, we introduce lexical chains to extract a set of semantically related words from texts, which can represent the semantic content of the texts. Although lexical chains have been extensively used in text summarization, their potential impact on text clustering problem has not been fully investigated. Our integrated way can identify the theme of documents based on the disambiguated core features extracted, and in parallel downsize the dimensions of feature space. The experimental results using the proposed framework on reuters-21578 show that clustering performance improves significantly compared to several classical methods. Yonghe Lu, HuiYou Chang, Xianyu Bao |
Expert Syst. Appl. | 5 |