Xinchuan Li

dblp:26/6582 · also XinChuan Li · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9465-0965ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Incorporating hydrological constraints with deep learning for streamflow prediction
Yilin Duan, Hong Yao, Xinchuan Li, Shengwen Li
Expert Syst. Appl.4
2024 A Clustering-Oriented Method for Open-Domain Named Entity Recognition
abstract
Named Entity Recognition (NER), as a fundamental task in natural language understanding, has garnered widespread attention. However, most existing research assumes that labeled data is available, limiting the ability of NER models to perform in open-domain scenarios. Especially in real-world applications, the absence of labeled data for novel entity classes is an inevitable scenario. To address this issue, this paper proposes CONER, a Clustering-Oriented Named Entity Recognition method, which mainly consists of two modules: Label Center Clustering (LCC) and Pseudo Label Learning (PLL). LCC uses label classes as clustering centers to guide the learning process of entity representations, thereby achieving high cohesion of entities of the same type in the embedding space. PLL generates pseudo labels based on the cluster-friendly embeddings generated by LCC and uses pairwise similarity learning for the discriminative representation of the novel classes. To balance the learning pace for both seen classes and novel classes, LCC is employed as a pre-training procedure to initialize the model, and then LCC is jointly optimized with PLL. The experimental results on OntoNotes and AnatEM datasets demonstrate that the proposed model outperforms the current zero-shot NER models in open-domain NER, which validates the effectiveness of the approach presented in this paper in open-domain scenarios.
Diange Zhou, Yilin Duan, Xinchuan Li, Hong Yao
CCGrid4
2023 Anchors-Based Incremental Embedding for Growing Knowledge Graphs
abstract
Knowledge graph embedding aims to transform the entities and relations of triplets into the low-dimensional vectors. Previous methods are oriented towards the static knowledge graphs, in which all entities and relations are assumed to be known and only some unknown triplets need to be predicted. However, the real-world knowledge graphs can grow dynamically, and some new knowledge are often added. To embed the new knowledge into the space of original knowledge graph, the classic models have to perform the entire re-embedding with including the new and original knowledge. This causes heavy computational burden for embedding. To address this problem, this study proposes a new model of anchors-based incremental embedding (ABIE) to implement the dynamical embedding for the growing knowledge graph. According to ABIE, every knowledge graph has some key entities, called anchors, which can fix the embedding space of knowledge graph. When some new knowledge is added into the graph, only a few updated entities and relations are embedded into the embedding space with the help of anchors, and the entire re-embedding on the whole graph is not necessary. By this way, the computational burden of embedding caused by the growth of knowledge graph is reduced significantly.
Lijun Dong, Dongyang Zhao, Xiaoai Zhang, Xinchuan Li, Xiaojun Kang, Hong Yao
IEEE Trans. Knowl. Data Eng.4
2023 Computable Access Control: Embedding Access Control Rules Into Euclidean Space
abstract
Access control is one of the most basic techniques to ensure the security of the information system. The traditional access controls of information systems are usually performed based on the traversals or queries of rules. However, with the increasing complexity of information systems, massive data, and open environments bring great workload and risk for the traditional methods. This study proposes a model of embedding-based computable access control (ECAC), by employing the idea of representation learning in artificial intelligence. According to ECAC, access control rules can be embedded into a Euclidean vector space, and the security of arbitrary behavior can be computed by numerical vector operations, without any traditional querying or traversing of rules, and thus the workload of access control is reduced. Furthermore, by the embedding-based computation, the security of unknown behaviors can be predicted. Potentially, due to the use of numerical vectors instead of traditional semantic symbols, the risk of privacy leakage via semantics can be reduced. Finally, as the first embedding-based access control model, the effectiveness of ECAC is evaluated and concluded by the experiment-based analyses and discussions.
Lijun Dong, Tiejun Wu, Xinchuan Li
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Location-aware neural graph collaborative filtering
abstract
Collaborative filtering (CF) is initiated by representing users and items as vectors and seeks to describe the relationship between users and items at a profound level, thus predicting users’ preferred behavior. To address the issue that previous research ignored higher-order geographical interactions hidden in users’ historical behaviors, this paper proposes a location-aware neural graph collaborative filtering model (LA-NGCF), which incorporates location information of items for improving prediction performance. The model characterizes the interactions between items based on spatial decay law from a graph perspective and designs two strategies to capture the interaction effects of users and items considering node heterogeneity. An optimized loss function with spatial distances of items is also developed in the model. Extensive experiments are conducted on three publicly available real-world datasets to examine the effectiveness of our model. Results show that LA-NGCF achieves competitive performances compared with several state-of-the-art models, which suggests that location information of items is beneficial for improving the performance of personalized recommendations. This paper offers an approach to incorporate weighted interactions between items into CF algorithms and enriches the methods of utilizing geographical information for artificial intelligence applications.
Shengwen Li, Chenpeng Sun, Renyao Chen, Xinchuan Li, Qingzhong Liang, Junfang Gong, Hong Yao
Int. J. Geogr. Inf. Sci.4
2022 Dynamic hypergraph neural networks based on key hyperedges
Xiaojun Kang, Xinchuan Li, Hong Yao, Xiaoyue Peng, Tiejun Wu, Shihua Qi, Lijun Dong
Inf. Sci.2
2022 Graph Neural Networks with Information Anchors for Node Representation Learning
Chao Liu 0007, Xinchuan Li, Dongyang Zhao, Shaolong Guo, Xiaojun Kang, Lijun Dong, Hong Yao
Mob. Networks Appl.2
2019 Distant-Supervised Relation Extraction with Hierarchical Attention Based on Knowledge Graph
abstract
Relation Extraction is concentrated on finding the unknown relational facts automatically from the unstructured texts. Most current methods, especially the distant supervision relation extraction (DSRE), have been successfully applied to achieve this goal. DSRE combines knowledge graph and text corpus to corporately generate plenty of labeled data without human efforts. However, the existing methods of DSRE ignore the noisy words within sentences and suffer from the noisy labelling problem; the additional knowledge is represented in a common semantic space and ignores the semantic-space difference between relations and entities. To address these problems, this study proposes a novel hierarchical attention model, named the Bi-GRU-based Knowledge Graph Attention Model (BG2KGA) for DSRE using the Bidirectional Gated Recurrent Unit (Bi-GRU) network. BG2KGA contains the word-level and sentence-level attentions with the guidance of additional knowledge graph, to highlight the key words and sentences respectively which can contribute more to the final relation representations. Further-more, the additional knowledge graph are embedded in the multi-semantic vector space to capture the relations in 1-N, N-1 and N-N entity pairs. Experiments are conducted on a widely used dataset for distant supervision. The experimental results have shown that the proposed model outperforms the current methods and can improve the Precision/Recall (PR) curve area by 8% to 16% compared to the state-of-the-art models; the AUC of BG2KGA can reach 0.468 in the best case.
Hong Yao, Lijun Dong, Shiqi Zhen, Xiaojun Kang, Xinchuan Li, Qingzhong Liang
ICTAI5
2019 Testing Scientific Software with Invariant Relations: A Case Study
abstract
Adequately testing scientific software is essential to the quality of the software. However, it is a grand challenge due to the oracle problem. Metamorphic testing has shown its effectiveness for alleviating the problem. But the effectiveness of metamorphic testing is highly dependent on the quality of metamorphic relations that are developed for testing the software. In this paper, we propose a framework for iteratively developing metamorphic relations for adequately testing scientific software. The basic idea is to refine metamorphic relations that are loosely defined to those that can be verified with only limited number of cases so that the relations can be accurately tested. We explain the framework through testing a scientific software system that is used for modeling light scattering of particles. Based on domain knowledge and general guidelines, a group of metamorphic relations are first identified and tested. According to testing results, the metamorphic relations are refined step by step until each of them can be accurately tested. In particular, an invariant transform is applied to the metamorphic relations for significantly reducing the number of cases that can satisfy the relations. Finally the relations are further transformed with a carefully defined hash function to ensure each of the metamorphic relations can be automatically verified. The proposed approach truly solves the oracle problem and greatly improves the effectiveness of metamorphic testing. Its effectiveness is evaluated by mutation testing and demonstrated by new problems found in the software.
Junhua Ding 0001, Xinchuan Li, Xin-Hua Hu
QRS2
2019 A-GNN: Anchors-Aware Graph Neural Networks for Node Embedding
Chao Liu 0007, Xinchuan Li, Dongyang Zhao, Shaolong Guo, Xiaojun Kang, Lijun Dong, Hong Yao
QSHINE2
2018 An Approach for Validating Quality of Datasets for Machine Learning
abstract
There are basically two ways for improving the accuracy of machine learning: building relevant machine learning models, and providing high quality datasets for training the models. Significant efforts have been made in designing powerful machine learning models. Furthermore, many open-source datasets have been created for machine learning research. However, research on assessing the impact of the quality of a dataset on the accuracy of a machine learning system has not received attention. In this paper, we present an experimental study to show how the quality of datasets impact the accuracy of machine learning models. We discovered a common problem in datasets that could greatly impact the accuracy of machine learning. This problem could also exist in many other machine learning systems, especially those that are developed using crowd-sourced datasets. This problem is difficult to detect using traditional validation approaches. We propose a novel technique based on metamorphic testing for validating a machine learning system together with its training and testing data. The key to metamorphic testing is to create tests that will adequately test the system. We propose an approach for creating such tests. The effectiveness of the proposed approach is demonstrated through a case study of automated classification of biological cell images.
Junhua Ding 0001, Xinchuan Li
IEEE BigData2
2017 Augmentation and evaluation of training data for deep learning
abstract
Deep learning is an important technique for extracting value from big data. However, the effectiveness of deep learning requires large volumes of high quality training data. In many cases, the size of training data is not large enough for effectively training a deep learning classifier. Data augmentation is a widely adopted approach for increasing the amount of training data. But the quality of the augmented data may be questionable. Therefore, a systematic evaluation of training data is critical. Furthermore, if the training data is noisy, it is necessary to separate out the noise data automatically. In this paper, we propose a deep learning classifier for automatically separating good training data from noisy data. To effectively train the deep learning classifier, the original training data need to be transformed to suit the input format of the classifier. Moreover, we investigate different data augmentation approaches to generate sufficient volume of training data from limited size original training data. We evaluated the quality of the training data through cross validation of the classification accuracy with different classification algorithms. We also check the pattern of each data item and compare the distributions of datasets. We demonstrate the effectiveness of the proposed approach through an experimental investigation of automated classification of massive biomedical images. Our approach is generic and is easily adaptable to other big data domains.
Junhua Ding 0001, Xinchuan Li, Venkat N. Gudivada
IEEE BigData2