Xiangyu Wang 0016

dblp:02/6128-16 · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0001-9843-5982ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Counterfactual-Driven Zero-Shot Classifier Expansion
abstract
Zero-shot classifier expansion aims to adapt existing model to new, unseen classes. It utilizes class attributes or textual descriptions to learn a mapping from the semantic space to the classifier's weight space, without requiring new visual training data. However, the learning process for this mapping relies solely on correlating semantic patterns with their corresponding classifier weights and lacks explicit modeling of inter-class differences. This makes it difficult for the model to capture the critical discriminative features required to define classification boundaries. To overcome this limitation, we reframe the problem from a causal perspective and introduce a novel framework driven by counterfactuals. Our method first generates factual descriptions alongside corresponding inter-class counterfactuals to pinpoint the causal attributes essential for classification, then refines these representations via a mutual purification process, and finally leverages a novel separation loss to explicitly push the factual and counterfactual classifier weights apart. This strategy forces the model to forge clearer and more discriminative classification boundaries, achieving more accurate and robust classification. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods.
Xiangyu Wang 0016, Yanze Gao, Changxin Rong, Lyuzhou Chen, Derui Lyu, Xiren Zhou, Taiyu Ban, Huanhuan Chen 0001
AAAI1
2026 Fault Diagnosis of Irregular Sequences by Adjoint Learning in Continuous-Time Model Space
abstract
Fault Diagnosis (FD) on sequential data suffers from irregular sampling (with missing values), limited training data, and varying underlying environments. In response, this paper proposes FD by adjoint learning in continuous-time model space. Model-Space Learning employs well-fitted models that capture data's dynamics (i.e., changing information) as more stable and concise representations of the original data. The Continuous-Time Reservoir Computing Network (CT-Res) is first introduced, which embeds Ordinary Differential Equation (ODE) within the reservoir-based hidden layer to govern continuous-time hidden-state evolution, naturally handling irregular sampling without relying on fixed time steps and effectively capturing intrinsic data dynamics. By fitting each sequence via CT-Res and representing it with the fitted model, the original sequences are mapped from the data space into the continuous-time model space. We further develop an adjoint learning strategy by incorporating a discrete-time "adjoint Echo State Network (ESN)" that shares structure and parameters with CT-Res, thus enabling efficient training by bypassing the computationally intensive ODE solver, with joint optimization of fitting accuracy and class discrimination in the model space. Experiments on multiple FD benchmarks highlight the effectiveness and efficiency of our study, particularly with missing values and scarce training data.
Xiren Zhou, Chuyang Wei, Ao Chen 0002, Shikang Liu, Xiangyu Wang 0016, Huanhuan Chen 0001
AAAI5
2026 Reliable Causal Mining via Knowledge Evaluation for Order-based Graph Constraints
Shunjie Wu, Xingjian Lin, Yanze Gao, Changxin Rong, Xiangyu Wang 0016
ICIC (4)5
2026 Conditional diffusion for causal inference with state space representation
Yanmin Li, Xiangyu Wang 0016, Weidong Bao 0001, Jibing Wu, Hang Zhang 0008, Lihua Liu 0002
Knowl. Based Syst.2
2026 Autonomous Causal Discovery: Evaluating LLMs' Priors and Constraint Strategies for Reliability
abstract
Expert-guided Causal Structure Learning (CSL) incorporates prior knowledge to improve the accuracy of causal discovery, yet the acquisition of such knowledge is often restricted by the availability of human experts. While Large Language Models (LLMs) provide an alternative source of causal priors, LLM-derived knowledge can be inconsistent with the true causal structure due to hallucinations or contextual misinterpretations. This paper introduces a structural constraint measurement framework, which defines constraint strength and constraint quality to describe reliability and effectiveness, enabling a systematic evaluation of LLM-derived constraints. Using this framework, we evaluate five categories of structural constraints: Edge Existence (EEC), Edge Forbidden (EFC), Path Existence (PEC), Path Forbidden (PFC), and Order Constraints (OC). Our theoretical and empirical analyses demonstrate that while EEC offers high constraint strength, it exhibits low quality when derived from LLMs; conversely, PFC and OC provide a balanced trade-off between search-space pruning and reliability. Building on these insights, we propose a two-level CSL optimization framework that partitions the search space by node order and refines the structure using global path constraints. The results show that this framework provides an effective way to incorporate noisy LLM-derived priors into CSL, particularly in settings where expert knowledge is limited.
Lyuzhou Chen, Xiangyu Wang 0016, Taiyu Ban, Derui Lyu, Qinrui Zhu, Xin Wang 0179, Huanhuan Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Continuous Structure Constraint Integration for Robust Causal Discovery
abstract
Causal discovery aims to infer a Directed Acyclic Graph (DAG) from observational data to represent causal relationships among variables. Traditional combinatorial methods search DAG spaces to identify optimal structures, while recent advances in continuous optimization improve this search process. However, integrating structural constraints informed by prior knowledge into these methods remains a substantial challenge. Existing methods typically integrate prior knowledge in a hard way, demanding precise information about causal relationships and struggling with erroneous priors. Such rigidity can lead to significant inaccuracies, especially when the priors are flawed. In response to these challenges, this work introduces the Edge Constraint Adaptive (ECA) method, a novel approach that softly represents the presence of edges, allowing for a differentiable representation of prior constraint loss. This soft integration can more flexibly adjust to both accurate and erroneous priors, enhancing both robustness and adaptability. Empirical evaluations demonstrate that our approach effectively leverages prior to improve causal structure accuracy while maintaining resilience against prior errors, thus offering significant advancements in the field of causal discovery.
Lyuzhou Chen, Taiyu Ban, Derui Lyu, Kangtao Hu, Xiangyu Wang 0016, Huanhuan Chen 0001
AISTATS6
2025 Instruction Tuning with Data Augmentation for Event Argument Extraction
Chengfei Wang, Xiangyu Wang 0016, Huanhuan Chen 0001
ICIC (23)2
2025 Differentiable Structure Learning with Ancestral Constraints
abstract
Differentiable structure learning of causal directed acyclic graphs (DAGs) is an emerging field in causal discovery, leveraging powerful neural learners. However, the incorporation of ancestral constraints, essential for representing abstract prior causal knowledge, remains an open research challenge. This paper addresses this gap by introducing a generalized framework for integrating ancestral constraints. Specifically, we identify two key issues: the non-equivalence of relaxed characterizations for representing path existence and order violations among paths during optimization. In response, we propose a binary-masked characterization method and an order-guided optimization strategy, tailored to address these challenges. We provide theoretical justification for the correctness of our approach, complemented by experimental evaluations on both synthetic and real-world datasets.
Taiyu Ban, Changxin Rong, Xiangyu Wang 0016, Lyuzhou Chen, Xin Wang 0179, Derui Lyu, Qinrui Zhu, Huanhuan Chen 0001
ICML3
2025 Expanding the Category of Classifiers with LLM Supervision
abstract
Zero-shot learning has shown significant potential for creating cost-effective and flexible systems to expand classifiers to new categories. However, existing methods still rely on manually created attributes designed by domain experts. Motivated by the widespread success of large language models (LLMs), we introduce an LLM-driven framework for class-incremental learning that removes the need for human intervention, termed Classifier Expansion with Multi-vIew LLM knowledge (CEMIL). In CEMIL, an LLM agent autonomously generates detailed textual multi-view descriptions for unseen classes, offering richer and more flexible class representations than traditional expert-constructed vectorized attributes. These LLM-derived textual descriptions are integrated through a contextual filtering attention mechanism to produce discriminative class embeddings. Subsequently, a weight injection module maps the class embeddings to classifier weights, enabling seamless expansion to new classes. Experimental results show that CEMIL outperforms existing methods using expert-constructed attributes, demonstrating its effectiveness for fully automated classifier expansion without human participation.
Derui Lyu, Xiangyu Wang 0016, Taiyu Ban, Lyuzhou Chen, Xiren Zhou, Huanhuan Chen 0001
IJCAI2
2025 Underground Diagnosis in 3D GPR Data by Learning in CuCoRes Model Space
abstract
Ground Penetrating Radar (GPR) provides detailed subterranean insights. Nevertheless, underground diagnosis via GPR is hindered by the fact that training data typically contain only normal samples, along with the complexity of GPR data’s wave-collection characteristics. This paper proposes subsurface anomaly detection within the Cubic Correlation Reservoir Network (CuCoRes) model space. CuCoRes incorporates three reservoirs with spatial correlation adjustment in each direction to adequately and accurately capture multi-directional dynamics (i.e., changing information) within GPR data. Fitting GPR data with CuCoRes and representing data with fitted models, the original GPR data is mapped into a category-discriminative CuCoRes model space, where anomalies could be efficiently identified and categorized based on model dissimilarities. Our approach leverages only limited normal GPR data, easily accessible, to support subsequent anomaly detection and categorization, enhancing its applicability in practical scenarios. Experiments on real-world data demonstrate its effectiveness, outperforming state-of-the-art.
Xiren Zhou, Shikang Liu, Xiangyu Wang 0016, Huanhuan Chen 0001
IJCAI4
2025 Fault Diagnosis in REDNet Model Space
abstract
Fault Diagnosis (FD) in time-varying data presents considerations such as limited training data, intra- and inter-dimensional correlations, and constraints of training time. In response, this paper introduces FD in the Reservoir-Embedded-Directional Network (REDNet) model space. Model-oriented methods utilize well-fitted networks or functions, denoted as "models" that capture data's changing information, as more stable and parsimonious representations of the data. Our approach employs REDNet for data fitting, wherein multiple reservoirs are organized along intrinsic correlation directions to establish intra- and inter-dimensional dependencies, thereby capturing multi-directional dynamics in high-dimensional data. Representing each data instance with an independently fitted REDNet model maps these instances into a class-separable REDNet model space, where FD could be performed on the models rather than the original data. Concentrating on the data-intrinsic dynamics, our method achieves rapid training speeds, and maintains robust performance even with minimal training data. Experiments on several datasets demonstrate its effectiveness.
Xiren Zhou, Ziyu Tang, Shikang Liu, Ao Chen 0002, Xiangyu Wang 0016, Huanhuan Chen 0001
IJCAI5
2025 Pattern-Guided Adaptive Prior for Structure Learning
abstract
Learning the causality between variables, known as DAG structure learning, is critical yet challenging due to issues such as insufficient data and noise. While prior knowledge can improve the learning process and refine the DAG structure, incorporating prior knowledge is not without pitfalls. In particular, we find that the gap between the imprecise prior knowledge and the exact weights modeled by existing methods may result in deviation in edge weights. Such deviation can subsequently cause significant inaccuracies when learning the DAG structure. This paper addresses this challenge by providing a theoretical analysis of the impact of deviation in edge weights during the optimization process of structure learning. We identify two special graph patterns that arise due to the deviation and show that their occurrence increases as the degree of deviation grows. Building on this analysis, we propose the Pattern-Guided Adaptive Prior (PGAP) framework. PGAP detects these patterns as structural signals during optimization and adaptively adjusts the structure learning process to counteract the identified weight deviation, thereby improving the integration of prior knowledge. Experiments verify the effectiveness and robustness of the proposed method.
Lyuzhou Chen, Yanze Gao, Xiangyu Wang 0016, Derui Lyu, Taiyu Ban, Xin Wang 0179, Xiren Zhou, Huanhuan Chen 0001
NeurIPS4
2025 Improving knowledge graphs via data-based causal structures
Derui Lyu, Xiangyu Wang 0016, Lyuzhou Chen, Taiyu Ban, Kaiming Xiao, Lihua Liu 0002
Knowl. Inf. Syst.2
2025 Mitigating Prior Errors in Causal Structure Learning: A Resilient Approach via Bayesian Networks
abstract
Causal structure learning (CSL), a prominent technique for encoding cause-and-effect relationships among variables, through Bayesian Networks (BNs). Although recovering causal structure solely from data is a challenge, the integration of prior knowledge, revealing partial structural truth, can markedly enhance learning quality. However, current methods based on prior knowledge exhibit limited resilience to errors in the prior, with hard constraint methods disregarding priors entirely, and soft constraints accepting priors based on a predetermined confidence level, which may require expert intervention. To address this issue, we propose a strategy resilient to edge-level prior errors for CSL, thereby minimizing human intervention. We classify prior errors into different types and provide their theoretical impact on the Structural Hamming Distance (SHD) under the presumption of sufficient data. Intriguingly, we discover and prove that the strong hazard of prior errors is associated with a unique acyclic closed structure, defined as "quasi-circle". Leveraging this insight, a post-hoc strategy is employed to identify the prior errors by its impact on the increment of "quasi-circles". Through empirical evaluation on both real and synthetic datasets, we demonstrate our strategy's robustness against prior errors. Specifically, we highlight its substantial ability to resist order-reversed errors while maintaining the majority of correct prior.
Lyuzhou Chen, Taiyu Ban, Xiangyu Wang 0016, Derui Lyu, Huanhuan Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 LLM-Driven Causal Discovery via Harmonized Prior
abstract
Traditional domain-specific causal discovery relies on expert knowledge to guide the data-based structure learning process, thereby improving the reliability of recovered causality. Recent studies have shown promise in using the Large Language Model (LLM) as causal experts to construct autonomous expert-guided causal discovery systems through causal reasoning between pairwise variables. However, their performance is hampered by inaccuracies in aligning LLM-derived causal knowledge with the actual causal structure. To address this issue, this paper proposes a novel LLM-driven causal discovery framework that limits LLM’s prior within a reliable range. Instead of pairwise causal reasoning that requires both precise and comprehensive output results, the LLM is directed to focus on each single aspect separately. By combining these distinct causal insights, a unified set of structural constraints is created, termed a harmonized prior, which draws on their respective strengths to ensure prior accuracy. On this basis, we introduce plug-and-play integrations of the harmonized prior into mainstream categories of structure learning methods, thereby enhancing their applicability in practical scenarios. Evaluations on real-world data demonstrate the effectiveness of our approach.
Taiyu Ban, Lyuzhou Chen, Derui Lyu, Xiangyu Wang 0016, Qinrui Zhu, Huanhuan Chen 0001
IEEE Trans. Knowl. Data Eng.4
2025 Large-Scale Hierarchical Causal Discovery via Weak Prior Knowledge
abstract
Causal discovery faces significant challenges as the number of hypotheses grows exponentially with the number of variables. This complexity becomes particularly daunting when dealing with large sets of variables. We introduce a novel divide-and-conquer method that uniquely handles this challenge. The existing division strategies often rely on conditional independency (CI) tests or data-driven clustering to split variables, which can suffer from the typical data scarcity in large-scale settings, thus leading to inaccurate division results. The proposed method overcomes this by implementing a data-independent division strategy, which constructs a prior structure, informed by potential causal relationships identified using a Large Language Model (LLM), to guide recursively dividing variables into sub-sets. This approach avoids the impact of data insufficiency and is robust against potential incompleteness in the prior structure. In the merging phase, we adopt a score-based refinement strategy to address fake causal links caused by hidden variables in sub-sets, which eliminates edges in the intersected parts of sub-sets to optimize the score of local structures. While maintaining both correctness and completeness under the faithfulness assumption, this novel merging approach demonstrates enhanced performance than the conventional CI-test based merging strategy in practical scenarios. Empirical evaluations on various large-scale datasets demonstrate the proposed approach's superior accuracy and efficiency compared to existing causal discovery methods.
Xiangyu Wang 0016, Taiyu Ban, Lyuzhou Chen, Derui Lyu, Qinrui Zhu, Huanhuan Chen 0001
IEEE Trans. Knowl. Data Eng.1
2024 Constructing a Knowledge-Guided Mental Health Chatbot with LLMs
Xi Fan, Lishan Yang 0004, Xiangyu Wang 0016, Derui Lyu, Huanhuan Chen 0001
ACML3
2024 Multi-Model Consistency for LLMs' Evaluation
abstract
This paper introduces an evaluation method for large language models (LLMs) based on multi-model factual cognition consistency. Traditional evaluation methods, especially in terms of factuality assessments, face challenges in constructing extensive domain-specific question sets and relying on specific model answers. These methods fall short in the face of dynamic and diverse model development. To overcome these limitations, the proposed approach does not depend on a fixed set of standard answers. Instead, it utilizes the responses of multiple models to construct a dynamic, relative evaluation benchmark. We first developed a framework to capture and compare the cognitive consistency of different models when addressing specific questions. Subsequently, a dynamic iterative algorithm was designed to evaluate models based on these sets of answers. Experiments across multiple domains demonstrated the effectiveness of this method. This innovative evaluation strategy not only provides a more comprehensive and flexible approach to understanding and assessing the performance of LLMs in various scenarios but also offers practical guidance for future model development and improvement.
Qinrui Zhu, Derui Lyu, Xi Fan, Xiangyu Wang 0016, Qiang Tu, Yibin Zhan, Huanhuan Chen 0001
IJCNN4
2024 Differentiable Structure Learning with Partial Orders
abstract
Differentiable structure learning is a novel line of causal discovery research that transforms the combinatorial optimization of structural models into a continuous optimization problem. However, the field has lacked feasible methods to integrate partial order constraints, a critical prior information typically used in real-world scenarios, into the differentiable structure learning framework. The main difficulty lies in adapting these constraints, typically suited for the space of total orderings, to the continuous optimization context of structure learning in the graph space. To bridge this gap, this paper formalizes a set of equivalent constraints that map partial orders onto graph spaces and introduces a plug-and-play module for their efficient application. This module preserves the equivalent effect of partial order constraints in the graph space, backed by theoretical validations of correctness and completeness. It significantly enhances the quality of recovered structures while maintaining good efficiency, which learns better structures using 90\% fewer samples than the data-based method on a real-world dataset. This result, together with a comprehensive evaluation on synthetic cases, demonstrates our method's ability to effectively improve differentiable structure learning with partial orders.
Taiyu Ban, Lyuzhou Chen, Xiangyu Wang 0016, Xin Wang 0179, Derui Lyu, Huanhuan Chen 0001
NeurIPS3
2024 Leveraging Multisource Label Learning for Underground Object Recognition
abstract
Currently, numerous deep learning (DL) methods have been proposed for the recognition of ground penetrating radar (GPR) B-scan images. Due to the sensitivity of GPR imaging to local underground conditions, DL models trained in other underground environment areas are likely to fail in new areas. Consequently, organizing and labeling new GPR images to train models have become a widely adopted approach in practical applications. However, expert annotation is costly, making it often difficult to collect large-scale datasets with high-quality annotations. Some studies have attempted to improve the quality of image annotations by integrating the efforts of multiple annotators. This process can be constrained by the varying levels of expertise and attention spans of the annotators, which may lead to the presence of errors or contradictions in the provided labels. To address these challenges, this article proposes a method for underground target recognition aimed at multisource annotation tasks. A probability multisource label aggregation (PMLA) module is designed to estimate the reliability of multisource labels, and a label-sensitive regularization (LSR) module is introduced to mitigate the negative impact of potentially erroneous labels on model training. Extensive experiments are conducted on multiple GPR B-scan datasets. The experimental results demonstrate the advantages of the proposed method in handling annotation conflicts and improving the accuracy of underground target recognition.
Derui Lyu, Lyuzhou Chen, Taiyu Ban, Xiangyu Wang 0016, Qinrui Zhu, Xiren Zhou, Huanhuan Chen 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Quality Evaluation of Triples in Knowledge Graph by Incorporating Internal With External Consistency
abstract
The evaluation of knowledge quality (KQ) in multisource knowledge graphs (KGs) is an essential step for many applications, such as fragmented knowledge fusion and knowledge base construction. Many existing quality evaluation methods for multisource knowledge are based on validation from high-quality knowledge bases or statistical analysis of knowledge related to a specific fact from multiple sources, named external consistency (EC)-based methods. However, high-quality KGs are difficult to obtain, and there might exist incorrect knowledge in multisource KGs interfering with KQ evaluation. To address the issue, this article refers to the internal structure of a KG to evaluate the degree to which the contained triples conform to the overall semantic pattern of the KG, such as KG embedding and logic inference-based approaches, defined as internal consistency (IC) evaluation. The IC is integrated with the EC to identify possible incorrect triples and reduce their influences on the KQ evaluation, thus alleviating the interference of incorrect knowledge. The proposed method is verified with multiple datasets, and the results demonstrate that the proposed method could significantly reduce wrong evaluations caused by incorrect knowledge and effectively improve the quality evaluation of triples.
Taiyu Ban, Xiangyu Wang 0016, Lyuzhou Chen, Qiuju Chen, Huanhuan Chen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Research Ideas Discovery via Hierarchical Negative Correlation
abstract
A new research idea may be inspired by the connections of keywords. Link prediction discovers potential nonexisting links in an existing graph and has been applied in many applications. This article explores a method of discovering new research ideas based on link prediction, which predicts the possible connections of different keywords by analyzing the topological structure of the keyword graph. The patterns of links between keywords may be diversified due to different domains and different habits of authors. Therefore, it is often difficult for a single learner to extract diverse patterns of different research domains. To address this issue, groups of learners are organized with negative correlation to encourage the diversity of sublearners. Moreover, a hierarchical negative correlation mechanism is proposed to extract subgraph features in different order subgraphs, which improves the diversity by explicitly supervising the negative correlation on each layer of sublearners. Experiments are conducted to illustrate the effectiveness of the proposed model to discover new research ideas. Under the premise of ensuring the performance of the model, the proposed method consumes less time and computational cost compared with other ensemble methods.
Lyuzhou Chen, Xiangyu Wang 0016, Taiyu Ban, Shikang Liu, Derui Lyu, Huanhuan Chen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Knowledge Verification From Data
abstract
Knowledge verification is an important task in the quality management of knowledge graphs (KGs). Knowledge is a summary of facts and events based on human cognition and experience. Due to the nature of knowledge, most knowledge quality (KQ) management methods are designed by human experts or the characteristics of existing knowledge, which may be limited by human cognition and the quality of existing knowledge. Numerical data contain a wealth of potential information that may be helpful in verifying knowledge, which is rarely explored. However, due to the implicit representation of numerical data to facts as well as the noise in the data, it is challenging to use data to verify the knowledge. Therefore, this article proposes a knowledge verification method, which discovers the correlation and causality from numerical data to validate knowledge and then evaluate the quality of knowledge. Moreover, to address the impact of noise, the method integrates multisource knowledge to jointly evaluate the KQ. Specifically, an iterative update method is designed to update KQ by utilizing the consistency between multisource knowledge while designing knowledge verification factors based on data causality and correlation to manage update process. The method is validated with multiple datasets, and the results demonstrate that the proposed method could evaluate KQ more accurately and has strong robustness to noise in the data.
Xiangyu Wang 0016, Taiyu Ban, Lyuzhou Chen, Derui Lyu, Huanhuan Chen 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Decentralised Knowledge Graph Evolution via Blockchain
abstract
In recent years, knowledge graphs (KGs) have been applied in various domains, where the construction and maintenance of the KGs are usually time- and labor-intensive. In this context, constructing shareable KG through multiple constructors is being attempted to reduce costs. In this collaborative process, security and quality issues are critical. The system for constructing shareable KGs should be capable to recover the KG from most malicious attack and to filter out wrong triples from dynamically submitted ones. Blockchain could naturally prevent malicious tampering with its record data, perfect for solving the security issue. However, the integration of multi-source KGs as well as the quality issue still lacks solutions. To address the issues, this paper proposes a blockchain-based high-quality KG collaborative construction framework to ensure the KG quality in its long-term evolution. The framework is built on the underlying consensus mechanism of the blockchain, adopted to an extensible data structure to store multi-source triples on the distributed ledger. A smart contract is implemented to publish triples, assess the contributor credibility and evaluate triple quality to keep the KG in high-quality. Anti-attack mechanisms are designed to defend against malicious triple submissions. Experiments are conducted demonstrating the effectiveness of the framework.
Xiangyu Wang 0016, Taiyu Ban, Lyuzhou Chen, Yifeng Guan, Derui Lyu, Jian Cheng 0004, Huanhuan Chen 0001, Cyril Leung, Chunyan Miao
IEEE Trans. Serv. Comput.1
2023 Temporal knowledge graph embedding via sparse transfer matrix
Xin Wang 0179, Shengfei Lyu, Xiangyu Wang 0016, Huanhuan Chen 0001
Inf. Sci.3
2023 Knowledge Extraction From National Standards for Natural Resources: A Method for Multi-Domain Texts
abstract
National standards for natural resources (NSNR) plays an important role in promoting efficient use of China's natural resources, which sets standards for many domains such as marine and land resources. Its revision is difficult since standards in different domains may overlap or conflict. To facilitate the revision of NSNR, this paper extracts structural knowledge from the NSNR files to assist its revision. NSNR files are in multi-domain texts, where the traditional knowledge extraction methods could fall short in recalling multi-domain entities. To address this issue, this paper proposes a knowledge extraction method for multi-domain texts, including sub-domain relation discovery (SRD) and domain semantic features fusion (DSFF) module. SRD splits NSNR into sub-domains to facilitate the relation discovery. DSFF integrates relation features in the conditional random field (CRF) model to improve the capability of multi-domain entity recognition. Experimental results demonstrate that the proposed method could effectively extract structural knowledge from NSNR.
Taiyu Ban, Xiangyu Wang 0016, Xin Wang 0179, Jiarun Zhu, Lvzhou Chen, Yizhan Fan
J. Database Manag.2
2023 A distribution-based representation of Knowledge Quality
Xiangyu Wang 0016, Taiyu Ban, Lyuzhou Chen, Qiuju Chen, Huanhuan Chen 0001
Knowl. Based Syst.1
2023 Accurate Label Refinement From Multiannotator of Remote Sensing Data
abstract
The remote sensing (RS) field has an increasing research interest in using deep learning (DL) models to recognize kinds of RS data, leading to a great demand for training data annotation. Due to the high cost of expertise, using nonexperts to label data has become an important way to improve labeling efficiency. Commonly, a single data sample is labeled by multiple annotators and the most voted label is accepted to promise accuracy. But in the RS context, the widely admitted strategy could lose effect. Usually RS data involve considerable classes on account of the complexity of surface environments, which is prone to interclass similarity difficult to distinguish. Annotators without expertise probably make mistakes on these indistinguishable classes, thus causing error voted labels. Although classification of different characteristics in RS data has been widely documented, the nonexpert annotators are unfamiliar with these expertise, and it is difficult to force them to handle specialized labeling skills. To address the issues, this article bases multiannotator label selection on the investigation of annotators’ own ability in distinguishing similar classes of images. A quality evaluation process is designed which weights the labels from capable annotators higher than those from weak ones. By a multi-round quality evaluation algorithm, correct labels could outcompete the wrong ones even disadvantaged in numbers. Experimental results demonstrate the advance of the proposed method on the RS datasets.
Xiangyu Wang 0016, Lyuzhou Chen, Taiyu Ban, Derui Lyu, Yifeng Guan, Xiren Zhou, Huanhuan Chen 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Feature Selection in the Data Stream Based on Incremental Markov Boundary Learning
abstract
Recent years have witnessed the proliferation of techniques for streaming data mining to meet the demands of many real-time systems, where high-dimensional streaming data are generated at high speed, increasing the burden on both hardware and software. Some feature selection algorithms for streaming data are proposed to tackle this issue. However, these algorithms do not consider the distribution shift due to nonstationary scenarios, leading to performance degradation when the underlying distribution changes in the data stream. To solve this problem, this article investigates feature selection in streaming data through incremental Markov boundary (MB) learning and proposes a novel algorithm. Different from existing algorithms focusing on prediction performance on off-line data, the MB is learned by analyzing conditional dependence/independence in data, which uncovers the underlying mechanism and is naturally more robust against the distribution shift. To learn MB in the data stream, the proposal transforms the learned information in previous data blocks to prior knowledge and employs them to assist MB discovery in current data blocks, where the likelihood of distribution shift and reliability of conditional independence test are monitored to avoid the negative impact from invalid prior information. Extensive experiments on synthetic and real-world datasets demonstrate the superiority of the proposed algorithm.
Bingbing Jiang 0001, Xiangyu Wang 0016, Taiyu Ban, Huanhuan Chen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Dynamic Link Prediction for Discovery of New Impactful COVID-19 Research Approaches
abstract
In fighting the COVID-19 pandemic, the main challenges include the lack of prior research and the urgency to find effective solutions. It is essential to accurately and rapidly summarize the relevant research work and explore potential solutions for diagnosis, treatment and prevention of COVID-19. It is a daunting task to summarize the numerous existing research works and to assess their effectiveness. This paper explores the discovery of new COVID-19 research approaches based on dynamic link prediction, which analyze the dynamic topological network of keywords to predict possible connections of research concepts. A dynamic link prediction method based on multi-granularity feature fusion is proposed. Firstly, a multi-granularity temporal feature fusion method is adopted to extract the temporal evolution of different order subgraphs. Secondly, a hierarchical feature weighting method is proposed to emphasize actively evolving nodes. Thirdly, a semantic repetition sampling mechanism is designed to avoid the negative effect of semantically equivalent medical entities on the real structure of the graph, and to capture the real topological structure features. Experiments are performed on the COVID-19 Open Research Dataset to assess the performance of the model. The results show that the proposed model performs significantly better than existing state-of-the-art models, thereby confirming the effectiveness of the proposed method for the discovery of new COVID-19 research approaches.
Xiangyu Wang 0016, Taiyu Ban, Jiarun Zhu, Lyuzhou Chen, Xin Wang 0179, Huanhuan Chen 0001, Cyril Leung, Chunyan Miao
IEEE J. Biomed. Health Informatics1
2021 Unsupervised Change Detection in Satellite Images With Generative Adversarial Network
abstract
Detecting changed regions in paired satellite images plays a key role in many remote sensing applications. The evolution of recent techniques could provide satellite images with very high spatial resolution (VHR) but made it challenging to apply image coregistration, and many change detection methods are dependent on its accuracy.Two images of the same scene taken at different time or from different angle would introduce unregistered objects and the existence of both unregistered areas and actual changed areas would lower the performance of many change detection algorithms in unsupervised condition.To alleviate the effect of unregistered objects in the paired images, we propose a novel change detection framework utilizing a special neural network architecture -- Generative Adversarial Network (GAN) to generate many better coregistered images. In this paper, we show that GAN model can be trained upon a pair of images through using the proposed expanding strategy to create a training set and optimizing designed objective functions. The optimized GAN model would produce better coregistered images where changes can be easily spotted and then the change map can be presented through a comparison strategy using these generated images explicitly.Compared to other deep learning-based methods, our method is less sensitive to the problem of unregistered images and makes most of the deep learning structure.Experimental results on synthetic images and real data with many different scenes could demonstrate the effectiveness of the proposed approach.
Caijun Ren, Xiangyu Wang 0016, Xiren Zhou, Huanhuan Chen 0001
IEEE Trans. Geosci. Remote. Sens.2