Linyao Yang

dblp:228/4749 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-0826-9453ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 From Instruction-Following to Structured Reasoning: Long Chain-of-Thought SFT for Geoscience QA
Shi Luo, Xu Jiao, Linyao Yang
ICIC3
2026 Retrieval augmented generation for open-set entity alignment with large language models
Linyao Yang, Xiao Wang 0002, Weiping Ding 0001, Hongyang Chen 0001, Long Chen 0001
Expert Syst. Appl.1
2026 Enhancing graph learning with quantum structural encoding
Gaoxiang Chen, Hongyang Chen 0001, Jingsong Lv, Linyao Yang
Neurocomputing5
2025 The Emperor's New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFT
abstract
Recent advances in large language models (LLMs) have yielded impressive gains on mathematical reasoning benchmarks via supervised fine-tuning (SFT).However, the brittleness of these models under input perturbations has cast doubt on whether such improvements reflect genuine reasoning abilities or merely superficial alignment with expected output formats.We investigate the mechanisms behind SFT improvements in small-scale LLMs, addressing four key questions: (1) Are performance gains primarily due to format alignment rather than reasoning?(2) Can high-quality supervision encourage genuine reasoning?(3) Does scaling data shift learning from format alignment to deeper reasoning?(4) Are format alignment gains consistent across model sizes and architectures?Through controlled experiments, we find that most performance improvements arise from format alignment rather than genuine reasoning enhancement.Moreover, SFT's effectiveness is strongly influenced by the alignment between the base model's inductive biases and the teacher model's output distribution, rather than the teacher's raw strength.Finally, scaling up training data offers diminishing returns and does not fundamentally alter the model's reasoning behavior.These findings suggest that current SFT practices may overestimate the reasoning abilities of LLMs and underscore the need for more rigorous evaluation methods.
Linyao Yang, Jian-Tao Huang, Yafei Lu, Zhenhui Jessie Li, Guirong Xue
EMNLP1
2025 TemPrompt: Multi-task prompt learning for temporal relation extraction in RAG-based crowdsourcing systems
Jing Yang 0044, Linyao Yang, Xiao Wang 0002, Long Chen 0005, Fei-Yue Wang 0001
Neurocomputing3
2025 Exploring Latent Transferability of feature components
Zhengshan Wang, Long Chen 0001, Juan He 0006, Linyao Yang, Fei-Yue Wang 0001
Pattern Recognit.4
2024 Open-Set Entity Alignment using Large Language Models with Retrieval Augmentation
abstract
Recent years have witnessed remarkable advance-ments in entity alignment, which endeavors to identify entities that represent the same real-world objects across different knowledge graphs (KGs). Nonetheless, prevailing approaches predominantly operate within closed-domain scenarios, rendering them inadequate for handling unmatchable entities. To address this challenge, we propose a retrieval augmented large language model framework (RALLM) to leverage the reasoning capacities of large language models (LLMs) to achieve open-set entity alignment, which not only enables the identification of equivalent entities for matchable entities but also addresses the identification of unmatchable ones. Specifically, we propose a novel retrieval augmentation method that leverages both textual and structural information of entities to retrieve potential equivalent candidates. Subsequently, we employ an iterative process to prompt the LLM to discern the equivalence between the retrieved candidate entity and the entity requiring alignment. To mitigate issues related to many-to-one alignment prediction and enhance alignment efficacy, we devise a memory mechanism to store highly confident aligned entity pairs and provide reminders to the LLM when a candidate entity has been matched. Our experimental findings underscore the superior performance of RALLM, highlighting the potential of LLMs in facilitating open-set entity alignment tasks.
Linyao Yang, Hongyang Chen 0001, Xiao Wang 0002, Yonglin Tian, Xingyuan Dai, Fei-Yue Wang 0001
SMC1
2024 Give us the Facts: Enhancing Large Language Models With Knowledge Graphs for Fact-Aware Language Modeling
abstract
Recently, ChatGPT, a representative large language model (LLM), has gained considerable attention. Due to their powerful emergent abilities, recent LLMs are considered as a possible alternative to structured knowledge bases like knowledge graphs (KGs). However, while LLMs are proficient at learning probabilistic language patterns and engaging in conversations with humans, they, like previous smaller pre-trained language models (PLMs), still have difficulty in recalling facts while generating knowledge-grounded contents. To overcome these limitations, researchers have proposed enhancing data-driven PLMs with knowledge-based KGs to incorporate explicit factual knowledge into PLMs, thus improving their performance in generating texts requiring factual knowledge and providing more informed responses to user queries. This paper reviews the studies on enhancing PLMs with KGs, detailing existing knowledge graph enhanced pre-trained language models (KGPLMs) as well as their applications. Inspired by existing studies on KGPLM, this paper proposes enhancing LLMs with KGs by developing knowledge graph-enhanced large language models (KGLLMs). KGLLM provides a solution to enhance LLMs’ factual reasoning ability, opening up new avenues for LLM research.
Linyao Yang, Hongyang Chen 0001, Zhao Li 0007, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2023 Parallel Reasoning Based on ACP Method for Power Grid Dispatching
abstract
Multi-source heterogeneous knowledge collaboration is the technical foundation for establishing a complete knowledge base. A parallel reasoning framework based on the ACP method is proposed to establish a complete knowledge base. The contribution of this framework is three-fold. First, it provides a virtual experimental platform for generating the artificial data needed for missing knowledge extraction by constructing an artificial system. Second, it generates artificial big data and organizes it into a knowledge graph to achieve structured representation and storage of system control knowledge by carrying out calculation experiments related to missing scene knowledge. Finally, it achieves unbiased application and update of knowledge through parallel execution, completing the optimization and control of the actual system. Parallel reasoning provides an effective technical means for multi-source knowledge collaboration and provides strong support for building knowledge-enhanced complex system control systems. The effectiveness of parallel reasoning is verified through experiments.
Yancai Xu, Linyao Yang, Fenghua Zhu, Xiao Wang 0002, Fei-Yue Wang 0001
SMC2
2023 HackGAN: Harmonious Cross-Network Mapping Using CycleGAN With Wasserstein-Procrustes Learning for Unsupervised Network Alignment
abstract
Network alignment (NA) that identifies equivalent nodes across networks is an effective tool for integrating knowledge from multiple networks. The state-of-the-art NA methods learn inter-network node similarities based on labeled anchor links, which are costly, time-consuming, and difficult to acquire. Therefore, a few unsupervised network alignment (UNA) methods propose solving NA problems without anchor links. However, most existing UNA methods rely on discriminative attributes to capture nodes’ similarities and are hard to obtain optimal one-to-one alignments. Toward these issues, this article proposes a novel method named HackGAN to solve the UNA problem solely based on the structural information. Specifically, HackGAN represents nodes with embeddings based on an unsupervised graph neural network (GNN) to capture their global and local structural features. After that, it initializes mapping functions to transform the embedding spaces of different networks into the same vector space by iteratively solving the Wasserstein–Procrustes problem. The mapping functions are then refined by an adversarial model with cycle-consistency and Sinkhorn distance losses to obtain optimized one-to-one mappings. Based on the distances between mapped embeddings, accurate and robust results are obtained with a collective alignment algorithm. Experimental comparisons on both synthetic and real-world datasets demonstrate the superiority of HackGAN.
Linyao Yang, Xiao Wang 0002, Jun Jason Zhang, Jun Yang 0019, Yancai Xu, Jiachen Hou, Kejun Xin, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2022 Fast and Progressive Misbehavior Detection in Internet of Vehicles Based on Broad Learning and Incremental Learning Systems
abstract
In recent years, deep learning (DL) has been widely used in vehicle misbehavior detection and has attracted great attention due to its powerful nonlinear mapping ability. However, because of the large number of network parameters, the training processes of these methods are time consuming. Besides, the existing detection methods lack scalability; thus, they are not suitable for Internet of Vehicles (IoV) where new data are constantly generated. In this article, the concept of the broad learning system (BLS) is innovatively introduced into vehicle misbehavior detection. In order to make better use of vehicle information, key features are first extracted from the collected raw data. Then, a BLS is established, which is able to calculate the connection weight of the network efficiently and effectively by ridge regression approximation. Finally, the system can be updated and refined by an incremental learning algorithm based on the newly generated data in IoV. The experimental results show that the proposed method performs much better than DL or traditional classifiers, and could update and optimize the old model fastly and progressively while improving the system’s misbehavior detection accuracy.
Xiao Wang 0002, Yushan Zhu, Shuangshuang Han, Linyao Yang, Haixia Gu, Fei-Yue Wang 0001
IEEE Internet Things J.4
2021 HackRL: Reinforcement learning with hierarchical attention for cross-graph knowledge fusion and collaborative reasoning
Linyao Yang, Xiao Wang 0002, Yuxin Dai, Kejun Xin, Xiaolong Zheng 0001, Weiping Ding 0001, Jun Jason Zhang, Fei-Yue Wang 0001
Knowl. Based Syst.1
2021 An IVC-Based Nuclear Emergency Parallel Evacuation System
abstract
Nuclear emergency evacuation is challenging and dangerous, with time constraints, resource limitations, and radiation exposure risks. The development of the Internet of Things (IoT) and artificial intelligence enables us to build an intelligent evacuation system to help mitigate this problem a great deal. In this article, we design the nuclear emergency parallel evacuation system based on the artificial systems (A), computational experiments (C), and parallel execution (P) approach and intelligent vehicle collaborative systems (IVCs). In this system, the evacuation risks of different regions under various possible scenarios are simulated and evaluated in the artificial systems. With data adversarially generated from the artificial systems and collected from sensors all over the area, an optimization model is proposed to find the optimal evacuation plans for emergent evacuation scenarios in the computational experiments. Eventually, the most suitable running strategy of autonomous buses will be selected and carried out based on the parallel execution in accordance with the real scene. A case study is conducted, and results indicate that the system can serve as an efficient tool for future nuclear emergency evacuation planning.
Linyao Yang, Xin Liu 0022, Yancai Xu, Jiazhen Lin, Xiao Wang 0002, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.2
2020 Parallel Internet of Vehicles: ACP-Based System Architecture and Behavioral Modeling
abstract
Vehicles in Internet of Vehicles (IoV) exchange information about location, environment, infotainment, as well as social information with other units via vehicular communication networks. This makes IoV with key social entities in the human-vehicle-infrastructure-roadside units (RSUs) as integrated intelligent transportation systems. Therefore, by identifying the cyber-physical-social features of IoV and presenting its complexity issues of both engineering and social dimensions, this article proposes and introduces the concept, architecture, and applications of parallel IoV (PIoV). Three main components of PIoV are demonstrated, which are artificial IoV to learn and describe the physical IoV, computation experiments to evaluate and predict the consequences and values of driving strategies, and parallel execution to prescribe the operation of the physical IoV. PIoV makes it possible to achieve safe, smart, effective, and efficient transportation management and control. The final objective of PIoV is to equip IoV with descriptive, predictive, and prescriptive intelligence based on the parallel intelligence approach.
Xiao Wang 0002, Shuangshuang Han, Linyao Yang, Lingxi Li 0001
IEEE Internet Things J.3
2020 Pedestrian Choice Modeling and Simulation of Staged Evacuation Strategies in Daya Bay Nuclear Power Plant
abstract
Considering the distances to exits, exits' capacities, the sizes of queues at exits, distances to the nuclear power plant, as well as individual characteristics, the exit choice model for pedestrians in the plume planning area is established based on a random forest model. This model is trained and verified with the survey data of residents around the Daya Bay Nuclear Power Plant collected from a serious game-based questionnaire system. Combining the pedestrian choice with the agent-based pedestrian behavior simulation model, the evacuation process of a nuclear accident is simulated. Based on the detailed evacuation simulation model, a comparative experiment is performed to evaluate the staged evacuation strategy in such scenarios. Simulation results indicate that staged evacuation may not be the best strategy all the time, and the number of groups highly impacts its performance.
Linyao Yang, Xiao Wang 0002, Jun Jason Zhang, Min Zhou 0003, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2018 LoRa on the Move: Performance Evaluation of LoRa in V2X Communications
abstract
Recent years have witnessed much interest in Low Power Wide Area (LPWA) technologies, which are gaining unprecedented momentum and commercial interest towards the realisation of the Internet of Things (IoT). Long Range (LoRa), as a representative LPWA technology, has the potential to satisfy the growing demand for the longer range and larger amount connectivities in vehicular communication networks. In this paper, LoRa is firstly applied into two typical vehicular networks, namely Vehicle-to-Infrastructure (V2I)and Vehicle-to-Vehicle (V2V), and performance of LoRa schemes with different parameter configurations are evaluated and compared. Further, Monte Carlo simulations indicate that the schemes equipped with higher bandwidth or lower spreading factor exhibit significant advantages in combating the fast fading caused by Doppler effect in networks.
Shuangshuang Han, Linyao Yang, Fei-Yue Wang 0001, Hui Zhang 0001
Intelligent Vehicles Symposium3