Xinkai Zhang

dblp:218/2818 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving
abstract
Trial-and-error is a fundamental strategy for humans to solve complex problems and a necessary capability for Artificial Intelligence (AI) systems operating in real-world environments. Although several trial-and-error AI techniques have recently been proposed, most of them rely on simple heuristics designed by researchers and achieve limited performance gains. The core issue is the absence of appropriate data: current models cannot learn from detailed records of how humans actually conduct trial-and-error in practice. To address this gap, we introduce a data annotation platform and a corresponding dataset, termed Trial-and-Error Collection (TEC). The platform records users' complete trajectories across multiple trials and collects their reflections after receiving error feedback. Using this platform, we record the problem-solving processes of 46 participants on 58 tasks, resulting in 5,370 trial trajectories along with error reflections across 41,229 webpages. With this dataset, we observe that humans achieve substantially higher accuracy compared to LLMs, which demonstrates that humans are more effective in trial-and-error than LLMs. We believe that the TEC platform and dataset provide a valuable foundation for understanding human trial-and-error behavior and for developing more capable AI systems. Platform and dataset are publicly available. https://github.com/Serendipity0429/TEC.
Xinkai Zhang, Jingtao Zhan, Yiqun Liu 0001, Qingyao Ai
SIGIR1
2025 A Deterministic Event-Driven Software Bus for Virtualized Industrial Edge Applications
abstract
To address the scalability limitations and maintenance challenges of traditional PLC-based systems, modern industrial deployments are increasingly adopting multi-core edge servers and virtualization technologies such as containers for resource isolation and dynamic deployment. In this context, high-concurrency local communication between containers must meet stringent real-time requirements, including low latency and predictability. This paper proposes a deterministic Software Bus tailored for virtualized industrial edge systems. The design integrates shared memory pooling, centralized event scheduling, and dynamic routing to enable zero-copy data exchange and event-driven communication across different industrial edge applications in containers. Experimental results demonstrate that the Software Bus incurs minimal overhead during connection initialization, supporting rapid deployment and dynamic module integration. In one-way message transmission, it significantly reduces average latency compared to traditional TCP and UDS mechanisms, while maintaining stable maximum latency under varying payloads.
Bintao Yu, Xinkai Zhang, Wenbin Dai
IECON3
2025 A High-Performance Fault-Tolerant Scheduling Mechanism for Virtualized Industrial Edge Applications
abstract
The core requirement of industrial automation systems is continuous and stable running. System design for the next-generation Industrial Internet must meet the high reliability requirements of industrial automation systems. This paper presents a design method for a dual-mode redundancy mechanism in virtualized runtime systems for the Industrial Internet. Through reliability analysis and comparison of multiple redundancy architectures, the applicability of the dual-mode redundancy mechanism in industrial scenarios is established. Compared with traditional hardware redundancy solutions, this method features lightweight deployment and fast iterative updates, significantly reducing system resource consumption and labor costs. Experimental data show that as the number of connected devices increases, the complexity growth curve of the virtualized runtime system is significantly lower than traditional hardware redundancy mechanisms, and performance indicators meet the stability and reliability requirements of next-generation industrial scenarios.
Xinkai Zhang, Bintao Yu, Wenbin Dai
IECON1
2025 Prioritized Deterministic Real-Time Execution Semantics for Industrial Edge Applications Based on IEC 61499 Event Types
abstract
Industrial edge computing brings intelligent applications to traditional industrial automation systems. The support of multi-core processors has dramatically enhanced the computational capabilities of these edge devices. An edge device can simultaneously execute multiple tasks with diverse real-time constraints, from motion control to machine learning-based fault prediction. The IEC 61499 function blocks (FBs) have been adopted for the system-level modeling language for industrial edge applications. To provide prioritized deterministic concurrent execution for real-time control, computational, and other tasks, event types are enhanced with different priorities and real-time attributes in accordance with the IEC 61499 standard. The concurrent scheduling algorithm with event types is proposed to ensure determinism for the real-time execution of IEC 61499 FB networks. Finally, the proposed execution semantics are proven to significantly reduce execution time while guaranteeing the determinism of both operations and information technology tasks with multicore support.
Kaiyun Qin, Xinkai Zhang, Wenbin William Dai
IEEE Trans. Ind. Informatics3
2024 Applying Embedded Multi-Core Control Technologies for the Next Generation Industrial Edge Applications
abstract
In the realm of industrial edge computing, higher-configured terminal processors are gradually replacing inefficient ones. Edge computing devices encounter challenges such as replicating rapid software development and software upgrade and maintenance. In response to these issues, an edge-cloud two-tier architecture has been proposed. In the previous article, while conducting research on the edge system architecture, we contemplated on how to effectively utilize the CPU virtualization capabilities of edge devices to achieve elastic expansion of device resources. This paper will commence from the perspective of efficient utilization of multi-core CPU resources. By conducting analyses of relevant research technology cases and identifying the deficiencies in the research, our edge computing multi-core control system architecture is re-engineered. Through comparative experiments of multi-core task testing, it is concluded that the multicore processor scheduling system designed in this paper can significantly enhance the task execution efficiency of edge computing systems.
Xinkai Zhang, Bintao Yu, Wenbin Dai
IECON1
2023 Applying Embedded Virtualization Technologies for the Next Generation Industrial Edge Applications
abstract
With the increasing availability of computing, storage, and network resources, there is a growing interest in efficiently utilizing these resources. Virtualization technology plays a crucial role in cloud computing as it provides flexibility and reliability in managing resources. In this paper, we conducted experiments with several virtualization technologies in the context of industrial edge computing. These technologies aim to reduce overall hardware costs. In addition to explaining the software architecture and describing various open-source solutions, we have also designed an architecture that incorporates embedded virtualization technology. This architecture simplifies and reconstructs the traditional industrial system architecture, enabling efficient utilization and maximizing the allocation of resources for edge-end devices. Finally, we implemented ACRN virtualization technologies in industrial edge-end devices and conducted experiments to validate their effectiveness.
Xinkai Zhang, Dali Yang, Wenbin Dai
IECON1
2021 Computing for Control and Control for Computing
abstract
Computing can be thought of as a service provided to a system to yield actionable tasks enacted by physical hardware. But rarely is control thought to be in the service of enhancing computation. Consideration of that perspective is what motivates co-regulation, our framework for holistic cyber-physical control of autonomous vehicles. In this paper we elaborate on how co-regulation will enable the next generation of autonomous vehicles precisely because it considers computation as an enabler and consumer of autonomous behavior. We report on the latest advances in this space showing how co-regulation exceeds results in event-triggered, self-triggered, and fixed-rate control strategies yielding more robustness and adaptivity to changing and uncertain conditions - a requirement for next-gen autonomous vehicles. We then describe a co-regulated decision making algorithm based on Markov Decision Processes showing how full consideration of computational resource allocation can increase decision-making capabilities in uncertain environments.
Xinkai Zhang, Justin M. Bradley
DATE1
2020 Trua: Efficient Task Replication for Flexible User-defined Availability in Scientific Grids
abstract
Failure is inevitable in scientific computing. As scientific applications and facilities increase their scales over the last decades, finding the root cause of a failure can be very complex or at times nearly impossible. Different scientific computing customers have varying availability demands as well as a diverse willingness to pay for availability. In contrast to existing solutions that try to provide higher and higher availability in scientific grids, we propose a model called Task Replication for Userdefined Availability (Trua). Trua provides flexible, user-defined, availability in scientific grids, allowing customers to express their desire for availability to computational providers. Trua differs from existing task replication approaches in two folds. First, it relies on the historic failure information collected from the virtual layer of the scientific grids. The reliability model for the failures can be represented with a bimodal Johnson distribution which is different from any existing distributions. Second, it adopts an anomaly detector to filter out anomalous failures; it additionally adopts novel selection algorithms to mitigate the effects of temporary and spatial correlations of the failures without knowing the root cause of the failures. We apply the Trua on real-world traces collected from the Open Science Grid (OSG). Our results show that the Trua can successfully meet user-defined availability demands.
Zhe Zhang 0003, Brian Bockelman, Derek Weitzel, Xinkai Zhang, Hamid Vakilzadian, David Swanson
CCGRID4