EDBT 2026 Demo / reviewers in the wild / expert
Xiaokun Yang
dblp:39/8513
· DBLP profile ↗
15ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwiftBot: A Decentralized Platform for LLM-Powered Federated Robotic Task Execution
YueMing Zhang, Zhengxiong Li, Fangtian Zhong, Xiaokun Yang, Hailu Xu |
CCGrid | 5 |
| 2026 | A Hierarchical Methodology for Hardware Design Comparison in HPC WorkloadsabstractAs Moore's law slows down, developers face difficult choices between low-level HDLs (Verilog, VHDL) offering fine-grained control and higher-level tools (HLS and Chisel) promising improved productivity. While high-level tools accelerate development, performance gaps persist compared to expert HDL implementations. Prior studies emphasize end-to-end performance, offering limited insight into why tools excel or where performance diverges in the design hierarchy. We introduce a hierarchical framework for comparing hardware generation tools by decomposing HPC kernels (FFT, GEMM, QR factorization) into reusable primitives (MAC arrays, butterflies, permutations, reduction trees). Across Verilog, Chisel, and Vivado HLS, we built an automated tool flow and synthesized ~1, 500 variants on AMD Alveo U250, measuring resource utilization and frequency. We derived theoretical bounds for validation. Verilog achieves the highest frequency and lowest resource usage; Chisel performs comparably (5--15% gap), while HLS shows a 20--40% gap. All tools operate within bounds for well-structured designs. Crucially, performance divergence arises during primitive assembly, indicating that high-level tools require better composition optimization. This reproducible framework provides actionable insights and is extensible to other tools, domains, and FPGA architectures. Doru-Thom Popovici, Mario Vega, Angelos Ioannou, Fabien Chaix, Dania Susanne Mosuli, Blair Reasoner, Tan Nguyen 0001, Xiaokun Yang, John Shalf |
FPGA | 8 |
| 2026 | Scalable Quantum Circuit Simulation via Circuit Cutting and FPGA Acceleration
Xiaokun Yang, Jeremy W. Turner, Cameron D. DiSomma, Yunhe Feng, Xuechen Zhang 0001, Vipin Chaudhary |
HPDC | 1 |
| 2024 | Hardware Generation on Trigonometric FunctionsabstractThis paper presents hardware generation for accelerating various floating-point (FP) trigonometric functions, including sine, cosine, and arctangent. The Chisel Hardware Construction Language (HCL) is used to develop parameterized and flexible designs for these functions. The hardware generator supports multiple design architectures with configurable parameters, such as precision (e.g., 16-bit, 32-bit, 64-bit, and 128-bit), iteration count, and pipeline depth, allowing for customization of hardware resource utilization, latency, speed, and accuracy. Paul Wong, Dania Susanne Mosuli, Xuechen Zhang 0001, Xiaokun Yang |
IEEE Big Data | 4 |
| 2022 | Research on the Potential Mechanism of Rhizoma Drynariae in the Treatment of Periodontitis Based on Network Pharmacology
Caixia Xu, Xiaokun Yang, Pengyong Han, Zhengwei Li 0001 |
ICIC (2) | 2 |
| 2021 | FPGA acceleration on a multi-layer perceptron neural network for digit recognition
Isaac Westby, Xiaokun Yang, Hailu Xu |
J. Supercomput. | 2 |
| 2020 | On Fundamental Principles for Thermal-Aware Design on Periodic Real-Time Multi-Core SystemsabstractWith the exponential rise of the transistor count in one chip, the thermal problem has become a pressing issue in computing system design. While there have been extensive methods and techniques published for design optimization with thermal awareness, there is a need for more rigorous and formal thermal analysis in designing real-time systems and applications that demand a strong exception guarantee. In this article, we analytically prove a series of fundamental properties and principles concerning the RC thermal model, peak temperature identification, and peak temperature reduction for periodic real-time systems, which are general enough to be applied on 2D and 3D multi-core platforms. These findings enhance the worst-case temperature predictability in runtime scenarios, as well as help to develop more effective thermal management policy, which is key to thermal-constrained periodic real-time system design. Shi Sha, Ajinkya S. Bankar, Xiaokun Yang, Wujie Wen, Gang Quan |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Improving AES Core Performance via an Advanced ASBUS ProtocolabstractSecurity is becoming a de-facto requirement of System-on-Chips (SoC), leading up to a significant share of circuit design cost. In this article, we propose an advanced SBUS protocol (ASBUS), to improve the data feeding efficiency of the Advanced Encryption Standard (AES) encrypted circuits. As a case study, the direct memory access (DMA) combined with AES engine and memory controller are implemented as our design-under-test (DUT) using field-programmable gate arrays (FPGA). The results show that our presented ASBUS structure outperforms the AXI-based design for cipher tests. As an example, the 32-bit ASBUS design costs less in terms of hardware resources and achieves higher throughput (1.30 ×) than the 32-bit AXI implementation, and the dynamic energy consumed by the ASBUS cipher test is reduced to 71.27% compared with the AXI test. Xiaokun Yang, Wujie Wen |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2017 | Design of a pre-scheduled data bus for advanced encryption standard encrypted system-on-chipsabstractThis paper proposes a high efficiency data bus (DBUS) for Advanced Encryption Standard (AES) encrypted system-on-chips (SoCs). Using DBUS, the data sequence can be pre-selected for AES encryption/decryption, so that the state buffering and rescheduling overhead can be reduced. FPGA results show that the DBUS based design lowers the dynamic energy to 66.93%, and achieves up to 1.30 times higher valid throughput compared with the Advanced eXensible Interface (AXI) based implementation. Xiaokun Yang, Wujie Wen |
ASP-DAC | 1 |
| 2017 | Energy minimization for on-line real-time scheduling with reliability awareness
Ming Fan 0001, Qiushi Han, Xiaokun Yang |
J. Syst. Softw. | 3 |
| 2016 | Exploiting cyclic features of walking for pedestrian dead reckoning with unconstrained smartphonesabstractPedestrian dead reckoning (PDR) is a promising complementary technique to balance the requirements on both accuracy and costs in outdoor and indoor positioning systems. In this paper, we propose a unified framework to comprehensively tackle the three sub problems involved in PDR, including step detection and counting, heading estimation and step length estimation, based on sequentially rotating the device (reference) frame to the Earth (reference) frame through sensor fusion. To be specific, a robust step detection and counting algorithm is devised according to vertical angular velocities and turns out to be tolerant of various smartphone placements; then, a zero velocity update (ZUPT) based algorithm is leveraged to calibrate the measurements in the Earth frame; on these grounds, the heading and step length are further estimated by exploiting the cyclic features of walking. A thorough and extensive experimental analysis is conducted and confirms the effectiveness and advantages of the proposed PDR framework as well as the corresponding algorithms. Baoqi Huang, Guodong Qi, Xiaokun Yang, Long Zhao 0004, Han Zou |
UbiComp | 3 |
| 2016 | A novel bus transfer mode (AS transfer) and a performance evaluation methodology
Xiaokun Yang, Nansong Wu, Jean Andrian |
Integr. | 1 |
| 2015 | A High-Performance On-Chip Bus (MSBUS) Design and VerificationabstractThis brief proposes a high-performance system-on-chip bus protocol termed the master-slave bus (MSBUS). Considering the inevitable tradeoff among area, throughput and energy efficiency, the control bus is developed as a low-cost and low-power bus, and the data bus is created as a high-throughput full-duplex bus with the feature of block data transfer. To evaluate the bus performance, we create four analytical models including transfer time consumption (TC), wire efficiency (WE), valid data bandwidth (VDB) and dynamic energy efficiency. Then, the advanced high-performance bus-, advanced eXensible interface (AXI)-, and MSBUS-based direct memory access (DMA) are developed as a case study of hardware implementation. It is observed that MSBUS DMA costs less hardware resources and achieves higher performance, especially in the block transfer mode. For instance, the results from both the analytical models and the practical tests show that the TC of MSBUS is close to 63% of the AXI, the WE and VDB of MSBUS are almost 2.3 and 1.6 times of the AXI respectively, and the energy consumption is half of AXI in the block transfer mode. Xiaokun Yang, Jean Andrian |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Co-Existence Analysis of LTE Micro Cell and LTE Out-Band BackhaulabstractIn this paper, a stand-alone LTE based out-band backhaul is designed for urban area in NLOS environment, and the interference and compatibility issues relating to co-existence of LTE micro cell and co-located LTE out-band backhaul are investigated by a static system level simulator. Feasibility and recommendation of installing out-band backhaul are analyzed according to the simulation results. Xinglin Wang, Xiaokun Yang |
VTC Fall | 2 |
| 2010 | Study on Co-Existence of Macro WCDMA Cell and Micro HSUPA CellabstractIn this paper, a static system level simulator is used to investigate the capacity loss in case of co-existence of macro WCDMA and micro HSUPA. The capacity loss under different frequency spacing or ACIRs is obtained. The simulation results show that the micro HSUPA cell has little impact on the macro WCDMA cell with frequency spacing of 5 MHz, while the macro WCDMA cell has larger impact on the micro HSUPA cell. Based on simulation, suggestions of frequency planning and parameter setting are given to avoid large capacity loss due to co-existence. Xinglin Wang, Xiaokun Yang, Xiaojin Zhang 0003 |
VTC Spring | 3 |