EDBT 2026 Demo / reviewers in the wild / expert
Yuan He 0002
dblp:11/1735-2
· DBLP profile ↗
18ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-4087-8905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EATS: Energy-Aware Adaptive Topology Switching for NoCs
Man Wu, Shaswot Shresthamali, Yuan He 0002 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | A2SHE: An anonymous authentication scheme for health emergencies in public venues
Xiao-han Yue, Haoran Si, Haibo Yang 0003, Fucai Zhou, Yuan He 0002 |
Inf. Sci. | 9 |
| 2025 | A face authentication-based searchable encryption scheme for mobile device
Xiao-han Yue, Gang Yi, Haoran Si, Haibo Yang 0003, Yuan He 0002 |
J. Supercomput. | 6 |
| 2024 | DAISM: Digital Approximate In-SRAM Multiplier-Based Accelerator for DNN Training and InferenceabstractDNNs are widely used but face significant computational costs due to matrix multiplications, especially from data movement between the memory and processing units. One promising approach is therefore Processing-in-Memory as it greatly reduces this overhead. However, most PIM solutions rely either on novel memory technologies that have yet to mature or bit-serial computations that have significant performance overhead and scalability issues. Our work proposes an in-SRAM digital multiplier, that uses a conventional memory to perform bit-parallel computations, leveraging multiple wordlines activation. We then introduce DAISM, an architecture leveraging this multiplier, which achieves up to two orders of magnitude higher area efficiency compared to the SOTA counterparts, with competitive energy efficiency. Lorenzo Sonnino, Shaswot Shresthamali, Yuan He 0002, Masaaki Kondo |
DATE | 3 |
| 2023 | Exploiting Data Parallelism in Graph-Based Simultaneous Localization and Mapping: A Case Study with GPU AccelerationsabstractGraph-based simultaneous localization and mapping (G-SLAM) is an intuitive SLAM implementation where graphs are used to represent poses, landmarks and sensor measurements when a mobile robot builds a map of the environment and locates itself in it. Being a very important application employed in many realistic scenarios, estimating the whole environment and all trajectories through solving graph problems for SLAM can incur a large amount of computation and consume a significant amount of energy. For the purpose of improving both performance and energy efficiency, we have unveiled the critical path of the G-SLAM algorithm in this paper and implemented a GPU-based solution to aid it. Furthermore, we have attempted to offload performance-critical components (such as matrix inversions when updating the trajectory) in the G-SLAM process into GPUs through CUDA to exploit data parallelism. With our solution, we observe a speed-up of up to 19.7x and an energy saving of up to 83.7% over a modern workstation class x86 CPU; while on a platform dedicated for edge computing (NVIDIA Jetson Nano), we achieve a speed-up of up to 2.5x and an energy saving of up to 6.4% with its integrated GPU, respectively. Junyuan Zheng, Yuan He 0002, Masaaki Kondo |
HPC Asia | 2 |
| 2022 | Memory Bandwidth Conservation for SpMV Kernels Through Adaptive Lossy Data Compression
Makiko Ito, Takahide Yoshikawa, Yuan He 0002, Masaaki Kondo |
PDCAT | 4 |
| 2022 | GraphDEAR: An Accelerator Architecture for Exploiting Cache Locality in Graph Analytics ApplicationsabstractData structure is the key in Edge Computing where various types of data are continuously generated by ubiquitous devices. Within all common data structures, graphs are used to express relationships and dependencies among human identities, objects, and locations; and they are expected to become one of the most important data infrastructure in the near future. Furthermore, as graph processing often requires random accesses to vast memory spaces, conventional memory hierarchies with caches cannot perform efficiently. To alleviate such memory access bottlenecks in graph processing, we present a solution through vertex accesses scheduling and edge array re-ordering, in parallel with the execution of graph processing application to improve both temporal and spatial locality of memory accesses, especially for edge-centric graphs which are popular means in handling dynamic graphs. Our proposed architecture is evaluated and tested through both trace-based cache simulations and cycle-accurate FPGA-based prototyping. Evaluation results show that our proposal has a potential of significantly reducing the quantity of Miss-Per-Kilo-Instructions (MPKI) for Last Level Cache (LLC) by 56.27% on average. Masaaki Kondo, Yuan He 0002, Ryuichi Sakamoto, Hiroshi Nakamura |
PDP | 3 |
| 2022 | A practical privacy-preserving communication scheme for CAMs in C-ITSabstractIn a Collaborative Intelligent Transportation System (C-ITS), vehicles send collaborative awareness messages (CAMs) carrying information such as speed, position, and identity, to interact with other vehicles and to obtain transportation services. Traditionally, CAMs are broadcasted in plaintext over unsafe networks, which may leak sensitive information of the vehicle. On one hand, many existing works attempting to solve this problem focused on anonymous authentication, but ignored the confidentiality of CAMs. Conversely, some other attempts trying to keep the privacy of CAMs have backward security problems in key management or vehicle revocation. Therefore, in this paper, we present a practical privacy-preserving communication scheme for CAMs in C-ITS. To ensure the privacy of a vehicle's identity, our scheme mainly uses the re-randomizable PS signature method to realize anonymous authentication; and to meet backward security, we use the complete subtree method to realize both key management and vehicle revocation. In more details, our proposal fulfills more security and privacy requirements, such as correctness, anonymity, traceability, revocability, backward security and anti-illegal message dissemination, while also hiding the content of CAMs. Xiao-han Yue, Shuaishuai Zeng, Xibo Wang, Lixin Yang 0002, Yuan He 0002 |
J. Inf. Secur. Appl. | 6 |
| 2021 | Local Traffic-Based Energy-Efficient Hybrid Switching for On-Chip NetworksabstractAdvanced flow control mechanisms employed by modern on-chip networks are the reasons of large energy footprint and long per-hop latency. On the other hand, dated and simpler flow controls such as circuit switching can draw far less power and offer an end-to-end latency analogous to wire delay. In this paper, we present a hybrid flow control mechanism, which mixes both virtual channels and circuit-switching, to provide a latency-competitive and energy-efficient on-chip network design. Contrary to existing hybrid switching designs, our proposal is based on local traffic so that circuits are formed without the knowledge of end-to-end traffic. When compared to on-chip networks with virtual channels, our proposal achieves a very competitive latency per flit, for up to 4% lower, while also dramatically suppressing the energy per flit by up to 18%. Yuan He 0002, Jinyu Jiao, Masaaki Kondo |
PDP | 1 |
| 2021 | An Aggregate Anonymous Credential Scheme in C-ITS for Multi-Service with RevocationabstractCooperative intelligent transport systems (C-ITS) are advanced applications providing various services to enable a more secure, efficient, and accessible transportation system. When requesting services in C-ITS, privacy-preserving is crucial to avoid attempts in tracking vehicles who send continuous messages containing their authentication information. In this paper, we propose a revocable privacy-preserving authentication scheme through attribute-based credentials with complete subtree method. Differing from the existing anonymous credential works in C-ITS, our scheme supports the aggregation of attributes which facilitates the construction of a constant-size credential for multi-service. Not only maintaining the privacy-preserving of identity, the aggregate credential is also able to improve the efficiency of credential presentation. For illegal vehicles, our scheme revokes and disallows them to generate valid signatures via using complete subtree method in the non-revoked epoch. Moreover, we present analyses on aspects of security and performance, respectively. The former focuses on the security goals including privacy-preserving of identity, unlinkability, unforgeability, revocability and anti-replay attacks; while the latter shows that our scheme makes more practical sense for C-ITS. Xiao-han Yue, Lixin Yang 0002, Xibo Wang, Shuaishuai Zeng, Jian Xu 0004, Yuan He 0002 |
TrustCom | 7 |
| 2021 | A Revocable Zone Encryption Scheme with Anonymous Authentication for C-ITSabstractIn Collaborative Intelligent Transportation System (C-ITS), vehicles send collaborative awareness messages (CAMs) carrying information such as speed, position, and identity, to interact with other vehicles, and thus, obtain transportation services. However, CAMs are broadcasted in plaintext over the unsafe network, which causes the leakage of vehicle's sensitive information when maliciously intercepted. At present, there is less scheme for encryption and anonymous authentication CAMs that is suitable for the actual scenario. In this paper, we propose a revocable zone encryption scheme with anonymous authentication. When vehicles enter the zone, the zone manager first completes the anonymous authentication according to the vehicle-generated join request containing a self-delegated certificate. Then, vehicles encrypt CAMs with the session key by symmetric encryption and wrap the session key with the zone key for secure communication. Besides, we use the complete subtree method to satisfy the revocation of the vehicle and zone key management respectively. Finally, simulation experiments and analysis show that our scheme provides stronger privacy-preserving compared with proposals that only support anonymous authentication. Differing from the similar scheme that tries to encrypt CAMs, our scheme not only satisfies more security requirements but also our key management is more feasible. Xiao-han Yue, Shuaishuai Zeng, Xibo Wang, Lixin Yang 0002, Jian Xu 0004, Yuan He 0002 |
TrustCom | 7 |
| 2021 | A Revocable Group Signatures Scheme to Provide Privacy-Preserving Authentications
Xiao-han Yue, Mengzhe Xi, Mingchao Gao, Yuan He 0002, Jian Xu 0004 |
Mob. Networks Appl. | 5 |
| 2020 | Energy-Efficient On-Chip Networks through Profiled Hybrid SwitchingabstractVirtual channel (VC) flow control is the de facto choice for modern networks-on-chip (NoCs) to allow better utilization of the link bandwidth through buffering and packet switching (PS), which are also the sources of large power footprint and long per-hop latency. However, bandwidth can be plentiful for parallel workloads under VC flow control. Thus, dated but simpler mechanisms, such as circuit switching (CS), can help improve the energy efficiency of modern NoCs. In this paper, we propose to apply CS to part of the link bandwidth so that a considerable amount of traffic can be transmitted bufferlessly without routing. Evaluations reveal that this proposal leads to a reduction of energy per flit by up to 32% while also provides very competitive latency when compared to networks under VC flow control. Yuan He 0002, Jinyu Jiao, Masaaki Kondo |
ACM Great Lakes Symposium on VLSI | 1 |
| 2017 | Cooling-Aware Job Scheduling and Node Allocation for Overprovisioned HPC SystemsabstractLimited power budget is becoming one of the most crucial challenges in developing supercomputer systems. Hardware overprovisioning which installs a larger number of nodes beyond the limitations of the power constraint is an attractive way to design next generation supercomputers. In air cooled HPC centers, about half of the total power is consumed by cooling facilities. Reducing cooling power and effectively utilizing power resource for computing nodes are important challenges. It is known that the cooling power depends on the hotspot temperature of the node inlets. Therefore, if we minimize the hotspot temperature, performance efficiency of the HPC system will be increased. One of the ways to reduce the hotspot temperature is to allocate power-hungry jobs to compute nodes whose effect on the hotspot temperature is small. It can be accomplished by optimizing job-to-node mapping in the job scheduler. In this paper, we propose a cooling and node location-aware job scheduling strategy which tries to optimize job-to-node mapping while improving the total system throughput under the constraint of total system (compute nodes and cooling facilities) power consumption. Experimental results with the job scheduling simulation show that our scheduling scheme achieves 1.49X higher total system throughput than the conventional scheme. Yuan He 0002, Masaaki Kondo |
IPDPS | 3 |
| 2016 | Demand-Aware Power Management for Power-Constrained HPC SystemsabstractAs limited power budget is becoming one of the most crucialchallenges in developing supercomputer systems, hardware overprovisioning which installs larger number of nodes beyond the limitations of the power constraint determinedby Thermal Design Power is an attractive way to design extreme-scale supercomputers. In this design, power consumption of each node should be controlled by power-knobs equipped in the hardware such as dynamic voltage and frequency scaling (DVFS) or power capping mechanisms. Traditionally, in supercomputer systems, schedulers determine when and where to allocate jobs. In overprovisioned systems, the schedulers also need to care about power allocation to each job. An easy way is to set a fixed power cap for each job so that the total power consumption is within the power constraint of the system. This fixed power capping does not necessarily provide good performance since the effective power usage of jobs changes throughout their execution. Moreover, because each job has its own performance requirement, fixed power cap may not work well for all the jobs. In this paper, we propose a demand-aware power management framework for overprovisioned and power-constrained high-performance computing (HPC) systems. The job scheduler selects a job to run based on available hardware and power resources. The power manager continuously monitors power usage, predicts performance of executing jobs and optimizes power cap of each CPU so that the required performance level of each job is satisfied while improving system throughput by making good use of available powerbudget. Experiments on a real HPC system and with simulation for a large scale system show that the power manager can successfully control power consumption of executing jobs while achieving 1.17x improvement in system throughput. Yuan He 0002, Masaaki Kondo |
CCGrid | 2 |
| 2016 | Opportunistic circuit-switching for energy efficient on-chip networksabstractModern on-chip networks (NoCs) rely on virtual channel (VC) flow control to allow effective utilization of link bandwidth at the cost of more power and longer per-hop latency. Despite many existing optimization techniques for NoCs under VC flow control, we take a further step on questioning its necessity. Our finding is, when the network is not busy, circuit-switching (CS) may already satisfy the performance requirements with much smaller power consumption and shorter per-hop latency. In this paper, we propose to opportunistically enable CS in NoCs under VC flow control. This allows us to effectively reduce the power consumption of NoCs through having less buffering and longer sleep intervals for power gating while retaining CS-like per-hop latency. Our evaluations reveal that this proposal leads to a reduction of network power by up to 70% while cutting the system energy footprint by up to 35%. Yuan He 0002, Masaaki Kondo |
VLSI-SoC | 1 |
| 2015 | Runtime multi-optimizations for energy efficient on-chip interconnections1abstractOn-chip interconnection (or NoC) is a major performance and power contributor to modern and future multicore processors. So far, many optimization techniques have been developed to improve its bandwidth, latency and power consumption. But it is not clear how energy efficiency is affected since an optimization technique normally comes with overheads. This paper thus attempts to address when and how such optimization techniques should be applied and tuned to help achieve better energy efficiency. We firstly model the performance and energy impacts of representative NoC optimization techniques. These models help us more easily understand the consequences when applying these optimization techniques and their combinations under different circumstances. Moreover, based on such modeling, we propose and implement an adaptive control over these NoC optimization techniques to improve both performance and energy efficiency of the network. Our results show that, this proposal can achieve an average improvement of 26% and 57% on network performance and energy delay product, respectively. Yuan He 0002, Masaaki Kondo, Takashi Nakada, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
ICCD | 1 |
| 2013 | McRouter: Multicast within a router for high performance network-on-chipsabstractThe inevitable advent of the multi-core era has driven an increasing demand for low latency on-chip inter-connection networks (or NoCs). Being a critical part of the memory hierarchy for modern chip multi-processors (CMPs), these networks face stringent design constraints to provide fast communication with tight power budget. Modern NoC's first-order concern is clearly its latency, while we also find that internal bandwidth of its routers is relatively plentiful; thus, we present a low latency router design utilizing a technique we call “multicast within a router” or McRouter, which allows productive utilization of remaining bandwidth inside a NoC router. McRouter allows a single cycle transfer of flits which shortens the communication latency when there is enough remaining bandwidth within the router. The key idea is to transmit a header flit to all possible output ports (multicast) so that it is always transmitted to the correct output port without relying on route computation. In addition, we find it is affordable with marginal power overhead while still being a stand-alone design by maintaining portability and modularity (unlike look-ahead routing based designs). Our evaluation with application traffic shows that McRouter helps achieving system speed-ups of 1.28, 1.17 and 1.05 over the conventional router (CR), the VSA router (VSAR) and the prediction router (PR), respectively. Yuan He 0002, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
PACT | 1 |