EDBT 2026 Demo / reviewers in the wild / expert
Taeweon Suh
dblp:63/677
· DBLP profile ↗
26ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-6377-5482ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Characterization and Optimization of LLM Inference on Tenstorrent AI Accelerators
Jangho Lim, Dongin Shin, Uichan Kim, Jinhyeok Choi, Sangwon Shin, Sangwoo Park 0005, Gunjae Koo, Taeweon Suh |
Euro-Par (2) | 9 |
| 2026 | Three Birds, One Stone: Fast, Accurate-aware and Cost-Efficient Accelerator for Ternary LLMabstractOn-device LLM inference is increasingly important for latency- and privacy-sensitive applications, yet it remains challenging due to the high compute and storage demands. Ternary-weight LLMs are a promising direction because they dramatically reduce model size and simplify arithmetic. In practice, deploying pretrained models on edge devices typically relies on post-training quantization (PTQ), but ternary PTQ often needs fine-grained scaling to preserve accuracy, which amplifies scale-metadata traffic and sub-byte decoding overhead that fits poorly with conventional NPU datapaths. This paper presents T-ACE, a Ternary Accuracy-aware Compute Engine that enables efficient ternary LLM inference under PTQ by jointly designing the data representation and execution pipeline. T-ACE co-packs 64 ternary weights and power-of-two scale metadata into a naturally aligned 16-byte block, eliminating separate scale fetches and preserving aligned memory access. To decode compact ternary packing efficiently, T-ACE proposes a compact two-stage 5-trit unpacker and integrates on-the-fly decoding and scaling directly into the ternary GEMM pipeline. The evaluation on an FPGA prototype shows that decoding and scaling are fully overlapped with GEMM execution, incurring no additional cycles over baseline. Moreover, the comparison against A100/H100 baselines in a normalized setting shows that T-ACE improves accuracy-adjusted compute density (ACD) by 66.8% and accuracy-adjusted energy efficiency (AEE) by 17.6% over the best GPU baseline. Wonseok Jung, Sangwon Shin, Hongjun Um, Jangho Lim, Yongjun Park 0001, Gunjae Koo, Sangwoo Park 0005, Taeweon Suh |
ICS | 9 |
| 2026 | SumcheckPIM: An Efficient HBM-Based PIM Architecture for Linear Complexity Zero Knowledge ProofsabstractZero-knowledge proofs (ZKPs) are emerging as a core technology for privacy-preserving computation. Despite steady progress in protocol and algorithm design, generating these proofs remains computationally intensive, driving growing interest in hardware acceleration for kernels such as number-theoretic transform (NTT) and multi-scalar multiplication (MSM). Among them, the sumcheck protocol offers a compelling alternative with O(n) prover complexity compared to O(nlog n) for NTT-based approaches, yet our analysis reveals its execution is fundamentally memory-bound, with severely underutilized compute resources. This characteristic demands a memory-centric acceleration strategy, in contrast to compute-centric approaches of prior work. Sunchae Kim, Taewoon Kang, Sangwon Shin, Taeweon Suh, Yibin Yang 0001, Gunjae Koo |
ICS | 4 |
| 2025 | SkipNZ: Non-zero Value Skipping for Efficient CNN Acceleration
Joonyup Kwon, Jinhyeok Choi, Ngoc-Son Pham, Sangwon Shin, Taeweon Suh |
Euro-Par (2) | 5 |
| 2025 | HBM-Aware Number Theoretic Transform Accelerator for Zero-Knowledge ProofabstractZero-Knowledge Proof (ZKP) cryptographic algorithms have garnered significant attention for their ability to enhance privacy. However, the practical deployment of these algorithms remains challenging because they demand extremely high computational effort and handle huge volumes of data, especially in the Number Theoretic Transform (NTT) step. In this work, we propose an HBM-aware dataflow that employs sub-tiling and row-shuffling techniques to overcome the nonuniform stride access problem and to maximize HBM bandwidth utilization. We also design the NTT accelerator to use minimal FPGA resources. In particular, we explore diverse design options for the 256-bit modular multiplier and adopt an efficient design that optimizes resource usage and performance. Experimental results demonstrate that the proposed accelerator achieves lower latency and enhanced resource utilization compared to state-of-the-art FPGA-based designs. Sangwon Shin, Ngoc-Son Pham, Lei Xu 0012, Larry Shi, Taeweon Suh |
ICCD | 5 |
| 2025 | SparsePIM: An Efficient HBM-Based PIM Architecture for Sparse Matrix-Vector MultiplicationsabstractSparse matrix-vector multiplication (SpMV) is a fundamental operation across diverse domains, including scientific computing, machine learning, and graph processing.However, its irregular memory access patterns necessitate frequent data retrieval from external memory, leading to significant inefficiencies on conventional processors such as CPUs and GPUs.Processing-in-memory (PIM) presents a promising solution to address these performance bottlenecks observed in memory-intensive workloads.However, existing PIM architectures are primarily optimized for dense matrix operations since conventional memory cell structures struggle with the challenges of indirect indexing and unbalanced data distributions inherent in sparse computations.In order to address these challenges, we propose SparsePIM, a novel PIM architecture designed to accelerate SpMV computations efficiently.SparsePIM introduces a DRAM row-aligned format (DRAF) to optimize memory access patterns.SparsePIM exploits K-means-based column group partitioning to achieve a balanced load distribution across memory banks.Furthermore, SparsePIM includes bank group (BG) accumulators to mitigate the performance burdens of accumulating partial sums in SpMV operations.By aggregating partial results across multiple banks, SparsePIM can significantly improve the throughput of sparse matrix computations.Leveraging a combination of hardware and software optimizations, SparsePIM can achieve significant performance gains over cuSPARSE-based SpMV kernels on the GPU.Our evaluation demonstrates that SparsePIM achieves up to 5.61× speedup over SpMV on GPUs. Taewoon Kang, Geonwoo Choi, Taeweon Suh, Gunjae Koo |
ICS | 3 |
| 2022 | CacheRewinder: Revoking Speculative Cache Updates Exploiting Write-Back BufferabstractTransient execution attacks are critical security threats since those attacks exploit speculative execution which is an essential architectural solution that can improve the performance of out-of-order processors significantly. Such attacks change cache state by accessing secret data during speculative executions, then the attackers leak the secret information exploiting cache timing side-channels. Even though software patches against transient execution attacks have been proposed, the software solutions significantly slow down the performance of a system. In this paper, we propose CacheRewinder, an efficient hardware-based defense mechanism against transient execution attacks. CacheRewinder prevents leakage of secret information by revoking the cache updates done by speculative executions. To restore the cache state efficiently, CacheRewinder exploits the underutilized write-back buffer space as the temporary storage for victimized cache blocks evicted during speculative executions. Hence, when speculation fails CacheRewinder can quickly restore the cache state using the victim blocks held in the write-back buffer. Our evaluation exhibits CacheRewinder can effectively defend against transient execution attacks. The performance overhead by CacheRewinder is only 0.6%, which is negligible compared to the unprotected baseline processor. CacheRewinder also requires minimal storage cost since it exploits unused write-back buffer entries as storage for evicted cache blocks. Jun-Yeon Lee, Taeweon Suh, Gunjae Koo |
DATE | 3 |
| 2020 | FPGA based Blockchain System for Industrial IoTabstractIndustrial IoT (IIoT) is critical for industrial infrastructure modernization and digitalization. Therefore, it is of utmost importance to provide adequate protection of the IIoT system. A modern IIoT system usually consists of a large number of devices that are deployed in multiple locations and owned/managed by different entities who do not fully trust each other. These features make it harder to manage the system in a coherent manner and utilize existing security mechanisms to offer adequate protection. The emerging blockchain technology provides a powerful tool for IIoT system management and protection because the IIoT nature of distributed deployment and involvement of multiple stakeholders fits the design philosophy of blockchain well. Most existing blockchain construction mechanisms are not scalable enough and too heavy for an IIoT system. One promising way to overcome these limitations is utilizing hardware based trusted execution environment (TEE) in blockchain construction. However, most of the existing works on this direction do not consider the characteristics of IIoT devices (e.g., fixed functionality and limited supply) and face several limitations when they are applied for IIoT system management and protection, such as high energy consumption, single root-of-trust, and low decentralization level. To mitigate these challenges, we propose a novel field programmable gate array (FPGA) based blockchain system. It leverages the FPGA to build a simple but efficient TEE for IIoT devices, and removes the single root-of-trust by allowing all stakeholders to participate in the management of the devices. The FPGA based blockchain system shifts the computation/storage intensive part of blockchain management to more powerful computers but still involves the IIoT devices in the block construction to achieve a high level of decentralization. We implement the major FPGA components of the design and evaluate the performance of the whole system with a simulation tool to demonstrate its feasibility for IIoT applications. Lei Xu 0012, Lin Chen 0009, Zhimin Gao, Han-Yee Kim, Taeweon Suh, Larry Shi |
TrustCom | 5 |
| 2020 | Blockchain based End-to-end Tracking System for Distributed IoT Intelligence Application Security EnhancementabstractIoT devices provide a rich data source that is not available in the past, which is valuable for a wide range of intelligence applications, especially deep neural network (DNN) applications that are data-thirsty. An established DNN model provides useful analysis results that can improve the operation of IoT systems in turn. The progress in distributed/federated DNN training further unleashes the potential of integration of IoT and intelligence applications. When a large number of IoT devices are deployed in different physical locations, distributed training allows training modules to be deployed to multiple edge data centers that are close to the IoT devices to reduce the latency and movement of large amounts of data. In practice, these IoT devices and edge data centers are usually owned and managed by different parties, who do not fully trust each other or have conflicting interests. It is hard to coordinate them to provide end-to-end integrity protection of the DNN construction and application with classical security enhancement tools. For example, one party may share an incomplete data set with others, or contribute a modified sub DNN model to manipulate the aggregated model and affect the decision-making process. To mitigate this risk, we propose a novel blockchain based end-to-end integrity protection scheme for DNN applications integrated with an IoT system in the edge computing environment. The protection system leverages a set of cryptography primitives to build a blockchain adapted for edge computing that is scalable to handle a large number of IoT devices. The customized blockchain is integrated with a distributed/federated DNN to offer integrity and authenticity protection services. Lei Xu 0012, Zhimin Gao, Xinxin Fan, Lin Chen 0009, Han-Yee Kim, Taeweon Suh, Larry Shi |
TrustCom | 6 |
| 2019 | SafeDB: Spark Acceleration on FPGA Clouds with Enclaved Data Processing and Bitstream ProtectionabstractThis paper proposes SafeDB: Spark Acceleration on FPGA Clouds with Enclaved Data Processing and Bitstream Protection. SafeDB provides a comprehensive and systematic hardware-based security framework from the bitstream protection to data confidentiality, especially for the cloud environment. The AES key shared between FPGA and client for the bitstream encryption is generated in hard-wired logic using PKI and ECC. The data security is assured by the enclaved processing with encrypted data, meaning that the encrypted data is processed inside the FPGA fabric. Thus, no one in the system is able to look into clients' data because plaintext data are not exposed to memory and/or memory-mapped space. SafeDB is resistant not only to the side channel attack but to the attacks from malicious insiders. We have constructed an 8-node cluster prototype with Zynq UltraScale+ FPGAs to demonstrate the security, performance, and practicability. Han-Yee Kim, Rohyoung Myung, Boeui Hong, Heon-Chang Yu, Taeweon Suh, Lei Xu 0012, Larry Shi |
CLOUD | 5 |
| 2019 | DQN-based OpenCL workload partition for performance optimization
Taeweon Suh |
J. Supercomput. | 2 |
| 2018 | Architectural Protection of Application Privacy against Software and Physical Attacks in Untrusted Cloud EnvironmentabstractIn cloud computing, it is often assumed that cloud vendors are trusted; the guest Operating System (OS) and the Virtual Machine Monitor (VMM, also called Hypervisor) are secure. However, these assumptions are not always true in practice and existing approaches cannot protect the data privacy of applications when none of these parties are trusted. We investigate how to cope with a strong threat model which is that the cloud vendors, the guest OS, or the VMM, or both of them are malicious or untrusted, and can launch attacks against privacy of trusted user applications. This model is relevant because applications may be small enough to be formally verified, while the guest OS and VMM are too complex to be formally verified. Specifically, we present the design and analysis of an architectural solution which integrates a set of components on-chip to protect the memory of trusted applications from potential software and hardware based attacks from untrusted cloud providers, compromised guest OS, or malicious VMM. Full-system performance evaluation results show that the design only incurs 9 percent overhead on average, which is a small performance price that is paid for the substantial security gain. Lei Xu 0012, Jong-Hyuk Lee, Qingji Zheng, Shouhuai Xu, Taeweon Suh, Won Woo Ro, Larry Shi |
IEEE Trans. Cloud Comput. | 6 |
| 2017 | Evaluating coherence-exploiting hardware TrojanabstractIncreasing complexity of integrated circuits and IP-based hardware designs have created the risk of hardware Trojans. This paper introduces a new type of threat, a coherence-exploiting hardware Trojan. This Trojan can be maliciously implanted in master components in a system, and continuously injects memory transactions onto the main interconnect. The injected traffic forces the eviction of cache lines, taking advantage of cache coherence protocols. This type of Trojans insidiously slows down the system performance, incurring Denial-of-Service (DoS) attack. We used a Xilinx Zynq-7000 device to implement the Trojan and evaluate its severity. Experiments revealed that the system performance can be severely degraded as much as 258% with the Trojan. A countermeasure to annihilate the Trojan attack is proposed in detail. We also found that AXI version 3.0 supports a seemingly irrelevant invalidation protocol through ACP, opening a door for the potential Trojan attack. Sunhee Kong, Boeui Hong, Lei Xu 0012, Larry Shi, Taeweon Suh |
DATE | 6 |
| 2014 | PFC: Privacy Preserving FPGA Cloud - A Case Study of MapReduceabstractPrivacy is one of the critical concerns that hinder the adoption of public cloud. For storage, encryption can be used to protect user's data. But for outsourced data processing, for example MapReduce, there is no satisfying solution. Users have to trust the cloud service providers totally. In this work, we propose PFC, a FPGA cloud for privacy preserving computation in the public cloud environment. PFC leverages the security feature of the existing FPGAs originally designed for bitstream IP protection and proxy re-encryption for preserving user data privacy. In PFC, cloud service providers are not necessarily trusted, and during outsourced computation, user's data is protected by a data encryption key only accessible by trusted FPGA devices. As an important application of cloud computing, we apply PFC to the popular MapReduce programming model and extend the FPGA based MapReduce pipeline with privacy protection capabilities. Proxy re-encryption is employed to support dynamic allocations of trusted FPGA devices as mappers and reducers. Finally, we conduct evaluation to demonstrate the effectiveness of PFC. Lei Xu 0012, Larry Shi, Taeweon Suh |
IEEE CLOUD | 3 |
| 2014 | Privacy preserving large scale DNA read-mapping in MapReduce framework using FPGAsabstractRead-mapping, i.e., finding certain patterns in a long DNA sequence, is an important operation for molecular biology. It is widely used in a variety of biological analyses including SNP discovery, genotyping and personal genomics. As next-generation DNA sequencing machines are generating an enormous amount of sequence data, it is a good choice to implement the read-mapping algorithm in the MapReduce framework and outsource the computation to the cloud. Data privacy becomes a big concern in this situation as DNA sequences are very sensitive. In response, encryption may be used to protect the data. However, it is very difficult for the cloud to process cipher texts. In the MapReduce framework, even if values (data to be processed) may be protected by encryption, keys cannot be encrypted using sematic secure encryption schemes as it will affect the MapReduce scheduling mechanism. But if no protection is utilized, attackers may extract useful information from unprotected keys. We propose a solution that can securely outsource read-mapping computations in the MapReduce framework by leveraging inherent tamper resistant properties of FPGAs. We also provide a method to protect the keys generated in this process. We implement our solution using FPGAs and apply it to some data sets. The security evaluation and experimental results show that with this method, DNA sequence privacy is well protected, and the extra cost is acceptable. Lei Xu 0012, Han-Yee Kim, Xi Wang 0011, Larry Shi, Taeweon Suh |
FPL | 5 |
| 2014 | A scheduling algorithm with dynamic properties in mobile grid
Jong-Hyuk Lee, SungJin Choi, Joon-Min Gil, Taeweon Suh, Heon-Chang Yu |
Frontiers Comput. Sci. | 4 |
| 2014 | Leveraging Process Variation for Performance and Energy: In the Perspective of OverclockingabstractProcess variation is one of the most important factors to be considered in recent microprocessor design, since it negatively affects performance, power, and yield of microprocessors. However, by leveraging process variation, overclocking techniques can improve performance. As microprocessors have substantial clock cycle time margin for yield, there is enough room for performance improvement by overclocking techniques. In this paper, we adopt the F-overclocking technique, which increases clock frequency without changing supply voltage. Our experimental results show that the F-overclocking technique significantly improves performance as well as energy consumption. In addition, the F-overclocking technique is superior to the conventional overclocking technique which increases clock frequency and supply voltage together in the perspective of energy efficiency and reliability, showing similar performance improvement. Furthermore, we propose an adaptive overclocking controller which dynamically applies the F-overclocking technique based on the application characteristics. By adopting our adaptive overclocking controller, we further minimize the reliability loss caused by the F-overclocking technique. Hyung Beom Jang, Junhee Lee 0004, Joonho Kong, Taeweon Suh, Sung Woo Chung |
IEEE Trans. Computers | 4 |
| 2010 | An Effective Job Replication Technique Based on Reliability and Performance in Mobile Grids
Daeyong Jung, Sung-Ho Chin, Kwang-Sik Chung, Taeweon Suh, Heon-Chang Yu, Joon-Min Gil |
GPC | 4 |
| 2010 | Adaptive service scheduling for workflow applications in Service-Oriented Grid
Sung-Ho Chin, Taeweon Suh, Heon-Chang Yu |
J. Supercomput. | 2 |
| 2009 | Balanced Scheduling Algorithm Considering Availability in Mobile Grid
Jong-Hyuk Lee, SungJin Song, Joon-Min Gil, Kwang-Sik Chung, Taeweon Suh, Heon-Chang Yu |
GPC | 5 |
| 2009 | A Potential Based Routing Protocol for Mobile Ad Hoc NetworksabstractIn this paper, we propose a novel proactive routing protocol, referred to as potential management based proactive routing (PMPR), for mobile ad hoc networks. Unlike other proactive routing protocols, PMPR performs request based routing recovery for proactive route maintenance. When a node has lost the routing information, it attempts a local route recovery by broadcasting a request message to neighbor nodes within a limited hop range. If the local recovery succeeds, the routing information is reconstructed by the interaction between the requesting node and the neighbor nodes. In this paper,we introduce a concept of potential and propose an efficient management method of potential. Potential is a value assigned to each node for each destination. Routes are determined based on the potential of each node. When a node is requested to perform a route recovery and the recovery is feasible, the node modifies its potential to a lower value to provide the requestor a new route. A potential management method determines the success rate of the local route recovery and consequent route optimality. In our simulation with a moderate node density and high node mobility, over 95% of broken routes are recovered with1 hop request. PMPR outperforms DVDS for all the simulated parameters. PMPR also outperforms AODV and DSR under high node mobility and high traffic load condition. Under a low node mobility or low data rate condition, PMPR provides comparable performance to AODV and DSR Dai Yong Kwon, Jaehwa Chung, Taeweon Suh, Won-Gyu Lee, Kyeong Hur |
HPCC | 3 |
| 2008 | A Desktop Computer with a Reconfigurable Pentium®abstractAdvancements in reconfigurable technologies, specifically FPGAs, have yielded faster, more power-efficient reconfigurable devices with enormous capacities. In our work, we provide testament to the impressive capacity of recent FPGAs by hosting a complete Pentium ® in a single FPGA chip. In addition we demonstrate how FPGAs can be used for microprocessor design space exploration while overcoming the tension between simulation speed, model accuracy, and model completeness found in traditional software simulator environments. Specifically, we perform preliminary experimentation/prototyping with an original Socket 7 based desktop processor system with typical hardware peripherals running modern operating systems such as Fedora Core 4 and Windows XP; however we have inserted a Xilinx Virtex-4 in place of the processor that should sit in the motherboard and have used the Virtex-4 to host a complete version of the Pentium ® microprocessor (which consumes less than half its resources). We can therefore apply architectural changes to the processor and evaluate their effects on the complete desktop system. We use this FPGA-based emulation system to conduct preliminary architectural experiments including growing the branch target buffer and the level 1 caches. In addition, we experimented with interfacing hardware accelerators such as DES and AES engines which resulted in a 27x speedup. Shih-Lien Lu, Peter Yiannacouras, Taeweon Suh, Rolf Kassa, Michael Konow |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2007 | An FPGA-based Pentium in a complete desktop systemabstractSoftware simulation has been the predominant method for architects to evaluate microprocessor research proposals. There are three tenets in modeling new designs with software models: simulation speed, model accuracy and model completeness. The increasing complexity of the processor and accelerated trend to have multiple processors on a chip are putting burden on simulators to achieve all tenets mentioned, including accurately capturing OS effects. In this work we perform preliminary experimentation/prototyping with an emulation system which overcomes the tension to satisfy all three requirements. The system is an original Socket-7 based desktop processor system with typical hardware peripherals running modern operating systems such as Fedora Core 4 and Windows XP; however we have inserted a Xilinx Virtex-4 in place of the processor that should sit in the motherboard and have used the Virtex-4 to host a complete version of the Pentium® microprocessor (which consumes less than half its resources). We can therefore apply architectural changes to the processor and evaluate their effects on the complete desktop system. We use this FPGA-based emulation system to conduct preliminary architectural experiments including growing the branch target buffer and the level 1 caches. In addition, we experimented with interfacing hardware accelerators such as DES and AES engines which resulted in 27x speedups. Shih-Lien Lu, Peter Yiannacouras, Rolf Kassa, Michael Konow, Taeweon Suh |
FPGA | 5 |
| 2007 | An FPGA Approach to Quantifying Coherence Traffic Efficiency on Multiprocessor SystemsabstractRecently, there is a surge of interests in using FPGAs for computer architecture research including applications from emulating and analyzing a new platform to accelerating microarchitecural simulation speed for design space exploration. This paper proposes and demonstrates a novel usage of FPGAs for measuring the efficiency of coherent traffic of an actual computer system. Our approach employs an FPGA acting as a bus agent, interacting with a real CPU in a dual processor system to measure the intrinsic delay of coherence traffic. This technique eliminates non-deterministic factors in the measurement, such as the arbitration delay and stall in the pipelined bus. It completely isolates the impact of pure coherence traffic delay on system performance while executing workloads natively. Our experiments show that the overall execution time of the benchmark programs on a system with coherence traffic was actually increased over one without coherent traffic. It indicates that cache-to-cache transfers are less efficient in an Intel-based server system, and there exists room for further improvement such as the inclusion of the O state and cache line buffers in the memory controller. Taeweon Suh, Shih-Lien Lu, Hsien-Hsin S. Lee |
FPL | 1 |
| 2005 | Cache coherence support for non-shared bus architecture on heterogeneous MPSoCsabstractWe propose two novel integration techniques ̬ bypass and bookkeeping̬in the memory controller to address the cache coherence compatibility issue of a non-shared bus heterogeneous MPSoC. The bypass approach is an inexpensive and efficient solution for computation-bound applications while the bookkeeping approach eliminating unnecessary forwarding traffic offers an alternative for bandwidth-limited applications. Our RTOS kernel simulations show up to 6.65x speedup over the conventional software solution. Taeweon Suh, Daehyun Kim 0001, Hsien-Hsin S. Lee |
DAC | 1 |
| 2004 | Supporting Cache Coherence in Heterogeneous Multiprocessor SystemsabstractIn embedded system-on-a-chip (SoC) applications, the demand for integrating heterogeneous processors onto a single chip is increasing. An important issue in integrating multiple heterogeneous processors on the same chip is to maintain the coherence of their data caches. In this paper, we propose a hardware/software methodology to make caches coherent in heterogeneous multiprocessor platforms with shared memory. Our approach works with any combination of processors that support invalidation-based protocols. As shown in our experiments, up to 58% performance improvement can be achieved with low miss penalty at the expense of adding simple hardware, compared to a pure software solution. Speedup can be improved even further as the miss penalty increases. In addition, our approach provides embedded system programmers a transparent view of shared data, removing the burden of software synchronization. Taeweon Suh, Douglas M. Blough, Hsien-Hsin S. Lee |
DATE | 1 |