VLDB 2026 Research / reviewers in the wild / expert
Hyunok Oh
dblp:59/2601
· DBLP profile ↗
38ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-9044-7441ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 5 first-authorSecurity and privacy · 12 · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tangram: Encryption-Friendly SNARK Framework Under Pedersen Committed Engines
Gweonho Jeong, Myeongkyun Moon, Geonho Yoon, Hyunok Oh, Jihye Kim 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | DUPLEX: Scalable Zero-Knowledge Lookup Arguments over RSA Group
Semin Han, Geonho Yoon, Hyunok Oh, Jihye Kim 0001 |
AsiaCCS | 3 |
| 2025 | zkMarket: Ensuring Fairness and Privacy in Decentralized Data ExchangeabstractEnsuring fairness in blockchain-based data trading presents significant challenges, as the transparency of blockchain can expose sensitive details and compromise fairness. Fairness ensures that the seller receives payment only if they provide the correct data, and the buyer gains access to the data only after making the payment. Existing approaches face limitations in efficiency, particularly when applied to large-scale data. Moreover, preserving privacy has also been a significant challenge in blockchain. In this paper, we introduce zkMarket, a privacy-preserving fair trade system on the blockchain. We ensure fairness by integrating encryption with zk-SNARKs, enabling verifiable proofs for fair trading. However, applying zk-SNARKs directly can be computationally expensive for the prover. To address this, we improve efficiency by leveraging our novel matrix-formed PRG (MatPRG) and commit-and-prove SNARK (CP-SNARK), making the data registration process more concise and significantly reducing the seller's proving time. To ensure transaction privacy, zkMarket is built upon an anonymous transfer protocol. Experimental results demonstrate that zkMarket significantly reduces the computational overhead associated with traditional blockchain solutions while maintaining robust security and privacy. Specifically, our evaluation quantifies this high efficiency: the seller can register 1 MB of data in 2.8 seconds, the buyer can generate the trade transaction in 0.2 seconds, and the seller can finalize the trade within 0.4 seconds. Seongho Park, Seungwoo Kim, Semin Han, Kyeongtae Lee, Jihye Kim 0001, Hyunok Oh |
SRDS | 6 |
| 2024 | zkLogis: Scalable, Privacy-Enhanced, and Traceable Logistics on Public BlockchainabstractDecentralized blockchain systems have provided a significant leap in autonomous logistics practices by enabling accurate product authentication through the tracking of ownership changes. Although public blockchains could provide open access to stored information, privacy and security concerns lead most existing logistics implementations to rely on private or consortium blockchains. This, however, limits or regulates the end-customer's ability to independently verify product provenance. Hyunok Oh, Jihye Kim 0001 |
AsiaCCS | 3 |
| 2024 | SAVER: SNARK-Compatible Verifiable Encryption
Jaekyoung Choi, Jihye Kim 0001, Hyunok Oh |
FC (2) | 4 |
| 2024 | vCNN: Verifiable Convolutional Neural Network Based on zk-SNARKsabstractIt is becoming important for the client to be able to check whether the AI inference services have been correctly calculated. Since the weight values in a CNN model are assets of service providers, the client should be able to check the correctness of the result without them. The Zero-knowledge Succinct Non-interactive Argument of Knowledge (zk-SNARK) allows verifying the result without input and weight values. However, the proving time in zk-SNARK is too slow to be applied to real AI applications. This article proposes a new efficient verifiable convolutional neural network (vCNN) framework that greatly accelerates the proving performance. We introduce a new efficient relation representation for convolution equations, reducing the proving complexity of convolution from O(ln) to O(l+n) compared to existing zero-knowledge succinct non-interactive argument of knowledge (zk-SNARK) approaches, where l and n denote the size of the kernel and the data in CNNs. Experimental results show that the proposed vCNN improves proving performance by 20-fold for a simple MNIST and 18,000-fold for VGG16. The security of the proposed scheme is formally proven. Seunghwa Lee, Hankyung Ko, Jihye Kim 0001, Hyunok Oh |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Efficient Transparent Polynomial Commitments for zk-SNARKs
Sungwook Kim 0001, Sungju Kim, Yulim Shin, Sunmi Kim, Jihye Kim 0001, Hyunok Oh |
ESORICS (3) | 6 |
| 2022 | Succinct Zero-Knowledge Batch Proofs for Set AccumulatorsabstractCryptographic accumulators are a common solution to proving information about a large set S. They allow one to compute a short digest of S and short certificates of some of its basic properties, notably membership of an element. Accumulators also allow one to track set updates: a new accumulator is obtained by inserting/deleting a given element. In this work we consider the problem of generating membership and update proofs for \em batches of elements so that we can succinctly prove additional properties of the elements (i.e., proofs are of constant size regardless of the batch size), and we can preserve privacy. Solving this problem would allow obtaining blockchain systems with improved privacy and scalability. Matteo Campanelli, Dario Fiore 0001, Semin Han, Jihye Kim 0001, Dimitris Kolonelos, Hyunok Oh |
CCS | 6 |
| 2021 | Efficient Verifiable Image Redacting based on zk-SNARKsabstractImage is a visual representation of a certain fact and can be used as proof of events. As the utilization of the image increases, it is required to prove its authenticity with the protection of its sensitive personal information. In this paper, we propose a new efficient verifiable image redacting scheme based on zk-SNARKs, a commitment, and a digital signature scheme. We adopt a commit-and-prove SNARK scheme which takes commitments as inputs, in which the authenticity can be quickly verified outside the circuit. We also specify relations between the original and redacted images to guarantee the redacting correctness. Our experimental results show that the proposed scheme is superior to the existing works in terms of the key size and proving time without sacrificing the other parameters. The security of the proposed scheme is proven formally. Hankyung Ko, Ingeun Lee, Seunghwa Lee, Jihye Kim 0001, Hyunok Oh |
AsiaCCS | 5 |
| 2019 | SA-SPM: an efficient compiler for security aware scratchpad memory (invited paper)abstractScratchpad memories (SPM) are often used to boost the performance of application-specific embedded systems. In embedded systems, main memories are vulnerable to external attacks such as bus snooping or memory extraction. Therefore it is desirable to guarantee the security of data in a main memory. In software-managed SPM, it is possible to provide security in main memory by performing software-assistant encryption. Thomas Haywood Dadzie, Jihye Kim 0001, Hyunok Oh |
LCTES | 4 |
| 2019 | Forward Secure Identity-Based Signature Scheme with RSA
Hankyung Ko, Gweonho Jeong, Jihye Kim 0001, Hyunok Oh |
SEC | 5 |
| 2019 | FAS: Forward secure sequential aggregate signatures for secure logging
Jihye Kim 0001, Hyunok Oh |
Inf. Sci. | 2 |
| 2019 | AuthCropper: Authenticated Image Cropper for Privacy Preserving Surveillance SystemsabstractAs surveillance systems are popular, the privacy of the recorded video becomes more important. On the other hand, the authenticity of video images should be guaranteed when used as evidence in court. It is challenging to satisfy both (personal) privacy and authenticity of a video simultaneously, since the privacy requires modifications (e.g., partial deletions) of an original video image while the authenticity does not allow any modifications of the original image. This paper proposes a novel method to convert an encryption scheme to support partial decryption with a constant number of keys and construct a privacy-aware authentication scheme by combining with a signature scheme. The security of our proposed scheme is implied by the security of the underlying encryption and signature schemes. Experimental results show that the proposed scheme can handle the UHD video stream with more than 17 fps on a real embedded system, which validates the practicality of the proposed scheme. Jihye Kim 0001, Hankyung Ko, Donghwan Oh, Semin Han, Gwonho Jeong, Hyunok Oh |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2018 | Scalable Wildcarded Identity-Based Encryption
Jihye Kim 0001, Seunghwa Lee, Hyunok Oh |
ESORICS (2) | 4 |
| 2018 | Forward-secure ID based digital signature scheme with forward-secure private key generator
Hyunok Oh, Jihye Kim 0001, Ji Sun Shin |
Inf. Sci. | 1 |
| 2018 | A hybrid performance analysis technique for distributed real-time embedded systems
Junchul Choi, Hyunok Oh, Soonhoi Ha |
Real Time Syst. | 2 |
| 2017 | PASS: Privacy aware secure signature scheme for surveillance systemsabstractIn a surveillance system, the privacy becomes important since those who are not relevant to an event may be recorded by many surveillance systems. On the other hand, the authenticity of video frames in surveillance systems should be guaranteed if a video is used as evidence. Hence a signature is attached for each frame. However, it is contradictory to provide both privacy and authenticity of a video since the privacy requires deletion of objects in an original video image while the authenticity disallows any modification of the original image. This paper devises a new novel privacy aware secure signature scheme for the surveillance system. In the proposed scheme, deletion (or masking) of objects in an image is allowed while a signature still remains valid for the modified image. The proposed scheme utilizes a chameleon hash so that deletion of objects in an image does not invalidate the signature. The proposed scheme provides forward security minimizing the damage from the secret key exposure. The deletion is executed in an authorized way. Experimental results show that the proposed scheme is practical in a real-time video surveillance system due to the high performance of signature generation (40ms per frame) and a small signature size overhead (1%). Jihye Kim 0001, Seunghwa Lee, Jungjun Yoon, Hankyung Ko, Seungri Kim, Hyunok Oh |
AVSS | 6 |
| 2017 | Hierarchical Dataflow Modeling of Iterative ApplicationsabstractEven though dataflow models are good at exploiting task-level parallelism of an application, it is difficult to exploit the parallelism of loop structures since they are not explicitly specified in existent dataflow models. To overcome this drawback, we propose a novel extension to the SDF model, called SDF/L graph, specifying the loop structures explicitly in a hierarchical fashion. With a given SDF/L graph specification and the mapping and scheduling information, an application can be automatically parallelized on a multicore system. The enhanced expression capability by the proposed extension is verified with two applications, k-means clustering and deep neural network application. Hyesun Hong, Hyunok Oh, Soonhoi Ha |
DAC | 2 |
| 2017 | Forward-Secure Digital Signature Schemes with Optimal Computation and Storage of Signers
Jihye Kim 0001, Hyunok Oh |
SEC | 2 |
| 2017 | Multiprocessor Scheduling of a Multi-Mode Dataflow Graph Considering Mode Transition DelayabstractThe Synchronous Data Flow (SDF) model is widely used for specifying signal processing or streaming applications. Since modern embedded applications become more complex with dynamic behavior changes at runtime, several extensions of the SDF model have been proposed to specify the dynamic behavior changes while preserving static analyzability of the SDF model. They assume that an application has a finite number of behaviors (or modes), and each behavior (mode) is represented by an SDF graph. They are classified as multi-mode dataflow models in this article. While there exist several scheduling techniques for multi-mode dataflow models, no one allows task migration between modes. By observing that the resource requirement can be additionally reduced if task migration is allowed, we propose a multiprocessor scheduling technique of a multi-mode dataflow graph considering task migration between modes. Based on a genetic algorithm, the proposed technique schedules all SDF graphs in all modes simultaneously to minimize the resource requirement. To satisfy the throughput constraint, the proposed technique calculates the actual throughput requirement of each mode and the output buffer size for tolerating throughput jitter. We compare the proposed technique with a method that analyzes SDF graphs in each execution mode separately, a method that does not allow task migration, and a method that does not allow mode-overlapped schedule for synthetic examples and five real applications: H.264 decoder, lane detection, vocoder, MP3 decoder, and printer pipeline. Hanwoong Jung, Hyunok Oh, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | In-storage processing of database scans and joins
Sungchan Kim, Hyunok Oh, Chanik Park, Sangyeun Cho, Sang-Won Lee 0001, Bongki Moon |
Inf. Sci. | 2 |
| 2015 | Optimization of multi-channel BCH error decoding for common casesabstractThis paper proposes a new method to optimize a BCH error correction decoder in multi-channel configurations. We break the BCH decoding process into its three basic blocks: syndrome calculation, the error locator polynomial generation, and the roots of the error locator polynomial computation. While an existing multi-channel BCH decoder consists of several single-channel BCH decoders operating in parallel, this paper utilizes a pooled group of shared decoding blocks. By considering the frequency of errors, the proposed pooled group approach requires fewer hardware blocks than in a traditional multi-channel configuration with a negligible impact on performance. Combined with a specialized root finding unit for blocks with only 1 error, our scheme reduces hardware area by 47%-71% and dynamic power by 44%-59% with 2% performance degradation in typical NAND flash systems. With a constant hardware area, the proposed scheme can improve throughput by 3x-5x or NAND flash lifetime by 1.4x-4.5x. Russ Dill, Aviral Shrivastava, Hyunok Oh |
CASES | 3 |
| 2015 | Optimal Checkpoint Selection with Dual-Modular Redundancy HardeningabstractWith the continuous scaling of semiconductor technology, failure rate is increasing significantly so that reliability becomes an important issue in multiprocessor system-on-chip (MPSoC) design. We propose an optimal checkpoint selection with task duplication hardening to tolerate transient faults. A target application is specified in a task graph, and the schedule/checkpoint placements are determined at design time. The proposed optimal algorithm minimizes the checkpoint overhead with a latency constraint. Experimental results show that the proposed algorithm effectively reduces the minimum end-to-end latency to perform a fault-tolerant schedule. In addition, the proposed algorithm dramatically decreases the checkpointing overhead on uniprocessor and multiprocessor systems compared with a greedy approach and an equidistant algorithm. Shin-Haeng Kang, Hae-woo Park, Sungchan Kim, Hyunok Oh, Soonhoi Ha |
IEEE Trans. Computers | 4 |
| 2014 | An Efficient Non-Linear Cost Compression Algorithm for Multi Level Cell MemoryabstractThis paper defines a non-linear cost compression problem, proposes an efficient algorithm, and applies it to a real application of multi level cell memory to minimize energy consumption and latency. The non-linear cost compression problem extends the traditional cost compression problem to allow a non-linear cost function of symbol frequencies, while it is a weighted linear combination of symbol frequencies in the cost compression problem. In order to solve the non-linear cost compression problem efficiently, we propose an encoding symbol frequency based approach. We first compute frequencies of encoding symbols to minimize a cost function. To achieve the computed frequencies of a cost-compressed message, we deploy existing size-decompression algorithms. The proposed algorithm is optimal and as fast as the existing size compression algorithms. Our experimental results show that it reduces the energy consumption and latency by 70 percent for a text file in multi level cell memory. Furthermore, it increases the lifetime of endurance limited memory. Hyunok Oh, Jihye Kim 0001 |
IEEE Trans. Computers | 1 |
| 2014 | Dynamic Behavior Specification and Dynamic Mapping for Real-Time Embedded Systems: HOPES ApproachabstractAs the number of processors in a chip increases and more functions are integrated, the system status will change dynamically due to various factors such as the workload variation, QoS requirement, and unexpected component failure. A typical method to deal with the dynamics of the system is to decide the mapping decision at runtime, based on the local information of the system status. It is very challenging to guarantee any real-time performance of a certain application in such a dynamically varying system. To solve this problem, we propose a hybrid specification of dataflow and FSM models to specify the dynamic behavior of a system distinguishing inter- and intra-application dynamism. At the top level, each application is specified by a dataflow task and the dynamic behavior is modeled as a control task that supervises the execution of applications. Inside a dataflow task, we specify the dynamic behavior using a similar way as FSM-based SADF in which an application is specified by a synchronous dataflow graph for each mode of operation. It enables us to perform compile-time scheduling of each graph to maximize the throughput varying the number of allocated processors, and store the scheduling information. When a change in system state is detected at runtime, the number of allocated processors to the active tasks is determined dynamically utilizing the stored scheduling information of those tasks in order to meet the real-time requirements. The proposed technique is implemented in the HOPES design environment. Through preliminary experiments with a simple smartphone example, we show the viability of the proposed methodology. Hanwoong Jung, Chanhee Lee 0002, Shin-Haeng Kang, Sungchan Kim, Hyunok Oh, Soonhoi Ha |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | Intelligent SSD: a turbo for big data miningabstractThis paper introduces the notion of intelligent SSDs. First, we present the design considerations of intelligent SSDs, and then examine their potential benefits under various settings in data mining applications. Duck-Ho Bae, Jin-Hyung Kim, Sang-Wook Kim, Hyunok Oh, Chanik Park |
CIKM | 4 |
| 2013 | A novel analytical method for worst case response time estimation of distributed embedded systemsabstractIn this paper, we propose a novel analytical method, called scheduling time bound analysis, to find a tight upper bound of the worst-case response time in a distributed real-time embedded system, considering execution time variations of tasks, jitter of input arrivals, and scheduling anomaly behavior in a multi-tasking system all together. By analyzing the graph topology and worst-case scheduling scenarios, we measure the conservative scheduling time bound of each task. The proposed method supports an arbitrary mixture of preemptive and non-preemptive processing elements. Its speed is comparable to compositional approaches while it gives a much tighter bound. The advantages of the proposed approach compared with related work were verified by experimental results with randomly generated task graphs and a real-life automotive application. Hyunok Oh, Junchul Choi, Hyojin Ha, Soonhoi Ha |
DAC | 2 |
| 2013 | Active disk meets flash: a case for intelligent SSDsabstractIntelligent solid-state drives (iSSDs) allow execution of limited application functions (e.g., data filtering or aggregation)on their internal hardware resources, exploiting SSD characteristics and trends to provide large and growing performance and energy efficiency benefits. Most notably, internal flash media bandwidth can be significantly (2-4x or more) higher than the external bandwidth with which the SSD is connected to a host system, and the higher internal bandwidth can be exploited within an iSSD. Also, SSD bandwidth is projected to increase rapidly over time, creating a substantial energy cost for streaming of data to an external CPU for processing, which can be avoided via iSSD processing. This paper makes a case for iSSDs by detailing these trends, quantifying the potential benefifits across a range of application activities, describing how SSD architectures could be extended cost-effectively, and demonstrating the concept with measurements of a prototype iSSD running simple data scan functions. Our analyses indicate that, with less than a 2% increase in hardware cost over a traditional SSD, an iSSD can provide 2-4x performance increases and 5-27x energy efficiency gains for a range of data-intensive computations. Sangyeun Cho, Chanik Park, Hyunok Oh, Sungchan Kim, Youngmin Yi, Gregory R. Ganger |
ICS | 3 |
| 2013 | A lifetime aware buffer assignment method for streaming applications on DRAM/PRAM hybrid memoryabstractThis article proposes a lifetime aware buffer assignment method for streaming applications like multimedia specified in a synchronous dataflow (SDF) graph on a DRAM/PRAM hybrid memory in which the endurance of PRAM is limited. We determine whether buffers are assigned to DRAM or PRAM to minimize the writing frequency of PRAM. To solve the problems, we formulate them using Answer Set Programming. Experimental results show that the proposed approach increases the PRAM lifetime by 63% compared with no optimization, and shows the tradeoff between PRAM and DRAM size to guarantee a lifetime constraint. Hyunok Oh |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Executing synchronous dataflow graphs on a SPM-based multicore architectureabstractIn this paper we are concerned about executing synchronous dataflow (SDF) applications on a multicore architecture where a core has a limited size of scratchpad memory (SPM). Unlike traditional multi-processor scheduling of SDF graphs, we consider the SPM size limitation that incurs code and data overlay overhead. Since the scheduling problem is intractable, we propose an EA(evolutionary algorithm)-based technique. To hide memory latency, prefetching is aggressively performed in the proposed technique. The experimental results show that our approach reduces the overlay overhead significantly compared to a non-optimized approach and the previous approach. Junchul Choi, Hyunok Oh, Sungchan Kim, Soonhoi Ha |
DAC | 2 |
| 2012 | An ILP-based Worst-case Performance Analysis Technique for Distributed Real-time Embedded SystemsabstractFinding a tight upper bound of the worst-case response time in a distributed real-time embedded system is a very challenging problem since we have to consider execution time variations of tasks, jitter of input arrivals, scheduling anomaly behavior in a multi-tasking system, all together. In this paper, we translate the problem as an optimization problem and propose a novel solution based on ILP (Integer Linear Programming). In the proposed technique, we formulate a set of ILP formulas in a compositional way for modeling flexibility, but solve the problem holistically to achieve tighter upper bounds. To mitigate the time complexity of the ILP method, we perform static analysis based on a scheduling heuristic to reduce the number of variables and confine the variable ranges. Preliminary experiments with the benchmarks used in the related work and a real-life example show promising results that give tight bounds in an affordable solution time. Hyunok Oh, Hyojin Ha, Shin-Haeng Kang, Junchul Choi, Soonhoi Ha |
RTSS | 2 |
| 2012 | A parallel and distributed meta-heuristic framework based on partially ordered knowledge sharing
Minyoung Kim 0002, Mark-Oliver Stehr, Hyunok Oh, Soonhoi Ha |
J. Parallel Distributed Comput. | 4 |
| 2011 | Minimizing buffer requirements for throughput constrained parallel execution of synchronous dataflow graphabstractThis paper concerns throughput-constrained parallel execution of synchronous data flow graphs. This paper assumes static mapping and dynamic scheduling of nodes, which has several benefits over static scheduling approaches. We determine the buffer size of all arcs to minimize the total buffer size while satisfying a throughput constraint. Dynamic scheduling is able to achieve the similar throughput performance as the static scheduling does by unfolding the given SDF graph. A key issue of dynamic scheduling is how to assign the priority to each node invocation, which is also discussed in this paper. Since the problem is NP-hard, we present a heuristic based on a genetic algorithm. The experimental results confirm the viability of the proposed technique. Tae-ho Shin, Hyunok Oh, Soonhoi Ha |
ASP-DAC | 2 |
| 2011 | Library Support in an Actor-Based Parallel Programming PlatformabstractActor model-based design is actively researched for parallel embedded SW design since the model exposes the potential parallelism explicitly in an architecture-neutral form. In most actor-oriented models, actors are self-contained and data channels are the only sharable object between actors, and they compose a system in a flat layer. In contrast, it is common to use shared library functions and construct vertically layered software for efficiency and modularity. To fill this gap between modeling and implementation, we propose a special actor, library task, with new types of ports: library master port and library slave port. It is a sharable and mappable object that defines a set of function interfaces inside. N:1 master-slave connection allows sharing a library task and the master-slave connection can specify vertically layered software and client-server applications naturally. To support the library task in our embedded software design environment, we develop an automatic mapping algorithm as well as an automatic code generator. The design environment with the library task is applied for two target platforms: IBM CELL Broad band Engine and an ARM-based multicore simulator. Preliminary experiments show that the special actor, or library task, extends the expression power of the previous actor model with efficiently generated codes. Hae-woo Park, Hanwoong Jung, Hyunok Oh, Soonhoi Ha |
IEEE Trans. Ind. Informatics | 3 |
| 2006 | Memory optimal single appearance schedule with dynamic loop count for synchronous dataflow graphsabstractIn this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size simultaneously. While a single appearance schedule promises only one appearance of each node definition in the generated code, it requires significant amount of data memory overhead compared with a buffer optimal schedule allowing multiple appearance. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that every buffer optimal schedule can be transformed to our single appearance schedule which requires optimal buffer size for arbitrary synchronous dataflow graphs. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. In order to minimize the overhead we propose optimization techniques. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms. Hyunok Oh, Nikil Dutt, Soonhoi Ha |
ASP-DAC | 1 |
| 2005 | Single appearance schedule with dynamic loop count for minimum data buffer from synchronous dataflow graphsabstractIn this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size at the same time. When the software code is automatically synthesized from the dataflow program graphs, a single appearance schedule promises only one appearance of each node definition in the generated code. While several heuristics have been developed to find a single appearance schedule, they all have to pay significant amount of data memory overhead compared with a buffer optimal schedule. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that the proposed scheduling technique is optimal for a chain-structured graph in terms of data memory requirement while maintaining the single appearance schedule. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms for CD2DAT and non uniform filter bank applications. Hyunok Oh, Nikil Dutt, Soonhoi Ha |
CASES | 1 |
| 2002 | Efficient code synthesis from extended dataflow graphs for multimedia applicationsabstractThis paper presents efficient automatic code synthesis techniques from dataflow graphs for multimedia applications. Since multimedia applications require large size buffers containing composite type data, we aim to reduce the buffer sizes with fractional rate dataflow extension and buffer sharing technique. In an H.263 encoder experiment, the FRDF extension and buffer sharing technique enable us to reduce the buffer size by 67%. The final buffer size is no more than in a manual reference code. Hyunok Oh, Soonhoi Ha |
DAC | 1 |
| 2000 | Data memory minimization by sharing large size buffersabstractThis paper presents software synthesis techniques to deal with non-primitive data type from graphical dataflow programs based on the synchronous dataflow (SDF) model.Non-primitive data types, often used in multimedia and graphics applications, require buffer memory of large size.To minimize the buffer requirement, we separate global data buffers and local pointer buffers.The proposed approach first allocates the minimum size of global buffers and next binds the local buffers to the global buffers by setting the pointers.Static binding and dynamic binding techniques are devised.Experimental results prove the significance of the proposed techniques. Hyunok Oh, Soonhoi Ha |
ASP-DAC | 1 |