Hyunok Oh

dblp:59/2601 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-9044-7441ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 5 first-authorSecurity and privacy · 12 · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Tangram: Encryption-Friendly SNARK Framework Under Pedersen Committed Engines
Gweonho Jeong, Myeongkyun Moon, Geonho Yoon, Hyunok Oh, Jihye Kim 0001
IEEE Trans. Dependable Secur. Comput.4
2025 DUPLEX: Scalable Zero-Knowledge Lookup Arguments over RSA Group
Semin Han, Geonho Yoon, Hyunok Oh, Jihye Kim 0001
AsiaCCS3
2025 zkMarket: Ensuring Fairness and Privacy in Decentralized Data Exchange
abstract
Ensuring fairness in blockchain-based data trading presents significant challenges, as the transparency of blockchain can expose sensitive details and compromise fairness. Fairness ensures that the seller receives payment only if they provide the correct data, and the buyer gains access to the data only after making the payment. Existing approaches face limitations in efficiency, particularly when applied to large-scale data. Moreover, preserving privacy has also been a significant challenge in blockchain. In this paper, we introduce zkMarket, a privacy-preserving fair trade system on the blockchain. We ensure fairness by integrating encryption with zk-SNARKs, enabling verifiable proofs for fair trading. However, applying zk-SNARKs directly can be computationally expensive for the prover. To address this, we improve efficiency by leveraging our novel matrix-formed PRG (MatPRG) and commit-and-prove SNARK (CP-SNARK), making the data registration process more concise and significantly reducing the seller's proving time. To ensure transaction privacy, zkMarket is built upon an anonymous transfer protocol. Experimental results demonstrate that zkMarket significantly reduces the computational overhead associated with traditional blockchain solutions while maintaining robust security and privacy. Specifically, our evaluation quantifies this high efficiency: the seller can register 1 MB of data in 2.8 seconds, the buyer can generate the trade transaction in 0.2 seconds, and the seller can finalize the trade within 0.4 seconds.
Seongho Park, Seungwoo Kim, Semin Han, Kyeongtae Lee, Jihye Kim 0001, Hyunok Oh
SRDS6
2024 zkLogis: Scalable, Privacy-Enhanced, and Traceable Logistics on Public Blockchain
abstract
Decentralized blockchain systems have provided a significant leap in autonomous logistics practices by enabling accurate product authentication through the tracking of ownership changes. Although public blockchains could provide open access to stored information, privacy and security concerns lead most existing logistics implementations to rely on private or consortium blockchains. This, however, limits or regulates the end-customer's ability to independently verify product provenance.
Hyunok Oh, Jihye Kim 0001
AsiaCCS3
2024 SAVER: SNARK-Compatible Verifiable Encryption
Jaekyoung Choi, Jihye Kim 0001, Hyunok Oh
FC (2)4
2024 vCNN: Verifiable Convolutional Neural Network Based on zk-SNARKs
abstract
It is becoming important for the client to be able to check whether the AI inference services have been correctly calculated. Since the weight values in a CNN model are assets of service providers, the client should be able to check the correctness of the result without them. The Zero-knowledge Succinct Non-interactive Argument of Knowledge (zk-SNARK) allows verifying the result without input and weight values. However, the proving time in zk-SNARK is too slow to be applied to real AI applications. This article proposes a new efficient verifiable convolutional neural network (vCNN) framework that greatly accelerates the proving performance. We introduce a new efficient relation representation for convolution equations, reducing the proving complexity of convolution from O(ln) to O(l+n) compared to existing zero-knowledge succinct non-interactive argument of knowledge (zk-SNARK) approaches, where l and n denote the size of the kernel and the data in CNNs. Experimental results show that the proposed vCNN improves proving performance by 20-fold for a simple MNIST and 18,000-fold for VGG16. The security of the proposed scheme is formally proven.
Seunghwa Lee, Hankyung Ko, Jihye Kim 0001, Hyunok Oh
IEEE Trans. Dependable Secur. Comput.4
2023 Efficient Transparent Polynomial Commitments for zk-SNARKs
Sungwook Kim 0001, Sungju Kim, Yulim Shin, Sunmi Kim, Jihye Kim 0001, Hyunok Oh
ESORICS (3)6
2022 Succinct Zero-Knowledge Batch Proofs for Set Accumulators
abstract
Cryptographic accumulators are a common solution to proving information about a large set S. They allow one to compute a short digest of S and short certificates of some of its basic properties, notably membership of an element. Accumulators also allow one to track set updates: a new accumulator is obtained by inserting/deleting a given element. In this work we consider the problem of generating membership and update proofs for \em batches of elements so that we can succinctly prove additional properties of the elements (i.e., proofs are of constant size regardless of the batch size), and we can preserve privacy. Solving this problem would allow obtaining blockchain systems with improved privacy and scalability.
Matteo Campanelli, Dario Fiore 0001, Semin Han, Jihye Kim 0001, Dimitris Kolonelos, Hyunok Oh
CCS6
2021 Efficient Verifiable Image Redacting based on zk-SNARKs
abstract
Image is a visual representation of a certain fact and can be used as proof of events. As the utilization of the image increases, it is required to prove its authenticity with the protection of its sensitive personal information. In this paper, we propose a new efficient verifiable image redacting scheme based on zk-SNARKs, a commitment, and a digital signature scheme. We adopt a commit-and-prove SNARK scheme which takes commitments as inputs, in which the authenticity can be quickly verified outside the circuit. We also specify relations between the original and redacted images to guarantee the redacting correctness. Our experimental results show that the proposed scheme is superior to the existing works in terms of the key size and proving time without sacrificing the other parameters. The security of the proposed scheme is proven formally.
Hankyung Ko, Ingeun Lee, Seunghwa Lee, Jihye Kim 0001, Hyunok Oh
AsiaCCS5
2019 SA-SPM: an efficient compiler for security aware scratchpad memory (invited paper)
abstract
Scratchpad memories (SPM) are often used to boost the performance of application-specific embedded systems. In embedded systems, main memories are vulnerable to external attacks such as bus snooping or memory extraction. Therefore it is desirable to guarantee the security of data in a main memory. In software-managed SPM, it is possible to provide security in main memory by performing software-assistant encryption.
Thomas Haywood Dadzie, Jihye Kim 0001, Hyunok Oh
LCTES4
2019 Forward Secure Identity-Based Signature Scheme with RSA
Hankyung Ko, Gweonho Jeong, Jihye Kim 0001, Hyunok Oh
SEC5
2019 FAS: Forward secure sequential aggregate signatures for secure logging
Jihye Kim 0001, Hyunok Oh
Inf. Sci.2
2019 AuthCropper: Authenticated Image Cropper for Privacy Preserving Surveillance Systems
abstract
As surveillance systems are popular, the privacy of the recorded video becomes more important. On the other hand, the authenticity of video images should be guaranteed when used as evidence in court. It is challenging to satisfy both (personal) privacy and authenticity of a video simultaneously, since the privacy requires modifications (e.g., partial deletions) of an original video image while the authenticity does not allow any modifications of the original image. This paper proposes a novel method to convert an encryption scheme to support partial decryption with a constant number of keys and construct a privacy-aware authentication scheme by combining with a signature scheme. The security of our proposed scheme is implied by the security of the underlying encryption and signature schemes. Experimental results show that the proposed scheme can handle the UHD video stream with more than 17 fps on a real embedded system, which validates the practicality of the proposed scheme.
Jihye Kim 0001, Hankyung Ko, Donghwan Oh, Semin Han, Gwonho Jeong, Hyunok Oh
ACM Trans. Embed. Comput. Syst.7
2018 Scalable Wildcarded Identity-Based Encryption
Jihye Kim 0001, Seunghwa Lee, Hyunok Oh
ESORICS (2)4
2018 Forward-secure ID based digital signature scheme with forward-secure private key generator
Hyunok Oh, Jihye Kim 0001, Ji Sun Shin
Inf. Sci.1
2018 A hybrid performance analysis technique for distributed real-time embedded systems
Junchul Choi, Hyunok Oh, Soonhoi Ha
Real Time Syst.2
2017 PASS: Privacy aware secure signature scheme for surveillance systems
abstract
In a surveillance system, the privacy becomes important since those who are not relevant to an event may be recorded by many surveillance systems. On the other hand, the authenticity of video frames in surveillance systems should be guaranteed if a video is used as evidence. Hence a signature is attached for each frame. However, it is contradictory to provide both privacy and authenticity of a video since the privacy requires deletion of objects in an original video image while the authenticity disallows any modification of the original image. This paper devises a new novel privacy aware secure signature scheme for the surveillance system. In the proposed scheme, deletion (or masking) of objects in an image is allowed while a signature still remains valid for the modified image. The proposed scheme utilizes a chameleon hash so that deletion of objects in an image does not invalidate the signature. The proposed scheme provides forward security minimizing the damage from the secret key exposure. The deletion is executed in an authorized way. Experimental results show that the proposed scheme is practical in a real-time video surveillance system due to the high performance of signature generation (40ms per frame) and a small signature size overhead (1%).
Jihye Kim 0001, Seunghwa Lee, Jungjun Yoon, Hankyung Ko, Seungri Kim, Hyunok Oh
AVSS6
2017 Hierarchical Dataflow Modeling of Iterative Applications
abstract
Even though dataflow models are good at exploiting task-level parallelism of an application, it is difficult to exploit the parallelism of loop structures since they are not explicitly specified in existent dataflow models. To overcome this drawback, we propose a novel extension to the SDF model, called SDF/L graph, specifying the loop structures explicitly in a hierarchical fashion. With a given SDF/L graph specification and the mapping and scheduling information, an application can be automatically parallelized on a multicore system. The enhanced expression capability by the proposed extension is verified with two applications, k-means clustering and deep neural network application.
Hyesun Hong, Hyunok Oh, Soonhoi Ha
DAC2
2017 Forward-Secure Digital Signature Schemes with Optimal Computation and Storage of Signers
Jihye Kim 0001, Hyunok Oh
SEC2
2017 Multiprocessor Scheduling of a Multi-Mode Dataflow Graph Considering Mode Transition Delay
abstract
The Synchronous Data Flow (SDF) model is widely used for specifying signal processing or streaming applications. Since modern embedded applications become more complex with dynamic behavior changes at runtime, several extensions of the SDF model have been proposed to specify the dynamic behavior changes while preserving static analyzability of the SDF model. They assume that an application has a finite number of behaviors (or modes), and each behavior (mode) is represented by an SDF graph. They are classified as multi-mode dataflow models in this article. While there exist several scheduling techniques for multi-mode dataflow models, no one allows task migration between modes. By observing that the resource requirement can be additionally reduced if task migration is allowed, we propose a multiprocessor scheduling technique of a multi-mode dataflow graph considering task migration between modes. Based on a genetic algorithm, the proposed technique schedules all SDF graphs in all modes simultaneously to minimize the resource requirement. To satisfy the throughput constraint, the proposed technique calculates the actual throughput requirement of each mode and the output buffer size for tolerating throughput jitter. We compare the proposed technique with a method that analyzes SDF graphs in each execution mode separately, a method that does not allow task migration, and a method that does not allow mode-overlapped schedule for synthetic examples and five real applications: H.264 decoder, lane detection, vocoder, MP3 decoder, and printer pipeline.
Hanwoong Jung, Hyunok Oh, Soonhoi Ha
ACM Trans. Design Autom. Electr. Syst.2
2016 In-storage processing of database scans and joins
Sungchan Kim, Hyunok Oh, Chanik Park, Sangyeun Cho, Sang-Won Lee 0001, Bongki Moon
Inf. Sci.2
2015 Optimization of multi-channel BCH error decoding for common cases
abstract
This paper proposes a new method to optimize a BCH error correction decoder in multi-channel configurations. We break the BCH decoding process into its three basic blocks: syndrome calculation, the error locator polynomial generation, and the roots of the error locator polynomial computation. While an existing multi-channel BCH decoder consists of several single-channel BCH decoders operating in parallel, this paper utilizes a pooled group of shared decoding blocks. By considering the frequency of errors, the proposed pooled group approach requires fewer hardware blocks than in a traditional multi-channel configuration with a negligible impact on performance. Combined with a specialized root finding unit for blocks with only 1 error, our scheme reduces hardware area by 47%-71% and dynamic power by 44%-59% with 2% performance degradation in typical NAND flash systems. With a constant hardware area, the proposed scheme can improve throughput by 3x-5x or NAND flash lifetime by 1.4x-4.5x.
Russ Dill, Aviral Shrivastava, Hyunok Oh
CASES3
2015 Optimal Checkpoint Selection with Dual-Modular Redundancy Hardening
abstract
With the continuous scaling of semiconductor technology, failure rate is increasing significantly so that reliability becomes an important issue in multiprocessor system-on-chip (MPSoC) design. We propose an optimal checkpoint selection with task duplication hardening to tolerate transient faults. A target application is specified in a task graph, and the schedule/checkpoint placements are determined at design time. The proposed optimal algorithm minimizes the checkpoint overhead with a latency constraint. Experimental results show that the proposed algorithm effectively reduces the minimum end-to-end latency to perform a fault-tolerant schedule. In addition, the proposed algorithm dramatically decreases the checkpointing overhead on uniprocessor and multiprocessor systems compared with a greedy approach and an equidistant algorithm.
Shin-Haeng Kang, Hae-woo Park, Sungchan Kim, Hyunok Oh, Soonhoi Ha
IEEE Trans. Computers4
2014 An Efficient Non-Linear Cost Compression Algorithm for Multi Level Cell Memory
abstract
This paper defines a non-linear cost compression problem, proposes an efficient algorithm, and applies it to a real application of multi level cell memory to minimize energy consumption and latency. The non-linear cost compression problem extends the traditional cost compression problem to allow a non-linear cost function of symbol frequencies, while it is a weighted linear combination of symbol frequencies in the cost compression problem. In order to solve the non-linear cost compression problem efficiently, we propose an encoding symbol frequency based approach. We first compute frequencies of encoding symbols to minimize a cost function. To achieve the computed frequencies of a cost-compressed message, we deploy existing size-decompression algorithms. The proposed algorithm is optimal and as fast as the existing size compression algorithms. Our experimental results show that it reduces the energy consumption and latency by 70 percent for a text file in multi level cell memory. Furthermore, it increases the lifetime of endurance limited memory.
Hyunok Oh, Jihye Kim 0001
IEEE Trans. Computers1
2014 Dynamic Behavior Specification and Dynamic Mapping for Real-Time Embedded Systems: HOPES Approach
abstract
As the number of processors in a chip increases and more functions are integrated, the system status will change dynamically due to various factors such as the workload variation, QoS requirement, and unexpected component failure. A typical method to deal with the dynamics of the system is to decide the mapping decision at runtime, based on the local information of the system status. It is very challenging to guarantee any real-time performance of a certain application in such a dynamically varying system. To solve this problem, we propose a hybrid specification of dataflow and FSM models to specify the dynamic behavior of a system distinguishing inter- and intra-application dynamism. At the top level, each application is specified by a dataflow task and the dynamic behavior is modeled as a control task that supervises the execution of applications. Inside a dataflow task, we specify the dynamic behavior using a similar way as FSM-based SADF in which an application is specified by a synchronous dataflow graph for each mode of operation. It enables us to perform compile-time scheduling of each graph to maximize the throughput varying the number of allocated processors, and store the scheduling information. When a change in system state is detected at runtime, the number of allocated processors to the active tasks is determined dynamically utilizing the stored scheduling information of those tasks in order to meet the real-time requirements. The proposed technique is implemented in the HOPES design environment. Through preliminary experiments with a simple smartphone example, we show the viability of the proposed methodology.
Hanwoong Jung, Chanhee Lee 0002, Shin-Haeng Kang, Sungchan Kim, Hyunok Oh, Soonhoi Ha
ACM Trans. Embed. Comput. Syst.5
2013 Intelligent SSD: a turbo for big data mining
abstract
This paper introduces the notion of intelligent SSDs. First, we present the design considerations of intelligent SSDs, and then examine their potential benefits under various settings in data mining applications.
Duck-Ho Bae, Jin-Hyung Kim, Sang-Wook Kim, Hyunok Oh, Chanik Park
CIKM4
2013 A novel analytical method for worst case response time estimation of distributed embedded systems
abstract
In this paper, we propose a novel analytical method, called scheduling time bound analysis, to find a tight upper bound of the worst-case response time in a distributed real-time embedded system, considering execution time variations of tasks, jitter of input arrivals, and scheduling anomaly behavior in a multi-tasking system all together. By analyzing the graph topology and worst-case scheduling scenarios, we measure the conservative scheduling time bound of each task. The proposed method supports an arbitrary mixture of preemptive and non-preemptive processing elements. Its speed is comparable to compositional approaches while it gives a much tighter bound. The advantages of the proposed approach compared with related work were verified by experimental results with randomly generated task graphs and a real-life automotive application.
Hyunok Oh, Junchul Choi, Hyojin Ha, Soonhoi Ha
DAC2
2013 Active disk meets flash: a case for intelligent SSDs
abstract
Intelligent solid-state drives (iSSDs) allow execution of limited application functions (e.g., data filtering or aggregation)on their internal hardware resources, exploiting SSD characteristics and trends to provide large and growing performance and energy efficiency benefits. Most notably, internal flash media bandwidth can be significantly (2-4x or more) higher than the external bandwidth with which the SSD is connected to a host system, and the higher internal bandwidth can be exploited within an iSSD. Also, SSD bandwidth is projected to increase rapidly over time, creating a substantial energy cost for streaming of data to an external CPU for processing, which can be avoided via iSSD processing. This paper makes a case for iSSDs by detailing these trends, quantifying the potential benefifits across a range of application activities, describing how SSD architectures could be extended cost-effectively, and demonstrating the concept with measurements of a prototype iSSD running simple data scan functions. Our analyses indicate that, with less than a 2% increase in hardware cost over a traditional SSD, an iSSD can provide 2-4x performance increases and 5-27x energy efficiency gains for a range of data-intensive computations.
Sangyeun Cho, Chanik Park, Hyunok Oh, Sungchan Kim, Youngmin Yi, Gregory R. Ganger
ICS3
2013 A lifetime aware buffer assignment method for streaming applications on DRAM/PRAM hybrid memory
abstract
This article proposes a lifetime aware buffer assignment method for streaming applications like multimedia specified in a synchronous dataflow (SDF) graph on a DRAM/PRAM hybrid memory in which the endurance of PRAM is limited. We determine whether buffers are assigned to DRAM or PRAM to minimize the writing frequency of PRAM. To solve the problems, we formulate them using Answer Set Programming. Experimental results show that the proposed approach increases the PRAM lifetime by 63% compared with no optimization, and shows the tradeoff between PRAM and DRAM size to guarantee a lifetime constraint.
Hyunok Oh
ACM Trans. Embed. Comput. Syst.2
2012 Executing synchronous dataflow graphs on a SPM-based multicore architecture
abstract
In this paper we are concerned about executing synchronous dataflow (SDF) applications on a multicore architecture where a core has a limited size of scratchpad memory (SPM). Unlike traditional multi-processor scheduling of SDF graphs, we consider the SPM size limitation that incurs code and data overlay overhead. Since the scheduling problem is intractable, we propose an EA(evolutionary algorithm)-based technique. To hide memory latency, prefetching is aggressively performed in the proposed technique. The experimental results show that our approach reduces the overlay overhead significantly compared to a non-optimized approach and the previous approach.
Junchul Choi, Hyunok Oh, Sungchan Kim, Soonhoi Ha
DAC2
2012 An ILP-based Worst-case Performance Analysis Technique for Distributed Real-time Embedded Systems
abstract
Finding a tight upper bound of the worst-case response time in a distributed real-time embedded system is a very challenging problem since we have to consider execution time variations of tasks, jitter of input arrivals, scheduling anomaly behavior in a multi-tasking system, all together. In this paper, we translate the problem as an optimization problem and propose a novel solution based on ILP (Integer Linear Programming). In the proposed technique, we formulate a set of ILP formulas in a compositional way for modeling flexibility, but solve the problem holistically to achieve tighter upper bounds. To mitigate the time complexity of the ILP method, we perform static analysis based on a scheduling heuristic to reduce the number of variables and confine the variable ranges. Preliminary experiments with the benchmarks used in the related work and a real-life example show promising results that give tight bounds in an affordable solution time.
Hyunok Oh, Hyojin Ha, Shin-Haeng Kang, Junchul Choi, Soonhoi Ha
RTSS2
2012 A parallel and distributed meta-heuristic framework based on partially ordered knowledge sharing
Minyoung Kim 0002, Mark-Oliver Stehr, Hyunok Oh, Soonhoi Ha
J. Parallel Distributed Comput.4
2011 Minimizing buffer requirements for throughput constrained parallel execution of synchronous dataflow graph
abstract
This paper concerns throughput-constrained parallel execution of synchronous data flow graphs. This paper assumes static mapping and dynamic scheduling of nodes, which has several benefits over static scheduling approaches. We determine the buffer size of all arcs to minimize the total buffer size while satisfying a throughput constraint. Dynamic scheduling is able to achieve the similar throughput performance as the static scheduling does by unfolding the given SDF graph. A key issue of dynamic scheduling is how to assign the priority to each node invocation, which is also discussed in this paper. Since the problem is NP-hard, we present a heuristic based on a genetic algorithm. The experimental results confirm the viability of the proposed technique.
Tae-ho Shin, Hyunok Oh, Soonhoi Ha
ASP-DAC2
2011 Library Support in an Actor-Based Parallel Programming Platform
abstract
Actor model-based design is actively researched for parallel embedded SW design since the model exposes the potential parallelism explicitly in an architecture-neutral form. In most actor-oriented models, actors are self-contained and data channels are the only sharable object between actors, and they compose a system in a flat layer. In contrast, it is common to use shared library functions and construct vertically layered software for efficiency and modularity. To fill this gap between modeling and implementation, we propose a special actor, library task, with new types of ports: library master port and library slave port. It is a sharable and mappable object that defines a set of function interfaces inside. N:1 master-slave connection allows sharing a library task and the master-slave connection can specify vertically layered software and client-server applications naturally. To support the library task in our embedded software design environment, we develop an automatic mapping algorithm as well as an automatic code generator. The design environment with the library task is applied for two target platforms: IBM CELL Broad band Engine and an ARM-based multicore simulator. Preliminary experiments show that the special actor, or library task, extends the expression power of the previous actor model with efficiently generated codes.
Hae-woo Park, Hanwoong Jung, Hyunok Oh, Soonhoi Ha
IEEE Trans. Ind. Informatics3
2006 Memory optimal single appearance schedule with dynamic loop count for synchronous dataflow graphs
abstract
In this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size simultaneously. While a single appearance schedule promises only one appearance of each node definition in the generated code, it requires significant amount of data memory overhead compared with a buffer optimal schedule allowing multiple appearance. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that every buffer optimal schedule can be transformed to our single appearance schedule which requires optimal buffer size for arbitrary synchronous dataflow graphs. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. In order to minimize the overhead we propose optimization techniques. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms.
Hyunok Oh, Nikil Dutt, Soonhoi Ha
ASP-DAC1
2005 Single appearance schedule with dynamic loop count for minimum data buffer from synchronous dataflow graphs
abstract
In this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size at the same time. When the software code is automatically synthesized from the dataflow program graphs, a single appearance schedule promises only one appearance of each node definition in the generated code. While several heuristics have been developed to find a single appearance schedule, they all have to pay significant amount of data memory overhead compared with a buffer optimal schedule. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that the proposed scheduling technique is optimal for a chain-structured graph in terms of data memory requirement while maintaining the single appearance schedule. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms for CD2DAT and non uniform filter bank applications.
Hyunok Oh, Nikil Dutt, Soonhoi Ha
CASES1
2002 Efficient code synthesis from extended dataflow graphs for multimedia applications
abstract
This paper presents efficient automatic code synthesis techniques from dataflow graphs for multimedia applications. Since multimedia applications require large size buffers containing composite type data, we aim to reduce the buffer sizes with fractional rate dataflow extension and buffer sharing technique. In an H.263 encoder experiment, the FRDF extension and buffer sharing technique enable us to reduce the buffer size by 67%. The final buffer size is no more than in a manual reference code.
Hyunok Oh, Soonhoi Ha
DAC1
2000 Data memory minimization by sharing large size buffers
abstract
This paper presents software synthesis techniques to deal with non-primitive data type from graphical dataflow programs based on the synchronous dataflow (SDF) model.Non-primitive data types, often used in multimedia and graphics applications, require buffer memory of large size.To minimize the buffer requirement, we separate global data buffers and local pointer buffers.The proposed approach first allocates the minimum size of global buffers and next binds the local buffers to the global buffers by setting the pointers.Static binding and dynamic binding techniques are devised.Experimental results prove the significance of the proposed techniques.
Hyunok Oh, Soonhoi Ha
ASP-DAC1