EDBT 2026 Demo / reviewers in the wild / expert
Hyunwoo Oh
dblp:205/0073
· DBLP profile ↗
17ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Symbolic Reasoning with Matrix-Based Brain-Inspired Representations and Vector-Space Acceleration
William Youngwoo Chung, Hyunwoo Oh, Hamza Errahmouni Barkam, Calvin Yeung 0002, Mohsen Imani |
DATE | 2 |
| 2026 | QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable AttentionabstractDeformable Transformers achieve state-of-the-art object detection, but deformable attention maps poorly to hardware due to irregular memory access and low arithmetic intensity. We present QUILL, a schedule-aware accelerator that makes MSDeformAttn cache-local and single-pass. QUILL’s Distance-based Out-of-Order Querying (DOOQ) reorders queries by spatial proximity, enabling a look-ahead, double-buffered prefetch that overlaps memory and compute. QUILL also uses a fused MSDeformAttn pipeline that performs interpolation, Softmax, aggregation, and output projection in one pass, avoiding intermediate spills and keeping small tensors on-chip. Implemented in RTL and evaluated end-to-end, QUILL achieves up to 7.29× higher throughput and 47.3× better energy efficiency than an RTX 4090, and improves throughput/energy efficiency over prior accelerators by 3.26–9.82× / 2.01–6.07×. With mixed precision, accuracy stays within ≤ 0.9 AP of FP32 across Deformable and Sparse DETR variants. By converting sparsity into locality and locality into utilization, QUILL delivers consistent end-to-end gains. Hyunwoo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Wenjun Huang 0001, Tamoghno Das, Suyeon Jang, Mohsen Imani |
DATE | 1 |
| 2026 | RIFT: A Single-Bitstream, Runtime-Adaptive FPGA-Based Accelerator for Multimodal AIabstractMultimodal models spanning ViTs, CNNs, GNNs, and NLP stress embedded systems because their heterogeneous compute and memory behaviors complicate resource allocation, load balancing, and real-time inference. We present RIFT, a single-bitstream FPGA accelerator and compiler for end-to-end multimodal inference. RIFT unifies layers as DDMM/SDDMM/SpMM kernels executed on a runtime mode-switchable engine that morphs among weight-/output-stationary systolic, 1×CSSIMD, and a routable adder tree (RADT) on a shared datapath. A two-stage hardware top-k unit, width-matched to the array, performs in-stream token pruning with minimal buffering, and dependency-aware scheduling overlaps independent kernels across multiple RPUs—achieving adaptation without bitstream reconfiguration. On Alveo U50 and ZCU104, RIFT reduces latency by up to 22.57× versus an RTX 4090 and 6.86× versus a Jetson Orin Nano at ∼20–21W; pruning alone yields up to 7.8× on ViT-heavy workloads. Hyunwoo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Suyeon Jang, Behnam Khaleghi, Fei Wen 0003, Mohsen Imani |
DATE | 1 |
| 2026 | T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU ReorganizationabstractRecent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. While ternary quantization enables significant resource savings, existing CPU solutions rely heavily on memory-based lookup tables (LUTs) which limit scalability, and FPGA or GPU accelerators remain impractical for edge use. This paper presents T-SAR, the first framework to achieve scalable ternary LLM inference on CPUs by repurposing the SIMD register file for dynamic, in-register LUT generation with minimal hardware modifications. T-SAR eliminates memory bottlenecks and maximizes data-level parallelism, delivering 5.6–24.5× and 1.1–86.2× improvements in GEMM latency and GEMV throughput, respectively, with only 3.2% power and 1.4% area overheads in SIMD units. T-SAR achieves up to 2.5–4.9× the energy efficiency of an NVIDIA Jetson AGX Orin, establishing a practical approach for efficient LLM inference on edge platforms. Hyunwoo Oh, KyungIn Nam, Rajat Bhattacharjya, Hanning Chen, Tamoghno Das, Sanggeon Yun, Suyeon Jang, Andrew Ding, Nikil Dutt, Mohsen Imani |
DATE | 1 |
| 2026 | DecoHD: Decomposed Hyperdimensional Classification under Extreme Memory BudgetsabstractDecomposition is a proven way to shrink deep networks without changing input-output dimensionality or interface semantics. We bring this idea to hyperdimensional computing (HDC), where footprint cuts usually shrink the feature axis and erode concentration and robustness. Prior HDC decompositions decode via fixed atomic hypervectors, which are ill-suited for compressing learned class prototypes. We introduce DecoHD, which learns directly in a decomposed HDC parameterization: a small, shared set of per-layer channels with multiplicative binding across layers and bundling at the end, yielding a large representational space from compact factors. DecoHD compresses along the class axis via a lightweight bundling head while preserving native bind-bundle-score; training is end-to-end, and inference remains pure HDC, aligning with in/near-memory accelerators. In evaluation, DecoHD attains extreme memory savings with only minor accuracy degradation under tight deployment budgets. On average it stays within about 0.1–0.15% of a strong non¬reduced HDC baseline (worst case 5.7%), is more robust to random bit-flip noise, reaches its accuracy plateau with up to ~ 97% fewer trainable parameters, and—in hardware—delivers roughly 277 × /35× energy/speed gains over a CPU (AMD Ryzen 9 9950X), 13.5 × /3.7× over a GPU (NVIDIA RTX 4090), and 2.0 × /2.4× over a baseline HDC ASIC. Sanggeon Yun, Hyunwoo Oh, Ryozo Masukawa, Mohsen Imani |
DATE | 2 |
| 2026 | LogHD: Robust Compression of Hyperdimensional Classifiers via Logarithmic Class-Axis ReductionabstractHyperdimensional computing (HDC) suits memory, energy, and reliability-constrained systems, yet the standard "one prototype per class" design requires $O(CD)$ memory (with $C$ classes and dimensionality $D$). Prior compaction reduces $D$ (feature axis), improving storage/compute but weakening robustness. We introduce LogHD, a logarithmic class-axis reduction that replaces the $C$ per-class prototypes with $n\!\approx\!\lceil\log_k C\rceil$ bundle hypervectors (alphabet size $k$) and decodes in an $n$-dimensional activation space, cutting memory to $O(D\log_k C)$ while preserving $D$. LogHD uses a capacity-aware codebook and profile-based decoding, and composes with feature-axis sparsification. Across datasets and injected bit flips, LogHD attains competitive accuracy with smaller models and higher resilience at matched memory. Under equal memory, it sustains target accuracy at roughly $2.5$-$3.0\times$ higher bit-flip rates than feature-axis compression; an ASIC instantiation delivers $498\times$ energy efficiency and $62.6\times$ speedup over an AMD Ryzen 9 9950X and $24.3\times$/$6.58\times$ over an NVIDIA RTX 4090, and is $4.06\times$ more energy-efficient and $2.19\times$ faster than a feature-axis HDC ASIC baseline. Sanggeon Yun, Hyunwoo Oh, Ryozo Masukawa, Pietro Mercati, Nathaniel D. Bastian, Mohsen Imani |
DATE | 2 |
| 2026 | Vector-Space Projection and Unified Complex-Valued Acceleration for Scaling Matrix-Based Brain-Inspired Representations
William Youngwoo Chung, Hyunwoo Oh, Calvin Yeung 0002, Hansen Jin Lillemark, Hamza Errahmouni Barkam, Mohsen Imani |
ISLPED | 2 |
| 2026 | FusionSense: Tri-Stage Near-Sensor Learning for Runtime-Adaptive Multimodal Edge IntelligenceabstractAutonomous systems and smart-industry deployments increasingly split computation across near-sensor, edge, and cloud resources, where tight energy, latency, and reliability budgets demand runtime adaptivity. In practice, deciding what to compute and transmit at each point is pivotal; yet as multimodal sensor suites (cameras, LiDAR/depth, etc.) proliferate at the edge, most prior approaches either (i) fuse modalities on powerful servers or (ii) apply uni-modal near-sensor filters that ignore cross-modal dependencies, leading to redundant transmissions or missed events. We present Fusion-Sense, a fusion-aware intelligent sensing framework for energy-constrained autonomous edge systems. Lightweight near-sensor classifiers are trained via a three-step procedure: (i) a server-side fusion model learns the downstream task, (ii) filter-out-safe (FoS) labels quantify each modality's necessity relative to the fused decision, and (iii) an edge-side fusion model is compacted by injecting near-sensor predictions as auxiliary signals. The result is a runtime decision layer that jointly reduces compute and communication while scaling linearly with sensor count. On a dual-modality (RGB+Depth/LiDAR) setup with SynDrone, FusionSense sustains task quality at substantially higher data-reduction rates than unimodal filters and delivers large end-to-end gains: up to 33× lower energy at 1% FoI prevalence, 11× at 10%, a 92.3% reduction in quality loss at a fixed 30% data reduction, and roughly 1.5× higher energy savings than the best prior filtering baseline. Sanggeon Yun, Ryozo Masukawa, Minhyoung Na, Hyunwoo Oh, Yoshiki Yamaguchi, Wenjun Huang 0001, Sungheon Jeong 0001, Mohsen Imani |
ISLPED | 4 |
| 2025 | VerbDiff: Text-Only Diffusion Models with Enhanced Interaction AwarenessabstractRecent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches. SeungJu Cha 0001, Kwanyoung Lee, Ye-Chan Kim, Hyunwoo Oh, Dong-Jin Kim 0003 |
CVPR | 4 |
| 2025 | iTaskSense: Task-Oriented Object Detection in Resource-Constrained EnvironmentsabstractTask-oriented object detection is increasingly essential for intelligent sensing applications, enabling AI systems to operate autonomously in complex, real-world environments such as autonomous driving, healthcare, and industrial automation. Conventional models often struggle with generalization, requiring vast datasets to accurately detect objects within diverse contexts. In this work, we introduce iTask, a taskoriented object detection framework that leverages large language models (LLMs) to generalize efficiently from limited samples by generating an abstract knowledge graph. This graph encapsulates essential task attributes, allowing iTask to identify objects based on high-level characteristics rather than extensive data, making it possible to adapt to complex mission requirements with minimal samples. iTask addresses the challenges of high computational cost and resource limitations in vision-language models by offering two configuration models: a distilled, task-specific vision transformer optimized for high accuracy in defined tasks, and a quantized version of the model for broader applicability across multiple tasks. Additionally, we designed a hardware acceleration circuit to support real-time processing, essential for edge devices that require low latency and efficient task execution. Our evaluations show that the task-specific configuration achieves a 15% higher accuracy over the quantized configuration in specific scenarios, while the quantized model provides robust multi-task performance. The hardware-accelerated iTask system achieves a $3.5 x$ speedup and a 40% reduction in energy consumption compared to GPU-based implementations. These results demonstrate that iTask’s dual-configuration approach and situational adaptability offer a scalable solution for task-specific object detection, providing robust and efficient performance in resourceconstrained environments. Sungheon Jeong 0001, Hamza Errahmouni Barkam, Hyunwoo Oh, Hanning Chen, Tamoghno Das, Mohsen Imani |
DAC | 3 |
| 2025 | Revisiting Reconfigurable Acceleration of Vision Transformer with Patch PruningabstractVision Transformers (ViTs) have become the backbone of numerous cutting-edge vision applications. The attention modules within ViTs play a crucial role in modeling spatial relationships between pixels. Although these attention modules enhance the accuracy of ViT models, they also increase computational demands, limiting the deployment of ViTs in edge computing environments. To address this issue, prior research has focused on optimizing ViTs from both software and hardware perspectives. A notable software optimization technique is reducing the image patches involved in attention computations. Two common methods to achieve this are window attention and patch pruning. However, they introduce new challenges for existing hardware platforms regarding attention computation. Therefore, it is essential to develop new hardware modules to simultaneously support pruned attention computations and efficient window shifts. In this study, we introduce an FPGA-based token reduction vision transformer accelerator called TRFPA. Experiments conducted on the Xilinx ZCU104 and Alveo U50 demonstrate that TRFPA outperforms previous FPGA-based ViT accelerators, achieving a 7× speedup and a 3× improvement in energy efficiency. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Tamoghno Das, Fei Wen 0003, Mohsen Imani |
ISLPED | 4 |
| 2025 | LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationabstractLarge Vision Language Models (LVLMs) have been widely adopted to guide vision foundation models in performing reasoning segmentation tasks, achieving impressive performance. However, the substantial computational overhead associated with LVLMs presents a new challenge. The primary source of this computational cost arises from processing hundreds of image tokens. Therefore, an effective strategy to mitigate such overhead is to reduce the number of image tokens-a process known as image token pruning. Previous studies on image token pruning for LVLMs have primarily focused on high-level visual understanding tasks, such as visual question answering and image captioning. In contrast, guiding vision foundation models to generate accurate visual masks based on textual queries demands precise semantic and spatial reasoning capabilities. Consequently, pruning methods must carefully control individual image tokens throughout the LVLM reasoning process. Our empirical analysis reveals that existing methods struggle to adequately balance reductions in computational overhead with the necessity to maintain high segmentation accuracy. In this work, we propose LVLM_CSP, a novel training-free visual token pruning method specifically designed for LVLM-based reasoning segmentation tasks. LVLM_CSP consists of three stages: clustering, scattering, and pruning. Initially, the LVLM performs coarse-grained visual reasoning using a subset of selected image tokens. Next, fine-grained reasoning is conducted, and finally, most visual tokens are pruned in the last stage. Extensive experiments demonstrate that LVLM_CSP achieves a 65% reduction in image token inference FLOPs with virtually no accuracy degradation, and a 70% reduction with only a minor 1% drop in accuracy on the 7B LVLM. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Yezi Liu, Tamoghno Das, Mohsen Imani |
ACM Multimedia | 4 |
| 2025 | CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image GenerationabstractWe propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining ) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector ). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment. Hyunwoo Oh, SeungJu Cha 0001, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim 0003 |
ACM Multimedia | 1 |
| 2025 | ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionabstractText-to-image diffusion models often exhibit degraded performance when generating images beyond their training resolution.
Recent training-free methods can mitigate this limitation, but they often require substantial computation or are incompatible with recent Diffusion Transformer models.
In this paper, we propose ScaleDiff, a model-agnostic and highly efficient framework for extending the resolution of pretrained diffusion models without any additional training.
A core component of our framework is Neighborhood Patch Attention (NPA), an efficient mechanism that reduces computational redundancy in the self-attention layer with non-overlapping patches.
We integrate NPA into an SDEdit pipeline and introduce Latent Frequency Mixing (LFM) to better generate fine details.
Furthermore, we apply Structure Guidance to enhance global structure during the denoising process.
Experimental results demonstrate that ScaleDiff achieves state-of-the-art performance among training-free methods in terms of both image quality and inference speed on both U-Net and Diffusion Transformer architectures. Sungho Koh, SeungJu Cha 0001, Hyunwoo Oh, Kwanyoung Lee, Dong-Jin Kim 0003 |
NeurIPS | 3 |
| 2025 | FlawMatch: Conditional defect image generation via flow matching for improved surface defect classification
Hyunwoo Oh, Seunghee Choi, Jinho Baek, Junegak Joung |
Adv. Eng. Informatics | 1 |
| 2018 | Learning to Identify Rush Strategies in StarCraft
Teguh Budianto, Hyunwoo Oh, Takehito Utsuro |
ICEC | 2 |
| 2017 | Identifying Rush Strategies Employed in StarCraft II Using Support Vector Machines
Teguh Budianto, Hyunwoo Oh, Zi Long, Takehito Utsuro |
ICEC | 2 |