EDBT 2026 Demo / reviewers in the wild / expert
Shinhyoung Jang
dblp:406/1432
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0008-0378-3801ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MeshHD: Near-Linear Encoding for Hyperdimensional Computing via Multi-Scale Bases and Kronecker FactorizationabstractHyperdimensional (HD) computing is attractive for low-power platforms, but common encoders flatten inputs and treat neighboring features as independent, discarding spatial structure and inflating the cost of a dense F×D apply. We present MESHHD, a spatially aware, relative and multi-scale base that maps 2D coordinates with random Fourier features to approximate a distance kernel; nearby locations receive similar hypervectors regardless of absolute position. We further introduce a compact Kronecker-structured apply that realizes the bundled base with three small GEMMs, reducing arithmetic and weight movement from O(FD) toward a near-linear form while preserving encoder semantics. Our experimental results show that MESHHD consistently improves accuracy over the state-of-the-art nonlinear HD encoders, especially at smaller D, and reduces per-batch encoding time by ~ 3×, with up to 10× savings in encoder MACs/state at D=10,000. Woongjae Han, Jiseung Kim 0005, Hyukjun Kwon, Hojeong Kim, Selim An, Shinhyoung Jang, Yeseong Kim |
DATE | 6 |
| 2026 | Million-Scale Text-to-Video Retrieval with Hyperdimensional ComputingabstractScalable video retrieval is increasingly challenging as datasets reach tens of millions of videos. Current text-to-video retrieval (T2VR) methods either compress videos into single dense vectors, losing segment-level detail, or expand them into multi-frame representations, incurring prohibitive storage and search costs. We propose a binary hyperdimensional representation that encodes each video into a compact 3,072-dimension hypervector, preserving semantic fidelity while reducing memory via bit-packing. To leverage the properties of hypervectors for sublinear search, we introduce Hypervector Retrieval (HVR), a frequency-aware inverted index that prioritizes rare informative positions and refines candidates using GPU-accelerated Hamming search. Experiments show that our approach matches or exceeds dense baselines for T2VR and surpasses state-of-the-art partially relevant video retrieval (PRVR) by over 5% Recall@ 10 on ActivityNet. At scale, HVR processes over 2,000 queries per second on 10M videos, maintains recall within 1% of exact search, and achieves 5.3× greater storage capacity than CLIP4Clip and over 2,116× over MS-SL. Hyunsei Lee, Jaewoo Gwak, Shinhyoung Jang, Yeseong Kim |
EuroSys | 3 |
| 2025 | Bit-Level Semantics: Scalable RAG Retrieval with Neurosymbolic Hyperdimensional ComputingabstractRetrieval-Augmented Generation (RAG) systems typically rely on dense floating-point embeddings to retrieve relevant documents, but this approach incurs significant memory and compute costs at scale. We propose a Hyperdimensional Computing (HDC) framework that projects transformer token embeddings into high-dimensional binary hypervectors, which are aggregated into compact document representations. To support sublinear search, we introduce HD-NSW, a graph-based index inspired by navigable small-world networks. HD-NSW clusters similar hypervectors into bundled centroids and connects them with sparse Hammingdistance edges, enabling efficient, beam-guided traversal entirely in the binary domain. Across 15 BEIR benchmarks and synthetic Gaussian mixture corpora, HD-NSW achieves over 99% of dense retrieval quality, reduces memory usage by $8 \times$, and supports over 860 queries per second at 10 million documents while maintaining over 80% throughput at 40 million documents. At five million documents, HD-NSW achieves $7.68 \times$ higher throughput compared to state of the art approximate nearest neighbor methods. Beyond this point, competing baselines encounter memory exhaustion, while HD-NSW continues scaling and maintains high throughput at larger corpus sizes. Hyunsei Lee, Shinhyoung Jang, Jaewoo Gwak, Yeseong Kim |
PACT | 2 |
| 2025 | PersonalizedHD: Hyperdimensional Online Learning with Scalable Personalization and Memory-Efficient Replay
Shinhyoung Jang, Hyunsei Lee, Ilhong Suh, Yeseong Kim |
IEEE Big Data | 1 |
| 2025 | Late Breaking Results: Dynamically Scalable Pruning for Transformer-Based Large Language ModelsabstractWe propose Matryoshka, a novel framework for transformer model pruning, enabling dynamic runtime controls while maintaining competitive accuracy to modern large language models (LLMs). Matryoshka incrementally constructs submodels with varying complexities, allowing runtime adaptation without maintaining separate models. Our evaluations on LLaMA-7B demonstrate that Matryoshka achieves up to 34% speedup and outperforms the quality of state-of-the-art pruning methods, providing a flexible solution for deploying LLMs. Shinhyoung Jang, Ilhong Suh, Hoon Sung Chwa, Yeseong Kim |
DATE | 2 |
| 2025 | Hyperdimensional Computing-Based Federated Learning in Mobile Robots Through Synthetic OversamplingabstractTraditional federated learning frameworks, often reliant on deep neural networks, face challenges related to computational demands and privacy risks. In this paper, we present a novel Hyperdimensional (HD) Computing-based federated learning framework designed for resource-constrained mobile robots. Unlike other HD-based learning, our approach introduces dynamic encoding, which improves both model accuracy and privacy by continuously updating hypervector representations. To further address the issue of imbalanced data, especially prevalent in robotics tasks, we propose a hypervector oversampling technique, enhancing model robustness. Extensive evaluations on LiDAR-equipped mobile robots demonstrate that our oversampling method outperforms state-of-the-art HD computing frameworks, achieving up to a 22.9% increase in accuracy while maintaining computational efficiency. Hyunsei Lee, Woongjae Han, Hojeong Kim, Hyukjun Kwon, Shinhyoung Jang, Ilhong Suh, Yeseong Kim |
ICRA | 5 |