Juno Kim

dblp:59/8200 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0003-1300-9875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 10 first-author · 12 since 2021Systems, architecture and hardware · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Kinematic Sickness: Understanding Cybersickness Through Body Kinematics
abstract
Postural Instability Theory (PIT) proposes that individuals who are naturally unstable on their feet are more susceptible to cybersickness. We hypothesize that this relationship extends to locomotive VR, such that people who exhibit greater instability when walking without VR will also be more susceptible to cybersickness in a locomotive VR setup. To test this, we analyzed participants' natural walking kinematics alongside their cybersickness responses and kinematic patterns during mobile VR use. Our results showed that vertical Center of Mass movement during pre-VR walking showed promise for identifying individuals susceptible to cybersickness. Spatial stability metrics emerged as stronger predictors of cybersickness than time-series measures, suggesting that spatial characteristics of gait may be more informative indicators of susceptibility in mobile VR contexts. These findings highlight the importance of accounting for baseline postural stability when designing and personalizing mobile VR experiences.
Carlos Alfredo Tirado Cortes, Yiheng Chi, Juno Kim, Hsiang-Ting Chen
IEEE Trans. Vis. Comput. Graph.3
2025 Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
abstract
We provide a convergence analysis of \emph{deep feature instrumental variable} (DFIV) regression (Xu et al., 2021), a nonparametric approach to IV regression using data-adaptive features learned by deep neural networks in two stages. We prove that the DFIV algorithm achieves the minimax optimal learning rate when the target structural function lies in a Besov space. This is shown under standard nonparametric IV assumptions, and an additional smoothness assumption on the regularity of the conditional distribution of the covariate given the instrument, which controls the difficulty of Stage 1. We further demonstrate that DFIV, as a data-adaptive algorithm, is superior to fixed-feature (kernel or sieve) IV methods in two ways. First, when the target function possesses low spatial homogeneity (i.e., it has both smooth and spiky/discontinuous regions), DFIV still achieves the optimal rate, while fixed-feature methods are shown to be strictly suboptimal. Second, comparing with kernel-based two-stage regression estimators, DFIV is provably more data efficient in the Stage 1 samples.
Juno Kim, Dimitri Meunier, Arthur Gretton, Taiji Suzuki
ICLR1
2025 Transformers Provably Solve Parity Efficiently with Chain of Thought
abstract
This work provides the first theoretical analysis of training transformers to solve complex problems by recursively generating intermediate states, analogous to fine-tuning for chain-of-thought (CoT) reasoning. We consider training a one-layer transformer to solve the fundamental $k$-parity problem, extending the work on RNNs by \citet{Wies23}. We establish three key results: (1) any finite-precision gradient-based algorithm, without intermediate supervision, requires substantial iterations to solve parity with finite samples. (2) In contrast, when intermediate parities are incorporated into the loss function, our model can learn parity in one gradient update when aided by \emph{teacher forcing}, where ground-truth labels of the reasoning chain are provided at each generation step. (3) Even without teacher forcing, where the model must generate CoT chains end-to-end, parity can be learned efficiently if augmented data is employed to internally verify the soundness of intermediate steps. Our findings, supported by numerical experiments, show that task decomposition and stepwise reasoning naturally arise from optimizing transformers with CoT; moreover, self-consistency checking can improve multi-step reasoning ability, aligning with empirical studies of CoT.
Juno Kim, Taiji Suzuki
ICLR1
2025 Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
abstract
A key paradigm to improve the reasoning capabilities of large language models (LLMs) is to allocate more inference-time compute to search against a verifier or reward model. This process can then be utilized to refine the pretrained model or distill its reasoning patterns into more efficient models. In this paper, we study inference-time computation by viewing chain-of-thought (CoT) generation as a metastable Markov process: easy reasoning steps (e.g., algebraic manipulations) form densely connected clusters, while hard reasoning steps (e.g., applying a relevant theorem) create sparse, low-probability edges between clusters, leading to phase transitions at longer timescales. Under this framework, we prove that implementing a search protocol that rewards sparse edges improves CoT by decreasing the expected number of steps to reach different clusters. In contrast, we establish a limit on reasoning capability when the model is restricted to local information of the pretrained graph. We also show that the information gained by search can be utilized to obtain a better reasoning model: (1) the pretrained model can be directly finetuned to favor sparse edges via policy gradient methods, and moreover (2) a compressed metastable representation of the reasoning dynamics can be distilled into a smaller, more efficient model.
Juno Kim, Denny Wu, Jason D. Lee, Taiji Suzuki
ICML1
2025 DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation
abstract
In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelfpicking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB-based methods tend to over-segment objects due to their reliance on texture, while depth-based methods often under-segment by focusing primarily on geometric features. To address these limitations, we propose DA-Fusion, a deformable attention-based RGB-D fusion Transformer designed for unseen object instance segmentation. DA-Fusion effectively combines the strengths of both RGB and depth data, enhancing segmentation accuracy in cluttered and multi-layered object environments. We also introduce the Object Clutter Bin Dataset (OCBD), a benchmark dataset specifically tailored for evaluating bin-picking scenarios in top-down views. Extensive evaluations demonstrate that DA-Fusion outperforms state-of-the-art methods across diverse environments, making it particularly suited for real-world logistics tasks.
Yesol Park, Hye Jung Yoon, Juno Kim, Byoung-Tak Zhang
ICRA3
2025 TOSS: Tiering of Serverless Snapshots for Memory-Efficient Serverless Computing
abstract
Serverless computing is an emerging cloud computing paradigm where users offload functions to serverless platforms that manage their own execution environments. Despite recent advancements on efficient serverless management, we find that cloud providers' solutions lead to unnecessary memory overheads, and significant memory cost by assuming a single tier of memory (DRAM). In this paper, we evaluate previous serverless works and offer insights about enabling memory tiering for serverless. Based on our insights, we introduce Tiering of Serverless Snapshots (TOSS), a heterogeneous memory mechanism that aims to reduce the total memory cost on serverless platforms, while achieving comparable performance to single-tier solutions. We show that TOSS achieves near optimal memory cost for most functions, while offloading on average 92% of the memory to the slow tier. In addition, TOSS achieves$52 \times$lower setup time and up to$4.2 \times$lower invocation time than the state-of-the-art DRAM-only mechanism.
Theodore Michailidis, Juno Kim, Linsong Guo, Steven Swanson, Jishen Zhao
IPDPS2
2025 CDIS : Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging
abstract
Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often breaks object identities across time and yields fragmented 3D instances. We introduce Cross-Dimensional Class-Agnostic 3D Instance Segmentation (CDIS), a zero-shot framework that explicitly tracks 2D instance masks across frames and associates them with 3D superpoints, creating a feedback loop between 2D and 3D. This cross-dimensional reasoning links temporally stable 2D tracks with spatially coherent 3D regions, producing globally consistent 3D instance labels without any 3D-specific training. Experiments on benchmark datasets demonstrate that CDIS achieves higher accuracy and consistency than state-of-the-art zero-shot methods, while remaining efficient and scalable to diverse real-world environments.
Juno Kim, Hye Jung Yoon, Yesol Park, Byoung-Tak Zhang
IROS1
2025 Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
abstract
Wasserstein gradient flow (WGF) is a common method to perform optimization over the space of probability measures. While WGF is guaranteed to converge to a first-order stationary point, for nonconvex functionals the converged solution does not necessarily satisfy the second-order optimality condition; i.e., it could converge to a saddle point. In this work, we propose a new algorithm for probability measure optimization, \emph{perturbed Wasserstein gradient flow} (PWGF), that achieves second-order optimality for general nonconvex objectives. PWGF enhances WGF by injecting noisy perturbations near saddle points via a Gaussian process-based scheme. By pushing the measure forward along a random vector field generated from a Gaussian process, PWGF helps the solution escape saddle points efficiently by perturbing the solution towards the smallest eigenvalue direction of the Wasserstein Hessian. We theoretically derive the computational complexity for PWGF to achieve a second-order stationary point. Furthermore, we prove that PWGF converges to a global optimum in polynomial time for strictly benign objectives.
Naoya Yamamoto, Juno Kim, Taiji Suzuki
NeurIPS2
2025 "Differences in Virtual and Physical Head Pose" Predict Cybersickness When Naturalistic Head-Movements are Made in VR
abstract
When we move during virtual reality (VR) display lag produces Differences in our Virtual and Physical head pose (DVP). Research suggests that DVP can be used to predict cybersickness during head-mounted display (HMD) based VR. However, these studies always had participants make unusual (continuous oscillatory) head-movements. This study examined whether DVP also predicts cybersickness during more typical VR conditions. After assessing their susceptibility to real-world motion sickness (using the MSSQ-Revised), 67 participants repeatedly moved their heads to “target” objects that appeared inside a virtual room (under different experimentally imposed display lags). We found that cybersickness was more likely and severe when: (1) participants had higher MSSQ scores; (2) the spatial magnitudes and the detrended fluctuation analysis α values of their DVP increased. Based on these findings we believe that real-time estimates of the DVP could be used to warn users about the imminent onset of sickness during consumer HMD VR.
Stephen A. Palmisano, Michael Mcfadyen, Sebastien Miellet, Robert S. Allison, Juno Kim
Int. J. Hum. Comput. Interact.5
2024 $t^3$-Variational Autoencoder: Learning Heavy-tailed Data with Student's t and Power Divergence
abstract
The variational autoencoder (VAE) typically employs a standard normal prior as a regularizer for the probabilistic latent encoder. However, the Gaussian tail often decays too quickly to effectively accommodate the encoded points, failing to preserve crucial structures hidden in the data. In this paper, we explore the use of heavy-tailed models to combat over-regularization. Drawing upon insights from information geometry, we propose $t^3$VAE, a modified VAE framework that incorporates Student's t-distributions for the prior, encoder, and decoder. This results in a joint model distribution of a power form which we argue can better fit real-world datasets. We derive a new objective by reformulating the evidence lower bound as joint optimization of KL divergence between two statistical manifolds and replacing with $\gamma$-power divergence, a natural alternative for power families. $t^3$VAE demonstrates superior generation of low-density regions when trained on heavy-tailed synthetic data. Furthermore, we show that $t^3$VAE significantly outperforms other models on CelebA and imbalanced CIFAR-100 datasets.
Juno Kim, Jaehyuk Kwon, Mincheol Cho, Joong-Ho Won
ICLR1
2024 Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
abstract
In this paper, we extend mean-field Langevin dynamics to minimax optimization over probability distributions for the first time with symmetric and provably convergent updates. We propose \emph{mean-field Langevin averaged gradient} (MFL-AG), a single-loop algorithm that implements gradient descent ascent in the distribution spaces with a novel weighted averaging, and establish average-iterate convergence to the mixed Nash equilibrium. We also study both time and particle discretization regimes and prove a new uniform-in-time propagation of chaos result which accounts for the dependency of the particle interactions on all previous distributions. Furthermore, we propose \emph{mean-field Langevin anchored best response} (MFL-ABR), a symmetric double-loop algorithm based on best response dynamics with linear last-iterate convergence. Finally, we study applications to zero-sum Markov games and conduct simulations demonstrating long-term optimality.
Juno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji Suzuki
ICLR1
2024 Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
abstract
Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of attention trained on linear regression tasks. In this paper, we study the optimization of a Transformer consisting of a fully connected layer followed by a linear attention layer. The MLP acts as a common nonlinear representation or feature map, greatly enhancing the power of in-context learning. We prove in the mean-field and two-timescale limit that the infinite-dimensional loss landscape for the distribution of parameters, while highly nonconvex, becomes quite benign. We also analyze the second-order stability of mean-field dynamics and show that Wasserstein gradient flow almost always avoids saddle points. Furthermore, we establish novel methods for obtaining concrete improvement rates both away from and near critical points. This represents the first saddle point analysis of mean-field dynamics in general and the techniques are of independent interest.
Juno Kim, Taiji Suzuki
ICML1
2024 OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
abstract
We introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance object recognition capabilities. A significant challenge arises when overlapping features from adjacent voxels reduce instance-level precision, as features spill over voxel boundaries, blending neighboring regions together. Our method overcomes this by employing a class-agnostic segmentation model to project 2D masks into 3D space, combined with a supplemented depth image created by merging raw and synthetic depth from point clouds. This approach, along with a 3D mask voting mechanism, enables accurate zero-shot 3D instance segmentation without relying on 3D supervised segmentation models. We assess the effectiveness of our method through comprehensive experiments on public datasets such as ScanNet200 and Replica, demonstrating superior zero-shot performance, robustness, and adaptability across diverse environments. Additionally, we conducted real-world experiments to demonstrate our method’s adaptability and robustness when applied to diverse real-world environments.
Juno Kim, Yesol Park, Hye Jung Yoon, Byoung-Tak Zhang
IROS1
2024 Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking
abstract
Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in unstructured environments. To overcome these limitations, we introduce Seg2Grasp, a modular pipeline designed for robust suction grasping in dynamic and cluttered bin scenarios. Seg2Grasp is built on a three-step process: Segmentation, Grasping, and Classification. The Segmentation module employs a Transformer-based model to generate class-agnostic object masks from RGB-D images, ensuring accurate detection across various conditions. The Grasping module uses surface normals and mask proposals to determine the optimal suction points, enhancing grasp success. Finally, the Classification module leverages fine-tuned open-vocabulary Mask-CLIP for precise object identification, enabling versatile handling of diverse objects. Real-world robotic experiments demonstrate that Seg2Grasp outperforms existing methods in success rates and adaptability, establishing it as a powerful tool for automated bin picking in industrial settings.
Hye Jung Yoon, Juno Kim, Yesol Park, Jun-Ki Lee, Byoung-Tak Zhang
IROS2
2024 Transformers are Minimax Optimal Nonparametric In-Context Learners
abstract
In-context learning (ICL) of large language models has proven to be a surprisingly effective method of learning a new task from only a few demonstrative examples. In this paper, we shed light on the efficacy of ICL from the viewpoint of statistical learning theory. We develop approximation and generalization error analyses for a transformer model composed of a deep neural network and one linear attention layer, pretrained on nonparametric regression tasks sampled from general function spaces including the Besov space and piecewise $\gamma$-smooth class. In particular, we show that sufficiently trained transformers can achieve -- and even improve upon -- the minimax optimal estimation risk in context by encoding the most relevant basis representations during pretraining. Our analysis extends to high-dimensional or sequential data and distinguishes the \emph{pretraining} and \emph{in-context} generalization gaps, establishing upper and lower bounds w.r.t. both the number of tasks and in-context examples. These findings shed light on the effectiveness of few-shot prompting and the roles of task diversity and representation learning for ICL.
Juno Kim, Tai Nakamaki, Taiji Suzuki
NeurIPS1
2024 Effects of Constant and Time-Varying Display Lag on DVP and Cybersickness When Making Head-Movements in Virtual Reality
abstract
When HMD users move their heads in virtual reality (VR), display lag creates differences between their virtual and physical head pose (DVP). This study examined whether objective estimates of DVP could predict experiences of cybersickness during simulations with three different types of added lag: (1) Constant lag (where the display was always delayed by 250 ms); (2) Predictable time-varying lag (where delays alternated between 0 and 250 ms every 5 s); and (3) Random time-varying lag (where delays alternated between 0 and a randomly determined value, up to 250 ms, every 1–5 s). Constant, Predictable, and Random added lag were found to generate similar levels of cybersickness—with all three conditions producing more severe sickness than the Baseline lag control. Consistent with our DVP hypothesis, the spatial magnitude and temporal dynamics of our participants’ DVP were both found to be reliable predictors of their cybersickness in all display lag conditions tested.
Stephen A. Palmisano, Robert S. Allison, Rodney G. Davies, Juno Kim
Int. J. Hum. Comput. Interact.5
2022 Blaze: Fast Graph Processing on Fast SSDs
abstract
Out-of-core graph processing is an attractive solution for processing very large graphs that do not fit in the memory of a single machine. The new class of ultra-low-latency SSDs should expand the impact and utility of out-of-core graph processing systems. However, current out-of-core systems cannot fully leverage the high IOPS these devices can deliver. We introduce Blaze, a new out-of-core graph processing system optimized for ultra-low-latency SSDs. Blaze offers high-performance out-of-core graph analytics by constantly saturating these fast SSDs with a new scatter-gather technique called online binning that allows value propagation among graph vertices without atomic synchronization. Blaze offers succinct APIs to allow programmers to write efficient out-of-core graph algorithms without the burden to manage complex IO executions. Our evaluation shows that Blaze outperforms current out-of-core systems by a wide margin on seven datasets and a set of representative graph queries on Intel Optane SSD.
Juno Kim, Steven Swanson
SC1
2021 Ayudante: A Deep Reinforcement Learning Approach to Assist Persistent Memory Programming
Hanxian Huang, Zixuan Wang 0027, Juno Kim, Steven Swanson, Jishen Zhao
USENIX ATC3
2020 An Empirical Guide to the Behavior and Use of Scalable Persistent Memory
Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz, Steven Swanson
FAST2
2019 Finding and Fixing Performance Pathologies in Persistent Memory Software Stacks
abstract
Emerging fast, non-volatile memories will enable systems with large amounts of non-volatile main memory (NVMM) attached to the CPU memory bus, bringing the possibility of dramatic performance gains for IO-intensive applications. This paper analyzes the impact of state-of-the-art NVMM storage systems on some of these applications and explores how those applications can best leverage the performance that NVMMs offer. Our analysis leads to several conclusions about how systems and applications should adapt to NVMMs. We propose FiLe Emulation with DAX (FLEX), a technique for moving file operations into user space, and show it and other simple changes can dramatically improve application performance. We examine the scalability of NVMM file systems in light of the rising core counts and pronounced NUMA effects in modern systems, and propose changes to Linux's virtual file system (VFS) to improve scalability. We also show that adding NUMA-aware interfaces to an NVMM file system can significantly improve performance.
Jian Xu 0012, Juno Kim, Amir Saman Memaripour, Steven Swanson
ASPLOS2
2019 Monocular Viewing Protects Against Cybersickness Produced by Head Movements in the Oculus Rift
abstract
We compared the cybersickness produced when a virtual environment (VE) was viewed binocularly and monocularly through an Oculus Rift CV1 head-mounted display (HMD). During each exposure to the VE participants made continuous yaw head movements in time with a computer-generated metronome. Across trials we also varied their head movement frequency (0.5 or 1.0 Hz) and motion-to-photon delays (from ∼5 - ∼212 ms). We found that: 1) cybersickness severity increased with added display lag; and 2) monocular viewing appeared to protect against these increases in cybersickness. We conclude that active binocular viewing with this HMD introduced artifacts that increased the likelihood of more severe sickness.
Stephen A. Palmisano, Luke Szalla, Juno Kim
VRST3
2018 The FuzzyLog: A Partially Ordered Shared Log
Joshua Lockerman, Jose M. Faleiro, Juno Kim, Soham Sankaran, Daniel J. Abadi, James Aspnes, Siddhartha Sen 0001, Mahesh Balakrishnan 0001
OSDI3
2018 Effects of head-display lag on presence in the oculus rift
abstract
We measured presence and perceived scene stability in a virtual environment viewed with different head-to-display lag (i.e., system lag) on the Oculus Rift (CV1). System lag was added on top of the measured benchmark system latency (22.3 ms) for our visual scene rendered in OpenGL Shading Language (GLSL). Participants made active head oscillations in pitch at 1.0Hz while viewing displays. We found that perceived scene instability increased and presence decreased when increasing system lag, which we attribute to the effect of multisensory visual-vestibular interactions on the interpretation of the visual information presented.
Juno Kim, Matthew Moroz, Benjamin Arcioni, Stephen A. Palmisano
VRST1
2010 Pilot gaze and glideslope control
abstract
We examined the eye movements of pilots as they carried out simulated aircraft landings under day and night lighting conditions. Our five students and five certified pilots were instructed to quickly achieve and then maintain a constant 3-degree glideslope relative to the runway. However, both groups of pilots were found to make significant glideslope control errors, especially during simulated night approaches. We found that pilot gaze was directed most often toward the runway and to the ground region located immediately in front of the runway, compared to other visual scene features. In general, their gaze was skewed toward the near half of the runway and tended to follow the runway threshold as it moved on the screen. Contrary to expectations, pilot gaze was not consistently directed at the aircraft's simulated aimpoint (i.e., its predicted future touchdown point based on scene motion). However, pilots did tend to fly the aircraft so that this point was aligned with the runway threshold. We conclude that the supplementary out-of-cockpit visual cues available during day landing conditions facilitated glideslope control performance. The available evidence suggests that these supplementary visual cues are acquired through peripheral vision, without the need for active fixation.
Juno Kim, Stephen A. Palmisano, April Ash, Robert S. Allison
ACM Trans. Appl. Percept.1