VLDB 2026 Research / reviewers in the wild / expert
Joonsoo Kim
dblp:02/777
· DBLP profile ↗
19ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-6470-0773ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has made significant strides in scene representation and neural rendering, with intense efforts focused on adapting it for dynamic scenes. Despite delivering remarkable rendering quality and speed, existing methods struggle with storage demands and the representation of complex real-world motions. To address these challenges, we propose MoDec-GS, a memory-efficient Gaussian splatting framework designed to reconstruct novel views in challenging scenarios with complex motions. We introduce Global-to-Local Motion Decomposition (GLMD) to effectively capture dynamic motions in a coarse-to-fine manner. This approach leverages Global Canonical Scaffolds (Global CS) and Local Canonical Scaffolds (Local CS), which extend static Scaffold representation to dynamic video reconstruction. For Global CS, we propose Global Anchor Deformation (GAD) to efficiently represent global dynamics along complex motions, by directly deforming the implicit Scaffold attributes which are anchor position, offset, and local context features. Next, we finely adjust local motions via the Local Gaussian Deformation (LGD) of Local CS explicitly. Additionally, we introduce Temporal Interval Adjustment (TIA) to automatically control the temporal coverage of each Local CS during training, enabling MoDec-GS to find optimal interval assignments based on the specified number of temporal segments. Extensive evaluations demonstrate that MoDec-GS achieves an average 70% reduction in model size over state-of-the-art methods for dynamic 3D Gaussians from real-world dynamic videos while maintaining or even improving rendering quality. Sangwoon Kwak, Joonsoo Kim, Jun Young Jeong, Won-Sik Cheong, Jihyong Oh, Munchurl Kim |
CVPR | 2 |
| 2025 | Fast Bounding Box HierarchyabstractWe introduce a fast algorithm designed for Bounding Box Hierarchy (BBH), a hierarchical tree structure wherein leaf nodes encapsulate sets of original bounding boxes, and non-leaf nodes represent merging operations of their children. A straightforward algorithm employs a bottom-up strategy in a brute force approach to construct the tree, entailing the iterative merging of candidate pairs with the minimum distance until reaching the root node. A pivotal challenge inherent to this brute force paradigm lies in its computational bottleneck, as determining the candidate pair with the minimum distance necessitates a global operation, rendering it highly computationally intensive. Our novel approach strategically circumvents this bottleneck by introducing the computation of approximate minimum distance pairs within local neighborhoods. By using the transformation of 2D bounding boxes into 1D space through Morton coding, the computational cost associated with identifying candidate bounding boxes for merging is significantly diminished, reducing it from O(N2) to O(logN). Compared with brute force approach which has the overall time complexity O(N3), our algorithm is only O(Nlog(N). The acceleration is critical for computation sensitive applications, particularly in embedded systems and mobile devices. Our approach also supports multi-class bounding box inputs, making it particularly useful for computer vision tasks where bounding boxes are inherently associated with class labels. We validate our algorithm on both real world data and simulated data. Zhe Zhu, Joonsoo Kim, Arshita Gupta, Tien C. Bau |
ICIP | 3 |
| 2024 | Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to detect action instances in untrimmed videos using only video-level annotations. Since many existing works optimize WTAL models based on action classification labels, they encounter the task discrepancy problem (i.e., localization-by-classification). To tackle this issue, recent studies have attempted to utilize action category names as auxiliary semantic knowledge through vision-language pre-training (VLP). However, there are still areas where existing research falls short. Previous approaches primarily focused on leveraging textual information from language models but overlooked the alignment of dynamic human action and VLP knowledge in a joint space. Furthermore, the deterministic representation employed in previous studies struggles to capture fine-grained human motions. To address these problems, we propose a novel framework that aligns human action knowledge and VLP knowledge in a probabilistic embedding space. Moreover, we propose intra- and inter-distribution contrastive learning to enhance the probabilistic embedding space based on statistical similarities. Extensive experiments and ablation studies reveal that our method significantly outperforms all previous state-of-the-art methods. Code is available at https://github.com/sejong-rcv/PVLR. Geuntaek Lim, Joonsoo Kim, Yukyung Choi |
ACM Multimedia | 3 |
| 2024 | Torque based Structured Pruning for Deep Neural NetworkabstractStructured pruning is a popular way of convolutional neural network (CNN) acceleration. However, current state of the art pruning techniques require modifications to the network architecture, implementation of complex gradient update rules or repetitive training and long fine-tuning stages. Our novel physics-inspired approach for structured pruning aims to solve these issues. Analogous to ‘Torque’ we apply a force that consolidates the weights of a convolutional layer around a selected pivot point during training. Using the distance-dependency nature of torque, we can encourage high density of weights in filters around this point and increase filter sparsity as we move away. Filters away from the pivot point can be pruned, resulting in a minimum loss of information. We can control the tightness of the weights by varying the hyper-parameters, thus assisting us in creating a more compact network. Our proposed technique is jointly able to perform both filter learning and filter importance sorting. Additionally, our method is easy to implement, requires no change to model architecture and needs very little to no fine-tuning. We show that our approach reaches competitive results with previous state-of-the-art by evaluating popular networks such as VGGNet and ResNet on multiple image classification tasks. Notably, our method can reduce the parameter count of VGGNet by 96% and still maintain the accuracy achieved by the full-size model without any fine-tuning. This makes our method both latency and memory efficient for hardware deployment. Arshita Gupta, Tien C. Bau, Joonsoo Kim, Zhe Zhu, Hrishikesh Garud |
WACV | 3 |
| 2024 | MosaicMVS: Mosaic-Based Omnidirectional Multi-View Stereo for Indoor ScenesabstractWe present MosaicMVS, a novel learning-based depth estimation framework for a mosaic-based omnidirectional multi-view stereo (MVS) camera setup. It uses a regular field of view (FOV) MVS network for an omnidirectional imaging setup with explicit consideration of hypothetical voxel-wise FOV overlaps. The resulting depth predictions are accurate and agree on the omnidirectional multi-view geometry. Unlike existing MVS setups, MosaicMVS camera setup can be easily applied to omnidirectional indoor scenes without having to account for constraints such as intricate epipolar constraints and the distortion of omnidirectional cameras. We validate the effectiveness of our framework on a new challenging indoor dataset in terms of depth estimation, reconstruction, and view synthesis. We also present new evaluation metric to check reconstruction performance using post-processed masks for accurate evaluation without any ground truth depth map or laser-scanned reconstructions. Experimental results show that our framework outperforms the state-of-the-art MVS methods in a large margin in all test scenes. Min-Jung Shin, Woojune Park, Minji Cho, Kyeongbo Kong, Hoseong Son, Joonsoo Kim, Kugjin Yun, Gwangsoon Lee, Suk-Ju Kang |
IEEE Trans. Multim. | 6 |
| 2024 | CMVDE: Consistent Multi-View Video Depth Estimation via Geometric-Temporal Coupling ApproachabstractIn the field of video depth estimation, significant strides have been made with deep learning-based multi-view stereo approaches. However, existing studies struggle to produce consistently accurate depth maps that account for both multi-view geometry and temporal consistency from monocular video contents. To overcome this limitation, we introduce CMVDE, an innovative video depth estimation framework that leverages a multi-view geometric-temporal coupling approach in an end-to-end manner. Our proposed geometric consistency module efficiently generates multi-view geometric features by employing mutual cross-view epipolar attention between adjacent video frames. Additionally, it compresses these features using the novel multi-scale feature compressor, producing an effective input tensor for the subsequent module. Moreover, our framework enhances temporal consistency across consecutive video frames with the temporal consistency module based on convolutional LSTM 1 leveraging previous depth information as geometric guidance. Compared to state-of-the-art models, our approach achieves superior performance in depth quality and consecutive consistency on the ScanNet 2 and 7-Scenes 3 datasets, surpassing previous multi-view video depth estimation methods. Min-Jung Shin, Minji Cho, Joonsoo Kim, Kugjin Yun, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2023 | Efficient-HDRTV: Efficient SDR to HDR Conversion for HDR TVabstractThe existing deep neural network (DNN) based SDR (Standard dynamic range) to HDR (High dynamic range) conversion methods outperform conventional methods, but they are either too large to implement on a device or with quantization artifacts generated on smooth regions on an image. We propose an efficient neural network for the SDR to HDR conversion, namely "Efficient-HDRTV". It consists of two efficient structures GIM (Global Inverse Mapping) and LIM (Local Inverse Mapping). The key features of GIM and LIM use the small series of a basis function with its coefficient function, which are implemented using small number of convolutions, logarithm and exponential functions. They are combined with other convolutional layers so that the entire network can be jointly trained for learning inverse tone, enhanced details and expanded color gamut from SDR to HDR. Thanks to the GIM and LIM, we can keep our network small with good performance. Our experimental results show that Efficient-HDRTV is much lighter but performs better than the state of the arts. Joonsoo Kim, Kamal Jnawali |
ICIP | 1 |
| 2021 | Structured Camera Pose Estimation for Mosaic-Based Omnidirectional ImagingabstractThis paper presents a novel structured camera pose estimation framework for mosaic-based omnidirectional imaging, i.e., producing a wide field of view (FoV) image that covers an entire sphere of the surroundings from a set of regular FoV images. With the effective utilization of geometric priors, the proposed framework exploits an individual image's connected structure while sequentially extracting correspondence between them. In the proposed framework, 2DSfM, a structure from motion method for multi-view images in structured 2D grids, is also proposed. Additionally, we propose a constraint term for rotation vectors in the bundle adjustment process that efficiently incorporates structural priors. We demonstrate our framework on structured omnidirectional image scenes and compare to existing frameworks. The experimental results show that our framework outperforms well-known conventional frameworks regarding both average reprojection error and reconstruction results. Woojune Park, Jung Hee Kim 0001, Suk-Ju Kang, Joonsoo Kim, Kugjin Yun, Won-Sik Cheong |
ISCAS | 4 |
| 2020 | Tri-level optimization-based image rectification for polydioptric cameras
Siyeong Lee, Gwon Hwan An, Joonsoo Kim, Kugjin Yun, Won-Sik Cheong, Suk-Ju Kang |
Signal Process. Image Commun. | 3 |
| 2017 | Spatial pyramid alignment for sparse coding based object classificationabstractThe bag of visual words (BOW) model is widely used for image representation and classification. Spatial pyramid based feature pooling utilizes the BOW model and is the most popular approach to capture the spatial distribution (layout) of local image features. It makes the assumption that the center of an object is aligned with the center of an image, which can lead to misalignment and degradation in performance. In this paper, we propose a method to utilize max pooled features to estimate objects centers and align the spatial pyramid accordingly. We also propose an image representation descriptor robust to misalignments and objects deformations. The experimental results demonstrate that our spatial pyramid alignment method is simple yet efficient in handling misalignments and achieves high object classification accuracy. Joonsoo Kim, Khalid Tahboub, Edward J. Delp |
ICIP | 1 |
| 2016 | Shape matching using a self similar affine invariant descriptorabstractIn this paper we introduce a shape descriptor known as Self Similar Affine Invariant (SSAI) descriptor for shape retrieval. The SSAI descriptor is based on the property that two sets of points are transformed by an affine transform, then subsets of each set of points are also related by the same affine transformation. Also, the SSAI descriptor is insensitive to local shape distortions. We use multiple SSAI descriptors based on different sets of neighbor points to improve shape recognition accuracy. We also describe an efficient image matching method for the multiple SSAI descriptors. Experimental results show that our approach achieves very good performance on two publicly available shape datasets. Joonsoo Kim, He Li 0002, Jiaju Yue, Edward J. Delp |
ICIP | 1 |
| 2015 | Robust local and global shape context for tattoo image matchingabstractTattoos can provide useful information related to criminal gang activity. Law enforcement can use the information embedded in tattoos to identify and track the criminal history of a suspect. For matching processes, tattoo images are difficult to use due to problems such as deformations and weak edge structures. In this paper we describe a tattoo image retrieval and matching system based on a combination of local and global image matching methods to improve matching accuracy. The proposed local shape context combined with SIFT descriptors are used for local features of a tattoo object and global shape is used for overall shape of a tattoo object. The contributions of this paper include the introduction of a multiple different sized-bin polar histograms based local shape context (MHLC) and a global shape descriptor combining the multiple different sized-bin polar histogram and 2D Fourier Transform for robustness of translation, scale, rotation and shape distortions. We also describe robust similarity for local descriptors and a weighted matching method based on local and global descriptors. Our experimental results show that our proposed method performs better than previously published tattoo image retrieval systems. Joonsoo Kim, Albert Parra Pozo, Jiaju Yue, He Li 0002, Edward J. Delp |
ICIP | 1 |
| 2011 | System accuracy estimation of SRAM-based device authenticationabstractIt is known that power-up values of embedded SRAM memory are unique for each individual chip. The uniqueness enables the power-up values to be considered as SRAM fingerprints used to verify device identities, which is a fundamental task in security applications. However, as the SRAM fingerprints are sensitive to environmental changes, there always exists a chance of error during the authentication process. Hence, the accuracy of a device authentication system with the SRAM fingerprints should be carefully estimated and verified in order to be implemented in practice. Consequently, a proper system evaluation method for the SRAM-based device authentication system should be provided. In this paper, we introduce tractable and computationally efficient system evaluation methods, which include novel parametric models for the distributions of matching distances among genuine and imposter devices. In addition, novel algorithms to calculate the confidence intervals of the estimates, which are crucial in system evaluation, are presented. Also, empirical results follow to validate the models and methods. Joonsoo Kim, Joonsoo Lee, Jacob A. Abraham |
ASP-DAC | 1 |
| 2010 | Toward reliable SRAM-based device identificationabstractDue to process variation, power-up values of embedded SRAM memory are unique for individual devices. They are used as SRAM fingerprints to identify integrated-circuits which is fundamental for security applications. The fingerprints, however, are sensitive to environmental changes. Consequently, during the identification process, errors may occur. To overcome this inherent nondeterminism, we provide a systematic approach to designing reliable SRAM-based identification system. We also discuss how to evaluate its system performance. We present a generic score-fusion-based matching recipe to identify devices with high confidence across a wide range of environmental conditions. Joonsoo Kim, Joonsoo Lee, Jacob A. Abraham |
ICCD | 1 |
| 2009 | QUICK: A flexible full-system functional modelabstractIn this paper, we introduce the concept of full-system complete-and-rollback functional simulators that make efficient functional models in functional/timing partitioned simulators. Complete-and-rollback functional simulators can efficiently drive simulators of resolutions ranging from functional-only to cycle-accurate for a wide range of simulated machines. Complete-and-rollback functional models achieve their capabilities by executing instructions to completion, enabling their execution to be highly optimized, but providing rollback capabilities to enable on-the-fly modifications to the functional execution. We also introduce QUICK, an implementation of a full-system complete-and-rollback functional model that supports the x86 and PowerPC ISAs, boots unmodified Windows XP and Linux, and runs unmodified applications such as YouTube on Internet Explorer while fully supporting rollbacks, including across I/O operations. We present various case studies using QUICK and conduct performance analyses to demonstrate its simulation performance. Dam Sunwoo, Joonsoo Kim, Derek Chiou |
ISPASS | 2 |
| 2008 | Parallelizing computer system simulatorsabstractThis paper describes NSF-supported work in parallelized computer system simulators being done in the Electrical and Computer Engineering Department at the University of Texas at Austin. Our work is currently following two paths: (i) the FAST simulation methodology[9, 11, 10] that is capable of simulating complex systems accurately and quickly (currently about 1.2MIPS executing the x86 ISA, modeling an out-of-order superscalar processor and booting Windows XP and Linux) and (ii) the RAMP-White (White)[1, 22] platform that will soon be capable of simulating very large systems of around 1000 cores. We plan to combine the projects to provide fast and accurate simulation of multicore systems. Derek Chiou, Dam Sunwoo, Hari Angepat, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson |
IPDPS | 4 |
| 2007 | The FAST methodology for high-speed SoC/computer simulationabstractThis paper describes the FAST methodology that enables a single FPGA to accelerate the performance of cycle-accurate computer system simulators modeling modem, realistic SoCs, embedded systems and standard desktop/laptop/server computer systems. The methodology partitions a simulator into (i) a functional model that simulates the functionality of the computer system and (ii) a predictive model that predicts performance and other metrics. The partitioning is crafted to map most of the parallel work onto a hardware-based predictive model, eliminating much of the complexity and difficulty of simulating parallel constructs on a sequential platform. FAST conventions and libraries have been designed to make creating, modifying, using and measuring such simulators straightforward. We describe a prototype FAST system: a full-system, RTL-level cycle-accurate-capable computer system simulator that executes the x86 ISA, boots unmodified Linux and executes unmodified x86 applications. The prototype runs two to three orders of magnitude faster than the fastest Intel and AMD RTL-level cycle-accurate x86 software-based simulators and about six to seven times faster than RTL simulation. Derek Chiou, Dam Sunwoo, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson, Zheng Xu 0004 |
ICCAD | 3 |
| 2007 | FPGA-Accelerated Simulation Technologies (FAST): Fast, Full-System, Cycle-Accurate SimulatorsabstractThis paper describes FAST, a novel simulation methodology that can produce simulators that (i) are orders of magnitude faster than comparable simulators, (ii) are cycle- accurate, (Hi) model the entire system running unmodified applications and operating systems, (iv) provide visibility with minimal simulation performance impact and (v) are capable of running current instruction sets such as x86. It achieves its capabilities by partitioning simulators into a speculative functional model component that simulates the instruction set architecture and a timing model component that predicts performance. The speculative functional model enables the simulator to be parallelized, implementing the timing model in FPGA hardware for speed and the functional model using a modified full-system simulators. We currently achieve an average simulation speed of 1.2MIPS running x86 applications on x86 Linux and Windows XP and expect to achieve 10MIPS over time. Such simulators are useful to virtually all computer system simulator users ranging from architects, through RTL designers and verifiers to software developers. Sharing a common simulation/design infrastructure couldfoster better communication between these groups, potentially resulting in better system designs. Derek Chiou, Dam Sunwoo, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson, Jebediah Keefe, Hari Angepat |
MICRO | 3 |
| 2006 | Towards formal probabilistic power-performance design space explorationabstractWe describe a formal probabilistic power-performance design space exploration technique. The technique aims at enabling hierarchical design space exploration based on a fully probabilistic description of power-performance tradeoffs. Probabilistic Pareto sets in power-performance space are proposed as canonical encodings of the power and delay tradeoffs in designs under any source of uncertainty. An algorithm to compute a composite probabilistic power-performance Pareto set for series or parallel connections of circuit blocks is also developed and validated. The algorithm is based on numerical convolution and is suitable for micro-architecture pipeline design exploration in the presence of process variability. Joonsoo Kim, Michael Orshansky |
ACM Great Lakes Symposium on VLSI | 1 |