VLDB 2026 Research / reviewers in the wild / expert
Kwang Moo Yi
dblp:30/5082
· DBLP profile ↗
68ranked-venue papers
9as first author
40since 2021 · last 2025
0000-0001-9036-3822ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 9 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 6 first-author · 33 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NoKSR: Kernel-Free Neural Surface Reconstruction via Point Cloud SerializationabstractWe present a novel approach to large-scale point cloud surface reconstruction by developing an efficient framework that converts an irregular point cloud into a signed distance field (SDF). Our backbone builds upon recent transformer-based architectures (i.e. PointTransformerV3), that serializes the point cloud into a locality-preserving sequence of tokens. We efficiently predict the SDF value at a point by aggregating nearby tokens, where fast approximate neighbors can be retrieved thanks to the serialization. We serialize the point cloud at different levels/scales, and non-linearly aggregate a feature to predict the SDF value. We show that aggregating across multiple scales is critical to overcome the approximations introduced by the serialization (i.e. false negatives in the neighborhood). Our frameworks sets the new state-of-the-art in terms of accuracy and efficiency (better or similar performance with half the latency of the best prior method, coupled with a simpler implementation), particularly on outdoor datasets where sparse-grid methods have shown limited performance. Zhen Li 0052, Shrisudhan Govindarajan, Shaobo Xia, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi |
3DV | 6 |
| 2025 | LSE-NeRF: Learning Sensor Modeling Errors for Deblured Neural Radiance Fields with RGB-Event StereoabstractWe present a method for reconstructing a clear Neural Radiance Field (NeRF) even with fast camera motions. To address blur artifacts, we leverage both (blurry) RGB images and event camera data captured in a binocular configuration. Importantly, when reconstructing our clear NeRF, we consider the camera modeling imperfections that arise from the simple pinhole camera model as learned embeddings for each camera measurement, and further learn a mapper that connects event camera measurements with RGB data. As no previous dataset exists for our binocular setting, we introduce an event camera dataset with captures from a 3D-printed stereo configuration between RGB and event cameras. Empirically, we evaluate our introduced dataset and EVIMOv2 and show that our method leads to improved reconstructions. Our code and dataset are available at https://github.com/ubc-vision/LSENeRF. Wei Zhi Tang, Daniel Rebain, Konstantinos G. Derpanis, Kwang Moo Yi |
3DV | 4 |
| 2025 | HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesabstractTo efficiently adapt large models or to train generative models of neural representations, Hypernetworks have drawn interest. While hypernetworks work well, training them is cumbersome, and often requires ground truth optimized weights for each sample. However, obtaining each of these weights is a training problem of its own—one needs to train, e.g., adaptation weights or even an entire neural field for hypernetworks to regress to. In this work, we propose a method to train hypernetworks, without the need for any per-sample ground truth. Our key idea is to learn a Hypernetwork ‘Field’ and estimate the entire trajectory of network weight training instead of simply its converged state. In other words, we introduce an additional input to the Hypernetwork, the convergence state, which then makes it act as a neural field that models the entire convergence pathway of a task network. A critical benefit in doing so is that the gradient of the estimated weights at any convergence state must then match the gradients of the original task—this constraint alone is sufficient to train the Hypernetwork Field. We demonstrate the effectiveness of our method through the task of personalized image generation and 3D shape reconstruction from images and point clouds, demonstrating competitive results without any per-sample ground truth. Eric Hedlin, Munawar Hayat, Fatih Porikli, Kwang Moo Yi, Shweta Mahajan |
CVPR | 4 |
| 2025 | Radiant Foam: Real-Time Differentiable Ray TracingabstractResearch on differentiable scene representations is consistently moving towards more efficient, real-time models. Recently, this has led to the popularization of splatting methods, which eschew the traditional ray-based rendering of radiance fields in favor of rasterization. This has yielded a significant improvement in rendering speeds due to the efficiency of rasterization algorithms and hardware, but has come at a cost: the approximations that make rasterization efficient also make implementation of light transport phenomena like reflection and refraction much more difficult. We propose a novel scene representation which avoids these approximations, but keeps the efficiency and reconstruction quality of splatting by leveraging a decades-old efficient volumetric mesh ray tracing algorithm which has been largely overlooked in recent computer vision research. The resulting model, which we name Radiant Foam, achieves rendering speed and quality comparable to Gaussian Splatting, without the constraints of rasterization. Unlike ray traced Gaussian models that use hardware ray tracing acceleration, our method requires no special hardware or APIs beyond the standard features of a programmable GPU. Shrisudhan Govindarajan, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi |
ICCV | 3 |
| 2025 | StochasticSplats: Stochastic Rasterization for Sorting-Free 3D Gaussian Splattingabstract3D Gaussian splatting (3DGS) is a popular radiance field method, with many application-specific extensions. Most variants rely on the same core algorithm: depth-sorting of Gaussian splats then rasterizing in primitive order. This ensures correct alpha compositing, but can cause rendering artifacts due to built-in approximations. Moreover, for a fixed representation, sorted rendering offers little control over render cost and visual fidelity. For example, and counter-intuitively, rendering a lower-resolution image is not necessarily faster. In this work, we address the above limitations by combining 3D Gaussian splatting with stochastic rasterization. Concretely, we leverage an unbiased Monte Carlo estimator of the volume rendering equation. This removes the need for sorting, and allows for accurate 3D blending of overlapping Gaussians. The number of Monte Carlo samples further imbues 3DGS with a way to trade off computation time and quality. We implement our method using OpenGL shaders, enabling efficient rendering on modern GPU hardware. At a reasonable visual quality, our method renders more than four times faster than sorted rasterization. Shakiba Kheradmand, Delio Vicini, Georgios Kopanas, Dmitry Lagun, Kwang Moo Yi, Mark J. Matthews, Andrea Tagliasacchi |
ICCV | 5 |
| 2025 | NESI: Neural Explicit-Shape-Intersection-Based Geometry RepresentationabstractCompressed representations of 3D shapes that are compact, accurate, and can be processed efficiently directly in compressed form, are extremely useful for digital media applications. Recent approaches in this space focus on learned implicit or parametric representations. While implicits are well suited for tasks such as in-out queries, they lack natural 2D parameterization, complicating tasks such as texture or normal mapping. Conversely, parametric representations support the latter tasks but are ill-suited for occupancy queries. We propose a novel learned alternative to these approaches, based on intersections of localized explicit , or height-field , surfaces. Since explicits can be trivially expressed both implicitly and parametrically, NESI directly supports a wider range of processing operations than implicit alternatives, including occupancy queries and parametric access. We represent input shapes using a collection of differently oriented height-field bounded half-spaces combined using volumetric Boolean intersections. We first tightly bound each input using a pair of oppositely oriented height-fields, forming a Double Height-Field (DHF) Hull . We refine this hull by intersecting it with additional localized height-fields (HFs) that capture surface regions in its interior. We minimize the number of HFs necessary to accurately capture each input and compactly encode both the DHF hull and the local HFs as neural functions defined over subdomains of \(\mathbb {R}^2\) . This reduced dimensionality encoding delivers high-quality compact approximations. Given similar parameter count, or storage capacity, NESI significantly reduces approximation error compared to the state-of-the-art, especially at lower parameter counts. Congyi Zhang 0001, Jinfan Yang, Eric Hedlin, Suzuran Takikawa, Nicholas Vining, Kwang Moo Yi, Wenping Wang 0001, Alla Sheffer |
ACM Trans. Graph. | 6 |
| 2024 | Unsupervised Keypoints from Pretrained Diffusion ModelsabstractUnsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures, but performance is yet to match the supervised counterpart, making their practicability questionable. We leverage the emergent knowledge within text-to-image diffusion models, towards more robust unsupervised keypoints. Our core idea is to find text embeddings that would cause the generative model to consistently attend to compact regions in images (i.e. keypoints). To do so, we simply optimize the text embedding such that the cross-attention maps within the denoising network are localized as Gaussians with small standard deviations. We validate our performance on multiple datasets: the CelebA, CUB-200-2011, Tai-Chi-HD, DeepFashion, and Human3.6m datasets. We achieve significantly improved accuracy, sometimes even outperforming supervised ones, particularly for data that is non-aligned and less curated. Our code is publicly available at the project page. Eric Hedlin, Gopal Sharma, Shweta Mahajan, Xingzhe He, Hossam Isack, Abhishek Kar, Helge Rhodin, Andrea Tagliasacchi, Kwang Moo Yi |
CVPR | 9 |
| 2024 | Accelerating Neural Field Training via Soft MiningabstractWe present an approach to accelerate Neural Field training by efficiently selecting sampling locations. While Neural Fields have recently become popular, it is often trained by uniformly sampling the training domain, or through handcrafted heuristics. We show that improved convergence and final training quality can be achieved by a soft mining technique based on importance sampling: rather than either considering or ignoring a pixel completely, we weigh the corresponding loss by a scalar. To implement our idea we use Langevin Monte-Carlo sampling. We show that by doing so, regions with higher error are being selected more frequently, leading to more than 2x improvement in convergence speed. The code and related resources for this study are publicly available at project page. Shakiba Kheradmand, Daniel Rebain, Gopal Sharma, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi |
CVPR | 7 |
| 2024 | ViVid-1-to-3: Novel View Synthesis with Video Diffusion ModelsabstractGenerating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and ren-dering high-quality, spatially consistent new views. While recent methods for view synthesis based on diffusion have shown great progress, achieving consistency among various view estimates and at the same time abiding by the desired camera pose remains a critical problem yet to be solved. In this work, we demonstrate a strikingly simple method, where we utilize a pre-trained video diffusion model to solve this problem. Our key idea is that synthesizing a novel view could be reformulated as synthesizing a video of a cam-era going around the object of interest-a scanning video-which then allows us to leverage the powerful priors that a video diffusion model would have learned. Thus, to perform novel-view synthesis, we create a smooth camera trajectory to the target view that we wish to render, and denoise using both a view-conditioned diffusion model and a video diffusion model. By doing so, we obtain a highly consistent novel view synthesis, outperforming the state of the art. Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, Kwang Moo Yi |
CVPR | 6 |
| 2024 | Prompting Hard or Hardly Prompting: Prompt Inversion for Text-to-Image Diffusion ModelsabstractThe quality of the prompts provided to text-to-image diffusion models determines how faithful the generated content is to the user's intent, often requiring 'prompt engineering ‘. To harness visual concepts from target images without prompt engineering, current approaches largely rely on embedding inversion by optimizing and then mapping them to pseudo-tokens. However, working with such highdimensional vector representations is challenging because they lack semantics and interpretability, and only allow simple vector operations when using them. Instead, this work focuses on inverting the diffusion model to obtain interpretable language prompts directly. The challenge of doing this lies in the fact that the resulting optimization problem is fundamentally discrete and the space of prompts is exponentially large; this makes using standard optimization techniques, such as stochastic gradient descent, difficult. To this end, we utilize a delayed projection scheme to optimize for prompts representative of the vocabulary space in the model. Further, we leverage the findings that different timesteps of the diffusion process cater to different levels of detail in an image. The later, noisy, timesteps of the forward diffusion process correspond to the semantic information, and therefore, prompt inversion in this range provides tokens representative of the image semantics. We show that our approach can identify semantically interpretable and meaningful prompts for a target image which can be used to synthesize diverse images with similar content. We further illustrate the application of the optimized prompts in evolutionary image generation and concept removal. Shweta Mahajan, Tanzila Rahman, Kwang Moo Yi, Leonid Sigal |
CVPR | 3 |
| 2024 | Neural Fields as Distributions: Signal Processing Beyond Euclidean SpaceabstractNeural fields have emerged as a powerful and broadly applicable method for representing signals. However, in contrast to classical discrete digital signal processing, the portfolio of tools to process such representations is still severely limited and restricted to Euclidean domains. In this paper, we address this problem by showing how a prob-abilistic re-interpretation of neural fields can enable their training and inference processes to become “filter-aware”. The formulation we propose not only merges training and filtering in an efficient way, but also generalizes beyond the familiar Euclidean coordinate spaces to the more general set of smooth manifolds and convolutions induced by the actions of Lie groups. We demonstrate how this framework can enable novel integrations of signal processing techniques for neural field applications on both Euclidean do-mains, such as images and audio, as well as non-Euclidean domains, such as rotations and rays. A noteworthy benefit of our method is its applicability. Our method can be sum-marized as primarily a modification of the loss function, and in most cases does not require changes to the network ar-chitecture or the inference process. Daniel Rebain, Soroosh Yazdani, Kwang Moo Yi, Andrea Tagliasacchi |
CVPR | 3 |
| 2024 | BANF: Band-Limited Neural Fields for Levels of Detail ReconstructionabstractLargely due to their implicit nature, neural fields lack a direct mechanism for filtering, as Fourier analysis from discrete signal processing is not directly applicable to these representations. Effective filtering of neural fields is critical to enable level-of-detail processing in downstream applications, and support operations that involve sampling the field on regular grids (e.g. marching cubes). Existing methods that attempt to decompose neural fields in the frequency domain either resort to heuristics or require extensive modifications to the neural field architecture. We show that via a simple modification, one can obtain neural fields that are low-pass filtered, and in turn show how this can be exploited to obtain a frequency decomposition of the entire signal. We demonstrate the validity of our technique by investigating level-of-detail reconstruction, and showing how coarser representations can be computed effectively. Source code is available at https://theialab.github.io/banf Akhmedkhan Ahan Shabanov, Shrisudhan Govindarajan, Cody Reading, Lily Goli, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi |
CVPR | 6 |
| 2024 | Salience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training
Hyesong Choi, Hyejin Park 0003, Kwang Moo Yi, Sungmin Cha, Dongbo Min |
ECCV (78) | 3 |
| 2024 | Lagrangian Hashing for Compressed Neural Field Representations
Shrisudhan Govindarajan, Zeno Sambugaro, Akhmedkhan Shabanov, Towaki Takikawa, Daniel Rebain, Nicola Conci, Kwang Moo Yi, Andrea Tagliasacchi |
ECCV (27) | 8 |
| 2024 | Volumetric Rendering with Baked Quadrature Fields
Gopal Sharma, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi |
ECCV (63) | 3 |
| 2024 | PointNeRF++: A Multi-scale, Point-Based Neural Radiance Field
Eduard Trulls, Yang-Che Tseng, Sneha Sambandam, Gopal Sharma, Andrea Tagliasacchi, Kwang Moo Yi |
ECCV (38) | 7 |
| 2024 | 3D Gaussian Splatting as Markov Chain Monte CarloabstractWhile 3D Gaussian Splatting has recently become popular for neural rendering, current methods rely on carefully engineered cloning and splitting strategies for placing Gaussians, which does not always generalize and may lead to poor-quality renderings. For many real-world scenes this leads to their heavy dependence on good initializations. In this work, we rethink the set of 3D Gaussians as a random sample drawn from an underlying probability distribution describing the physical representation of the scene—in other words, Markov Chain Monte Carlo (MCMC) samples. Under this view, we show that the 3D Gaussian updates can be converted as Stochastic Gradient Langevin Dynamics (SGLD) update by simply introducing noise. We then rewrite the densification and pruning strategies in 3D Gaussian Splatting as simply a deterministic state transition of MCMC samples, removing these heuristics from the framework. To do so, we revise the ‘cloning’ of Gaussians into a relocalization scheme that approximately preserves sample probability. To encourage efficient use of Gaussians, we introduce an L1-regularizer on the Gaussians. On various standard evaluation scenes, we show that our method provides improved rendering quality, easy control over the number of Gaussians, and robustness to initialization. The project website is available at https://3dgs-mcmc.github.io/. Shakiba Kheradmand, Daniel Rebain, Gopal Sharma, Yang-Che Tseng, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi |
NeurIPS | 9 |
| 2023 | Pointersect: Neural Rendering with Cloud-Ray IntersectionabstractWe propose a novel method that renders point clouds as if they are surfaces. The proposed method is differentiable and requires no scene-specific optimization. This unique capability enables, out-of-the-box, surface normal estimation, rendering room-scale point clouds, inverse rendering, and ray tracing with global illumination. Unlike existing work that focuses on converting point clouds to other representations—e.g., surfaces or implicit functions—our key idea is to directly infer the intersection of a light ray with the underlying surface represented by the given point cloud. Specifically, we train a set transformer that, given a small number of local neighbor points along a light ray, provides the intersection point, the surface normal, and the material blending weights, which are used to render the outcome of this light ray. Localizing the problem into small neighborhoods enables us to train a model with only 48 meshes and apply it to unseen point clouds. Our model achieves higher estimation accuracy than state-of-the-art surface reconstruction and point-cloud rendering methods on three test sets. When applied to room-scale point clouds, without any scene-specific optimization, the model achieves competitive quality with the state-of-the-art novel-view rendering methods. Moreover, we demonstrate ability to render and manipulate Lidar-scanned point clouds such as lighting control and object insertion. Jen-Hao Rick Chang, Wei-Yu Chen, Anurag Ranjan, Kwang Moo Yi, Oncel Tuzel |
CVPR | 4 |
| 2023 | BlendFields: Few-Shot Example-Driven Facial ModelingabstractGenerating faithful visualizations of human faces requires capturing both coarse and fine-level details of the face geometry and appearance. Existing methods are either data-driven, requiring an extensive corpus of data not publicly accessible to the research community, or fail to capture fine details because they rely on geometric face models that cannot represent fine-grained details in texture with a mesh discretization and linear deformation designed to model only a coarse face geometry. We introduce a method that bridges this gap by drawing inspiration from traditional computer graphics techniques. Unseen expressions are modeled by blending appearance from a sparse set of extreme poses. This blending is performed by measuring local volumetric changes in those expressions and locally reproducing their appearance whenever a similar expression is performed at test time. We show that our method generalizes to unseen expressions, adding fine-grained effects on top of smooth volumetric deformations of a face, and demonstrate how it generalizes beyond faces. Kacper Kania, Stephan J. Garbin, Andrea Tagliasacchi, Virginia Estellers, Kwang Moo Yi, Julien Valentin, Tomasz Trzcinski, Marek Kowalski |
CVPR | 5 |
| 2023 | FaceLit: Neural 3D Relightable FacesabstractWe propose a generative framework, FaceLit, capable of generating a 3D face that can be rendered at various user-defined lighting conditions and views, learned purely from 2D images in-the-wild without any manual annotation. Unlike existing works that require careful capture setup or human labor, we rely on off-the-shelf pose and illumination estimators. With these estimates, we incorporate the Phong reflectance model in the neural volume rendering framework. Our model learns to generate shape and material properties of a face such that, when rendered according to the natural statistics of pose and illumination, produces photorealistic face images with multiview 3D and illumination consistency. Our method enables photorealistic generation of faces with explicit illumination and view controls on multiple datasets - FFHQ, MetFaces and CelebA-HQ. We show state-of-the-art photorealism among 3D aware GANs on FFHQ dataset achieving an FID score of 3.5. Anurag Ranjan, Kwang Moo Yi, Jen-Hao Rick Chang, Oncel Tuzel |
CVPR | 2 |
| 2023 | Neural Fourier Filter BankabstractWe present a novel method to provide efficient and highly detailed reconstructions. Inspired by wavelets, we learn a neural field that decompose the signal both spatially and frequency-wise. We follow the recent grid-based paradigm for spatial decomposition, but unlike existing work, encourage specific frequencies to be stored in each grid via Fourier features encodings. We then apply a multi-layer perceptron with sine activations, taking these Fourier encoded features in at appropriate layers so that higher-frequency components are accumulated on top of lower-frequency components sequentially, which we sum up to form the final output. We demonstrate that our method outperforms the state of the art regarding model compactness and convergence speed on multiple tasks: 2D image fitting, 3D shape reconstruction, and neural radiance fields. Our code is available at https://github.com/ubc-vision/NFFB. Yuhe Jin, Kwang Moo Yi |
CVPR | 3 |
| 2023 | Unsupervised Semantic Correspondence Using Stable DiffusionabstractText-to-image diffusion models are now capable of generating images that are often indistinguishable from real images. To generate such images, these models must understand the semantics of the objects they are asked to generate. In this work we show that, without any training, one can leverage this semantic knowledge within diffusion models to find semantic correspondences – locations in multiple images that have the same semantic meaning. Specifically, given an image, we optimize the prompt embeddings of these models for maximum attention on the regions of interest. These optimized embeddings capture semantic information about the location, which can then be transferred to another image. By doing so we obtain results on par with the strongly supervised state of the art on the PF-Willow dataset and significantly outperform (20.9% relative for the SPair-71k dataset) any existing weakly- or unsupervised method on PF-Willow, CUB-200 and SPair-71k datasets. Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi |
NeurIPS | 7 |
| 2023 | Efficient Multi-purpose Video Annotation for Fast LabelingabstractWe propose an efficient method to label the frames of a video consistently and quickly. Our method is motivated by welding videos, where molten metal moves continuously, and some frames may be noisy due to the high intensity of arc and spatters that can obscure the desired point to be annotated. The traditional annotation methods which pause the videos and annotate each frame individually are time-consuming and may result in inconsistent labeling between different annotators for noisy frames of a video. Our new proposed method benefits from the fact that video sequences are contiguous. We track the location of the cursor while the video is being played, with the user controlling the playback speed, including the playback direction. We process the recorded cursor locations to convert them to the final annotation. Compared to frame-to-frame annotation methods, we show that our proposed interface speeds up the annotating process by 21 times while maintaining consistency of the labels. Mobina Mobaraki, Soodeh Ahani, Kwang Moo Yi, Mahyar Asadi, Klaske van Heusden, Guy Albert Dumont |
VCIP | 3 |
| 2023 | NeuralBF: Neural Bilateral Filtering for Top-down Instance Segmentation on Point CloudsabstractWe introduce a method for instance proposal generation for 3D point clouds. Existing techniques typically directly regress proposals in a single feed-forward step, leading to inaccurate estimation. We show that this serves as a critical bottleneck, and propose a method based on iterative bilateral filtering with learned kernels. Following the spirit of bilateral filtering, we consider both the deep feature embeddings of each point, as well as their locations in the 3D space. We show via synthetic experiments that our method brings drastic improvements when generating instance proposals for a given point of interest. We further validate our method on the challenging ScanNet benchmark, achieving the best instance segmentation performance amongst the sub-category of top-down methods. Daniel Rebain, Renjie Liao 0001, Vladimir Tankovich, Soroosh Yazdani, Kwang Moo Yi, Andrea Tagliasacchi |
WACV | 6 |
| 2022 | Bootstrapping Human Optical Flow and Pose
Aritro Roy Arko, James J. Little, Kwang Moo Yi |
BMVC | 3 |
| 2022 | Kubric: A scalable dataset generatorabstractData is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification. Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J. Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam H. Laradji, Hsueh-Ti Derek Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, A. Cengiz Öztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Tianhao Wu 0003, Kwang Moo Yi, Fangcheng Zhong, Andrea Tagliasacchi |
CVPR | 32 |
| 2022 | CoNeRF: Controllable Neural Radiance FieldsabstractWe extend neural 3D representations to allow for intu-itive and interpretable user control beyond novel view ren-dering (i. e. camera control). We allow the user to annotate which part of the scene one wishes to control with just a small number of mask annotations in the training images. Our key idea is to treat the attributes as latent variables that are regressed by the neural network given the scene en-coding. This leads to afew-shot learning framework, where attributes are discovered automatically by the framework, when annotations are not provided. We apply our method to various scenes with different types of controllable attributes (e.g. expression control on human faces, or state control in movement of inanimate objects). Overall, we demonstrate, to the best of our knowledge, for the first time novel view and novel attribute re-rendering of scenes from a single video. Kacper Kania, Kwang Moo Yi, Marek Kowalski, Tomasz Trzcinski, Andrea Tagliasacchi |
CVPR | 2 |
| 2022 | LOLNeRF: Learn from One LookabstractWe present a method for learning a generative 3D model based on neural radiance fields, trained solely from data with only single views of each object. While generating realistic images is no longer a difficult task, producing the corresponding 3D structure such that they can be rendered from different views is non-trivial. We show that, unlike existing methods, one does not need multi-view data to achieve this goal. Specifically, we show that by reconstructing many images aligned to an approximate canonical pose with a single network conditioned on a shared latent space, you can learn a space of radiance fields that models shape and appearance for a class of objects. We demonstrate this by training models to reconstruct object categories using datasets that contain only one view of each subject without depth or geometry information. Our experiments show that we achieve state-of-the-art results in novel view synthesis and high-quality results for monocular depth prediction. https://lolnerf.github.io. Daniel Rebain, Mark J. Matthews, Kwang Moo Yi, Dmitry Lagun, Andrea Tagliasacchi |
CVPR | 3 |
| 2022 | Layered Controllable Video Generation
Yuhe Jin, Kwang Moo Yi, Leonid Sigal |
ECCV (16) | 3 |
| 2022 | NeuMan: Neural Human Radiance Field from a Single Video
Wei Jiang 0034, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel, Anurag Ranjan |
ECCV (32) | 2 |
| 2022 | TUSK: Task-Agnostic Unsupervised KeypointsabstractExisting unsupervised methods for keypoint learning rely heavily on the assumption that a specific keypoint type (e.g. elbow, digit, abstract geometric shape) appears only once in an image. This greatly limits their applicability, as each instance must be isolated before applying the method—an issue that is never discussed or evaluated. We thus propose a novel method to learn Task-agnostic, UnSupervised Keypoints (TUSK) which can deal with multiple instances. To achieve this, instead of the commonly-used strategy of detecting multiple heatmaps, each dedicated to a specific keypoint type, we use a single heatmap for detection, and enable unsupervised learning of keypoint types through clustering. Specifically, we encode semantics into the keypoints by teaching them to reconstruct images from a sparse set of keypoints and their descriptors, where the descriptors are forced to form distinct clusters in feature space around learned prototypes. This makes our approach amenable to a wider range of tasks than any previous unsupervised keypoint method: we show experiments on multiple-instance detection and classification, object discovery, and landmark detection—all unsupervised—with performance on par with the state of the art, while also being able to deal with multiple instances. Yuhe Jin, Jan Hosang, Eduard Trulls, Kwang Moo Yi |
NeurIPS | 5 |
| 2022 | Novel-View Synthesis of Human Tourist PhotosabstractWe present a novel framework for performing novel-view synthesis on human tourist photos. Given a tourist photo from a known scene, we reconstruct the photo in 3D space through modeling the human and the background independently. We generate a deep buffer from a novel viewpoint of the reconstruction and utilize a deep network to translate the buffer into a photo-realistic rendering of the novel view. We additionally present a method to relight the renderings, allowing for relighting of both human and background to match either the provided input image or any other. The key contributions of our paper are: 1) a framework for performing novel view synthesis on human tourist photos, 2) an appearance transfer method for relighting of humans to match synthesized backgrounds, and 3) a method for estimating lighting properties from a single human photo. We demonstrate the proposed framework on photos from two different scenes of various tourists. Jonathan Freer, Kwang Moo Yi, Wei Jiang 0034, Jongwon Choi 0002, Hyung Jin Chang |
WACV | 2 |
| 2022 | Repurposing existing deep networks for caption and aesthetic-guided image cropping
Nora Horanyi, Kedi Xia, Kwang Moo Yi, Abhishake Kumar Bojja, Ales Leonardis, Hyung Jin Chang |
Pattern Recognit. | 3 |
| 2021 | MIST: Multiple Instance Spatial TransformerabstractWe propose a deep network that can be trained to tackle image reconstruction and classification problems that involve detection of multiple object instances, without any supervision regarding their whereabouts. The network learns to extract the most significant K patches, and feeds these patches to a task-specific network – e.g., auto-encoder or classifier – to solve a domain specific problem. The challenge in training such a network is the non-differentiable top-K selection process. To address this issue, we lift the training optimization problem by treating the result of top-K selection as a slack variable, resulting in a simple, yet effective, multi-stage training. Our method is able to learn to detect recurring structures in the training dataset by learning to reconstruct images. It can also learn to localize structures when only knowledge on the occurrence of the object is provided, and in doing so it outperforms the state-of-the-art. Code available at https://github.com/ubc-vision/mist Baptiste Angles, Yuhe Jin, Simon Kornblith, Andrea Tagliasacchi, Kwang Moo Yi |
CVPR | 5 |
| 2021 | VaB-AL: Incorporating Class Imbalance and Difficulty With Variational Bayes for Active LearningabstractActive Learning for discriminative models has largely been studied with the focus on individual samples, with less emphasis on how classes are distributed or which classes are hard to deal with. In this work, we show that this is harmful. We propose a method based on the Bayes’ rule, that can naturally incorporate class imbalance into the Active Learning framework. We derive that three terms should be considered together when estimating the probability of a classifier making a mistake for a given sample; i) probability of mislabelling a class, ii) likelihood of the data given a predicted class, and iii) the prior probability on the abundance of a predicted class. Implementing these terms requires a generative model and an intractable likelihood estimation. Therefore, we train a Variational Auto Encoder (VAE) for this purpose. To further tie the VAE with the classifier and facilitate VAE training, we use the classifiers’ deep feature representations as input to the VAE. By considering all three probabilities, among them, especially the data imbalance, we can substantially improve the potential of existing methods under limited data budget. We show that our method can be applied to classification tasks on multiple different datasets – including one that is a real-world dataset with heavy data imbalance – significantly outperforming the state of the art. Jongwon Choi 0002, Kwang Moo Yi, Jinho Choo, Byoungjip Kim, Jin-Yeop Chang, Youngjune Gwon, Hyung Jin Chang |
CVPR | 2 |
| 2021 | DeRF: Decomposed Radiance FieldsabstractWith the advent of Neural Radiance Fields (NeRF), neural networks can now render novel views of a 3D scene with quality that fools the human eye. Yet, generating these images is very computationally intensive, limiting their applicability in practical scenarios. In this paper, we propose a technique based on spatial decomposition capable of mitigating this issue. Our key observation is that there are diminishing returns in employing larger (deeper and/or wider) networks. Hence, we propose to spatially decompose a scene and dedicate smaller networks for each decomposed part. When working together, these networks can render the whole scene. This allows us near-constant inference time regardless of the number of decomposed parts. Moreover, we show that a Voronoi spatial decomposition is preferable for this purpose, as it is provably compatible with the Painter’s Algorithm for efficient and GPU-friendly rendering. Our experiments show that for real-world scenes, our method provides up to 3× more efficient inference than NeRF (with the same rendering quality), or an improvement of up to 1.0 dB in PSNR (for the same inference cost). Daniel Rebain, Wei Jiang 0034, Soroosh Yazdani, Ke Li 0011, Kwang Moo Yi, Andrea Tagliasacchi |
CVPR | 5 |
| 2021 | COTR: Correspondence Transformer for Matching Across ImagesabstractWe propose a novel framework for finding correspondences in images based on a deep neural network that, given two images and a query point in one of them, finds its correspondence in the other. By doing so, one has the option to query only the points of interest and retrieve sparse correspondences, or to query all points in an image and obtain dense mappings. Importantly, in order to capture both local and global priors, and to let our model relate between image regions using the most relevant among said priors, we realize our network using a transformer. At inference time, we apply our correspondence network by recursively zooming in around the estimates, yielding a multiscale pipeline able to provide highly-accurate correspondences. Our method significantly outperforms the state of the art on both sparse and dense correspondence problems on multiple datasets and tasks, ranging from wide-baseline stereo to optical flow, without any retraining for a specific dataset. We commit to releasing data, code, and all the tools necessary to train from scratch and ensure reproducibility. Wei Jiang 0034, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, Kwang Moo Yi |
ICCV | 5 |
| 2021 | Canonical Capsules: Self-Supervised Capsules in Canonical PoseabstractWe propose a self-supervised capsule architecture for 3D point clouds. We compute capsule decompositions of objects through permutation-equivariant attention, and self-supervise the process by training with pairs of randomly rotated objects. Our key idea is to aggregate the attention masks into semantic keypoints, and use these to supervise a decomposition that satisfies the capsule invariance/equivariance properties. This not only enables the training of a semantically consistent decomposition, but also allows us to learn a canonicalization operation that enables object-centric reasoning. To train our neural network we require neither classification labels nor manually-aligned training datasets. Yet, by learning an object-centric representation in a self-supervised manner, our method outperforms the state-of-the-art on 3D point cloud reconstruction, canonicalization, and unsupervised classification. Andrea Tagliasacchi, Boyang Deng, Sara Sabour, Soroosh Yazdani, Geoffrey E. Hinton, Kwang Moo Yi |
NeurIPS | 7 |
| 2021 | Image Matching Across Wide Baselines: From Paper to Practice
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk, Jiri Matas, Pascal Fua, Kwang Moo Yi, Eduard Trulls |
Int. J. Comput. Vis. | 6 |
| 2021 | Eigendecomposition-Free Training of Deep Networks for Linear Least-Square ProblemsabstractMany classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be tackled by solving a linear least-square problem, which can be done by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning frameworks would allow us to explicitly encode known notions of geometry, instead of having the network implicitly learn them from data. However, performing eigendecomposition within a network requires the ability to differentiate this operation. While theoretically doable, this introduces numerical instability in the optimization process in practice. In this paper, we introduce an eigendecomposition-free approach to training a deep network whose loss depends on the eigenvector corresponding to a zero eigenvalue of a matrix predicted by the network. We demonstrate that our approach is much more robust than explicit differentiation of the eigendecomposition using two general tasks, outlier rejection and denoising, with several practical examples including wide-baseline stereo, the perspective-n-point problem, and ellipse fitting. Empirically, our method has better convergence properties and yields state-of-the-art results. Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | ACNe: Attentive Context Normalization for Robust Permutation-Equivariant LearningabstractMany problems in computer vision require dealing with sparse, unordered data in the form of point clouds. Permutation-equivariant networks have become a popular solution - they operate on individual data points with simple perceptrons and extract contextual information with global pooling. This can be achieved with a simple normalization of the feature maps, a global operation that is unaffected by the order. In this paper, we propose Attentive Context Normalization (ACN), a simple yet effective technique to build permutation-equivariant networks robust to outliers. Specifically, we show how to normalize the feature maps with weights that are estimated within the network, excluding outliers from this normalization. We use this mechanism to leverage two types of attention: local and global - by combining them, our method is able to find the essential data points in high-dimensional space in order to solve a given task. We demonstrate through extensive experiments that our approach, which we call Attentive Context Networks (ACNe), provides a significant leap in performance compared to the state-of-the-art on camera pose estimation, robust fitting, and point cloud classification under noise and outliers. Source code: https://github.com/vcg-uvic/acne. Wei Jiang 0034, Eduard Trulls, Andrea Tagliasacchi, Kwang Moo Yi |
CVPR | 5 |
| 2020 | Improving Music Transcription by Pre-Stacking A U-NetabstractWe propose to pre-stack a U-Net as a way of improving the polyphonic music transcription performance of various baseline Convolutional Neural Networks (CNNS). The U-Net, a network architecture based on skip-connections between layers acts as a transformation network followed by a transcription network. Notably, we do not introduce any additional loss terms specific to the transformation network, but instead, jointly train the entire combined model with the original loss function that was designed for the back-end transcription network. We argue that this U-Net network transforms the input signal into a representation that is more effective for transcription, and thus enables the observed improvements in accuracy. We empirically confirm with several experiments using the MusicNet dataset, that the proposed configuration consistently improves the accuracy of transcription networks. This enhancement cannot be achieved by simply introducing more neurons or more layers to the baseline CNNs. Moreover, we show that using the proposed architecture we can go beyond general music transcription and perform transcription in an instrument-specific fashion. By doing so, the original general transcription performance is also increased. Fabrizio Pedersoli, George Tzanetakis, Kwang Moo Yi |
ICASSP | 3 |
| 2020 | Optimizing Through Learned Errors for Accurate Sports Field RegistrationabstractWe propose an optimization-based framework to register sports field templates onto broadcast videos. For accurate registration we go beyond the prevalent feed-forward paradigm. Instead, we propose to train a deep network that regresses the registration error, and then register images by finding the registration parameters that minimize the regressed error. We demonstrate the effectiveness of our method by applying it to real-world sports broadcast videos, outperforming the state of the art. We further apply our method on a synthetic toy example and demonstrate that our method brings significant gains even when the problem is simplified and unlimited training data is available1. Wei Jiang 0034, Juan Camilo Gamboa Higuera, Baptiste Angles, Mehrsan Javan Roshtkhari, Kwang Moo Yi |
WACV | 6 |
| 2019 | Beyond Cartesian Representations for Local DescriptorsabstractThe dominant approach for learning local patch descriptors relies on small image regions whose scale must be properly estimated a priori by a keypoint detector. In other words, if two patches are not in correspondence, their descriptors will not match. A strategy often used to alleviate this problem is to “pool” the pixel-wise features over log-polar regions, rather than regularly spaced ones. By contrast, we propose to extract the “support region” directly with a log-polar sampling scheme. We show that this provides us with a better representation by simultaneously oversampling the immediate neighbourhood of the point and undersampling regions far away from it. We demonstrate that this representation is particularly amenable to learning descriptors with deep networks. Our models can match descriptors across a much wider range of scales than was possible before, and also leverage much larger support regions without suffering from occlusions. We report state-of-the-art results on three different datasets. Patrick Ebel 0002, Eduard Trulls, Kwang Moo Yi, Pascal Fua, Anastasiia Mishchuk |
ICCV | 3 |
| 2019 | Linearized Multi-Sampling for Differentiable Image TransformationabstractWe propose a novel image sampling method for differentiable image transformation in deep neural networks. The sampling schemes currently used in deep learning, such as Spatial Transformer Networks, rely on bilinear interpolation, which performs poorly under severe scale changes, and more importantly, results in poor gradient propagation. This is due to their strict reliance on direct neighbors. Instead, we propose to generate random auxiliary samples in the vicinity of each pixel in the sampled image, and create a linear approximation with their intensity values. We then use this approximation as a differentiable formula for the transformed image. We demonstrate that our approach produces more representative gradients with a wider basin of convergence for image alignment, which leads to considerable performance improvements when training networks for registration and classification tasks. This is not only true under large downsampling, but also when there are no scale changes. We compare our approach with multi-scale sampling and show that we outperform it. We then demonstrate that our improvements to the sampler are compatible with other tangential improvements to Spatial Transformer Networks and that it further improves their performance. Wei Jiang 0034, Andrea Tagliasacchi, Eduard Trulls, Kwang Moo Yi |
ICCV | 5 |
| 2018 | Learning to Find Good CorrespondencesabstractWe develop a deep architecture to learn to find good correspondences for wide-baseline stereo. Given a set of putative sparse matches and the camera intrinsics, we train our network in an end-to-end fashion to label the correspondences as inliers or outliers, while simultaneously using them to recover the relative pose, as encoded by the essential matrix. Our architecture is based on a multi-layer perceptron operating on pixel coordinates rather than directly on the image, and is thus simple and small. We introduce a novel normalization technique, called Context Normalization, which allows us to process each data point separately while embedding global information in it, and also makes the network invariant to the order of the correspondences. Our experiments on multiple challenging datasets demonstrate that our method is able to drastically improve the state of the art with little training data. Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, Pascal Fua |
CVPR | 1 |
| 2018 | Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses
Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
ECCV (5) | 2 |
| 2018 | LF-Net: Learning Local Features from ImagesabstractWe present a novel deep architecture and a training strategy to learn a local feature pipeline from scratch, using collections of images without the need for human supervision. To do so we exploit depth and relative camera pose cues to create a virtual target that the network should achieve on one image, provided the outputs of the network for the other image. While this process is inherently non-differentiable, we show that we can optimize the network in a two-branch setup by confining it to one branch, while preserving differentiability in the other. We train our method on both indoor and outdoor datasets, with depth data from 3D sensors for the former, and depth estimates from an off-the-shelf Structure-from-Motion solution for the latter. Our models outperform the state of the art on sparse feature matching on both datasets, while running at 60+ fps for QVGA images. Yuki Ono, Eduard Trulls, Pascal Fua, Kwang Moo Yi |
NeurIPS | 4 |
| 2018 | Robust 3D Object Tracking from Monocular Images Using Stable PartsabstractWe present an algorithm for estimating the pose of a rigid object in real-time under challenging conditions. Our method effectively handles poorly textured objects in cluttered, changing environments, even when their appearance is corrupted by large occlusions, and it relies on grayscale images to handle metallic environments on which depth cameras would fail. As a result, our method is suitable for practical Augmented Reality applications including industrial environments. At the core of our approach is a novel representation for the 3D pose of object parts: We predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible; when several parts are visible, we can easily combine them to compute a better pose of the object; the 3D pose we obtain is usually very accurate, even when only few parts are visible. We show how to use this representation in a robust 3D tracking framework. In addition to extensive comparisons with the state-of-the-art, we demonstrate our method on a practical Augmented Reality application for maintenance assistance in the ATLAS particle detector at CERN. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Learning Lightprobes for Mixed Reality IlluminationabstractThis paper presents the first photometric registration pipeline for Mixed Reality based on high quality illumination estimation using convolutional neural networks (CNNs). For easy adaptation and deployment of the system, we train the CNNs using purely synthetic images and apply them to real image data. To keep the pipeline accurate and efficient, we propose to fuse the light estimation results from multiple CNN instances and show an approach for caching estimates over time. For optimal performance, we furthermore explore multiple strategies for the CNN training. Experimental results show that the proposed method yields highly accurate estimates for photo-realistic augmentations. David Mandl, Kwang Moo Yi, Peter Mohr, Peter M. Roth, Pascal Fua, Vincent Lepetit, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 2 |
| 2016 | Learning to Assign Orientations to Feature PointsabstractWe show how to train a Convolutional Neural Network to assign a canonical orientation to feature points given an image patch centered on the feature point. Our method improves feature point matching upon the state-of-the art and can be used in conjunction with any existing rotation sensitive descriptors. To avoid the tedious and almost impossible task of finding a target orientation to learn, we propose to use Siamese networks which implicitly find the optimal orientations during training. We also propose a new type of activation function for Neural Networks that generalizes the popular ReLU, maxout, and PReLU activation functions. This novel activation performs better for our task. We validate the effectiveness of our method extensively with four existing datasets, including two non-planar datasets, as well as our own dataset. We show that we outperform the state-of-the-art without the need of retraining for each dataset. Kwang Moo Yi, Yannick Verdie, Pascal Fua, Vincent Lepetit |
CVPR | 1 |
| 2016 | LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua |
ECCV (6) | 1 |
| 2015 | TILDE: A Temporally Invariant Learned DEtectorabstractWe introduce a learning-based approach to detect repeatable keypoints under drastic imaging changes of weather and lighting conditions to which state-of-the-art keypoint detectors are surprisingly sensitive. We first identify good keypoint candidates in multiple training images taken from the same viewpoint. We then train a regressor to predict a score map whose maxima are those points so that they can be found by simple non-maximum suppression. As there are no standard datasets to test the influence of these kinds of changes, we created our own, which we will make publicly available. We will show that our method significantly outperforms the state-of-the-art methods in such challenging conditions, while still achieving state-of-the-art performance on untrained standard datasets. Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
CVPR | 2 |
| 2015 | A Novel Representation of Parts for Accurate 3D Object Detection and Tracking in Monocular ImagesabstractWe present a method that estimates in real-time and under challenging conditions the 3D pose of a known object. Our method relies only on grayscale images since depth cameras fail on metallic objects, it can handle poorly textured objects, and cluttered, changing environments, the pose it predicts degrades gracefully in presence of large occlusions. As a result, by contrast with the state-of-the-art, our method is suitable for practical Augmented Reality applications even in industrial environments. To be robust to occlusions, we first learn to detect some parts of the target object. Our key idea is to then predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible, when several parts are visible, we can combine them easily to compute a better pose of the object, the 3D pose we obtain is usually very accurate, even when only few parts are visible. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ICCV | 4 |
| 2015 | Category Attentional Search for Fast Object Detection by Mimicking Human Visual PerceptionabstractIn this paper, we propose a novel selective search method to speed up the object detection via category-based attention scheme. The proposed attentional searching strategy is designed to focus on a small set of selected regions where the object category is expected to exist. The selected regions are estimated by mimicking three properties of the attentional scheme of human visual perception: spotlighting interest regions with low-level saliency (saliency attention), focusing on distinctive features for an object category (feature attention), and estimating potential object position by following human gaze path (gaze attention). Also, the time complexity of each attentional scheme is implemented to be low so that it can hardly affect the computational time. To validate the performance of our method, experiments were conducted on the challenging PASCAL VOC dataset. Experimental results show that our method efficiently generates a small number of candidate boxes for object detection (less than 10ms=image), and the combined object detection system achieves more than 2 times faster performance than the baseline with comparable average precision. Hawook Jeong, Sangdoo Yun, Kwang Moo Yi, Jin Young Choi 0002 |
WACV | 3 |
| 2015 | Visual tracking of non-rigid objects with partial occlusion through elastic structure of local patches and hierarchical diffusion
Kwang Moo Yi, Hawook Jeong, Soo Wan Kim, Shimin Yin, Songhwai Oh, Jin Young Choi 0002 |
Image Vis. Comput. | 1 |
| 2015 | Visual tracking in complex scenes through pixel-wise tri-modeling
Kwang Moo Yi, Hawook Jeong, Byeongju Lee, Jin Young Choi 0002 |
Mach. Vis. Appl. | 1 |
| 2015 | Erratum to: Visual tracking in complex scenes through pixel-wise tri-modeling
Kwang Moo Yi, Hawook Jeong, Byeongju Lee, Jin Young Choi 0002 |
Mach. Vis. Appl. | 1 |
| 2014 | Scale preserving PTZ tracking with size estimation using tilt sensory dataabstractWe propose an pan-tilt-zoom (PTZ) tracking method to keep the target object at the center of the image with a predefined size in image. For this purpose, we develop an efficient method for object size estimation using only the tilt sensory data. First we investigate the relational function between the object size in image and the tilt sensory data, and find its approximated function to estimate the object size. Using the approximated function, the object size in each frame is estimated recursively from the known initial size. The proposed method requires smaller amount of computation and shows higher accuracy to keep the size preserving PTZ tracking than the existing state-of-the-art method. Byeong Ju Lee, Kwang Moo Yi |
AVSS | 2 |
| 2014 | Motion Interaction Field for Accident Detection in Traffic Surveillance VideoabstractThis paper presents a novel method for modeling of interaction among multiple moving objects to detect traffic accidents. The proposed method to model object interactions is motivated by the motion of water waves responding to moving objects on water surface. The shape of the water surface is modeled in a field form using Gaussian kernels, which is referred to as the Motion Interaction Field (MIF). By utilizing the symmetric properties of the MIF, we detect and localize traffic accidents without solving complex vehicle tracking problems. Experimental results show that our method outperforms the existing works in detecting and localizing traffic accidents. Kimin Yun, Hawook Jeong, Kwang Moo Yi, Soo Wan Kim, Jin Young Choi 0002 |
ICPR | 3 |
| 2014 | Tracking texture-less, shiny objects with descriptor fieldsabstractOur demo demonstrates the method we published at CVPR this year for tracking specular and poorly textured objects, and lets the visitors experiment with it and with their own patterns. Our approach only requires a standard monocular camera (no need for a depth sensor), and can be easily integrated within existing systems to improve their robustness and accuracy. Alberto Crivellaro, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ISMAR | 3 |
| 2014 | Two-stage online inference model for traffic pattern analysis and anomaly detection
Hawook Jeong, Young Joon Yoo, Kwang Moo Yi, Jin Young Choi 0002 |
Mach. Vis. Appl. | 3 |
| 2013 | Action Chart: A Representation for Efficient Recognition of Complex ActivityabstractIn this paper we propose an efficient method for the recognition of long and complex action streams. First, we design a new motion feature flow descriptor by composing low-level local features. Then a new data embedding method is developed in order to represent the motion flow as an one-dimensional sequence, whilst preserving useful motion information for recognition. Finally attentional motion spots (AMSs) are defined to automatically detect meaningful motion changes from the embedded one-dimensional sequence. An unsupervised learning strategy based on expectation maximization and a weighted Gaussian mixture model is then applied to the AMSs for each action class, resulting in an action representation which we refer to as Action Chart. The Action Chart is then used efficiently for recognizing each action class. Through comparison with the state-of-the-art methods, experimental results show that the Action Chart gives promising recognition performance with low computational load and can be used for abstracting long video sequences. Hyung Jin Chang, Jungchan Cho, Songhwai Oh, Kwang Moo Yi, Jin Young Choi 0002 |
BMVC | 5 |
| 2013 | Initialization-Insensitive Visual Tracking through Voting with Salient Local FeaturesabstractIn this paper we propose an object tracking method in case of inaccurate initializations. To track objects accurately in such situation, the proposed method uses "motion saliency" and "descriptor saliency" of local features and performs tracking based on generalized Hough transform (GHT). The proposed motion saliency of a local feature emphasizes features having distinctive motions, compared to the motions which are not from the target object. The descriptor saliency emphasizes features which are likely to be of the object in terms of its feature descriptors. Through these saliencies, the proposed method tries to "learn and find" the target object rather than looking for what was given at initialization, giving robust results even with inaccurate initializations. Also, our tracking result is obtained by combining the results of each local feature of the target and the surroundings with GHT voting, thus is robust against severe occlusions as well. The proposed method is compared against nine other methods, with nine image sequences, and hundred random initializations. The experimental results show that our method outperforms all other compared methods. Kwang Moo Yi, Hawook Jeong, Byeongho Heo, Hyung Jin Chang, Jin Young Choi 0002 |
ICCV | 1 |
| 2013 | Detection of moving objects with a moving camera using non-panoramic background model
Soo Wan Kim, Kimin Yun, Kwang Moo Yi, Sun Jung Kim, Jin Young Choi 0002 |
Mach. Vis. Appl. | 3 |
| 2010 | Recovery Video Stabilization Using MRF-MAP OptimizationabstractIn this paper, we propose a novel approach for video stabilization using Markov random field (MRF) modeling and maximum a posteriori (MAP) optimization. We build an MRF model describing a sequence of unstable images and find joint pixel matchings over all image sequences with MAP optimization via Gibbs sampling. The resulting displacements of matched pixels in consecutive frames indicate the camera motion between frames and can be used to remove the camera motion to stabilize image sequences. The proposed method shows robust performance even when a scene has moving foreground objects and brings more accurate stabilization results. The performance of our algorithm is evaluated on outdoor scenes. Soo Wan Kim, Kwang Moo Yi, Songhwai Oh, Jin Young Choi 0002 |
ICPR | 2 |
| 2009 | Orientation and Scale Invariant Kernel-Based Object Tracking with Probabilistic Emphasizing
Kwang Moo Yi, Soo Wan Kim, Jin Young Choi 0002 |
ACCV (2) | 1 |
| 2008 | Orientation and scale invariant mean shift using object mask-based kernelabstractIn this paper, we propose a new method for object tracking based on mean shift algorithm using a kernel which has the shape of the target object, and with probabilistic estimation of the orientation change and scale adaptation. The proposed method uses an object mask to construct a kernel which has the shape of the actual object for tracking. Orientation is adjusted using probabilistic estimation of orientation and scale is adapted using a newly proposed descriptor for scale. Tests results show that the proposed method is robust to background clutter and tracks objects very accurately. Kwang Moo Yi, Ho Seok Ahn, Jin Young Choi 0002 |
ICPR | 1 |