EDBT 2026 Demo / reviewers in the wild / expert
Hansung Kim 0001
dblp:44/4871-1
· DBLP profile ↗
46ranked-venue papers
20as first author
10since 2021 · last 2026
0000-0003-4907-0491ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 18 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion GenerationabstractRecent advances in transformer-based text-to-motion generation have significantly improved motion quality. However, achieving both real-time performance and long-horizon scalability remains an open challenge. In this paper, we present MOGO (Motion Generation with One-pass), a novel autoregressive framework for efficient and scalable 3D human motion generation. MOGO consists of two key components. First, we introduce MoSA-VQ, a motion scale-adaptive residual vector quantization module that hierarchically discretizes motion sequences through learnable scaling parameters, enabling dynamic allocation of representation capacity and producing compact yet expressive multi-level representations. Second, we design the RQHC-Transformer, a residual quantized hierarchical causal transformer that decodes motion tokens in a single forward pass. Each transformer block aligns with one quantization level, allowing hierarchical abstraction and temporally coherent generation with strong semantic flow. Compared to diffusion- and LLM-based approaches, MOGO achieves lower inference latency while preserving high motion fidelity. Moreover, its hierarchical latent design enables seamless and controllable infinite-length motion generation, with stable transitions and the ability to adaptively incorporate updated control signals at arbitrary points in time. To further enhance generalization and interpretability, we introduce Textual Condition Alignment (TCA), which leverages large language models with Chain-of-Thought reasoning to bridge the gap between real-world prompts and training data. TCA not only improves zero-shot performance on unseen datasets but also enriches motion comprehension for in-distribution prompts through explicit intent decomposition. Extensive experiments on HumanML3D, KIT-ML, and the unseen CMP dataset demonstrate that MOGO outperforms prior methods in generation quality, inference efficiency, and temporal scalability. Tengjiao Sun, Pengcheng Fang, Xiaohao Cai, Hansung Kim 0001 |
AAAI | 5 |
| 2025 | MatSpectNet: Material Segmentation Network with Domain-Aware and Physically-Constrained Hyperspectral ReconstructionabstractAchieving accurate material segmentation for 3-channel RGB images is challenging due to the considerable variation in material appearance. Hyperspectral images, which are sets of spectral measurements sampled at multiple wavelengths, theoretically offer distinct information for material identification, as variations in the intensity of electromagnetic radiation reflected by a surface depend on the material composition of a scene. However, existing hyperspectral datasets are impoverished in terms of the number of images and material categories for the dense material segmentation task, and collecting and annotating hyperspectral images with a spectral camera is prohibitively expensive. To address this, we propose a novel model, the MatSpectNet to segment materials with recovered hyperspectral images from RGB images. The network leverages the principles of colour perception in modern cameras to constrain the reconstructed hyperspectral images and employs the domain adaptation method to generalise the hyperspectral reconstruction capability from a spectral recovery dataset to material segmentation datasets. The reconstructed hyperspectral images are further filtered using learnt response curves and enhanced with human perception. The performance of MatSpectNet is evaluated on the LMD dataset as well as the OpenSurfaces dataset. Our experiments demonstrate that MatSpectNet attains a 1.60% increase in average pixel accuracy and a 3.42% improvement in mean class accuracy compared with the most recent publication. Additional experiments and the project code are published on https://github.com/heng-yuwen/MatSpectNet. Yuwen Heng, Yihong Wu 0004, Srinandan Dasmahapatra, Hansung Kim 0001 |
WACV | 4 |
| 2024 | 3D Semantic Scene Completion From A Depth Map With Unsupervised Learning For Semantics PrioritisationabstractThe Semantic Scene Completion (SSC) problem entails generating a comprehensive 3D voxel representation of a scene from a partial view, while simultaneously predicting volumetric occupancy and object category. A significant challenge in SSC is evaluating occluded regions in 3D space and accurately predicting object categories within an imbalanced setting. In addressing this challenge, our study explores SSC literature and introduces a simple, innovative class-balancing re-weighting technique rooted in an unsupervised clustering, leading to balanced learning and generalised representation. This method modulates the penalty on dataset classes during the CNN learning, emphasizing infrequent classes while moderately de-prioritizing the dominant ones, combining the strengths of both re-sampling and cost-sensitive learning enhancing the performance for both scene completion and scene semantics tasks. Our design, which relies on a single depth input without any RGB information, has shown to significantly outperform comparable baseline models. Our results are also competitively matched with other multi-input methods. Mona Alawadh, Mahesan Niranjan, Hansung Kim 0001 |
ICIP | 3 |
| 2024 | LInKs "Lifting Independent Keypoints" - Partial Pose Lifting for Occlusion Handling with Improved Accuracy in 2D-3D Human Pose EstimationabstractWe present LInKs, a novel unsupervised learning method to recover 3D human poses from 2D kinematic skeletons obtained from a single image, even when occlusions are present. Our approach follows a unique two-step process, which involves first lifting the occluded 2D pose to the 3D domain, followed by filling in the occluded parts using the partially reconstructed 3D coordinates. This lift-then-fill approach leads to significantly more accurate results compared to models that complete the pose in 2D space alone. Additionally, we improve the stability and likelihood estimation of normalising flows through a custom sampling function replacing PCA dimensionality reduction used in prior work. Furthermore, we are the first to investigate if different parts of the 2D kinematic skeleton can be lifted independently which we find by itself reduces the error of current lifting approaches. We attribute this to the reduction of long-range keypoint correlations. In our detailed evaluation, we quantify the error under various realistic occlusion scenarios, showcasing the versatility and applicability of our model. Our results consistently demonstrate the superiority of handling all types of occlusions in 3D space when compared to others that complete the pose in 2D space. Our approach also exhibits consistent accuracy in scenarios without occlusion, as evidenced by a 7.9% reduction in reconstruction error compared to prior works on the Human3.6M dataset. Furthermore, our method excels in accurately retrieving complete 3D poses even in the presence of occlusions, making it highly applicable in situations where complete 2D pose information is unavailable. Peter Hardy, Hansung Kim 0001 |
WACV | 2 |
| 2023 | Depth Estimation for a Single Omnidirectional Image with Reversed-Gradient Warming-up Thresholds DiscriminatorabstractDepth estimation for single image using deep learning requires a large labelled depth dataset with various scenes for training. However, currently published omnidirectional depth datasets cover limited types of scenes and are not suitable for depth estimation for various real-world scenes. With the challenge of labelled real-world datasets generation and stability of the performance, we propose an architecture with the Reverse-gradient Warming-up Threshold Discriminator (RWTD) to estimate real-world depth maps from the synthetic ground truth. It takes labelled synthetic scenes of a source domain and unlabelled real-world scenes of a target domain as inputs to predict the corresponding depth maps. Compared with state-of-the-art encoder-decoder models, the proposed architecture shows an 11% points improvement on the testing dataset for depth accuracy. Yihong Wu 0004, Yuwen Heng, Mahesan Niranjan, Hansung Kim 0001 |
ICASSP | 4 |
| 2023 | Spatial Audio Reconstruction for VR Applications Using a Combined Method Based on SIRR and RSAO ApproachesabstractIn order to recreate a sound field as realistic and immersive as the original setting, it is important to preserve the properties of the room acoustics. The room acoustics are usually captured by room impulse responses (RIRs) which can be employed to generate sound that could lead to the same audio perception. In this paper, we compare and combine two main parametric approaches, Spatial Impulse Response Rendering (SIRR) and Reverberant Spatial Audio Object (RSAO) to encode and render the RIRs. We first discuss that SIRR synthesises the early reflections more precisely whereas RSAO is a better approach to render the reverberation. For our proposed combined method, each RIR is divided into three parts: The direct sound, early reflections, and late reverberation, as in RSAO. To estimate the boundary time between the early reflections and late reverberation, we employ the diffuseness factor as in SIRR and set the mixing time as the time when the diffuseness factor reaches a threshold. Then the early part is analysed and synthesised by SIRR while the reverberant tail is encoded and rendered using RSAO. We show that in general, the combined method benefits the positive aspects of the two baseline methods. Atiyeh Alinaghi, Luca Remaggi, Hansung Kim 0001 |
MMSP | 3 |
| 2022 | Enhancing Material Features Using Dynamic Backward Attention on Cross-Resolution Patches
Yuwen Heng, Yihong Wu 0004, Srinandan Dasmahapatra, Hansung Kim 0001 |
BMVC | 4 |
| 2021 | Temporally Consistent 3D Human Pose Estimation Using Dual 360° CamerasabstractThis paper presents a 3D human pose estimation system that uses a stereo pair of 360° sensors to capture the complete scene from a single location. The approach combines the advantages of omnidirectional capture, the accuracy of multiple view 3D pose estimation and the portability of monocular acquisition. Joint monocular belief maps for joint locations are estimated from 360° images and are used to fit a 3D skeleton to each frame. Temporal data association and smoothing is performed to produce accurate 3D pose estimates throughout the sequence. We evaluate our system on the Panoptic Studio dataset, as well as real 360° video for tracking multiple people, demonstrating an average Mean Per Joint Position Error of 12.47cm with 30cm baseline cameras. We also demonstrate improved capabilities over perspective and 360° multi-view systems when presented with limited camera views of the subject. Matthew Shere, Hansung Kim 0001, Adrian Hilton 0001 |
WACV | 2 |
| 2021 | Temporally Coherent General Dynamic Scene ReconstructionabstractAbstract Existing techniques for dynamic scene reconstruction from multiple wide-baseline cameras primarily focus on reconstruction in controlled environments, with fixed calibrated cameras and strong prior constraints. This paper introduces a general approach to obtain a 4D representation of complex dynamic scenes from multi-view wide-baseline static or moving cameras without prior knowledge of the scene structure, appearance, or illumination. Contributions of the work are: an automatic method for initial coarse reconstruction to initialize joint estimation; sparse-to-dense temporal correspondence integrated with joint multi-view segmentation and reconstruction to introduce temporal coherence; and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes by introducing shape constraint. Comparison with state-of-the-art approaches on a variety of complex indoor and outdoor scenes, demonstrates improved accuracy in both multi-view segmentation and dense reconstruction. This paper demonstrates unsupervised reconstruction of complete temporally coherent 4D scene models with improved non-rigid object segmentation and shape reconstruction and its application to various applications such as free-view rendering and virtual reality. Armin Mustafa, Marco Volino, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Acoustic Room Modelling Using 360 Stereo CamerasabstractIn this paper we propose a pipeline for estimating acoustic 3D room structure with geometry and attribute prediction using spherical 360$^{\circ }$cameras. Instead of setting microphone arrays with loudspeakers to measure acoustic parameters for specific rooms, a simple and practical single-shot capture of the scene using a stereo pair of 360 cameras can be used to simulate those acoustic parameters. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). The scene is captured as a stereo pair using off-the-shelf consumer spherical 360 cameras. A cuboid-based 3D room geometry model is estimated by correspondence matching between captured images and semantic labelling using a convolutional neural network (SegNet). The estimated geometry is used to produce frequency-dependent acoustic predictions of the scene. This is, to our knowledge, the first attempt in the literature to use visual geometry estimation and object classification algorithms to predict acoustic properties. Results are compared to measurements through calculated reverberant spatial audio object parameters used for reverberation reproduction customized to the given loudspeaker set up. Hansung Kim 0001, Luca Remaggi, Sam Fowler, Philip J. B. Jackson, Adrian Hilton 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | 3D Multi Person Tracking With Dual 360° CamerasabstractPerson tracking is an often studied facet of computer vision, with applications in security, automated driving and entertainment. However, despite the advantages they offer, few current solutions work for 360° cameras, due to projection distortion. This paper presents a simple yet robust method for 3D tracking of multiple people in a scene from a pair of 360° cameras. By using 2D pose information, rather than potentially unreliable 3D position or repeated colour information, we create a tracker that is both appearance independent as well as capable of operating at narrow baseline. Our results demonstrate state of the art performance on 360° scenes, as well as the capability to handle vertical axis rotation. Matthew Shere, Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 2 |
| 2020 | EdgeNet: Semantic Scene Completion from a Single RGB- D ImageabstractSemantic scene completion is the task of predicting a complete 3D representation of volumetric occupancy with corresponding semantic labels for a scene from a single point of view. In this paper, we present EdgeNet, a new end-to-end neural network architecture that fuses information from depth and RGB, explicitly representing RGB edges in 3D space. Previous works on this task used either depth-only or depth with colour by projecting 2D semantic labels generated by a 2D segmentation network into the 3D volume, requiring a two step training process. Our EdgeNet representation encodes colour information in 3D space using edge detection and flipped truncated signed distance, which improves semantic completion scores especially in hard to detect classes. We achieved state-of-the-art scores on both synthetic and real datasets with a simpler and a more computationally efficient training pipeline than competing approaches. Aloisio Dourado, Teófilo Emídio de Campos, Hansung Kim 0001, Adrian Hilton 0001 |
ICPR | 3 |
| 2019 | Super Long Interval Time-Lapse Image Generation for Proactive Preservation of Cultural Heritage Using CrowdsourcingabstractTo establish advanced analytical methods for preserving cultural heritage, this research proposes a method to generate a time-lapse image with a super-long temporal interval. The key issue is to realize an image collection method using crowdsourcing and a method to improve the matching accuracy between images of cultural heritage buildings captured 50 to 100 years ago and current images. As degradation and damage to the appearance of cultural heritage buildings occurs due to ageing, rebuilding, and renovation, image features of the timed images are changed. This decreases the accuracy of the matching process that uses the appearance of patch-region. In addition, we need to give more consideration to incorrect feature correspondence that is prominent in buildings with considerable symmetry. We aim to solve these difficulties by applying an Autoencoder and a guided matching method. Our method involves utilizing the function of crowdsourcing, which can easily obtain the current image captured at the same position and orientation as the past image. We propose this method to address the inability to obtain the correspondence points between two images when observation times are significantly different. Hidehiko Shishido, Hansung Kim 0001, Itaru Kitahara |
IEEE BigData | 2 |
| 2019 | Immersive Spatial Audio Reproduction for VR/AR Using Room Acoustic Modelling from 360° ImagesabstractRecent progresses in Virtual Reality (VR) and Augmented Reality (AR) allow us to experience various VR/AR applications in our daily life. In order to maximise the immersiveness of user in VR/AR environments, a plausible spatial audio reproduction synchronised with visual information is essential. In this paper, we propose a simple and efficient system to estimate room acoustic for plausible reproducton of spatial audio using 360° cameras for VR/AR applications. A pair of 360° images is used for room geometry and acoustic property estimation. A simplified 3D geometric model of the scene is estimated by depth estimation from captured images and semantic labelling using a convolutional neural network (CNN). The real environment acoustics are characterised by frequency-dependent acoustic predictions of the scene. Spatially synchronised audio is reproduced based on the estimated geometric and acoustic properties in the scene. The reconstructed scenes are rendered with synthesised spatial audio as VR/AR content. The results of estimated room geometry and simulated spatial audio are evaluated against the actual measurements and audio calculated from ground-truth Room Impulse Responses (RIRs) recorded in the rooms. Hansung Kim 0001, Luca Hernaggi, Philip J. B. Jackson, Adrian Hilton 0001 |
VR | 1 |
| 2019 | OCEAN: Object-centric arranging network for self-supervised visual representations learning
Changjae Oh, Bumsub Ham, Hansung Kim 0001, Adrian Hilton 0001, Kwanghoon Sohn |
Expert Syst. Appl. | 3 |
| 2019 | MSFD: Multi-Scale Segmentation-Based Feature Detection for Wide-Baseline Scene ReconstructionabstractA common problem in wide-baseline matching is the sparse and non-uniform distribution of correspondences when using conventional detectors, such as SIFT, SURF, FAST, A-KAZE, and MSER. In this paper, we introduce a novel segmentation-based feature detector (SFD) that produces an increased number of accurate features for wide-baseline matching. A multi-scale SFD is proposed using bilateral image decomposition to produce a large number of scale-invariant features for wide-baseline reconstruction. All input images are over-segmented into regions using any existing segmentation technique, such as Watershed, Mean-shift, and simple linear iterative clustering. Feature points are then detected at the intersection of the boundaries of three or more regions. The detected feature points are local maxima of the image function. The key advantage of feature detection based on segmentation is that it does not require global threshold setting and can, therefore, detect features throughout the image. A comprehensive evaluation demonstrates that SFD gives an increased number of features that are accurately localized and matched between wide-baseline camera views; the number of features for a given matching error increases by a factor of 3-5 compared with SIFT; feature detection and matching performance are maintained with increasing baseline between views; multi-scale SFD improves matching performance at varying scales. Application of SFD to sparse multi-view wide-baseline reconstruction demonstrates a factor of 10 increases in the number of reconstructed points with improved scene coverage compared with SIFT/MSER/A-KAZE. Evaluation against ground-truth shows that SFD produces an increased number of wide-baseline matches with a reduced error. Armin Mustafa, Hansung Kim 0001, Adrian Hilton 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Human-Centric Scene Understanding from Single View 360 VideoabstractIn this paper, we propose an approach to indoor scene understanding from observation of people in single view spherical video. As input, our approach takes a centrally located spherical video capture of an indoor scene, estimating the 3D localisation of human actions performed throughout the long term capture. The central contribution of this work is a deep convolutional encoder-decoder network trained on a synthetic dataset to reconstruct regions of affordance from captured human activity. The predicted affordance segmentation is then applied to compose a reconstruction of the complete 3D scene, integrating the affordance segmentation into 3D space. The mapping learnt between human activity and affordance segmentation demonstrates that omnidirectional observation of human activity can be applied to scene understanding tasks such as 3D reconstruction. We show that our approach using only observation of people performs well against previous approaches, allowing reconstruction of occluded regions and labelling of scene affordances. Sam Fowler, Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 2 |
| 2018 | Acoustic Reflector Localization and ClassificationabstractThe process of understanding acoustic properties of environments is important for several applications, such as spatial audio, augmented reality and source separation. In this paper, multichannel room impulse responses are recorded and transformed into their direction of arrival (DOA)-time domain, by employing a superdirective beamformer. This domain can be represented as a 2D image. Hence, a novel image processing method is proposed to analyze the DOA-time domain, and estimate the reflection times of arrival and DOAs. The main acoustically reflective objects are then localized. Recent studies in acoustic reflector localization usually assume the room to be free from furniture. Here, by analyzing the scattered reflections, an algorithm is also proposed to binary classify reflectors into room boundaries and interior furniture. Experiments were conducted in four rooms. The classification algorithm showed high quality performance, also improving the localization accuracy, for non-static listener scenarios. Luca Remaggi, Hansung Kim 0001, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
ICASSP | 2 |
| 2018 | AVSU: Workshop on Audio-Visual Scene Understanding for Immersive MultimediaabstractThis workshop aims to provide a forum to exchange ideas in scene understanding techniques researched in audio and visual communities, and to ultimately unlock the creative potential of joint audio-visual signal processing to deliver a step change in various multimedia applications. Papers and talks presented in this workshop will contribute to the emerging technology for audio and visual information that can improve traditional approaches for multimedia content production and reproduction. The goals of this workshop are to (1) present and discuss the latest trends in audio and computer vision fields for the common research goals, (2) understand state-of-the-art techniques and bottlenecks in the other's discipline for the common topics, (3) investigate research opportunities of joint audio-visual scene understandings in multimedia content production. This workshop will be a good opportunity to bring together leading experts in audio processing and computer vision, and will bridge the gap between two research fields in multimedia content production and reproduction. Adrian Hilton 0001, Hong-Goo Kang, Hansung Kim 0001, Kwanghoon Sohn |
ACM Multimedia | 3 |
| 2018 | Multimodal Visual Data Registration for Web-Based Visualization in Media ProductionabstractRecent developments of video and sensing technology have led to large volumes of digital media data. Current media production relies on videos from the principal camera together with a wide variety of heterogeneous source of supporting data [photos, light detection and ranging point clouds, witness video camera, high dynamic range imaging, and depth imagery]. Registration of visual data acquired from various 2D and 3D sensing modalities is challenging because current matching and registration methods are not appropriate due to differences in structure, format, and noise characteristics for multimodal data. A combined 2D/3D visualization of this registered data allows an integrated overview of the entire data set. For such a visualization, a Web-based context presents several advantages. In this paper, we propose a unified framework for registration and visualization of this type of visual media data. A new feature description and matching method is proposed, adaptively considering local geometry, semiglobal geometry, and color information in the scene for more robust registration. The resulting registered 2D/3D multimodal visual data are too large to be downloaded and viewed directly via the Web browser, while maintaining an acceptable user experience. Thus, we employ hierarchical techniques for compression and restructuring to enable efficient transmission and visualization over the Web, leading to interactive visualization as registered point clouds, 2D images, and videos in the browser, improving on the current state-of-the-art techniques for Web-based visualization of big media data. This is the first unified 3D Web-based visualization of multimodal visual media production data sets. The proposed pipeline is tested on big multimodal data set typical of film and broadcast production, which are made publicly available. The proposed feature description method shows two times higher precision of feature matching and more stable registration performance than existing 3D feature descriptors. Hansung Kim 0001, Alun Evans, Josep Blat, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | 3D Room Geometry Reconstruction Using Audio-Visual SensorsabstractIn this paper we propose a cuboid-based air-tight indoor room geometry estimation method using combination of audio-visual sensors. Existing vision-based 3D reconstruction methods are not applicable for scenes with transparent or reflective objects such as windows and mirrors. In this work we fuse multi-modal sensory information to overcome the limitations of purely visual reconstruction for reconstruction of complex scenes including transparent and mirror surfaces. A full scene is captured by 360$^{\circ}$ cameras and acoustic room impulse responses (RIRs) recorded by a loudspeaker and compact microphone array. Depth information of the scene is recovered by stereo matching from the captured images and estimation of major acoustic reflector locations from the sound. The coordinate systems for audio-visual sensors are aligned into a unified reference frame and plane elements are reconstructed from audio-visual data. Finally cuboid proxies are fitted to the planes to generate a complete room model. Experimental results show that the proposed system generates complete representations of the room structures regardless of transparent windows, featureless walls and shiny surfaces. Hansung Kim 0001, Luca Remaggi, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
3DV | 1 |
| 2017 | Towards Complete Scene Reconstruction from Single-View Depth and Human Motion
Sam Fowler, Hansung Kim 0001, Adrian Hilton 0001 |
BMVC | 2 |
| 2016 | Room Layout Estimation with Object and Material Attributes Information Using a Spherical CameraabstractIn this paper we propose a pipeline for estimating 3D room layout with object and material attribute prediction using a spherical stereo image pair. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). A spherical stereo alignment algorithm is proposed to align two spherical images to the global world coordinate system. Depth information of the scene is estimated by stereo matching between images. Cubic projection images of the spherical RGB and estimated depth are used for object and material attribute detection. A single Convolutional Neural Network is designed to assign object and attribute labels to geometrical elements built from the spherical image. Finally simplified room layout is reconstructed by cuboid fitting. The reconstructed cuboid-based model shows the structure of the scene with object information and material attributes. Hansung Kim 0001, Teófilo Emídio de Campos, Adrian Hilton 0001 |
3DV | 1 |
| 2016 | Temporally Coherent 4D Reconstruction of Complex Dynamic ScenesabstractThis paper presents an approach for reconstruction of 4D temporally coherent models of complex dynamic scenes. No prior knowledge is required of scene structure or camera calibration allowing reconstruction from multiple moving cameras. Sparse-to-dense temporal correspondence is integrated with joint multi-view segmentation and reconstruction to obtain a complete 4D representation of static and dynamic objects. Temporal coherence is exploited to overcome visual ambiguities resulting in improved reconstruction of complex scenes. Robust joint segmentation and reconstruction of dynamic objects is achieved by introducing a geodesic star convexity constraint. Comparative evaluation is performed on a variety of unstructured indoor and outdoor dynamic scenes with hand-held cameras and multiple people. This demonstrates reconstruction of complete temporally coherent 4D scene models with improved nonrigid object segmentation and shape reconstruction. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
CVPR | 2 |
| 2016 | 4D Match Trees for Non-rigid Surface Alignment
Armin Mustafa, Hansung Kim 0001, Adrian Hilton 0001 |
ECCV (1) | 2 |
| 2016 | Big Data Analysis for Media ProductionabstractA typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields. Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas |
Proc. IEEE | 3 |
| 2015 | Segmentation Based Features for Wide-Baseline Multi-view ReconstructionabstractA common problem in wide-baseline stereo is the sparse and non-uniform distribution of correspondences when using conventional detectors such as SIFT, SURF, FAST and MSER. In this paper we introduce a novel segmentation based feature detector SFD that produces an increased number of 'good' features for accurate wide-baseline reconstruction. Each image is segmented into regions by over-segmentation and feature points are detected at the intersection of the boundaries for three or more regions. Segmentation-based feature detection locates features at local maxima giving a relatively large number of feature points which are consistently detected across wide-baseline views and accurately localised. A comprehensive comparative performance evaluation with previous feature detection approaches demonstrates that: SFD produces a large number of features with increased scene coverage, detected features are consistent across wide-baseline views for images of a variety of indoor and outdoor scenes, and the number of wide-baseline matches is increased by an order of magnitude compared to alternative detector-descriptor combinations. Sparse scene reconstruction from multiple wide-baseline stereo views using the SFD feature detector demonstrates at least a factor six increase in the number of reconstructed points with reduced error distribution compared to SIFT when evaluated against ground-truth and similar computational cost to SURF/FAST. Armin Mustafa, Hansung Kim 0001, Evren Imre, Adrian Hilton 0001 |
3DV | 2 |
| 2015 | General Dynamic Scene Reconstruction from Multiple View VideoabstractThis paper introduces a general approach to dynamic scene reconstruction from multiple moving cameras without prior knowledge or limiting constraints on the scene structure, appearance, or illumination. Existing techniques or dynamic scene reconstruction from multiple wide-baseline camera views primarily focus on accurate reconstruction in controlled environments, where the cameras are fixed and calibrated and background is known. These approaches are not robust for general dynamic scenes captured with sparse moving cameras. Previous approaches for outdoor dynamic scene reconstruction assume prior knowledge of the static background appearance and structure. The primary contributions of this paper are twofold: an automatic method for initial coarse dynamic scene segmentation and reconstruction without prior knowledge of background appearance or structure, and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes from multiple wide-baseline static or moving cameras. Evaluation is performed on a variety of indoor and outdoor scenes with cluttered backgrounds and multiple dynamic non-rigid objects such as people. Comparison with state-of-the-art approaches demonstrates improved accuracy in both multiple view segmentation and dense reconstruction. The proposed approach also eliminates the requirement for prior knowledge of scene structure and appearance. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
ICCV | 2 |
| 2015 | Multi-modal big-data management for film productionabstractModern digital film production uses large quantities of data from videos, digital photographs, LIDAR scans, spherical photography and many other sources to create the final film frames. The processing and management of this massive amount of heterogeneous data consumes enormous resources. We propose an integrated pipeline for 2D/3D data registration for film production. We present the prototype application Jigsaw, which allows users to efficiently manage and process various data from digital photographs to 3D point clouds. A key requirement in the use of multi-modal 2D/3D data for content production is the registration into a common coordinate frame. 3D geometric information is reconstructed from 2D data and registered to the reference 3D models using 3D feature matching. We provide a public multi-modal database captured with a wide variety of devices in different environments to assist further research. An order of magnitude gain in efficiency is achieved with the proposed approach. Hansung Kim 0001, Simon Pabst, Justin Sneddon, Ted Waine, Jeff Clifford, Adrian Hilton 0001 |
ICIP | 1 |
| 2015 | Block world reconstruction from spherical stereo image pairs
Hansung Kim 0001, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 1 |
| 2014 | Influence of Colour and Feature Geometry on Multi-modal 3D Point Clouds Data RegistrationabstractWith the current transition of various digital contents from 2D to 3D, the problem of 3D data matching and registration is increasingly important. Registration of multi-modal 3D data acquired from different sensors remains a challenging problem due to the difference in types and characteristics of the data. In this paper, we evaluate the registration performance of 3D feature descriptors with different domains on datasets from various environments and modalities. Datasets are acquired in indoor and outdoor environments with 2D and 3D sensing devices including LIDAR, spherical imaging, digital camera and RGBD camera. FPFH, PFH and SHOT feature descriptors are applied to the 3D point clouds generated from the multi-modal datasets. Local neighbouring point distribution, key points distribution, colour information and their combinations are used for feature description. Finally we analyse their influences on the multi-modal 3D point clouds data registration. Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 1 |
| 2014 | Hybrid 3D feature description and matching for multi-modal data registrationabstractWe propose a robust 3D feature description and registration method for 3D models reconstructed from various sensor devices. General 3D feature detectors and descriptors generally show low distinctiveness and repeatability for matching between different data modalities due to differences in noise and errors in geometry. The proposed method considers not only local 3D points but also neighbouring 3D keypoints to improve keypoint matching. The proposed method is tested on various multi-modal datasets including LIDAR scans, multiple photos, spherical images and RGBD videos to evaluate the performance against existing methods. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 1 |
| 2013 | Evaluation of 3D Feature Descriptors for Multi-modal Data RegistrationabstractWe propose a framework for 2D/3D multi-modal data registration and evaluate 3D feature descriptors for registration of 3D datasets from different sources. 3D datasets of outdoor environments can be acquired using a variety of active and passive sensor technologies. Registration of these datasets into a common coordinate frame is required for subsequent modelling and visualisation. 2D images are converted into 3D structure by stereo or multiview reconstruction techniques and registered to a unified 3D domain with other datasets in a 3D world. Multi-modal datasets have different density, noise, and types of errors in geometry. This paper provides a performance benchmark for existing 3D feature descriptors across multi-modal datasets. This analysis highlights the limitations of existing 3D feature detectors and descriptors which need to be addressed for robust multi-modal data registration. We analyse and discuss the performance of existing methods in registering various types of datasets then identify future directions required to achieve robust multi-modal data registration. Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 1 |
| 2013 | 3D Scene Reconstruction from Multiple Spherical Stereo Pairs
Hansung Kim 0001, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 1 |
| 2012 | Outdoor Dynamic 3-D Scene ReconstructionabstractExisting systems for 3-D reconstruction from multiple view video use controlled indoor environments with uniform illumination and backgrounds to allow accurate segmentation of dynamic foreground objects. In this paper, we present a portable system for 3-D reconstruction of dynamic outdoor scenes that require relatively large capture volumes with complex backgrounds and nonuniform illumination. This is motivated by the demand for 3-D reconstruction of natural outdoor scenes to support film and broadcast production. Limitations of existing multiple view 3-D reconstruction techniques for use in outdoor scenes are identified. Outdoor 3-D scene reconstruction is performed in three stages: 1) 3-D background scene modeling using spherical stereo image capture; 2) multiple view segmentation of dynamic foreground objects by simultaneous video matting across multiple views; and 3) robust 3-D foreground reconstruction and multiple view segmentation refinement in the presence of segmentation and calibration errors. Evaluation is performed on several outdoor productions with complex dynamic scenes including people and animals. Results demonstrate that the proposed approach overcomes limitations of previous indoor multiple view reconstruction approaches enabling high-quality free-viewpoint rendering and 3-D reference models for production. Hansung Kim 0001, Jean-Yves Guillemaut, Takeshi Takai, Muhammad Sarim, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | PDE-based disparity estimation with occlusion and texture handling for accurate depth recovery from a stereo image pairabstractThis paper presents a novel PDE-based method for floating-point disparity estimation which produces smooth disparity fields with sharp object boundaries for surface reconstruction. In order to avoid the over-segmentation problem of image-driven structure tensor and the blurred boundary problem of field-driven tensor, we propose a new anisotropic diffusivity function controlled by image and disparity gradients. We also embed a bi-directional disparity matching term to control the data term in occluded regions. We evaluate the proposed method on data sets from the Middlebury benchmarking site and real data sets with ground-truth models scanned by a LIDAR sensor. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 1 |
| 2010 | Natural image matting for multiple wide-baseline viewsabstractIn this paper we present a novel approach to estimate the alpha mattes of a foreground object captured by a wide-baseline circular camera rig provided a single key frame trimap. Bayesian inference coupled with camera calibration information are used to propagate high confidence trimaps labels across the views. Recent techniques have been developed to estimate an alpha matte of an image using multiple views but they are limited to narrow baseline views with low foreground variation. The proposed wide-baseline trimap propagation is robust to inter-view foreground appearance changes, shadows and similarity in foreground/background appearance for cameras with opposing views enabling high quality alpha matte extraction using any state-of-the-art image matting algorithm. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Takeshi Takai, Hansung Kim 0001 |
ICIP | 5 |
| 2009 | Graph-based foreground extraction in extended color spaceabstractWe propose a region-based method to extract semantic foreground regions from color video sequences with static backgrounds. First, we introduce a new distance measure for background subtraction which is robust against shadows. Then the foreground region is extracted with a graph-based region segmentation method considering background difference and spatial homogeneity. For efficient computation, the graph structure is optimized by the minimum spanning tree before segmentation. The main contribution is that the proposed algorithm improves on conventional approaches especially in strong shadow regions and does not require manual initialization. We have verified through experiments and comparison to state of the art methods that the proposed algorithm works well with various cameras and environment. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 1 |
| 2009 | Non-parametric natural image mattingabstractNatural image matting is an extremely challenging image processing problem due to its ill-posed nature. It often requires skilled user interaction to aid definition of foreground and background regions. Current algorithms use these predefined regions to build local foreground and background colour models. In this paper we propose a novel approach which uses non-parametric statistics to model image appearance variations. This technique overcomes the limitations of previous parametric approaches which are purely colour-based and thereby unable to model natural image structure. The proposed technique consists of three successive stages: (i) background colour estimation, (ii) foreground colour estimation, (iii) alpha estimation. Colour estimation uses patch-based matching techniques to efficiently recover the optimum colour by comparison against patches from the known regions. Quantitative evaluation against ground truth demonstrates that the technique produces better results and successfully recovers fine details such as hair where many other algorithms fail. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Hansung Kim 0001 |
ICIP | 4 |
| 2009 | Toward cinematizing our daily lives
Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure |
Multim. Tools Appl. | 1 |
| 2007 | Robust Foreground Extraction Technique Using Gaussian Family Model and Multiple Thresholds
Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure |
ACCV (1) | 1 |
| 2007 | Reliability-based 3D reconstruction in real environmentabstractWe present a practical 3D reconstruction method that guarantees robust visual hull construction in real environments where segmentation errors and occlusion exist. The proposed method consists of foreground extraction and reliability-based shape-from-silhouette, and they are connected by the intra-/inter-silhouette reliabilities. In foreground extraction, all regions are classified into four categories based on their intra-reliabilities. Then the reliability-based shape-from-silhouette technique reconstructs a visual hull by carving a 3D space based on the intra-/inter-silhouette reliabilities. The proposed method provides a reliable visual hull in real environments without much increment of the system complexity compared with conventional systems. Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure |
ACM Multimedia | 1 |
| 2006 | Edge-preserving joint motion-disparity estimation in stereo image sequences
Dongbo Min, Hansung Kim 0001, Kwanghoon Sohn |
Signal Process. Image Commun. | 2 |
| 2005 | 3D reconstruction from stereo images for interactions between real and virtual objects
Hansung Kim 0001, Kwanghoon Sohn |
Signal Process. Image Commun. | 1 |
| 2003 | Hierarchical disparity estimation with energy-based regularizationabstractIn this paper, we propose a hierarchical disparity estimation algorithm with energy-based regularization. Initial disparity vectors are obtained from downsampled stereo images using a feature-based region-dividing disparity estimation technique. Dense disparities are estimated from these initial vectors with shape-adaptive windows in full resolution images. Finally, the vector fields are regularized with the minimization of the energy functional which considers both fidelity and smoothness of the fields. The first two steps provide highly reliable disparity vectors, so that local minimum problem can be avoided in regularization step. The proposed algorithm generates accurate disparity map which is smooth inside objects while preserving its discontinuities in boundaries. Experimental results are presented to illustrate the capabilities of the proposed disparity estimation technique. Hansung Kim 0001, Kwanghoon Sohn |
ICIP (1) | 1 |
| 2003 | 3D Reconstruction of Stereo Images for Interaction between Real and Virtual WorldsabstractMixed reality is different from the virtual reality in that users can feel immersed in a space which is composed of not only virtual but also real objects. Thus, it is essential to realize seamless integration and interaction of the virtual and real worlds. We need depth information of the real scene to synthesize the real and virtual objects. We propose a two-stage algorithm to find smooth and precise disparity vector fields with sharp object boundaries in a stereo image pair for depth estimation. Hierarchical region-dividing disparity estimation increases the efficiency and the reliability of the estimation process, and a shape-adaptive window provides high reliability of the fields around the object boundary region. At the second stage, the vector fields are regularized with a energy model which produces smooth fields while preserving their discontinuities resulting from the object boundaries. The vector fields are used to reconstruct 3D surface of the real scene. Simulation results show that the proposed algorithm provides accurate and spatially correlated disparity vector fields in various kinds of images, and synthesized 3D models produce natural space where the virtual objects interact with the real world as if they are in the same world. Hansung Kim 0001, Seung-Jun Yang, Kwanghoon Sohn |
ISMAR | 1 |