VLDB 2026 Research / reviewers in the wild / expert
Mariko Isogawa
dblp:149/1494
· DBLP profile ↗
27ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0001-9560-0276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EventEgoHands: Event-Based Egocentric 3D Hand Mesh ReconstructionabstractReconstructing 3D hand mesh is challenging but an important task for human-computer interaction and AR/VR applications. In particular, RGB and/or depth cameras have been widely used in this task. However, methods using these conventional cameras face challenges in low-light environments and during motion blur. Thus, to address these limitations, event cameras have been attracting attention in recent years for their high dynamic range and high temporal resolution. Despite their advantages, event cameras are sensitive to background noise or camera motion, which has limited existing studies to static backgrounds and fixed cameras. In this study, we propose EventEgoHands, a novel method for event-based 3D hand mesh reconstruction in an egocentric view. Our approach introduces a Hand Segmentation Module that extracts hand regions, effectively mitigating the influence of dynamic background events. We evaluated our approach and demonstrated its effectiveness on the N-HOT3D dataset, improving MPJPE by approximately more than 4.5 cm (43%). Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa |
ICIP | 4 |
| 2025 | Event-Based Egocentric Human Pose Estimation in Dynamic EnvironmentabstractEstimating human pose using a front-facing egocentric camera is essential for applications such as sports motion analysis, VR/AR, and AI for wearable devices. However, many existing methods rely on RGB cameras and do not account for low-light environments or motion blur. Event-based cameras have the potential to address these challenges. In this work, we introduce a novel task of human pose estimation using a front-facing event-based camera mounted on the head and propose D-EventEgo, the first framework for this task. The proposed method first estimates the head poses, and then these are used as conditions to generate body poses. However, when estimating head poses, the presence of dynamic objects mixed with background events may reduce head pose estimation accuracy. Therefore, we introduce the Motion Segmentation Module to remove dynamic objects and extract background information. Extensive experiments on our synthetic event-based dataset derived from EgoBody, demonstrate that our approach outperforms our baseline in four out of five evaluation metrics in dynamic environments. Wataru Ikeda, Masashi Hatano, Ryosei Hara, Mariko Isogawa |
ICIP | 4 |
| 2025 | Occlusion-Free 4D Gaussians for Open Surgery Videos Using Multi-camera Shadowless Lamps
Yuna Kato, Shohei Mori, Hideo Saito 0001, Yoshifumi Takatsume, Hiroki Kajita, Mariko Isogawa |
MICCAI (10) | 6 |
| 2025 | Dense Depth from Event Focal StackabstractWe propose a method for dense depth estimation from an event stream generated when sweeping the focal plane of the driving lens attached to an event camera. In this method, a depth map is inferred from an “event focal stack” composed of the event stream using a convolutional neural network trained with synthesized event focal stacks. The synthesized event stream is created from a focal stack generated by Blender for any arbitrary 3D scene. This allows for training on scenes with diverse structures. Additionally, we explored methods to eliminate the domain gap between real event streams and synthetic event streams. Our method demonstrates superior performance over a depth-from-defocus method in the image domain on synthetic and real datasets. Kenta Horikawa, Mariko Isogawa, Hideo Saito 0001, Shohei Mori |
WACV | 2 |
| 2025 | EventPointMesh: Human Mesh Recovery Solely From Event Point CloudsabstractHow much can we infer about human shape using an event camera that only detects the pixel position where the luminance changed and its timestamp? This neuromorphic vision technology captures changes in pixel values at ultra-high speeds, regardless of the variations in environmental lighting brightness. Existing methods for human mesh recovery (HMR) from event data need to utilize intensity images captured with a generic frame-based camera, rendering them vulnerable to low-light conditions, energy/memory constraints, and privacy issues. In contrast, we explore the potential of solely utilizing event data to alleviate these issues and ascertain whether it offers adequate cues for HMR, as illustrated in Fig. 1. This is a quite challenging task due to the substantially limited information ensuing from the absence of intensity images. To this end, we propose EventPointMesh, a framework which treats event data as a three-dimensional (3D) spatio-temporal point cloud for reconstructing the human mesh. By employing a coarse-to-fine pose feature extraction strategy, we extract both global features and local features. The local features are derived by processing the spatio-temporally dispersed event points into groups associated with individual body segments. This combination of global and local features allows the framework to achieve a more accurate HMR, capturing subtle differences in human movements. Experiments demonstrate that our method with only sparse event data outperforms baseline methods. Ryosuke Hori, Mariko Isogawa, Dan Mikami, Hideo Saito 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Acoustic-based 3D Human Pose Estimation Robust to Human Position
Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa |
BMVC | 6 |
| 2024 | Efficient Circular and Confocal Non-Line-Of-Sight Imaging With Transient Sinogram Super ResolutionabstractNon-line-of-sight (NLOS) imaging techniques use light that diffusely reflects off of visible surfaces (e.g., walls) to see around corners. It has many potential applications, including autonomous driving, search and rescue, and medical imaging. The efficient utilization of NLOS imaging in these applications requires quick measurement and imaging. So far, unlike traditional 2D raster scanning, circular and confocal non-line-of-sight ($\mathrm{C}^{2} \mathrm{NLOS}$) utilizes 1D circular scanning for faster measurements and memory-efficient imaging. However, this technique’s limitation to circular measurements causes spatial bias in the gathered data, affecting the reconstruction result quality. To address this issue, we propose introducing learning-based image enhancement to $\mathrm{C}^{2}$ NLOS. To this end, we propose E-C ${ }^{2} \mathrm{NLOS}$ (Enhanced $\mathrm{C}^{2} \mathrm{NLOS}$), a framework in which we directly enhance the measured information. In our framework, we propose the Cropping module and Stitching module for faster and more effective image processing. Furthermore, considering that existing learning-based image enhancement methods are trained only on natural images, our work involves generating a synthetic dataset specifically designed for transient images, which is used for fine-tuning. Our experiments suggest that the framework works effectively, and our method outperforms the baseline method. Dixin Yang, Mariko Isogawa |
ICIP | 2 |
| 2023 | Listening Human Behavior: 3D Human Pose Estimation with Acoustic SignalsabstractGiven only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available. Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki |
CVPR | 3 |
| 2023 | Adaptive and Robust Mmwave-Based 3D Human Mesh Estimation for Diverse PosesabstractThis paper proposes a three-dimensional (3D) human mesh estimation framework with only a single commercial portable millimeter-wave device. Perceiving a 3D human mesh that includes poses and body shapes of a person with such a simple setting has remarkable potential for various applications such as daily activity monitoring and motion analysis for sports enhancement. Due to estimation difficulties, given a noisy input that includes signals reflected from the other person or objects in addition to the target person, most existing studies implicitly assume that the person stands almost vertically or that there is nothing else to be observed. Since such situations are unlikely to occur in real life, there is an urgent need for a more practical method. Therefore, we propose a framework that has the ability to extract only signals reflected from the target person and obtain local features that flexibly fit various poses that humans can take, including horizontal postures, such as lying. Our experiments suggest that the framework works effectively and that our method outperforms the baseline method. Kotaro Amaya, Mariko Isogawa |
ICIP | 2 |
| 2023 | Scapegoat Generation for Privacy Protection from DeepfakeabstractTo protect privacy and prevent malicious use of deepfake, current studies propose methods that interfere with the generation process, such as detection and destruction approaches. However, these methods suffer from sub-optimal generalization performance to unseen models and add undesirable noise to the original image. To address these problems, we propose a new problem formulation for deepfake prevention: generating a "scapegoat image" by modifying the style of the original input in a way that is recognizable as an avatar by the user, but impossible to reconstruct the real face. Even in the case of malicious deepfake, the privacy of the users is still protected. To achieve this, we introduce an optimization-based editing method that utilizes GAN inversion to discourage deepfake models from generating similar scapegoats. We validate the effectiveness of our proposed method through quantitative and user studies. Gido Kato, Yoshihiro Fukuhara, Mariko Isogawa, Hideki Tsunashima, Hirokatsu Kataoka, Shigeo Morishima |
ICIP | 3 |
| 2023 | High-Quality Virtual Single-Viewpoint Surgical Video: Geometric Autocalibration of Multiple Cameras in Surgical Lights
Yuna Kato, Mariko Isogawa, Shohei Mori, Hideo Saito 0001, Hiroki Kajita, Yoshifumi Takatsume |
MICCAI (9) | 2 |
| 2022 | Bilateral Video Magnification FilterabstractEulerian video magnification (EVM) has progressed to magnify subtle motions with a target frequency even under the presence of large motions of objects. However, existing EVM methods often fail to produce desirable results in real videos due to (1) misextracting subtle motions with a non-target frequency and (2) collapsing results when large de/acceleration motions occur (e.g., objects suddenly start, stop, or change direction). To enhance EVM performance on real videos, this paper proposes a bilateral video magnification filter (BVMF) that offers simple yet robust temporal filtering. BVMF has two kernels; (I) one kernel performs temporal bandpass filtering via a Laplacian of Gaussian whose passband peaks at the target frequency with unity gain and (II) the other kernel excludes large motions outside the magnitude of interest by Gaussian filtering on the intensity of the input signal via the Fourier shift theorem. Thus, BVMF extracts only subtle motions with the target frequency while excluding large motions outside the magnitude of interest, regardless of motion dynamics. In addition, BVMF runs the two kernels in the temporal and intensity domains simultaneously like the bilateral filter does in the spatial and intensity domains. This simplifies implementation and, as a secondary effect, keeps the memory usage low. Experiments conducted on synthetic and real videos show that BVMF outperforms state-of-the-art methods. Shoichiro Takeda, Kenta Niwa, Mariko Isogawa, Shinya Shimizu, Kazuki Okami, Yushi Aono |
CVPR | 3 |
| 2021 | Silhouette-Based Synthetic Data Generation For 3D Human Pose Estimation With A Single Wrist-Mounted 360° CameraabstractIn this paper, we propose a framework for 3D human pose estimation with a single 360° camera mounted on the user’s wrist. Perceiving a 3D human pose with such a simple setting has remarkable potential for various applications (e.g., daily-living activity monitoring, motion analysis for sports enhancement). However, no existing work has tackled this task due to the difficulty of estimating a human pose from a single camera image in which only a part of the human body is captured and the lack of training data. Therefore, we propose an effective method for translating wrist-mounted 360° camera images into 3D human poses. We also propose silhouette-based synthetic data generation dedicated to this task, which enables us to bridge the domain gap between real-world data and synthetic data. We achieved higher estimation accuracy quantitatively and qualitatively compared with other baseline methods. Ryosuke Hori, Ryo Hachiuma, Hideo Saito 0001, Mariko Isogawa, Dan Mikami |
ICIP | 4 |
| 2020 | Optical Non-Line-of-Sight Physics-Based 3D Human Pose EstimationabstractWe describe a method for 3D human pose estimation from transient images (i.e., a 3D spatio-temporal histogram of photons) acquired by an optical non-line-of-sight (NLOS) imaging system. Our method can perceive 3D human pose by 'looking around corners' through the use of light indirectly reflected by the environment. We bring together a diverse set of technologies from NLOS imaging, human pose estimation and deep reinforcement learning to construct an end-to-end data processing pipeline that converts a raw stream of photon measurements into a full 3D human pose sequence estimate. Our contributions are the design of data representation process which includes (1) a learnable inverse point spread function (PSF) to convert raw transient images into a deep feature vector; (2) a neural humanoid control policy conditioned on the transient image feature and learned from interactions with a physics simulator; and (3) a data synthesis and augmentation strategy based on depth data that can be transferred to a real-world NLOS imaging system. Our preliminary experiments suggest that our method is able to generalize to real-world NLOS measurement to estimate physically-valid 3D human poses. Mariko Isogawa, Ye Yuan 0007, Matthew O'Toole, Kris Makoto Kitani |
CVPR | 1 |
| 2020 | Efficient Non-Line-of-Sight Imaging from Transient Sinograms
Mariko Isogawa, Dorian Chan, Ye Yuan 0007, Kris Makoto Kitani, Matthew O'Toole |
ECCV (7) | 1 |
| 2019 | VR-based Batter Training System with Motion Sensing and Performance VisualizationabstractThis paper aims to establish a novel VR system for evaluating the performance of baseball batters. Existing VR systems for sports have been utilized as a tool for image training. In order to move such VR systems to the next stage, we introduce functions that sense the users' reaction to the VR stimulus. Our VR system has three features; (a) it synthesizes highly realistic VR video from the data captured in actual games, (b) it estimates the reaction of the user to the VR stimulus by capturing the 3D positions of full body parts, and (c) it consists of off-the-shelf devices and is easy to use. Our demonstration provides users with a chance to experience our VR system and give them some quick feedback by visualizing the estimated 3D positions of their body parts. Kosuke Takahashi, Dan Mikami, Mariko Isogawa, Yoshinori Kusachi, Naoki Saijo |
VR | 3 |
| 2019 | Which is the Better Inpainted Image?Training Data Generation Without Any Manual OperationsabstractThis paper proposes a learning-based quality evaluation framework for inpainted results that does not require any subjectively annotated training data. Image inpainting, which removes and restores unwanted regions in images, is widely acknowledged as a task whose results are quite difficult to evaluate objectively. Thus, existing learning-based image quality assessment (IQA) methods for inpainting require subjectively annotated data for training. However, subjective annotation requires huge cost and subjects’ judgment occasionally differs from person to person in accordance with the judgment criteria. To overcome these difficulties, the proposed framework generates and uses simulated failure results of inpainted images whose subjective qualities are controlled as the training data. We also propose a masking method for generating training data towards fully automated training data generation. These approaches make it possible to successfully estimate better inpainted images, even though the task is quite subjective. To demonstrate the effectiveness of our approach, we test our algorithm with various datasets and show it outperforms existing IQA methods for inpainting. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Daisuke Iwai, Kosuke Sato, Hideaki Kimata |
Int. J. Comput. Vis. | 1 |
| 2019 | Image quality assessment for inpainted images via learning to rankabstractThis paper proposes an image quality assessment (IQA) method for image inpainting, aiming at selecting the best one from a plurality of results. It is known that inpainting results vary largely with the method used for inpainting and the parameters set. Thus, in a typical use case, users need to manually select the inpainting method and the parameters that yield the best result. This manual selection takes a great deal of time and thus there is a great need for a way to automatically estimate the best result. Unlike existing IQA methods for inpainting, our method solves this problem as a learning-based ordering task between inpainted images. This approach makes it possible to introduce auto-generated training sets for more effective learning, which has been difficult for existing methods because judging inpainting quality is quite subjective. Our method focuses on the following three points: (1) the problem can be divided into a set of “pairwise preference order estimation” elemental problems, (2) this pairwise ordering approach enables a training set to be generated automatically, and (3) effective feature design is enabled by investigating actually measured human gazes for order estimation. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Hideaki Kimata |
Multim. Tools Appl. | 1 |
| 2018 | What Can VR Systems Tell Sports Players? Reaction-Based Analysis of Baseball Batters in Virtual and Real WorldsabstractThis study aims at ascertaining the applicability of a virtual reality environment (VRE) to sports training. Hitting an incoming object is one of the most common actions in various ball games, in which players are required to move to a suitable position and hit the object in a split second; this is a complicated task requiring spatio-temporal reaction to the object. Due to this complexity, how a VRE can serve as a training environment still remains an open question. In the work reported in this paper, we investigated the idea of substituting a VRE for an actual environment for training on the task of hitting a baseball. By focusing on the batter's temporal behavior with real and virtual environments, we clarified factors that contribute to the batter's reaction. This helped us understand how training VREs can be effectively utilized and the VRE requirements needed for sports training. Mariko Isogawa, Dan Mikami, Takehiro Fukuda, Naoki Saijo, Kosuke Takahashi, Hideaki Kimata, Makio Kashino |
VR | 1 |
| 2018 | Extrinsic Camera Calibration Without Visible Corresponding Points Using Omnidirectional CamerasabstractThis paper proposes a novel algorithm that calibrates multiple cameras scattered across a broad area. The key idea of the proposed method is “using the position of an omnidirectional camera as a reference point.” The common approach to calibrating multiple cameras assumes that the cameras capture at least some common points. This means calibration becomes quite difficult if there are no shared points in each camera's field of view (FOV). The proposed method uses the position of an omnidirectional camera to determine point correspondence. The position of an omnidirectional camera relative to the calibrated camera is estimated by the theory of epipolar geometry, even if the omnidirectional camera is placed outside the camera's FOV. This property makes our method applicable to multiple cameras scattered across a broad area. Qualitative and quantitative evaluations using synthesized and real data, e.g., a sports field, demonstrate the advantages of the proposed method. Shogo Miyata, Hideo Saito 0001, Kosuke Takahashi, Dan Mikami, Mariko Isogawa, Akira Kojima |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Which is the better inpainted image? Learning without subjective annotation
Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Hideaki Kimata |
BMVC | 1 |
| 2017 | Image and video completion via feature reduction and compensationabstractThis paper proposes a novel framework for image and video completion that removes and restores unwanted regions inside them. Most existing works fail to carry out the completion processing when similar regions do not exist in undamaged regions. To overcome this, our approach creates similar regions by projecting a low dimensional space from the original space. The approach comprises three stages. First, input images/videos are converted to a lower dimensional feature space. Second, a damaged region is restored in the converted feature space. Finally, inverse conversion is performed from the lower dimensional space to the original space. This generates two advantages: (1) it enhances the possibility of applying patches dissimilar to those in the original color space and (2) it enables the use of many existing restoration methods, each having various advantages, because the feature space for retrieving the similar patches is the only extension. The framework’s effectiveness was verified in experiments using various methods, the feature space for restoration in the second stage, and inverse conversion methods. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Akira Kojima |
Multim. Tools Appl. | 1 |
| 2016 | Eye gaze analysis and learning-to-rank to obtain the most preferred result in image inpaintingabstractThis paper proposes a method that blindly predicts preference order between inpainted images, aiming at selecting the best one from a plurality of results. Image inpainting, which removes unwanted regions and restores them, has attracted recent attention. However, it is known that the inpainting result varies largely with the method used for inpainting and the parameters set. Thus, in a typical use case, users need to manually select the inpainting method and the parameter that yields the best one. This manual selection takes a great deal of time and thus there is a great need for a way to automatically estimate the best result. Although some methods, such as estimating perceptual preference score from image features, have been proposed in recent years, none of them are considered very promising approaches. Our method focuses on the following two points: (1) what we essentially need is a preference order relation rather than an absolute score, and (2) we consider that image features for order estimation can be effectively designed by using actually measured human visual attention. Comparison with other image quality assessment methods shows that our method estimates the preference order with high accuracy. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Akira Kojima |
ICIP | 1 |
| 2015 | Content Completion in Lower Dimensional Feature Space through Feature Reduction and CompensationabstractA novel framework for image/video content completion comprising three stages is proposed. First, input images/videos are converted to a lower dimensional feature space, which is done to achieve effective restoration even in cases where a damaged region includes complex structures and changes in color. Second, a damaged region is restored in the converted feature space. Finally, an inverse conversion from the lower dimensional feature space to the original feature space is performed to generate the completed image in the original feature space. This three-step solution generates two advantages. First, it enhances the possibility of applying patches dissimilar to those in the original color space. Second, it enables the use of many existing restoration methods, each having various advantages, because the feature space for retrieving the similar patches is the only extension. Experiments verify the effectiveness of the proposed framework. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Akira Kojima |
ISMAR | 1 |
| 2015 | Toward Enhancing Robustness of DSystem: Ranking Model for Background InpaintingabstractA method for blindly predicting inpainted image quality is proposed for enhancing the robustness of diminished reality (DR), which uses inpainting to remove unwanted objects by replacing them with background textures in real time. The method maps from inpainted image features to subjective image quality scores without the need for reference images. It enables more complex background textures to be applied to DR. Mariko Isogawa, Dan Mikami, Kosuke Takahashi, Akira Kojima |
ISMAR | 1 |
| 2015 | Automatic Visual Feedback from Multiple Views for Motor LearningabstractA system providing visual feedback of a trainee's motions for effectively enhancing motor learning is presented. It provides feedback in synchronization with a reference motion from multiple view angles automatically with only a few seconds delay. Because the feedback is provided automatically, a trainee can obtain it without performing any operations while the memory of the motion is still clear. By employing features with low computational cost, the system achieves synchronized video feedback with four cameras connected to a consumer tablet PC. Dan Mikami, Mariko Isogawa, Kosuke Takahashi, Akira Kojima |
ISMAR | 2 |
| 2014 | Making Graphical Information Visible in Real Shadows on Interactive TabletopsabstractWe introduce a shadow-based interface for interactive tabletops. The proposed interface allows a user to browse graphical information by casting the shadow of his/her body, such as a hand, on a tabletop surface. Central to our technique is a new optical design that utilizes polarization in addition to the additive nature of light so that the desired graphical information is displayed only in a shadow area on a tabletop surface. In other words, our technique conceals the graphical information on surfaces other than the shadow area, such as the surface of the occluder and non-shadow areas on the tabletop surface. We combine the proposed shadow-based interface with a multi-touch detection technique to realize a novel interaction technique for interactive tabletops. We implemented a prototype system and conducted proof-of-concept experiments along with a quantitative evaluation to assess the feasibility of the proposed optical design. Finally, we showed implemented application systems of the proposed shadow-based interface. Mariko Isogawa, Daisuke Iwai, Kosuke Sato |
IEEE Trans. Vis. Comput. Graph. | 1 |