EDBT 2026 Demo / reviewers in the wild / expert
Photchara Ratsamee
dblp:14/11052
· DBLP profile ↗
19ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-3081-2232ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSystems, architecture and hardware · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HybridSphere: Enhancing Hybrid Meetings with Avatar-Based VR EnvironmentsabstractWith recent advances in information and communication technologies, Hybrid meetings, where local attendees are physically present and remote participants join virtually, have become increasingly common. However, remote participants often experience reduced contextual awareness and a sense of isolation. To address these issues, we propose HybridSphere, a hybrid meeting system that reconstructs a shared virtual reality environment for remote participants. The system employs a 360-degree camera and pose estimation to generate real-time avatar representations of local attendees to allow remote users with head-mounted displays to experience the meeting as if all participants are in the same virtual space. We conducted a user study comparing HybridSphere to a baseline condition in which remote participants viewed an unmodified 360-degree video. Although the avatars were rated lower in perceived trustworthiness and likability due to limited visual fidelity, participants appreciated the seated-6DoF function. These results suggest that immersive and spatially flexible VR representations can enhance remote engagement in hybrid meetings. Koji Momota, Shizuka Shirai, Masato Kobayashi 0001, Naoya Chiba, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Detective Networks: Enhancing Disaster Recognition in Images Through Attention Shifting Using Optimal MaskingabstractAerial investigation is used for surveying damage and identifying post-disaster events through imagery data. However, the challenge lies in detecting disaster-related areas within aerial or shipborne images, as these can appear as minor regions, making recognition difficult. To address this challenge, we introduce the Detective Network (DeNet), designed to optimally mask images, thereby shifting the attention of machine learning models towards these small yet crucial regions. Utilizing the concepts of patch and anchor box, DeNet incorporates a masking candidate layer and a masking layer to facilitate optimal masking. Our experimental findings are compelling; by preprocessing images with DeNet before analysis using an image captioning model, we achieved a remarkable accuracy of 92.91% in landslide detection from side-view image captions and 87.50% for shipborne view detection. The result demonstrates the efficacy of DeNet in enhancing the recognition of disaster-related areas in challenging imaging conditions. Narongthat Thanyawet, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura |
WACV | 2 |
| 2025 | User-Centric Locomotion Techniques for Virtual Reality Games: A Survey of User Needs and IssuesabstractVirtual reality (VR) video games that are played on a VR headset are becoming increasingly common in households, and though many games require players to navigate vast virtual spaces, most homes cannot provide a large enough physical space to encompass the entire virtual space. Thus, VR video games that require locomotion often provide users with alternative locomotion techniques. While teleportation or steering is typically used as a standard, new techniques can overcome remaining problems, such as motion sickness. However, a holistic perspective of user needs and issues regarding these techniques in practical situations has not been studied on a broad basis. To address this gap in the literature and contribute to future VR video game development and research, we conducted 16 semi-structured interviews and surveyed 88 participants to help explore issues regarding existing locomotion techniques. Our results revealed preferences related to teleportation versus steering and the postures that users adopt while playing VR video games, along with user needs for locomotion techniques in each posture. Daichi Hirobe, Shizuka Shirai, Jason Orlosky, Mehrasa Alizadeh, Masato Kobayashi 0001, Yuuki Uranishi, Photchara Ratsamee, Haruo Takemura |
IEEE Trans. Games | 7 |
| 2024 | Solving Multi-Robot Task Allocation and Planning in Trans-media ScenariosabstractTrans-media robots, capable of operating across diverse environments, add significant complexity for multi-robot task allocation and planning problems. This paper introduces a novel approach to plan missions for such multi-robot systems, that addresses the associated specific complexities and constraints. It streamlines the overall mission planning process by decomposing it into tractable sub-problems, and addresses the issues of coalition formation, path planning, and task scheduling. It provides mission plans in very little computation time and allows to tackle large missions intractable by global planners, with negligible loss in plan optimality. Virgile De La Rochefoucauld, Simon Lacroix, Photchara Ratsamee, Haruo Takemura |
IROS | 3 |
| 2024 | Panoptic-Level Image-to-Image Translation for Object Recognition and Visual Odometry EnhancementabstractImage-to-image translation methods have progressed from only considering the image-level information to integrating the global- and instance-level information. However, only the foreground instances are refined, and the background semantics are taken as an entire feature, which causes a substantial loss of the semantic information in the translation. Additionally, the insufficient quality of the translated semantic regions also leads to an unsatisfactory performance of the object recognition or visual odometry tasks in which the translated images/videos are further used. In this paper, we propose a novel generative adversarial network for panoptic-level image-to-image translation (PanopticGAN). The proposed method has three advantages: 1) the extracted panoptic perception (i.e., the foreground instances and background semantic regions) as content codes are aligned with the sampled panoptic style codes, which considers the panoptic-level information to avoid the semantic information loss, and the latent space of each object has a rich fusion of content and style codes to generate the higher-fidelity results; 2) a feature masking module is proposed to extract the representations within each object contour by masks for sharpening the object boundaries; 3) the improved fidelity of the translated semantic regions further contributes to enhancing the performance of the object recognition or visual odometry tasks that the translated images/videos are used in. In this paper, we also annotate a compact panoptic segmentation dataset for the thermal-to-color translation task. Extensive experiments are conducted to demonstrate the effectiveness of our PanopticGAN over the latest methods. Photchara Ratsamee, Zhaojie Luo, Yuuki Uranishi, Manabu Higashida, Haruo Takemura |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Panoptic-aware Image-to-Image TranslationabstractDespite remarkable progress in image translation, the complex scene with multiple discrepant objects remains a challenging problem. The translated images have low fidelity and tiny objects in fewer details causing unsatisfactory performance in object recognition. Without thorough object perception (i.e., bounding boxes, categories, and masks) of images as prior knowledge, the style transformation of each object will be difficult to track in translation. We propose panoptic-aware generative adversarial networks (PanopticGAN) for image-to-image translation together with a compact panoptic segmentation dataset. The panoptic perception (i.e., foreground instances and background semantics of the image scene) is extracted to achieve alignment between object content codes of the input domain and panoptic-level style codes sampled from the target style space, then refined by a proposed feature masking module for sharping object boundaries. The image-level combination between content and sampled style codes is also merged for higher fidelity image generation. Our proposed method was systematically compared with different competing methods and obtained significant improvement in both image quality and object recognition performance. Photchara Ratsamee, Bowen Wang 0002, Zhaojie Luo, Yuuki Uranishi, Manabu Higashida, Haruo Takemura |
WACV | 2 |
| 2023 | Mitigation of VR Sickness During Locomotion With a Motion-Based Dynamic Vision ModulatorabstractIn virtual reality, VR sickness resulting from continuous locomotion via controllers or joysticks is still a significant problem. In this article, we present a set of algorithms to mitigate VR sickness that dynamically modulate the user's field of view by modifying the contrast of the periphery based on movement, color, and depth. In contrast with previous work, this vision modulator is a shader that is triggered by specific motions known to cause VR sickness, such as acceleration, strafing, and linear velocity. Moreover, the algorithm is governed by delta velocity, delta angle, and average color of the view. We ran two experiments with different washout periods to investigate the effectiveness of dynamic modulation on the symptoms of VR sickness, in which we compared this approach against a baseline and pitch-black field-of-view restrictors. Our first experiment made use of a just-noticeable-sickness design, which can be useful for building experiments with a short washout period. Guanghan Zhao, Jason Orlosky, Steven K. Feiner, Photchara Ratsamee, Yuuki Uranishi |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | UAV Target-Selection: 3D Pointing Interface System for Large-Scale EnvironmentabstractThis paper presents a 3D pointing interface application to signal a UAV’s target in a large-scale environment. This system enables UAVs equipped with a monocular camera to determine which window of a building is selected by a human user in large-scale indoor or outdoor environments. The 3D pointing interface consists of three parts: YOLO, Open- Pose, and ORB-SLAM. YOLO detects the target objects, e.g., windows, OpenPose extracts the user pose, and ORB-SLAM builds a scale-dependent 3D map, a set of 3D sparse feature points. To obtain the visual scale, it performs a calibration step with the user standing in front of the UAV at a certain distance. We detail how we chose the gesture, localize and detect objects, and transform between coordinate systems. The real- world experiment results showed that the 3D pointing interface obtained a 0.73 F1-score average and a 0.58 F1-Score at the maximum distance of 25 meters between UAV and building. Anna Medeiros, Photchara Ratsamee, Jason Orlosky, Yuuki Uranishi, Manabu Higashida, Haruo Takemura |
ICRA | 2 |
| 2021 | Spherical Magnetic Joint for Inverted Locomotion of Multi-Legged RobotabstractIn this paper, we present a spherical magnetic joint for the inverted locomotion of a multi-legged robot. The permanent magnet’s spherical shape allows the robot to attach its foot to a steel surface without energy consumption. However, the robot’s inverted locomotion requires foot flexibility for placement and gait construction of the robot. Therefore, the spherical magnetic joint mechanism was designed and implemented for the robot feet to deal with angular placement. For decoupling the foot from the steel surface, the attractive force is adjusted by tilting the adjustable sleeve mechanism at an adequate angle between the surface and foot tip. Experimental results show that the spherical magnetic joint can maintain the attractive force at any angle, and the sleeve mechanism can reduce 20% of the reaction force for pulling the legs from the steel surfaces. Furthermore, the designed gait for inverted locomotion with a spherical magnetic joint was tested and compared to prove the concept of the spherical magnetic joint and sleeve mechanism. Harn Sison, Photchara Ratsamee, Manabu Higashida, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura |
ICRA | 2 |
| 2019 | Evaluation of Pointing Interfaces with an AR Agent for Multi-section Information GuidanceabstractIn educational settings such as art galleries or museums, Augmented Reality (AR) has the potential to provide detailed information about exhibits. However, dealing with items that contain information in multiple sections or areas is still a significant challenge. For example, a large painting may contain many minute details, which requires a system that can explain its broader features rather than just a generic description. To address this challenge, we introduce an AR guidance system that uses an embodied agent to point out items and explain each piece and part of exhibit items in detail. We also designed and tested 3 different pointing interfaces for the embodied agent: gesture only, gesture with a dot laser, and gesture with line laser. To evaluate this interface, we conducted a user experiment simulating painting guidance to test interest and exhibit memory. During the experiment, the agent pointed to various areas of interest in the painting and provided a detailed description to participants. The result shows that the search times for target positions were the fastest with the line laser. However, no particular interface outperformed others in memory recall of exhibit content. Nattaon Techasarntikul, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura, Jason Orlosky, Kiyoshi Kiyokawa |
VR | 3 |
| 2019 | A Comparison of Adaptive View Techniques for Exploratory 3D Drone TeleoperationabstractDrone navigation in complex environments poses many problems to teleoperators. Especially in three dimensional (3D) structures such as buildings or tunnels, viewpoints are often limited to the drone’s current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and provide smooth operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D point-cloud information into account to modify the user’s viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera, and we use the resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first- and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation. John Thomason, Photchara Ratsamee, Jason Orlosky, Kiyoshi Kiyokawa, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2018 | IntelliPupil: Pupillometric Light Modulation for Optical See-Through Head-Mounted DisplaysabstractIn practical use of optical see-through head-mounted displays, users often have to adjust the brightness of virtual content to ensure that it is at the optimal level. Automatic adjustment is still a challenging problem, largely due to the bidirectional nature of the structure of the human eye, complexity of real world lighting, and user perception. Allowing the right amount of light to pass through to the retina requires a constant balance of incoming light from the real world, additional light from the virtual image, pupil contraction, and feedback from the user. While some automatic light adjustment methods exist, none have completely tackled this complex input-output system. As a step towards overcoming this issue, we introduce IntelliPupil, an approach that uses eye tracking to properly modulate augmentation lighting for a variety of lighting conditions and real scenes. We first take the data from a small form factor light sensor and changes in pupil diameter from an eye tracking camera as passive inputs. This data is coupled with user-controlled brightness selections, allowing us to fit a brightness model to user preference using a feed-forward neural network. Using a small amount of training data, both scene luminance and pupil size are used as inputs into the neural network, which can then automatically adjust to a user's personal brightness preferences in real time. Experiments in a high dynamic range AR scenario with varied lighting show that pupil size is just as important as environment light for optimizing brightness and that our system outperforms linear models. Chang Liu 0081, Alexander Plopski, Kiyoshi Kiyokawa, Photchara Ratsamee, Jason Orlosky |
ISMAR | 4 |
| 2017 | Social Drone Companion for the Home Environment: a User-Centric ExplorationabstractRecent research has focused on how to facilitate interaction between humans and robots, giving rise to the field of human robot interaction. A related research area is human-drone interaction (HDI), investigating how interaction between humans and drones can be expanded in novel and meaningful ways. In this work, we explore the use of drones as companions in a home environment. We present three consecutive studies addressing user requirements and design space of companion drones. Following a user-centered approach, the three stages include online questionnaire, design workshops, and simulated virtual reality (VR) home environment. Our results show that participants preferred the idea of a drone companion at home, particularly for tasks such as fetching items and cleaning. The participants were also positive towards a drone companion that featured anthropomorphic features. Kari Daniel Karjalainen, Anna Elisabeth Sofia Romell, Photchara Ratsamee, Asim Evren Yantaç, Morten Fjeld, Mohammad Obaid |
HAI | 3 |
| 2017 | Exploring Proxemics for Human-Drone InteractionabstractWe present a human-centered designed social drone aiming to be used in a human crowd environment. Based on design studies and focus groups, we created a prototype of a social drone with a social shape, face and voice for human interaction. We used the prototype for a proxemic study, comparing the required distance from the drone humans could comfortably accept compared with what they would require for a nonsocial drone. The social shaped design with greeting voice added decreased the acceptable distance markedly, as did present or previous pet ownership, and maleness. We also explored the proximity sphere around humans with a social shaped drone based on a validation study with variation of lateral distance and heights. Both lateral distance and the higher height of 1.8 m compared to the lower height of 1.2 m decreased the required comfortable distance as it approached. Alexander Yeh, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi, Tomohiro Mashita, Haruo Takemura, Morten Fjeld, Mohammad Obaid |
HAI | 2 |
| 2017 | VisMerge: Light Adaptive Vision Augmentation via Spectral and Temporal Fusion of Non-visible LightabstractLow light situations pose a significant challenge to individuals working in a variety of different fields such as firefighting, rescue, maintenance and medicine. Tools like flashlights and infrared (IR) cameras have been used to augment light in the past, but they must often be operated manually, provide a field of view that is decoupled from the operator's own view, and utilize color schemes that can occlude content from the original scene. To help address these issues, we present VisMerge, a framework that combines a thermal imaging head mounted display (HMD) and algorithms that temporally and spectrally merge video streams of different light bands into the same field of view. For temporal synchronization, we first develop a variant of the time warping algorithm used in virtual reality (VR), but redesign it to merge video see-through (VST) cameras with different latencies. Next, using computer vision and image compositing we develop five new algorithms designed to merge non-uniform video streams from a standard RGB camera and small form-factor infrared (IR) camera. We then implement six other existing fusion methods, and conduct a series of comparative experiments, including a system level analysis of the augmented reality (AR) time warping algorithm, a pilot experiment to test perceptual consistency across all eleven merging algorithms, and an in-depth experiment on performance testing the top algorithms in a VR (simulated AR) search task. Results showed that we can reduce temporal registration error due to inter-camera latency by an average of 87.04%, that the wavelet and inverse stipple algorithms were perceptually rated the highest, that noise modulation performed best, and that freedom of user movement is significantly increased with visualizations engaged. Jason Orlosky, Peter Kim, Kiyoshi Kiyokawa, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura |
ISMAR | 5 |
| 2017 | Adaptive View Management for Drone Teleoperation in Complex 3D StructuresabstractDrone navigation in complex environments poses many problems to teleoperators. Especially in 3D structures like buildings or tunnels, viewpoints are often limited to the drone's current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and smooth user operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D pointcloud information into account to modify user-viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera and we use resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation. John Thomason, Photchara Ratsamee, Kiyoshi Kiyokawa, Pakpoom Kriengkomol, Jason Orlosky, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura |
IUI | 2 |
| 2013 | Lifelogging keyframe selection using image quality measurements and physiological excitement featuresabstractKeyframe selection is the process of finding a representative frame in an image sequence. Although mostly known from video processing, keyframe selection faces new challenges in the lifelog domain. To obtain a keyframe that is close to a user-selected frame, we propose a keyframe selection method based on image quality measurements and excitement features. Image quality measurements such as contrast, color variance, sharpness, noise and saliency are used to filter high quality images. However, high quality images are not necessarily keyframes because humans also use emotions in the selection process. In this study, we employ a biosensor to measure the excitement of humans. In previous investigation, keyframe selection using only image quality measurements yielded an acceptance rate of 79.70%. Our proposed method achieves an acceptance rate of 84.45%. Photchara Ratsamee, Yasushi Mae, Amornched Jinda-Apiraksa, Jana Machajdik, Kenichi Ohara, Masaru Kojima, Robert Sablatnig, Tatsuo Arai |
IROS | 1 |
| 2013 | Social navigation model based on human intention analysis using face orientationabstractWe propose a social navigation model that allows a robot to navigate in a human environment according to human intentions, in particular during a situation where the human encounters a robot and he/she wants to avoid, unavoid (maintain his/her course), or approach the robot. Avoiding, unavoiding, and approaching trajectories of humans are classified based on the face orientation on a social force model and their predicted motion. The proposed model is developed based on human motion and behavior (especially face orientation and overlapping personal space) analysis in preliminary experiments. Our experimental evidence demonstrates that the robot is able to adapt its motion by preserving personal distance from passers-by, and approaching persons who want to interact with the robot. This work contributes to the future development of a human-robot socialization environment. Photchara Ratsamee, Yasushi Mae, Kenichi Ohara, Masaru Kojima, Tatsuo Arai |
IROS | 1 |
| 2012 | Modified social force model with face pose for human collision avoidanceabstractIn order for robots to be a part of human society, their social accceptance is an important issue if smooth interaction with humans is to be achieved. We propose a modified social force model that allows robots to move naturally like humans, based on estimated human motion and face pose. We add to the previous model the effect of the force due to face pose, in order to predict human motion and compute the robot motion itself. Our approach was implemented and tested on a real humanoid robot in a situation in which a human is confronted with a robot in an indoor environment. Experimental results illustrate that the robot is able to perform human-like navigation by avoiding the human in a face-to-face confrontation. Our system provides accurate face pose tracking that allows a robot to have a more realistic behaviour compared to the original social force model. Photchara Ratsamee, Yasushi Mae, Kenichi Ohara, Tomohito Takubo, Tatsuo Arai |
HRI | 1 |