VLDB 2026 Research / reviewers in the wild / expert
Ryo Yonetani
dblp:29/8618
· DBLP profile ↗
36ranked-venue papers
12as first author
14since 2021 · last 2025
0000-0002-2724-6233ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 2 since 2021Systems, architecture and hardware · 10 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language GuidanceabstractThis paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IMU sensor noise that causes trajectory drift over time. The diversity of human actions further complicates IMU signal processing by introducing various motion patterns. Nevertheless, we observe that some actions captured by the head-mounted IMU correlate with spatial environmental structures (e.g., bending down to look inside an oven, washing dishes next to a sink), thereby serving as spatial anchors to compensate for the localization drift. The proposed EAIL framework learns such correlations via hierarchical multi-modal alignment with vision-language guidance. By assuming that the 3D point cloud of the environment is available, it contrastively learns modality encoders that align short-term egocentric action cues in IMU signals with local environmental features in the point cloud. The learning process is enhanced using concurrently collected vision and language signals to improve multimodal alignment. The learned encoders are then used in reasoning the IMU data and the point cloud over time and space to perform inertial localization. Interestingly, these encoders can further be utilized to recognize the corresponding sequence of actions as a by-product. Extensive experiments demonstrate the effectiveness of the proposed framework over state-of-the-art inertial localization and inertial action recognition baselines. Mingfang Zhang 0002, Ryo Yonetani, Yifei Huang 0002, Liangyang Ouyang, Ruicong Liu, Yoichi Sato 0001 |
ICCV | 2 |
| 2025 | Path Planning Using Instruction-Guided Probabilistic RoadmapsabstractThis work presents a novel data-driven path planning algorithm named Instruction-Guided Probabilistic Roadmap (IG-PRM). Despite the recent development and widespread use of mobile robot navigation, the safe and effective travels of mobile robots still require significant engineering effort to take into account the constraints of robots and their tasks. With IG-PRM, we aim to address this problem by allowing robot operators to specify such constraints through natural language instructions, such as “aim for wider paths” or “mind small gaps”. The key idea is to convert such instructions into embedding vectors using large-language models (LLMs) and use the vectors as a condition to predict instruction-guided cost maps from occupancy maps. By constructing a roadmap based on the predicted costs, we can find instruction-guided paths via the standard shortest path search. Experimental results demonstrate the effectiveness of our approach on both synthetic and real-world indoor navigation environments. Jiaqi Bao, Ryo Yonetani |
ICRA | 2 |
| 2025 | TSPDiffuser: Diffusion Models as Learned Samplers for Traveling Salesperson Path Planning ProblemsabstractThis paper presents TSPDiffuser, a novel data-driven path planner for traveling salesperson path planning problems (TSPPPs) in environments rich with obstacles. Given a set of destinations within obstacle maps, our objective is to efficiently find the shortest possible collision-free path that visits all the destinations. In TSPDiffuser, we train a diffusion model on a large collection of TSPPP instances and their respective solutions to generate plausible paths for unseen problem instances. The model can then be employed as a learned sampler to construct a roadmap that contains potential solutions with a small number of nodes and edges. This approach enables efficient and accurate estimation of travel costs between destinations, effectively addressing the primary computational challenge in solving TSPPPs. Experimental evaluations with diverse synthetic and real-world indoor/outdoor environments demonstrate the effectiveness of TSPDiffuser over existing methods in terms of the trade-off between solution quality and computational time requirements. Ryo Yonetani |
ICRA | 1 |
| 2025 | GSplatVNM: Point-of-View Synthesis for Visual Navigation Models Using Gaussian SplattingabstractThis paper presents a novel approach to image-goal navigation by integrating 3D Gaussian Splatting (3DGS) with Visual Navigation Models (VNMs), a method we refer to as GSplatVNM. VNMs offer a promising paradigm for image-goal navigation by guiding a robot through a sequence of point-of-view images without requiring metrical localization or environment-specific training. However, constructing a dense and traversable sequence of target viewpoints from start to goal remains a central challenge, particularly when the available image database is sparse. To address these challenges, we propose a 3DGS-based viewpoint synthesis framework for VNMs that synthesizes intermediate viewpoints to seamlessly bridge gaps in sparse data while significantly reducing storage overhead. Experimental results in a photorealistic simulator demonstrate that our approach not only enhances navigation efficiency but also exhibits robustness under varying levels of image database sparsity. Kohei Honda 0002, Takeshi Ishita, Yasuhiro Yoshimura, Ryo Yonetani |
IROS | 4 |
| 2025 | Opt-in Camera: Person Identification in Video via UWB Localization and Its Application to Opt-in SystemsabstractThis paper presents opt-in camera, a concept of privacy-preserving camera systems capable of recording only specific individuals in a crowd who explicitly consent to be recorded. Our system utilizes a mobile wireless communication tag attached to personal belongings as proof of opt-in and as a means of localizing tag carriers in video footage. Specifically, the on-ground positions of the wireless tag are first tracked over time using the unscented Kalman filter (UKF). The tag trajectory is then matched against visual tracking results for pedestrians found in videos to identify the tag carrier. Technically, we devise a dedicated trajectory matching technique based on constrained linear optimization, as well as a novel calibration technique that handles wireless tag-camera calibration and hyperparameter tuning for the UKF, which mitigates the non-lineof-sight (NLoS) issue in wireless localization. We implemented the proposed opt-in camera system using ultra-wideband (UWB) devices and an off-the-shelf webcam. Experimental results demonstrate that our system can perform opt-in recording of individuals in real-time at 10 fps, with reliable identification accuracy in crowds of 8–23 people in a confined space. Matthew Ishige, Yasuhiro Yoshimura, Ryo Yonetani |
IROS | 3 |
| 2024 | CXSimulator: A User Behavior Simulation using LLM Embeddings for Web-Marketing Campaign AssessmentabstractThis paper presents the Customer Experience (CX) Simulator, a novel framework designed to assess the effects of untested web-marketing campaigns through user behavior simulations. The proposed framework leverages large language models (LLMs) to represent various events in a user's behavioral history, such as viewing an item, applying a coupon, or purchasing an item, as semantic embedding vectors. We train a model to predict transitions between events from their LLM embeddings, which can even generalize to unseen events by learning from diverse training data. In web-marketing applications, we leverage this transition prediction model to simulate how users might react differently when new campaigns or products are presented to them. This allows us to eliminate the need for costly online testing and enhance the marketers' abilities to reveal insights. Our numerical evaluation and user study, utilizing BigQuery Public Datasets from the Google Merchandise Store, demonstrate the effectiveness of our framework. Akira Kasuga, Ryo Yonetani |
CIKM | 2 |
| 2024 | When to Replan? An Adaptive Replanning Strategy for Autonomous Navigation using Deep Reinforcement LearningabstractThe hierarchy of global and local planners is one of the most commonly utilized system designs in autonomous robot navigation. While the global planner generates a reference path from the current to goal locations based on the pre-built map, the local planner produces a kinodynamic trajectory to follow the reference path while avoiding perceived obstacles. To account for unforeseen or dynamic obstacles not present on the pre-built map, "when to replan" the reference path is critical for the success of safe and efficient navigation. However, determining the ideal timing to execute replanning in such partially unknown environments still remains an open question. In this work, we first conduct an extensive simulation experiment to compare several common replanning strategies and confirm that effective strategies are highly dependent on the environment as well as the global and local planners. Based on this insight, we then derive a new adaptive replanning strategy based on deep reinforcement learning, which can learn from experience to decide appropriate replanning timings in the given environment and planning setups. Our experimental results show that the proposed replanner can perform on par or even better than the current best-performing strategies in multiple situations regarding navigation robustness and efficiency. Kohei Honda 0002, Ryo Yonetani, Mai Nishimura, Tadashi Kozuno |
ICRA | 2 |
| 2024 | Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement TrajectoriesabstractThis paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores.Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management.The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map.Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects.Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics. Hikaru Asano, Ryo Yonetani, Taiki Sekii, Hiroki Ouchi |
INLG | 2 |
| 2023 | Periodic Multi-Agent Path PlanningabstractMulti-agent path planning (MAPP) is the problem of planning collision-free trajectories from start to goal locations for a team of agents. This work explores a relatively unexplored setting of MAPP where streams of agents have to go through the starts and goals with high throughput. We tackle this problem by formulating a new variant of MAPP called periodic MAPP in which the timing of agent appearances is periodic. The objective with periodic MAPP is to find a periodic plan, a set of collision-free trajectories that the agent streams can use repeatedly over periods, with periods that are as small as possible. To meet this objective, we propose a solution method that is based on constraint relaxation and optimization. We show that the periodic plans once found can be used for a more practical case in which agents in a stream can appear at random times. We confirm the effectiveness of our method compared with baseline methods in terms of throughput in several scenarios that abstract autonomous intersection management tasks. Kazumi Kasaura, Ryo Yonetani, Mai Nishimura |
AAAI | 2 |
| 2023 | Risk-aware Path Planning via Probabilistic Fusion of Traversability Prediction for Planetary Rovers on Heterogeneous TerrainsabstractMachine learning (ML) plays a crucial role in assessing traversability for autonomous rover operations on deformable terrains but suffers from inevitable prediction errors. Especially for heterogeneous terrains where the geological features vary from place to place, erroneous traversability prediction can become more apparent, increasing the risk of unrecoverable rover's wheel slip and immobilization. In this work, we propose a new path planning algorithm that explicitly accounts for such erroneous prediction. The key idea is the probabilistic fusion of distinctive ML models for terrain type classification and slip prediction into a single distribution. This gives us a multimodal slip distribution accounting for heterogeneous terrains and further allows statistical risk assessment to be applied to derive risk-aware traversing costs for path planning. Extensive simulation experiments have demonstrated that the proposed method is able to generate more feasible paths on heterogeneous terrains compared to existing methods. Masafumi Endo, Tatsunori Taniai, Ryo Yonetani, Genya Ishigami |
ICRA | 3 |
| 2021 | Path Planning using Neural A* SearchabstractWe present Neural A*, a novel data-driven search method for path planning problems. Despite the recent increasing attention to data-driven path planning, machine learning approaches to search-based planning are still challenging due to the discrete nature of search algorithms. In this work, we reformulate a canonical A* search algorithm to be differentiable and couple it with a convolutional encoder to form an end-to-end trainable neural network planner. Neural A* solves a path planning problem by encoding a problem instance to a guidance map and then performing the differentiable A* search with the guidance map. By learning to match the search results with ground-truth paths provided by experts, Neural A* can produce a path consistent with the ground truth accurately and efficiently. Our extensive experiments confirmed that Neural A* outperformed state-of-the-art data-driven planners in terms of the search optimality and efficiency trade-off. Furthermore, Neural A* successfully predicted realistic human trajectories by directly performing search-based planning on natural image inputs. Ryo Yonetani, Tatsunori Taniai, Mohammadamin Barekatain, Mai Nishimura, Asako Kanezaki |
ICML | 1 |
| 2021 | Precise Multi-Modal In-Hand Pose Estimation using Low-Precision Sensors for Robotic AssemblyabstractIn industrial assembly tasks, the in-hand pose of grasped objects needs to be known with high precision for subsequent manipulation tasks such as insertion. This problem (in-hand-pose estimation) has traditionally been addressed using visual recognition or tactile sensing. On the one hand, while visual recognition can provide efficient pose estimates, it tends to suffer from low precision due to noise, occlusions and calibration errors. On the other hand, tactile fingertip sensors can provide precise complementary information, but their low durability significantly limits their use in real-world applications. To get the best of both worlds, we propose an efficient method for in-hand pose estimation using off-the-shelf cameras and robot wrist force sensors, which requires no precise camera calibration. The key idea is to utilize visual and contact information adaptively to maximally reduce the uncertainty about the in-hand object pose in a Bayesian state estimation framework. As most of the uncertainty can be resolved from visual observations, our approach reduces the number of physical environment interactions while keeping a high pose estimation accuracy. Our experimental evaluation demonstrates that our approach can estimate object poses with sub-mm precision with an off-the-shelf camera and force-torque sensor. Felix von Drigalski, Kennosuke Hayashi, Yifei Huang 0002, Ryo Yonetani, Masashi Hamaya, Kazutoshi Tanaka, Yoshihisa Ijiri |
ICRA | 4 |
| 2021 | TRANS-AM: Transfer Learning by Aggregating Dynamics Models for Soft Robotic AssemblyabstractPractical industrial assembly scenarios often require robotic agents to adapt their skills to unseen tasks quickly. While transfer reinforcement learning (RL) could enable such quick adaptation, much prior work has to collect many samples from source environments to learn target tasks in a model-free fashion, which still lacks sample efficiency on a practical level. In this work, we develop a novel transfer RL method named TRANSfer learning by Aggregating dynamics Models (TRANS-AM). TRANS-AM is based on model-based RL (MBRL) for its high-level sample efficiency, and only requires dynamics models to be collected from source environments. Specifically, it learns to aggregate source dynamics models adaptively in an MBRL loop to better fit the state-transition dynamics of target environments and execute optimal actions there. As a case study to show the effectiveness of this proposed approach, we address a challenging contact-rich peg-in-hole task with variable hole orientations using a soft robot. Our evaluations with both simulation and real-robot experiments demonstrate that TRANS-AM enables the soft robot to accomplish target tasks with fewer episodes compared when learning the tasks from scratch. Kazutoshi Tanaka, Ryo Yonetani, Masashi Hamaya, Robert Lee, Felix von Drigalski, Yoshihisa Ijiri |
ICRA | 2 |
| 2021 | Learning Robotic Contact JugglingabstractRobotic contact juggling is a challenging task in which robots must control the movement of a ball rapidly and indirectly without holding it while keeping the ball in and sometimes out of contact with the robot’s body. In this work, we address the problem of learning such robotic contact juggling from trial and error via model-based reinforcement learning (MBRL). The key insight is that complex robot-ball interactions of the contact juggling actually consist of a small set of simple dynamics that each corresponds to a distinct interaction "primitive" such as touching and releasing the ball. Accordingly, we develop a tailored MBRL method that incrementally fits a set of simple dynamics models to the movements of a robot and a ball while also learning a switching model that can select a proper dynamics model depending on the current state and action. The learned model can then be used in an MBRL framework to seek optimal juggling control. We demonstrated the effectiveness of our approach on a simulator of contact juggling performed by a robotic arm. Kazutoshi Tanaka, Masashi Hamaya, Devwrat Joshi, Felix von Drigalski, Ryo Yonetani, Takamitsu Matsubara, Yoshihisa Ijiri |
IROS | 5 |
| 2020 | Support Strategies for Remote Guides in Assisting People with Visual Impairments for Effective Indoor NavigationabstractPeople with visual impairments often require mobility assistance of sighted guides but they are not always available. Recent technological strides have opened up new directions for sighted guidance services, assigning guides from a network of remote workers to provide real-time assistance via audio/video communication. However, little has been known regarding desirable support characteristics of remote guides or challenges experienced in guide practices without the requisite expertise. To recommend support strategies that contribute to facilitating a successful platform for remote sighted guidance, this paper presents a comparative study of the performance of trained and untrained sighted guides who are recruited for a remote scenario in assisting people with visual impairments in indoor navigation. As an outcome of this research, we provide a deeper understanding of design opportunities for HCI to scaffold requirements of remote guides, such that their collaborative efforts and environmental knowledge influence the user experience. Based on our empirical insights, we suggest to develop the expertise of remote guides through: a) preliminary guidance cooperation awareness b) guidelines for verbal description methods, and c) approaches to compensate for the lack of environmental knowledge. Rie Kamikubo, Naoya Kato, Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001 |
CHI | 4 |
| 2020 | Hybrid-FL for Wireless Networks: Cooperative Learning Mechanism Using Non-IID DataabstractThis paper proposes a cooperative mechanism for mitigating the performance degradation due to non-independent and-identically-distributed (non-IID) data in collaborative machine learning (ML), namely federated learning (FL), which trains an ML model using the rich data and computational resources of mobile clients without gathering their data to central systems. The data of mobile clients is typically non-IID owing to diversity among mobile clients' interests and usage, and FL with non-IID data could degrade the model performance. Therefore, to mitigate the degradation induced by non-IID data, we assume that a limited number (e.g., less than 1%) of clients allow their data to be uploaded to a server, and we propose a hybrid learning mechanism referred to as Hybrid-FL, wherein the server updates the model using the data gathered from the clients and aggregates the model with the models trained by clients. The HybridFL solves both client- and data-selection problems via heuristic algorithms, which try to select the optimal sets of clients who train models with their own data, clients who upload their data to the server, and data uploaded to the server. The algorithms increase the number of clients participating in FL and make more data gather in the server IID, thereby improving the prediction accuracy of the aggregated model. Evaluations, which consist of network simulations and ML experiments, demonstrate that the proposed scheme achieves a 13.5% higher classification accuracy than those of the previously proposed schemes for the non-IID case. Naoya Yoshida, Takayuki Nishio, Masahiro Morikura, Koji Yamamoto 0001, Ryo Yonetani |
ICC | 5 |
| 2020 | Adaptive Distillation for Decentralized Learning from Heterogeneous ClientsabstractThis paper addresses the problem of decentralized learning to achieve a high-performance global model by asking a group of clients to share local models pre-trained with their own data resources. We are particularly interested in a specific case where both the client model architectures and data distributions are diverse, which makes it nontrivial to adopt conventional approaches such as Federated Learning and network co-distillation. To this end, we propose a new decentralized learning method called Decentralized Learning via Adaptive Distillation (DLAD). Given a collection of client models and a large number of unlabeled distillation samples, the proposed DLAD 1) aggregates the outputs of the client models while adaptively emphasizing those with higher confidence in given distillation samples and 2) trains the global model to imitate the aggregated outputs. Our extensive experimental evaluation on multiple public datasets (MNIST, CIFAR-10, and CINIC-10) demonstrates the effectiveness of the proposed method. Ryo Yonetani |
ICPR | 2 |
| 2020 | MULTIPOLAR: Multi-Source Policy Aggregation for Transfer Reinforcement Learning between Diverse Environmental DynamicsabstractTransfer reinforcement learning (RL) aims at improving the learning efficiency of an agent by exploiting knowledge from other source agents trained on relevant tasks. However, it remains challenging to transfer knowledge between different environmental dynamics without having access to the source environments. In this work, we explore a new challenge in transfer RL, where only a set of source policies collected under diverse unknown dynamics is available for learning a target task efficiently. To address this problem, the proposed approach, MULTI-source POLicy AggRegation (MULTIPOLAR), comprises two key techniques. We learn to aggregate the actions provided by the source policies adaptively to maximize the target task performance. Meanwhile, we learn an auxiliary network that predicts residuals around the aggregated actions, which ensures the target policy's expressiveness even when some of the source policies perform poorly. We demonstrated the effectiveness of MULTIPOLAR through an extensive experimental evaluation across six simulated environments ranging from classic control problems to challenging robotics simulations, under both continuous and discrete action spaces. The demo videos and code are available on the project webpage: https://omron-sinicx.github.io/multipolar/. Mohammadamin Barekatain, Ryo Yonetani, Masashi Hamaya |
IJCAI | 2 |
| 2020 | L2B: Learning to Balance the Safety-Efficiency Trade-off in Interactive Crowd-aware Robot NavigationabstractThis work presents a deep reinforcement learning framework for interactive navigation in a crowded place. Our proposed approach, Learning to Balance (L2B) framework enables mobile robot agents to steer safely towards their destinations by avoiding collisions with a crowd, while actively clearing a path by asking nearby pedestrians to make room, if necessary, to keep their travel efficient. We observe that the safety and efficiency requirements in crowd-aware navigation have a trade-off in the presence of social dilemmas between the agent and the crowd. On the one hand, intervening in pedestrian paths too much to achieve instant efficiency will result in collapsing a natural crowd flow and may eventually put everyone, including the self, at risk of collisions. On the other hand, keeping in silence to avoid every single collision will lead to the agent's inefficient travel. With this observation, our L2B framework augments the reward function used in learning an interactive navigation policy to penalize frequent active path clearing and passive collision avoidance, which substantially improves the balance of the safety-efficiency trade-off. We evaluate our L2B framework in a challenging crowd simulation and demonstrate its superiority, in terms of both navigation success and collision rate, over a state-of-the-art navigation approach. Mai Nishimura, Ryo Yonetani |
IROS | 2 |
| 2019 | Client Selection for Federated Learning with Heterogeneous Resources in Mobile EdgeabstractWe envision a mobile edge computing (MEC) framework for machine learning (ML) technologies, which leverages distributed client data and computation resources for training high-performance ML models while preserving client privacy. Toward this future goal, this work aims to extend Federated Learning (FL), a decentralized learning framework that enables privacy-preserving training of models, to work with heterogeneous clients in a practical cellular network. The FL protocol iteratively asks random clients to download a trainable model from a server, update it with own data, and upload the updated model to the server, while asking the server to aggregate multiple client updates to further improve the model. While clients in this protocol are free from disclosing own private data, the overall training process can become inefficient when some clients are with limited computational resources (i.e., requiring longer update time) or under poor wireless channel conditions (longer upload time). Our new FL protocol, which we refer to as FedCS, mitigates this problem and performs FL efficiently while actively managing clients based on their resource conditions. Specifically, FedCS solves a client selection problem with resource constraints, which allows the server to aggregate as many client updates as possible and to accelerate performance improvement in ML models. We conducted an experimental evaluation using publicly-available large-scale image datasets to train deep neural networks on MEC environment simulations. The experimental results show that FedCS is able to complete its training process in a significantly shorter time compared to the original FL protocol. Takayuki Nishio, Ryo Yonetani |
ICC | 2 |
| 2019 | Assisting group activity analysis through hand detection and identification in multiple egocentric videosabstractResearch in group activity analysis has put attention to monitor the work and evaluate group and individual performance, which can be reflected towards potential improvements in future group interactions. As a new means to examine individual or joint actions in the group activity, our work investigates the potential of detecting and disambiguating hands of each person in first-person points-of-view videos. Based on the recent developments in automated hand-region extraction from videos, we develop a new multiple-egocentric-video browsing interface that gives easy access to the frames of 1) individual action when only the hands of the viewer are detected, 2) joint action when collective hands are detected, and 3) the viewer checking the others' action as only their hands are detected. We take the evaluation process to explore the effectiveness of our interface with proposed hand-related features which can help perceive actions of interests in the complex analysis of videos involving co-occurred behaviors of multiple people. Nathawan Charoenkulvanich, Rie Kamikubo, Ryo Yonetani, Yoichi Sato 0001 |
IUI | 3 |
| 2018 | Future Person Localization in First-Person VideosabstractWe present a new task that predicts future locations of people observed in first-person videos. Consider a first-person video stream continuously recorded by a wearable camera. Given a short clip of a person that is extracted from the complete stream, we aim to predict that person's location in future frames. To facilitate this future person localization ability, we make the following three key observations: (a) First-person videos typically involve significant ego-motion which greatly affects the location of the target person in future frames; (b) Scales of the target person act as a salient cue to estimate a perspective effect in first-person videos; (c) First-person videos often capture people up-close, making it easier to leverage target poses (e.g., where they look) for predicting their future locations. We incorporate these three observations into a prediction framework with a multi-stream convolution-deconvolution architecture. Experimental results reveal our method to be effective on our new dataset as well as on a public social interaction dataset. Takuma Yagi, Karttikeya Mangalam, Ryo Yonetani, Yoichi Sato 0001 |
CVPR | 3 |
| 2018 | Browsing Group First-Person Videos with 3D VisualizationabstractThis work presents a novel user interface applying 3D visualization to understand complex group activities from multiple first-person videos. The proposed interface is designed to assist video viewers to easily understand the collaborative relationships of group activity based on where the individual worker is located in a workspace and how multiple workers are positioned to one another during the group activity. More specifically, the interface not only shows all recorded first-person videos but also visualizes the 3D position and orientation of each view point (i.e., the 3D position of each worker wearing a head-mounted camera) with a reconstructed 3D model of the workspace. Our user study confirms that the 3D visualization helps video viewers to understand geometric information of a worker and collaborative relationships of group activity easily and accurately. Yuki Sugita, Keita Higuchi, Ryo Yonetani, Rie Kamikubo, Yoichi Sato 0001 |
ISS | 3 |
| 2018 | Ego-Surfing: Person Localization in First-Person Videos Using Ego-Motion SignaturesabstractWe envision a future time when wearable cameras are worn by the masses and recording first-person point-of-view videos of everyday life. While these cameras can enable new assistive technologies and novel research challenges, they also raise serious privacy concerns. For example, first-person videos passively recorded by wearable cameras will necessarily include anyone who comes into the view of a camera-with or without consent. Motivated by these benefits and risks, we developed a self-search technique tailored to first-person videos. The key observation of our work is that the egocentric head motion of a target person (i.e., the self) is observed both in the point-of-view video of the target and observer. The motion correlation between the target person's video and the observer's video can then be used to identify instances of the self uniquely. We incorporate this feature into the proposed approach that computes the motion correlation over densely-sampled trajectories to search for a target individual in observer videos. Our approach significantly improves self-search performance over several well-known face detectors and recognizers. Furthermore, we show how our approach can enable several practical applications such as privacy filtering, target video retrieval, and social group clustering. Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Rapid Prototyping of Accessible Interfaces With Gaze-Contingent Tunnel Vision SimulationabstractActive involvement of users with disabilities is difficult to employ during the iterative stages of the design process due to high costs and effort associated with user studies. This research proposes a user centered design (UCD) strategy to incorporate the use of gaze-contingent tunnel vision simulation with sighted individuals to facilitate rapid prototyping of accessible interfaces. Through three types of validation studies, we examined how our simulation techniques can provide the opportunity for continued evaluation and refinement of the design. Our simulation approach was effective in emulating scanning behaviors caused by tunnel vision along with grasping user feedback to recognize user interface and usability criteria early in the design cycle. Rie Kamikubo, Keita Higuchi, Ryo Yonetani, Hideki Koike, Yoichi Sato 0001 |
ASSETS | 3 |
| 2017 | EgoScanning: Quickly Scanning First-Person Videos with Egocentric Elastic TimelinesabstractThis work presents EgoScanning, a novel video fast-forwarding interface that helps users to find important events from lengthy first-person videos recorded with wearable cameras continuously. This interface is featured by an elastic timeline that adaptively changes playback speeds and emphasizes egocentric cues specific to first-person videos, such as hand manipulations, moving, and conversations with people, based on computer-vision techniques. The interface also allows users to input which of such cues are relevant to events of their interests. Through our user study, we confirm that users can find events of interests quickly from first-person videos thanks to the following benefits of using the EgoScanning interface: 1) adaptive changes of playback speeds allow users to watch fast-forwarded videos more easily; 2) Emphasized parts of videos can act as candidates of events actually significant to users; 3) Users are able to select relevant egocentric cues depending on events of their interests. Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001 |
CHI | 2 |
| 2017 | Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic EncryptionabstractWe propose a privacy-preserving framework for learning visual classifiers by leveraging distributed private image data. This framework is designed to aggregate multiple classifiers updated locally using private data and to ensure that no private information about the data is exposed during and after its learning procedure. We utilize a homomorphic cryptosystem that can aggregate the local classifiers while they are encrypted and thus kept secret. To overcome the high computational cost of homomorphic encryption of high-dimensional classifiers, we (1) impose sparsity constraints on local classifier updates and (2) propose a novel efficient encryption scheme named doublypermuted homomorphic encryption (DPHE) which is tailored to sparse high-dimensional data. DPHE (i) decomposes sparse data into its constituent non-zero values and their corresponding support indices, (ii) applies homomorphic encryption only to the non-zero values, and (iii) employs double permutations on the support indices to make them secret. Our experimental evaluation on several public datasets shows that the proposed approach achieves comparable performance against state-of-the-art visual recognition methods while preserving privacy and significantly outperforms other privacy-preserving methods. Ryo Yonetani, Vishnu Naresh Boddeti, Kris Makoto Kitani, Yoichi Sato 0001 |
ICCV | 1 |
| 2016 | Can Eye Help You?: Effects of Visualizing Eye Fixations on Remote Collaboration Scenarios for Physical TasksabstractIn this work, we investigate how remote collaboration between a local worker and a remote collaborator will change if eye fixations of the collaborator are presented to the worker. We track the collaborator's points of gaze on a monitor screen displaying a physical workspace and visualize them onto the space by a projector or through an optical see-through head-mounted display. Through a series of user studies, we have found the followings: 1) Eye fixations can serve as a fast and precise pointer to objects of the collaborator's interest. 2) Eyes and other modalities, such as hand gestures and speech, are used differently for object identification and manipulation. 3) Eyes are used for explicit instructions only when they are combined with speech. 4) The worker can predict some intentions of the collaborator such as his/her current interest and next instruction. Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001 |
CHI | 2 |
| 2016 | Recognizing Micro-Actions and Reactions from Paired Egocentric VideosabstractWe aim to understand the dynamics of social interactions between two people by recognizing their actions and reactions using a head-mounted camera. Our work will impact several first-person vision tasks that need the detailed understanding of social interactions, such as automatic video summarization of group events and assistive systems. To recognize micro-level actions and reactions, such as slight shifts in attention, subtle nodding, or small hand actions, where only subtle body motion is apparent, we propose to use paired egocentric videos recorded by two interacting people. We show that the first-person and second-person points-of-view features of two people, enabled by paired egocentric videos, are complementary and essential for reliably recognizing micro-actions and reactions. We also build a new dataset of dyadic (two-persons) interactions that comprises more than 1000 pairs of egocentric videos to enable systematic evaluations on the task of micro-action and reaction recognition. Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001 |
CVPR | 1 |
| 2016 | Visual Motif Discovery via First-Person Vision
Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001 |
ECCV (2) | 1 |
| 2015 | Ego-surfing first person videosabstractWe envision a future time when wearable cameras (e.g., small cameras in glasses or pinned on a shirt collar) are worn by the masses and record first-person point-of-view (POV) videos of everyday life. While these cameras can enable new assistive technologies and novel research challenges, they also raise serious privacy concerns. For example, first-person videos passively recorded by wearable cameras will necessarily include anyone who comes into the view of a camera - with or without consent. Motivated by these benefits and risks, we develop a self-search technique tailored to first-person POV videos. The key observation of our work is that the egocentric head motions of a target person (i.e., the self) are observed both in the POV video of the target and observer. The motion correlation between the target person's video and the observer's video can then be used to uniquely identify instances of the self. We incorporate this feature into our proposed approach that computes the motion correlation over supervoxel hierarchies to localize target instances in observer videos. Our proposed approach significantly improves self-search performance over several well-known face detectors and recognizers. Furthermore, we show how our approach can enable several practical applications such as privacy filtering, automated video collection and social group discovery. Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001 |
CVPR | 1 |
| 2013 | Predicting where we look from spatiotemporal gapsabstractWhen we are watching videos, there exist spatiotemporal gaps between where we look and what we focus on, which result from temporally delayed responses and anticipation in eye movements. We focus on the underlying structures of those gaps and propose a novel method to predict points of gaze from video data. In the proposed methods, we model the spatiotemporal patterns of salient regions that tend to be focused on and statistically learn which types of the patterns strongly appear around the points of gaze with respect to each type of eye movements. It allows us to exploit the structures of gaps affected by eye movements and salient motions for the gaze-point prediction. The effectiveness of the proposed method is confirmed with several public datasets. Ryo Yonetani, Hiroaki Kawashima, Takashi Matsuyama |
ICMI | 1 |
| 2012 | Single Image Segmentation with Estimated DepthabstractObject segmentation is a fundamental problem in computer vision. Although many segmentation methods have been proposed, most of them still rely on the appearances of images (i.e., colors or textures) [1, 2, 3, 4, 6, 8]. Consequently, they have a difficulty in distinguishing an object from the background with a similar appearance to the object. To overcome this difficulty, we employ a depth map of an input image as an additional cue to the object segmentation. The main contribution of this work is to introduce a novel segmentation framework that utilizes the depth map combined with a color image to describe the features of objects and backgrounds, where the depth map is estimated from the color image. While a depth map has great potential for use in segmentation, finding a way of integrating two completely different physical quantities, namely the color and depth, has remained unclear. We introduce an integration of the color and depth likelihood on objectness and backgroundness, which simply and effectively extends a traditional segmentation framework based on the Markov random fields (MRF) [2]. By refining the likelihood with the depth information, our proposed method can suppress the incorrect detection of misleading backgrounds. A single image is expressed by K, where K includes color information C = {Cx ∈R}x∈Ω, and in our case, depth informationZ = {Zx ∈R}x∈Ω (x is a position in the image domain Ω ⊂ N2). Object segmentation is the problem of assigning the label A = {Ax}x∈Ω, which gives a label Ax = {0,1} to each pixel, where the labels 1 and 0 at x respectively correspond to the object and background. The statistical relationship between K and A can be described by an MRF, and the appropriate configuration of the labels can be derived by minimizing the following energy function E: Ryo Yonetani, Akisato Kimura, Hitoshi Sakano, Ken Fukuchi |
BMVC | 1 |
| 2012 | Multi-mode saliency dynamics model for analyzing gaze and attentionabstractWe present a method to analyze a relationship between eye movements and saliency dynamics in videos for estimating attentive states of users while they watch the videos. The multi-mode saliency-dynamics model (MMSDM) is introduced to segment spatio-temporal patterns of the saliency dynamics into multiple sequences of primitive modes underlying the saliency patterns. The MMSDM enables us to describe the relationship by the local saliency dynamics around gaze points, which is modeled by a set of distances between gaze points and salient regions characterized by the extracted modes. Experimental results show the effectiveness of the proposed model to classify the attentive states of users by learning the statistical difference of the local saliency dynamics on gaze-paths at each level of attentiveness. Ryo Yonetani, Hiroaki Kawashima, Takashi Matsuyama |
ETRA | 1 |
| 2012 | Modeling video viewing behaviors for viewer state estimationabstractHuman gaze behaviors when watching videos reflect their cognitive states as well as characteristics of the video scenes being watched. Our goal is to establish a method to estimate the viewer states from his/her eye movements toward general videos, such as TV news and commercials. The proposed method is based on a novel model of video viewing behaviors, which takes into account structural and statistical relationships between video dynamics, gaze dynamics and viewer states. This model realizes statistical learning of gaze information while considering dynamic characteristics of video scenes to achieve viewer-state estimation. In this paper, we present an overview of the viewer-state estimation method based on the model of video-viewing behaviors, including several past work done by the author's team. Ryo Yonetani |
ACM Multimedia | 1 |
| 2010 | Gaze Probing: Event-Based Estimation of Objects Being Focused OnabstractWe propose a novel method to estimate the object that a user is focusing on by using the synchronization between the movements of objects and a user's eyes as a cue. We first design an event as a characteristic motion pattern, and we then embed it within the movement of each object. Since the user's ocular reactions to these events are easily detected using a passive camera-based eye tracker, we can successfully estimate the object that the user is focusing on as the one whose movement is most synchronized with the user's eye reaction. Experimental results obtained from the application of this system to dynamic content (consisting of scrolling images) demonstrate the effectiveness of the proposed method over existing methods. Ryo Yonetani, Hiroaki Kawashima, Takatsugu Hirayama, Takashi Matsuyama |
ICPR | 1 |