Ryoma Yataka

dblp:199/9334 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0004-7311-6431ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Indoor Multi-View Radar Object Detection via 3D Bounding Box Diffusion
abstract
Multi-view indoor radar perception has drawn attention due to its cost-effectiveness and low privacy risks. Existing methods often rely on implicit cross-view radar feature association, such as proposal pairing in RFMask or query-to-feature cross-attention in RETR, which can lead to ambiguous feature matches and degraded detection in complex indoor scenes. To address these limitations, we propose REXO (multi-view Radar object dEtection with 3D bounding boX diffusiOn), which lifts the 2D bounding box (BBox) diffusion process of DiffusionDet into the 3D radar space. REXO utilizes these noisy 3D BBoxes to guide an explicit cross-view radar feature association, enhancing the cross-view radar-conditioned denoising process. By accounting for prior knowledge that the person is in contact with the ground, REXO reduces the number of diffusion parameters by determining them from this prior. Evaluated on two open indoor radar datasets, our approach surpasses state-of-the-art methods by a margin of +4.22 AP on the HIBER dataset and +11.02 AP on the MMVR dataset. Our implementation is available at https://github.com/merlresearch/radar-bbox-diffusion.
Ryoma Yataka, Pu Wang 0004, Petros Boufounos, Ryuhei Takahashi
AAAI1
2025 Multi-View Radar Detection Transformer with Differentiable Positional Encoding
abstract
The Radar dEtection TRansformer (RETR) has recently been introduced to fuse multi-view millimeter-wave radar heatmaps by leveraging the detection transformer architecture and a geometric learning framework for indoor radar perception. A notable feature of RETR is its tunable positional encoding (TPE), which allows for adjusting the significance of depth positional embedding across multiple views to promote depth-prioritized feature association. However, the TPE ratio is predetermined, rather than being optimized during the training process. In this paper, we propose a differentiable positional encoding (DiPE) scheme for RETR by automatically adjusting the TPE ratio during the training for enhanced performance and avoiding exhaustive grid value search. DiPE can be applied along with either pre-fixed (e.g., sinusoidal) or learnable positional embeddings, achieved by multiplying dual differentiable masks over the depth and angular positional embedding vectors. Comprehensive evaluations on the open MMVR dataset demonstrate that the proposed DiPE not only simplifies the determination of the TPE ratio but also enhances the overall detection performance.
Ryoma Yataka, Pu Wang 0004, Petros Boufounos, Ryuhei Takahashi
ICASSP1
2025 RAPTR: Radar-based 3D Pose Estimation using Transformer
abstract
Radar-based indoor 3D human pose estimation typically relied on fine-grained 3D keypoint labels, which are costly to obtain especially in complex indoor settings involving clutter, occlusions, or multiple people. In this paper, we propose \textbf{RAPTR} (RAdar Pose esTimation using tRansformer) under weak supervision, using only 3D BBox and 2D keypoint labels which are considerably easier and more scalable to collect. Our RAPTR is characterized by a two-stage pose decoder architecture with a pseudo-3D deformable attention to enhance (pose/joint) queries with multi-view radar features: a pose decoder estimates initial 3D poses with a 3D template loss designed to utilize the 3D BBox labels and mitigate depth ambiguities; and a joint decoder refines the initial poses with 2D keypoint labels and a 3D gravity loss. Evaluated on two indoor radar datasets, RAPTR outperforms existing methods, reducing joint position error by $34.3$\% on HIBER and $76.9$\% on MMVR. Our implementation is available at \url{https://github.com/merlresearch/radar-pose-transformer}.
Sorachi Kato, Ryoma Yataka, Pu Wang 0004, Pedro Miraldo, Takuya Fujihashi, Petros Boufounos
NeurIPS2
2024 SIRA: Scalable Inter-Frame Relation and Association for Radar Perception
abstract
Conventional radar feature extraction faces limitations due to low spatial resolution, noise, multipath reflection, the presence of ghost targets, and motion blur. Such limitations can be exacerbated by nonlinear object motion, particularly from an ego-centric viewpoint. It becomes evident that to address these challenges, the key lies in exploiting temporal feature relation over an extended horizon and enforcing spatial motion consistency for effective association. To this end, this paper proposes SIRA (Scalable Inter-frame Relation and Association) with two designs. First, inspired by Swin Transformer, we introduce extended temporal relation, generalizing the existing temporal relation layer from two consecutive frames to multiple inter-frames with temporally regrouped window attention for scalability. Second, we propose motion consistency track with the concept of a pseudo-tracklet generated from observational data for better trajectory prediction and subsequent object association. Our approach achieves 58.11 [email protected] for oriented object detection and 47.79 MOTA for multiple object tracking on the Radiate dataset, surpassing previous state-of-the-art by a margin of +4.11 [email protected] and +9.94 MOTA, respectively.
Ryoma Yataka, Pu Wang 0004, Petros Boufounos, Ryuhei Takahashi
CVPR1
2024 MMVR: Millimeter-Wave Multi-view Radar Dataset and Benchmark for Indoor Perception
Mohammad Mahbubur Rahman, Ryoma Yataka, Sorachi Kato, Pu Wang 0004, Peizhao Li, Adriano Cardace, Petros Boufounos
ECCV (79)2
2024 Radar Perception with Scalable Connective Temporal Relations for Autonomous Driving
abstract
Due to the noise and low spatial resolution in automotive radar data, exploring temporal relations of learnable features over consecutive 2 radar frames has shown performance gain on downstream tasks (e.g., object detection and tracking) in our previous study [1]. In this paper, we further enhance radar perception by significantly extending the time horizon of temporal relations. To this end, we propose a scalable connective temporal radar (SCTR) method that consists of 1) a standard temporal relation layer (TRL), 2) a connective TRL with shifted window attention, and 3) a window merging operation, to facilitate feature connectivity between radar frames over an extended time interval. Our complexity analysis and comprehensive evaluation of the Radiate dataset demonstrate that the SCTR achieves a great tradeoff between the complexity and downstream detection performance.
Ryoma Yataka, Pu Wang 0004, Petros Boufounos, Ryuhei Takahashi
ICASSP1
2024 RETR: Multi-View Radar Detection Transformer for Indoor Perception
abstract
Indoor radar perception has seen rising interest due to affordable costs driven by emerging automotive imaging radar developments and the benefits of reduced privacy concerns and reliability under hazardous conditions (e.g., fire and smoke). However, existing radar perception pipelines fail to account for distinctive characteristics of the multi-view radar setting. In this paper, we propose Radar dEtection TRansformer (RETR), an extension of the popular DETR architecture, tailored for multi-view radar perception. RETR inherits the advantages of DETR, eliminating the need for hand-crafted components for object detection and segmentation in the image plane. More importantly, RETR incorporates carefully designed modifications such as 1) depth-prioritized feature similarity via a tunable positional encoding (TPE); 2) a tri-plane loss from both radar and camera coordinates; and 3) a learnable radar-to-camera transformation via reparameterization, to account for the unique multi-view radar setting. Evaluated on two indoor radar perception datasets, our approach outperforms existing state-of-the-art methods by a margin of 15.38+ AP for object detection and 11.91+ IoU for instance segmentation, respectively. Our implementation is available at https://github.com/merlresearch/radar-detection-transformer.
Ryoma Yataka, Adriano Cardace, Pu Wang 0004, Petros Boufounos, Ryuhei Takahashi
NeurIPS1
2023 Grassmann Manifold Flows for Stable Shape Generation
abstract
Recently, studies on machine learning have focused on methods that use symmetry implicit in a specific manifold as an inductive bias. Grassmann manifolds provide the ability to handle fundamental shapes represented as shape spaces, enabling stable shape analysis. In this paper, we present a novel approach in which we establish the theoretical foundations for learning distributions on the Grassmann manifold via continuous normalization flows, with the explicit goal of generating stable shapes. Our approach facilitates more robust generation by effectively eliminating the influence of extraneous transformations, such as rotations and inversions, through learning and generating within a Grassmann manifold designed to accommodate the essential shape information of the object. The experimental results indicated that the proposed method could generate high-quality samples by capturing the data structure. Furthermore, the proposed method significantly outperformed state-of-the-art methods in terms of the log-likelihood or evidence lower bound. The results obtained are expected to stimulate further research in this field, leading to advances for stable shape generation and analysis.
Ryoma Yataka, Kazuki Hirashima, Masashi Shiraishi
NeurIPS1
2017 Three-dimensional Object Recognition via Subspace Representation on a Grassmann Manifold
Ryoma Yataka, Kazuhiro Fukui
ICPRAM1