Taiki Sekii

dblp:191/4650 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-1895-3075ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 60% Trustworthy machine learning · 23% 3D vision · 18%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
0.712023
Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.712023
Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.712023
Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.712023
Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023
Computer vision › 3D vision
pose estimation
0.312018
Pose Proposal Networks · ECCV (13) 2018
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.212016
Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016
Computer vision › 3D vision
3d reconstruction
0.212016
Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016
Computer vision › Video understanding and tracking
multi-camera tracking
0.212016
Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016
Computer vision › Video understanding and tracking
multi-object tracking
0.212016
Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016
Computer vision › 3D vision › 3d shape reconstruction
object shape reconstruction
0.212016
Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016
Computer vision › 3D vision › point cloud analysis
point cloud learning
0.212023
Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023

Methods — techniques the papers use, named apart from their topics

point cloud deep learning · 1.3zero-shot learning · 0.7weakly supervised learning · 0.7structured keypoint pooling · 0.7prompt-guided learning · 0.7pooling-switching trick · 0.7track graph · 0.2mixture model · 0.2maximum a posteriori expectation-maximization · 0.2
YearPublicationVenuePosition
2025 Flashback: Memory Mechanism for Enhancing Memory Efficiency and Speed in Deep Sequential Models
abstract
In this study, we tackle three main challenges of deep sequential processing models in previous research: (1) memory degradation, (2) inaccurate gradient backpropagation, and (3) compatibility with next-token prediction. Specifically, to address (1-2), we define a Flashback property in which memory is preserved perfectly as an identity mapping of its stored value in a memory region until it is overwritten by a hidden state at a different time step. We propose a Flashback mechanism that satisfies this property in a fully differentiable, end-to-end manner. Further, to tackle (3), we propose architectures that incorporate the Flashback mechanism into Transformers and Mamba, enabling next-token prediction for language modeling tasks. In experiments, we trained on The Pile dataset, which includes diverse texts, to evaluate tradeoffs between commonsense reasoning accuracy, processing speed, and memory usage after introducing the Flashback mechanism into existing methods. The evaluations confirmed the effectiveness of the Flashback mechanism.
Taiki Sekii
COLING1
2025 Frozen Network Few-Shot Object Detection
abstract
In this paper, we propose a method that addresses two challenges faced by existing generalized few-shot object detection methods, limiting their scalability in practical applications: (i) the high training cost required for few-shot learning of a small number of target object samples, and (ii) the lack of variation in these samples. Leveraging the generalization ability of a deep neural network (DNN) trained with vision-language pretraining, we introduce a frozen network few-shot object detection framework in which a common DNN, shared with the pretraining phase, is used to model the appearance features of target objects without needing to be updated during few-shot learning. In experiments, we employed publicly available datasets with diverse object appearances for pretraining and few-shot learning. By comparing our proposed method’s detection accuracy with existing methods, we demonstrate that our method effectively addresses the challenges described above.
Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
ICIP5
2024 Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories
abstract
This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores.Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management.The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map.Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects.Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics.
Hikaru Asano, Ryo Yonetani, Taiki Sekii, Hiroki Ouchi
INLG3
2023 Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling
abstract
This paper simultaneously addresses three limitations associated with conventional skeleton-based action recognition; skeleton detection and tracking errors, poor variety of the targeted actions, as well as person-wise and framewise action recognition. A point cloud deep-learning paradigm is introduced to the action recognition, and a unified framework along with a novel deep neural network architecture called Structured Keypoint Pooling is proposed. The proposed method sparsely aggregates keypoint features in a cascaded manner based on prior knowledge of the data structure (which is inherent in skeletons), such as the instances and frames to which each keypoint belongs, and achieves robustness against input errors. Its less constrained and tracking-free architecture enables time-series keypoints consisting of human skeletons and nonhuman object contours to be efficiently treated as an input 3D point cloud and extends the variety of the targeted action. Furthermore, we propose a Pooling-Switching Trick inspired by Structured Keypoint Pooling. This trick switches the pooling kernels between the training and inference phases to detect person-wise and framewise actions in a weakly supervised manner using only video-level action labels. This trick enables our training scheme to naturally introduce novel data augmentation, which mixes multiple point clouds extracted from different videos. In the experiments, we comprehensively verify the effectiveness of the proposed method against the limitations, and the method outperforms state-of-the-art skeleton-based action recognition and spatio-temporal action localization methods.
Ryo Hachiuma, Fumiaki Sato, Taiki Sekii
CVPR3
2023 Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features
abstract
This study investigates unsupervised anomaly action recognition, which identifies video-level abnormal-human-behavior events in an unsupervised manner without abnormal samples, and simultaneously addresses three limitations in the conventional skeleton-based approaches: target domain-dependent DNN training, robustness against skeleton errors, and a lack of normal samples. We present a unified, user prompt-guided zero-shot learning framework using a target domain-independent skeleton feature extractor, which is pretrained on a large-scale action recognition dataset. Particularly, during the training phase using normal samples, the method models the distribution of skeleton features of the normal actions while freezing the weights of the DNNs and estimates the anomaly score using this distribution in the inference phase. Additionally, to increase robustness against skeleton errors, we introduce a DNN architecture inspired by a point cloud deep learning paradigm, which sparsely propagates the features between Joints. Furthermore, to prevent the unobserved normal actions from being misidentified as abnormal actions, we incorporate a similarity score between the user prompt embeddings and skeleton features aligned in the common space into the anomaly score, which indirectly supplements normal actions. On two publicly available datasets, we conduct experiments to test the effectiveness of the proposed method with respect to above-mentioned limitations.
Fumiaki Sato, Ryo Hachiuma, Taiki Sekii
CVPR3
2018 Pose Proposal Networks
Taiki Sekii
ECCV (13)1
2016 Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances
abstract
This paper proposes a novel method for tracking multiple moving objects and recovering their three-dimensional (3D) models separately using multiple calibrated cameras. For robustly tracking objects with similar appearances, the proposed method uses geometric information regarding 3D scene structure rather than appearance. A major limitation of previous techniques is foreground confusion, in which the shapes of objects and/or ghosting artifacts are ignored and are hence not appropriately specified in foreground regions. To overcome this limitation, our method classifies foreground voxels into targets (objects and artifacts) in each frame using a novel, probabilistic two-stage framework. This is accomplished by step-wise application of a track graph describing how targets interact and the maximum a posteriori expectation-maximization algorithm for the estimation of target parameters. We introduce mixture models with semiparametric component distributions regarding 3D target shapes. In order to not confuse artifacts with objects of interest, we automatically detect and track artifacts based on a closed-world assumption. Experimental results show that our method outperforms state-of-the-art trackers on seven public sequences while achieving real-time performance.
Taiki Sekii
CVPR1