EDBT 2026 Demo / reviewers in the wild / expert
Taiki Sekii
dblp:191/4650
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-1895-3075ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Video understanding and tracking · 60% Trustworthy machine learning · 23% 3D vision · 18% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
0.7 | 1 | 2023 | Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023 |
Machine learning › Trustworthy machine learning › robustness
adversarial attack |
0.7 | 1 | 2023 | Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023 |
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
0.7 | 1 | 2023 | Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023 |
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization |
0.7 | 1 | 2023 | Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling · CVPR 2023 |
Computer vision › 3D vision
pose estimation |
0.3 | 1 | 2018 | Pose Proposal Networks · ECCV (13) 2018 |
Computer vision › Video understanding and tracking › object tracking
3d object tracking |
0.2 | 1 | 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016 |
Computer vision › Video understanding and tracking
multi-camera tracking |
0.2 | 1 | 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.2 | 1 | 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016 |
Computer vision › 3D vision › 3d shape reconstruction
object shape reconstruction |
0.2 | 1 | 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar Appearances · CVPR 2016 |
Computer vision › 3D vision › point cloud analysis
point cloud learning |
0.2 | 1 | 2023 | Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton Features · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
point cloud deep learning · 1.3zero-shot learning · 0.7weakly supervised learning · 0.7structured keypoint pooling · 0.7prompt-guided learning · 0.7pooling-switching trick · 0.7track graph · 0.2mixture model · 0.2maximum a posteriori expectation-maximization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flashback: Memory Mechanism for Enhancing Memory Efficiency and Speed in Deep Sequential ModelsabstractIn this study, we tackle three main challenges of deep sequential processing models in previous research: (1) memory degradation, (2) inaccurate gradient backpropagation, and (3) compatibility with next-token prediction. Specifically, to address (1-2), we define a Flashback property in which memory is preserved perfectly as an identity mapping of its stored value in a memory region until it is overwritten by a hidden state at a different time step. We propose a Flashback mechanism that satisfies this property in a fully differentiable, end-to-end manner. Further, to tackle (3), we propose architectures that incorporate the Flashback mechanism into Transformers and Mamba, enabling next-token prediction for language modeling tasks. In experiments, we trained on The Pile dataset, which includes diverse texts, to evaluate tradeoffs between commonsense reasoning accuracy, processing speed, and memory usage after introducing the Flashback mechanism into existing methods. The evaluations confirmed the effectiveness of the Flashback mechanism. Taiki Sekii |
COLING | 1 |
| 2025 | Frozen Network Few-Shot Object DetectionabstractIn this paper, we propose a method that addresses two challenges faced by existing generalized few-shot object detection methods, limiting their scalability in practical applications: (i) the high training cost required for few-shot learning of a small number of target object samples, and (ii) the lack of variation in these samples. Leveraging the generalization ability of a deep neural network (DNN) trained with vision-language pretraining, we introduce a frozen network few-shot object detection framework in which a common DNN, shared with the pretraining phase, is used to model the appearance features of target objects without needing to be updated during few-shot learning. In experiments, we employed publicly available datasets with diverse object appearances for pretraining and few-shot learning. By comparing our proposed method’s detection accuracy with existing methods, we demonstrate that our method effectively addresses the challenges described above. Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii |
ICIP | 5 |
| 2024 | Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement TrajectoriesabstractThis paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores.Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management.The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map.Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects.Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics. Hikaru Asano, Ryo Yonetani, Taiki Sekii, Hiroki Ouchi |
INLG | 3 |
| 2023 | Unified Keypoint-Based Action Recognition Framework via Structured Keypoint PoolingabstractThis paper simultaneously addresses three limitations associated with conventional skeleton-based action recognition; skeleton detection and tracking errors, poor variety of the targeted actions, as well as person-wise and framewise action recognition. A point cloud deep-learning paradigm is introduced to the action recognition, and a unified framework along with a novel deep neural network architecture called Structured Keypoint Pooling is proposed. The proposed method sparsely aggregates keypoint features in a cascaded manner based on prior knowledge of the data structure (which is inherent in skeletons), such as the instances and frames to which each keypoint belongs, and achieves robustness against input errors. Its less constrained and tracking-free architecture enables time-series keypoints consisting of human skeletons and nonhuman object contours to be efficiently treated as an input 3D point cloud and extends the variety of the targeted action. Furthermore, we propose a Pooling-Switching Trick inspired by Structured Keypoint Pooling. This trick switches the pooling kernels between the training and inference phases to detect person-wise and framewise actions in a weakly supervised manner using only video-level action labels. This trick enables our training scheme to naturally introduce novel data augmentation, which mixes multiple point clouds extracted from different videos. In the experiments, we comprehensively verify the effectiveness of the proposed method against the limitations, and the method outperforms state-of-the-art skeleton-based action recognition and spatio-temporal action localization methods. Ryo Hachiuma, Fumiaki Sato, Taiki Sekii |
CVPR | 3 |
| 2023 | Prompt-Guided Zero-Shot Anomaly Action Recognition using Pretrained Deep Skeleton FeaturesabstractThis study investigates unsupervised anomaly action recognition, which identifies video-level abnormal-human-behavior events in an unsupervised manner without abnormal samples, and simultaneously addresses three limitations in the conventional skeleton-based approaches: target domain-dependent DNN training, robustness against skeleton errors, and a lack of normal samples. We present a unified, user prompt-guided zero-shot learning framework using a target domain-independent skeleton feature extractor, which is pretrained on a large-scale action recognition dataset. Particularly, during the training phase using normal samples, the method models the distribution of skeleton features of the normal actions while freezing the weights of the DNNs and estimates the anomaly score using this distribution in the inference phase. Additionally, to increase robustness against skeleton errors, we introduce a DNN architecture inspired by a point cloud deep learning paradigm, which sparsely propagates the features between Joints. Furthermore, to prevent the unobserved normal actions from being misidentified as abnormal actions, we incorporate a similarity score between the user prompt embeddings and skeleton features aligned in the common space into the anomaly score, which indirectly supplements normal actions. On two publicly available datasets, we conduct experiments to test the effectiveness of the proposed method with respect to above-mentioned limitations. Fumiaki Sato, Ryo Hachiuma, Taiki Sekii |
CVPR | 3 |
| 2018 | Pose Proposal Networks
Taiki Sekii |
ECCV (13) | 1 |
| 2016 | Robust, Real-Time 3D Tracking of Multiple Objects with Similar AppearancesabstractThis paper proposes a novel method for tracking multiple moving objects and recovering their three-dimensional (3D) models separately using multiple calibrated cameras. For robustly tracking objects with similar appearances, the proposed method uses geometric information regarding 3D scene structure rather than appearance. A major limitation of previous techniques is foreground confusion, in which the shapes of objects and/or ghosting artifacts are ignored and are hence not appropriately specified in foreground regions. To overcome this limitation, our method classifies foreground voxels into targets (objects and artifacts) in each frame using a novel, probabilistic two-stage framework. This is accomplished by step-wise application of a track graph describing how targets interact and the maximum a posteriori expectation-maximization algorithm for the estimation of target parameters. We introduce mixture models with semiparametric component distributions regarding 3D target shapes. In order to not confuse artifacts with objects of interest, we automatically detect and track artifacts based on a closed-world assumption. Experimental results show that our method outperforms state-of-the-art trackers on seven public sequences while achieving real-time performance. Taiki Sekii |
CVPR | 1 |