Yusuke Sekikawa

dblp:148/8805 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-1111-5949ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
abstract
Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decoder (DISCODE), a novel finetuning-free method that generates robust evaluation scores better aligned with human judgments across diverse domains. The core idea behind DISCODE lies in its test-time adaptive evaluation approach, which introduces the Adaptive Test-Time (ATT) loss, leveraging a Gaussian prior distribution to improve robustness in evaluation score estimation. This loss is efficiently minimized at test time using an analytical solution that we derive. Furthermore, we introduce the Multi-domain Caption Evaluation (MCEval) benchmark, a new image captioning evaluation benchmark covering six distinct domains, designed to assess the robustness of evaluation metrics. In our experiments, we demonstrate that DISCODE achieves state-of-the-art performance as a reference-free evaluation metric across MCEval and four representative existing benchmarks.
Nakamasa Inoue, Kanoko Goto, Masanari Oi, Martyna Gruszka, Mahiro Ukai, Takumi Hirose, Yusuke Sekikawa
AAAI7
2026 CoL2A: Convolution-free Local Linear Attention for SpatioTemporal Event Processing
abstract
Linear attention is sparse, recurrent, and GPU-parallel; these are essential features for processing sparse data from event-based cameras. We argue that locality is missing to efficiently model event-to-event relationships for continuous spatiotemporal perception. We propose CoL2A by introducing locality into linear attention without using a computationally demanding convolution operation. The key idea for the convolution-free formulation is restricting the positional embedding local convolutional kernel into the special class which can be decomposed into two global positional embeddings that can be absorbed into query and key; this replaces convolution with a local sum. To the best of our knowledge, CoL2A is the first to equip sparsity, recurrence, GPU parallelism and locality, simultaneously. We demonstrate CoL2A’s effectiveness on dense, high-temporal-resolution (> 1000 fps ) prediction task from events, demonstrating real-time capability while maintaining competitive results over the conventional method. https://github.com/DensoITLab/CoLA
Yusuke Sekikawa, Jun Nagata, Itsumi Araki, Andreu Girbau-Xalabarder
WACV1
2026 DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
abstract
Modeling daily hand interactions often struggles with severe occlusions, such as when two hands overlap, which highlights the need for robust feature learning in 3D hand pose estimation (HPE). To handle such occluded hand images, it is vital to effectively learn the relationship between local image features (e.g., for occluded joints) and global context (e.g., cues from inter-joints, inter-hands, or the scene). However, most current 3D HPE methods still rely on ResNet for feature extraction, and such CNN’s inductive bias may not be optimal for 3D HPE due to its limited capability to model the global context. To address this limitation, we propose an effective and efficient framework for visual feature extraction in 3D HPE using recent state space modeling (i.e., Mamba), dubbed Deformable Mamba (DF-Mamba). DF-Mamba is designed to capture global context cues beyond standard convolution through Mamba’s selective state modeling and the proposed deformable state scanning. Specifically, for local features after convolution, our deformable scanning aggregates these features within an image while selectively preserving useful cues that represent the global context. This approach significantly improves the accuracy of structured 3D HPE, with comparable inference speed to ResNet-50. Our experiments involve extensive evaluations on five divergent datasets including single-hand and two-hand scenarios, hand-only and hand-object interactions, as well as RGB and depth-based estimation. DF-Mamba outperforms the latest image backbones, including VMamba and Spatial-Mamba, on all datasets and achieves state-of-the-art performance.
Takehiko Ohkawa, Guwenxiao Zhou, Kanoko Goto, Takumi Hirose, Yusuke Sekikawa, Nakamasa Inoue
WACV6
2025 Masked Gated Linear Unit
abstract
Gated Linear Units (GLUs) have become essential components in the feed-forward networks of state-of-the-art Large Language Models (LLMs). However, they require twice as many memory reads compared to feed-forward layers without gating, due to the use of separate weight matrices for the gate and value streams. To address this bottleneck, we introduce Masked Gated Linear Units (MGLUs), a novel family of GLUs with an efficient kernel implementation. The core contribution of MGLUs include: (1) the Mixture of Element-wise Gating (MoEG) architecture that learns multiple binary masks, each determining gate or value assignments at the element level on a single shared weight matrix resulting in reduced memory transfer, and (2) FlashMGLU, a hardware-friendly kernel that yields up to a 19.7$\times$ inference-time speed-up over a na\"ive PyTorch MGLU and is 47\% more memory-efficient and 34\% faster than standard GLUs despite added architectural complexity on an RTX5090 GPU. In LLM experiments, the Swish-activated variant SwiMGLU preserves its memory advantages while matching—or even surpassing—the downstream accuracy of the SwiGLU baseline.
Yukito Tajima, Nakamasa Inoue, Yusuke Sekikawa, Ikuro Sato, Rio Yokota
NeurIPS3
2024 Gumbel-NeRF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
abstract
We propose Gumbel-NeRF, a mixture-of-expert (MoE) neural radiance fields (NeRF) model with a hindsight expert selection mechanism for synthesizing novel views of unseen objects. Previous studies have shown that the MoE structure provides high-quality representations of a given large-scale scene consisting of many objects. However, we observe that such a MoE NeRF model often produces low-quality representations in the vicinity of experts’ boundaries when applied to the task of novel view synthesis of an unseen object from one/few-shot input. We find that this deterioration is primarily caused by the foresight expert selection mechanism, which may leave an unnatural discontinuity in the object shape near the experts’ boundaries. Gumbel-NeRF adopts a hindsight expert selection mechanism, which guarantees continuity in the density field even near the experts’ boundaries. Experiments using the SRN cars dataset demonstrate the superiority of Gumbel-NeRF over the baselines in terms of various image quality metrics. The code will be available upon acceptance.
Yusuke Sekikawa, Chingwei Hsu, Satoshi Ikehata, Rei Kawakami, Ikuro Sato
ICIP1
2024 SAS: Structured Activation Sparsification
abstract
Wide networks usually yield better accuracy than their narrower counterpart at the expense of the massive $\texttt{mult}$ cost. To break this tradeoff, we advocate a novel concept of $\textit{Structured Activation Sparsification}$, dubbed SAS, which boosts accuracy without increasing computation by utilizing the projected sparsity in activation maps with a specific structure. Concretely, the projected sparse activation is allowed to have N nonzero value among M consecutive activations. Owing to the local structure in sparsity, the wide $\texttt{matmul}$ between a dense weight and the sparse activation is executed as an equivalent narrow $\texttt{matmul}$ between a dense weight and dense activation, which is compatible with NVIDIA's $\textit{SparseTensorCore}$ developed for the N:M structured sparse weight. In extensive experiments, we demonstrate that increasing sparsity monotonically improves accuracy (up to 7% on CIFAR10) without increasing the $\texttt{mult}$ count. Furthermore, we show that structured sparsification of $\textit{activation}$ scales better than that of $\textit{weight}$ given the same computational budget.
Yusuke Sekikawa, Shingo Yashima
ICLR1
2023 Tangentially Elongated Gaussian Belief Propagation for Event-Based Incremental Optical Flow Estimation
abstract
Optical flow estimation is a fundamental functionality in computer vision. An event-based camera, which asynchronously detects sparse intensity changes, is an ideal device for realizing low-latency estimation of the optical flow owing to its low-latency sensing mechanism. An existing method using local plane fitting of events could utilize the sparsity to realize incremental updates for low-latency estimation; however, its output is merely a normal component of the full optical flow. An alternative approach using a frame-based deep neural network could estimate the full flow; however, its intensive non-incremental dense operation prohibits the low-latency estimation. We propose tangentially elongated Gaussian (TEG) belief propagation (BP) that realizes incremental full-flow estimation. We model the probability of full flow as the joint distribution of TEGs from the normal flow measurements, such that the marginal of this distribution with correct prior equals the full flow. We formulate the marginalization using a message-passing based on the BP to realize efficient incremental updates using sparse measurements. In addition to the theoretical justification, we evaluate the effectiveness of the TEGBP in real-world datasets; it outperforms SOTA incremental quasi-full flow method by a large margin. (The code is available at https://github.com/DensoITLab/tegbp/).
Jun Nagata, Yusuke Sekikawa
CVPR2
2023 Bit-Pruning: A Sparse Multiplication-Less Dot-Product
Yusuke Sekikawa, Shingo Yashima
ICLR1
2022 Implicit Neural Representations for Variable Length Human Motion Generation
Pablo Cervantes, Yusuke Sekikawa, Ikuro Sato, Koichi Shinoda
ECCV (17)2
2022 Neural Implicit Event Generator for Motion Tracking
abstract
We present a novel framework of motion tracking from event data using implicit expression. Our framework uses pre-trained event generation MLP called the implicit event generator (IEG) and carries out motion tracking by updating its state (position and velocity) based on the difference between the observed event and generated event from the current state estimation. The difference is computed implicitly by the IEG. Unlike the conventional explicit approach, which requires dense computation to evaluate the difference, our implicit approach realizes the update of the efficient state directly from sparse event data. Our sparse algorithm is especially suitable for mobile robotics applications in which computational resources and battery life are limited. To verify the effectiveness of our method on real-world data, we applied it to the AR marker tracking application. We have confirmed that our framework works well in real-world environments in the presence of noise and background clutter.
Mana Masuda, Yusuke Sekikawa, Ryo Fujii, Hideo Saito 0001
ICRA2
2021 Learning to Sparsify Differences of Synaptic Signal for Efficient Event Processing
Yusuke Sekikawa, Keisuke Uto
BMVC1
2021 Toward Unsupervised 3d Point Cloud Anomaly Detection Using Variational Autoencoder
abstract
In this paper, we present an end-to-end unsupervised anomaly detection framework for 3D point clouds. To the best of our knowledge, this is the first work to tackle the anomaly detection task on a general object represented by a 3D point cloud. We propose a deep variational autoencoder based unsupervised anomaly detection network adapted to the 3D point cloud and an anomaly score specifically for 3D point clouds. To verify the effectiveness of the model, we conducted extensive experiments on ShapeNet dataset. Through quantitative and qualitative evaluation, we demonstrate that the proposed method outperforms the baseline method.
Mana Masuda, Ryo Hachiuma, Ryo Fujii, Hideo Saito 0001, Yusuke Sekikawa
ICIP5
2021 Snapshot Multispectral Image Completion Via Self-Dictionary Transformed Tensor Nuclear Norm Minimization With Total Variation
abstract
Snapshot multispectral imaging suffers from severely low spatial resolution and degraded signals due to mosaic rearrangement. In order to recover a signal of full bands and full sensor size from a single snapshot, we propose a self-dictionary-transformed tensor nuclear norm and develop a joint optimization with total variation regularization as a convex completion problem. The proposed nuclear norm is designed specifically for the intrinsic structure of snapshot multispectral data, reflects the parsimony of tensors that conventional approaches ignore, as well as incorporates inter-axis correlation unlike matrix-based optimization. We show increased accuracy with our self-dictionary throughout simulation experiments and demonstrate quality enhancement in recovering real snapshot multispectral images.
Keisuke Ozawa, Shinichi Sumiyoshi, Yusuke Sekikawa, Keisuke Uto, Yuichi Yoshida, Mitsuru Ambai
ICIP3
2020 Rethinking PointNet Embedding for Faster and Compact Model
abstract
PointNet, which is the widely used point-wise embedding method and known as a universal approximator for continuous set functions, can process one million points per second. Nevertheless, real-time inference for the recent development of high-performing sensors is still challenging with existing neural network-based methods, including PointNet. In ordinary cases, the embedding function of PointNet behaves like a soft-indicator function that is activated when the input points exist in a certain local region of the input space. Leveraging this property, we reduce the computational costs of point-wise embedding by replacing the embedding function of PointNet with the soft-indicator function by Gaussian kernels. Moreover, we show that the Gaussian kernels also satisfy the universal approximation theorem that PointNet satisfies. In experiments, we verify that our model using the Gaussian kernels achieves comparable results to baseline methods, but with much fewer floating-point operations per sample up to 92% reduction from PointNet.
Teppei Suzuki, Keisuke Ozawa, Yusuke Sekikawa
3DV3
2020 QR-code Reconstruction from Event Data via Optimization in Code Subspace
abstract
We propose an image reconstruction method from event data, assuming the target images belong to a prespecified class like QR codes. Instead of solving the reconstruction problem in the image space, we introduce a code space that covers all the noiseless target class images and solves the reconstruction problem on it. This restriction enormously reduces the number of optimizing parameters and makes the reconstruction problem well posed and robust to noise. We demonstrate fast and robust QR-code scanning in difficult, high-speed scenes with industrial high-speed cameras and other reconstruction methods.
Jun Nagata, Yusuke Sekikawa, Kosuke Hara, Teppei Suzuki, Yoshimitsu Aoki
WACV2
2019 EventNet: Asynchronous Recursive Event Processing
abstract
Event cameras are bio-inspired vision sensors that mimic retinas to asynchronously report per-pixel intensity changes rather than outputting an actual intensity image at regular intervals. This new paradigm of image sensor offers significant potential advantages; namely, sparse and non-redundant data representation. Unfortunately, however, most of the existing artificial neural network architectures, such as a CNN, require dense synchronous input data, and therefore, cannot make use of the sparseness of the data. We propose EventNet, a neural network designed for real-time processing of asynchronous event streams in a recursive and event-wise manner. EventNet models dependence of the output on tens of thousands of causal events recursively using a novel temporal coding scheme. As a result, at inference time, our network operates in an event-wise manner that is realized with very few sum-of-the-product operations---look-up table and temporal feature aggregation---which enables processing of 1 mega or more events per second on standard CPU. In experiments using real data, we demonstrated the real-time performance and robustness of our framework.
Yusuke Sekikawa, Kosuke Hara, Hideo Saito 0001
CVPR1
2018 Constant Velocity 3D Convolution
abstract
We propose a novel three-dimensional (3D)-convolution method, cv3dconv, for detecting spatiotemporal features from videos. It reduces the number of sum-of-products of 3D convolution by thousands of times by assuming the constant moving velocity of the camera. We observed that a specific class of video sequences, such as those captured by an in-vehicle camera, can be well approximated with piece-wise linear movements of 2D features in the temporal dimension. Our principal finding is that the 3D kernel, represented by the constant-velocity, can be decomposed into a convolution of a 2D kernel representing the shapes and a 3D kernel representing the velocity. We derived the efficient recursive algorithm for this class of 3D convolution which is exceptionally suited for sparse data, and this parameterized decomposed representation imposes a structured regularization along the temporal direction. We experimentally verified the validity of our approximation using a controlled dataset, and we also showed the effectiveness of cv3dconv for the visual odometry estimation task using real event camera data captured in urban road scene.
Yusuke Sekikawa, Kohta Ishikawa, Kosuke Hara, Yuichi Yoshida, Koichiro Suzuki, Ikuro Sato, Hideo Saito 0001
3DV1
2016 Fast Eigen Matching
Yusuke Sekikawa, Koichiro Suzuki, Yuichi Yoshida, Kosuke Hara, Ikuro Sato
BMVC1