VLDB 2026 Research / reviewers in the wild / expert
Satoshi Komorita
dblp:85/8088
· DBLP profile ↗
12ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0003-1526-5514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Computer networks · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Rate-Distortion Optimized Variable-Node-size Trisoup for Point Cloud CodingabstractTriangle soup (Trisoup) is being studied as a new coding tool for Geometry-based Point Cloud Compression (G-PCC) stan-dardized in the Moving Picture Experts Group (MPEG). Outside of MPEG, a variable-node-size extension of Trisoup is studied to increase the flexibility of G-PCC. A primary advantage of variable node size is to achieve better coding performance by selecting appropriate node size according to local geometric complexity and required bits. However, the node size is not optimized in terms of bit rate and distortion in the conventional extension. To maximize the coding performances of the variable-node-size method, we propose a new cost function considering both bit rates and distortions. The experimental results show that the proposed method provides -1.5 % coding performance improvement in point-to-point PSNR versus bit rate against the conventional extension. Kyohei Unno, Kohei Matsuzaki, Satoshi Komorita, Kei Kawamura |
ICASSP | 3 |
| 2023 | Consistent and Diverse Human Motion Prediction Using Conditional Variational Autoencoder with Context-Aware Latent Spaceabstract3D human motion prediction is the task of forecasting future motions from past observations. Previous work predicts diverse motions to cover uncertain future possibilities over a longer time. However, since most of them do not consider the surrounding context and use only the target person’s motion as the past observation, they generate overly diverse predictions rather than actual possibilities. Therefore, we propose a model that narrows down the possibilities by considering the context, especially interactions with nearby objects. Our proposed model predicts motions based on latent codes diversely sampled from a context-aware latent space. The experimental results show that our model improves accuracy and generates diverse predictions by focusing on possibilities that are consistent with object locations. Chihiro Nakatsuka, Satoshi Komorita |
ICIP | 2 |
| 2022 | Stable 3D Human Pose Estimation in Low- Resolution Videos with a Few ViewsabstractWe discuss the problem of 3D pose estimation for multi-view videos. With previous frame-by-frame multi-view methods, it has been difficult to achieve stable estimation under challenging settings such as low-resolution or with only a few views. Temporal approaches are effective ways of addressing such problems, but enforcing temporal consistency with neigh-boring frames sometimes damages the precision of the results. We propose a temporal approach with selective corrections based on the observation that errors in the frame-by-frame approach are concentrated under certain adverse conditions. Our method evaluates the confidence of the frame-by-frame results and compensates for the inaccurate keypoints with temporal information while retaining the accurate keypoints. In our experiments on the CMU Panoptic dataset customized for low-resolution and a few views, we reported 32.98 mm for MPJPE and 98.64% for 3D-PCK@150. Compared to the state-of-the-art method, our method improved MPJPE by 1.14 mm and corrected 16 % of incorrect keypoints. Chihiro Nakatsuka, Satoshi Komorita |
ICRA | 2 |
| 2022 | Efficient Linear Attention for Fast and Accurate Keypoint MatchingabstractRecently Transformers have provided state-of-the-art performance in sparse matching, crucial to realize high-performance 3D vision applications. Yet, these Transformers lack efficiency due to the quadratic computational complexity of their attention mechanism. To solve this problem, we employ an efficient linear attention for the linear computational complexity. Then, we propose a new attentional aggregation that achieves high accuracy by aggregating both the global and local information from sparse keypoints. To further improve the efficiency, we propose the joint learning of feature matching and description. Our learning enables simpler and faster matching than Sinkhorn, often used in matching the learned descriptors from Transformers. Our method achieves competitive performance with only 0.84M learnable parameters against the bigger SOTAs, SuperGlue (12M parameters) and SGMNet (30M parameters), on three benchmarks, HPatch, ETH, Aachen Day-Night. Suwichaya Suwanwimolkul, Satoshi Komorita |
ICMR | 2 |
| 2021 | Denoising 3D Human Poses from Low-Resolution Video using Variational AutoencoderabstractWe tackle the problem of refining and denoising a series of 3D human poses estimated from a low-resolution video. Low-resolution often causes the wrong pose estimation, e.g., left-right switching and the absence of keypoints. We propose to use the variational autoencoder (VAE) to remove these challenging noises. The VAE model utilizes time-series information and motion priors in denoising. From our experiments, the VAE model can reduce the pose estimation error (MPJPE) for poor-quality images by 24.37mm, from the original 105.53mm. This improves about 6.5 times over the traditional DCT approach. In addition, it removes jitters and generates smooth movements, which is helpful in recognition of human behaviors. Chihiro Nakatsuka, Satoshi Komorita |
IROS | 2 |
| 2021 | Facial Action Unit-based Deep Learning Framework for Spotting Macro- and Micro-expressions in Long Video SequencesabstractIn this paper, we utilize facial action units (AUs) detection to construct an end-to-end deep learning framework for the macro- and micro-expressions spotting task in long video sequences. The proposed framework focuses on individual components of facial muscle movement rather than processing the whole image, which eliminates the influence of image change caused by noises, such as body or head movement. Compared with existing models deploying deep learning methods with classical Convolutional Neural Network (CNN) models, the proposed framework utilizes Gated Recurrent Unit (GRU) or Long Short-term Memory (LSTM) or our proposed Concat-CNN models to learn the characteristic correlation between AUs of distinctive frames. The Concat-CNN uses three convolutional kernels with different sizes to observe features of different duration and emphasizes both local and global mutation features by changing dimensionality (max-pooling size) of the output space. Our proposal achieves state-of-the-art performance from the aspect of overall F1-scores: 0.2019 on CAS(ME)2-cropped, 0.2736 on SAMM Long Video, and 0.2118 on CAS(ME)2, which not only outperforms the baseline but is also ranked the 3rd of FME challenge 2021 for combined datasets of CAS(ME)2-cropped and SAMM-LV. Zhiguang Zhou, Megumi Komiya, Koki Kishimoto, Keisuke Nonaka, Toshiharu Horiuchi, Satoshi Komorita, Gen Hattori, Sei Naito, Yasuhiro Takishima |
ACM Multimedia | 9 |
| 2021 | Learning of low-level feature keypoints for accurate and robust detectionabstractJoint learning of feature descriptor and detector has offered promising 3D reconstruction results; however, they often lack the low-level feature awareness, which causes low accuracy in matched keypoint locations. The others employed fixed operations to select the keypoints, but the selected keypoints may not correspond to the descriptor matching. To address these problems, we propose the supervised learning of keypoint detection with low-level features. Our detector is a single CNN layer extended from the descriptor backbone, which can be jointly learned with the descriptor for maximizing the descriptor matching. This results in a state-of-the-art 3D reconstruction, especially on improving reprojection error, and the highest accuracy in keypoint detection and matching on benchmark datasets. We also present a dedicated study on evaluation metrics to measure the accuracy of keypoint detection and matching. Suwichaya Suwanwimolkul, Satoshi Komorita, Kazuyuki Tasaka |
WACV | 2 |
| 2018 | Depth Representation of LiDAR Point Cloud with Adaptive Surface Patching for Object Classification
Kanokphan Lertniphonphan, Satoshi Komorita, Kazuyuki Tasaka, Hiromasa Yanagihara |
MMM (2) | 2 |
| 2013 | Congestion-based automatic calling for improving call establishment in VoLTEabstractThrough the introduction of LTE (Long Term Evolution), which is a next-generation radio access network, the traditional circuit switching system in cellular telephony networks is migrating to VoLTE (Voice over LTE). Although LTE can provide high quality and high bandwidth, the call processing capacity of a VoLTE system is designed for normal traffic density, as is the traditional cellular network. However, its performance, measured by the successful call rate, deteriorates when overwhelmed by the large number of calls generated by the special circumstances such as large gatherings and disasters. At such times, users continue to retry calls until the call goes through, resulting in additional degradation of system performance. In this paper, we propose a novel method to improve performance in such circumstances by placing a hold on calls that are likely to fail, and automatically making those calls in IMS (IP Multimedia Subsystem), which is the main call-control function in VoLTE. Finally, we present a simulation of our proposal that demonstrates its effectiveness when users are persistent with their calls. Satoshi Komorita, Yoshinori Kitatsuji, Hidetoshi Yokota |
CCNC | 1 |
| 2011 | User-transparent reconfiguration method for self-organizing IP multimedia subsystemabstractThe NGN (Next Generation Network), which can provide advanced multimedia services over an all-IP based network, has been the subject of considerable attention for several years. While there have been tremendous efforts to develop its architecture and protocols, especially for IMS, a key technology of the NGN, its wide deployment is still a long way off. However, efforts to create an advanced signaling infrastructure able to meet many requirements have resulted in a large number of functional components and interactions between these components. Thus, the carriers are trying to explore effective ways to deploy IMS while offering value-added services. As one such approach, we have proposed a self-organizing IMS. A self-organizing IMS enables IMS functional components and corresponding physical nodes to adapt dynamically and automatically based on actual conditions such as network load and available system resources while continuing IMS operation. To realize this, service continuity for users is an important requirement when a reconfiguration occurs during operation. In this paper, we propose a mechanism that will provide service continuity to users without any impact on clients by extending the SBC (Session Border Controller). Furthermore, we implement it, show its behavior and evaluate its processing time. Satoshi Komorita, Hidetoshi Yokota, Ashutosh Dutta, Christian Makaya, Subir Das, Dana Chee, Fuchun Joseph Lin, Henning Schulzrinne |
ISCC | 1 |
| 2010 | Proposal for Distributed IMS by Hierarchical Administrative DomainabstractIn recent years, IMS architectures standardized by 3GPP have gained importance as the key technologies for the Next Generation Network that will provide high quality services over an IP network. These systems are being put to practical use and are mainly maintained and operated by telecommunication operators. In the future, deployment of new IMS services such as an IMS-based IPTV service and IMS-based office network will substantially increase the number of IMS-compatible devices. Such an increase raises concerns about overloading the telecommunication operator's IMS facilities, which are maintained in a centralized manner. On the other hand, in-house communications, such as an extension line via a PBX, can be managed internally. In this paper, we propose a method to distribute the maintenance of subscribers and localize the control traffic by separating the IMS into hierarchical and internally-controllable IMSs based on administrative domains in order to reduce the load on the telecommunication operator's facilities. Further, we implement and evaluate the proposed method, and show its benefits. Satoshi Komorita, Takeshi Kubo, Hidetoshi Yokota |
WCNC | 1 |
| 2009 | Proxy-based MOBIKE fast handoff for femtocell integrated mobile networkabstractRecently, small base stations such as femtocell and picocell are drawing attention as new components of wireless access networks. These small base stations are used not only to expand the wireless area within a large building, but also to increase bandwidth capacity and reduce the load on the core network. If multiple femtocells are deployed in an office building to expand the communication area, a handoff mechanism between femtocells is required to support the continuity of data communication. Especially, real-time applications such as VoIP require a seamless handoff mechanism. In this paper, we propose Proxy MOBIKE mechanisms that can realize an efficient handoff between femtocells and mitigate the difficulty of femtocell implementation resulting from supporting multiple protocols. We also implement our proposed mechanisms to verify and compare their effectiveness. Tsunehiko Chiba, Satoshi Komorita, Hidetoshi Yokota |
LCN | 2 |