Takahiro Komamizu

dblp:60/10773 · DBLP profile ↗
← Back
43ranked-venue papers
19as first author
26since 2021 · last 2026
0000-0002-3041-4330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 16 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 View-aware Cross-modal Distillation for Multi-view Action Recognition
abstract
The widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
WACV4
2026 MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environments
abstract
Multi-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
Pattern Recognit.4
2026 Hierarchical Global-Local Fusion for One-stage Open-vocabulary Temporal Action Detection
abstract
Open-vocabulary Temporal Action Detection (Open-vocab TAD) extends the detection scope of Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) to unseen action classes specified by vocabularies not included in the training data, within untrimmed video. Typical Open-vocab TAD methods adopt a two-stage approach that first proposes candidate action intervals and then identifies those actions. However, errors in the first stage can affect the subsequent stage and the final detection results. Moreover, conventional methods for temporal context analyses tend to focus solely on either global or local context. Focusing solely on the global context can lead to lack of momentary detail, making it difficult to distinguish one action from another. Conversely, focusing only on the local context makes it challenging to determine the start and end timings of action intervals. To address these challenges, we introduce a one-stage approach named Hierarchical Open-vocab TAD (HOTAD), consisting of two branches: Temporal Context Analysis (TCA) and Video–Text Alignment (VTA). The former utilizes Hierarchical Encoder (HE) to fuse global and local temporal features, enabling a comprehensive capture of temporal actions, while the latter branch exploits the synergy between visual and textual modalities for precisely detecting unseen actions in the Open-vocab setting. Experiments and in-depth analysis using the widely recognized datasets THUMOS14 and ActivityNet-1.3 are performed to show the effectiveness of HOTAD. The results highlight remarkable accuracy in detecting a wide range of unseen actions. Furthermore, HOTAD significantly reduces wrong labels and localizes action instances with high precision, showcasing its robustness in complex and dynamic video settings.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
abstract
Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
FG4
2025 ICDAR 25: Intelligent Cross-Data Analysis and Retrieval
abstract
The sixth edition of the Intelligent Cross-Data Analysis and Retrieval (ICDAR) workshop continues to serve as a forum for researchers and practitioners addressing the integration, analysis, and retrieval of heterogeneous data sources. While individual modalities such as wearable sensors, lifelogging cameras, and social media have been well studied, analyzing cross-data that incorporates multiple perspectives remains a crucial yet challenging task for advancing human-centered applications. In 2025, the workshop received 19 submissions, of which 7 were accepted following a careful peer-review process, resulting in an acceptance rate of 37%. The accepted papers covered a wide range of topics, including zero-shot composed image retrieval, vision-language scene understanding, adaptive modality fusion, lightweight fine-tuning with truncated SVD, and real-world federated split learning on mobile devices. By fostering interdisciplinary collaboration across domains such as well-being, disaster mitigation, mobility, food computing, and smart cities, the workshop continues to highlight emerging challenges and solutions for building intelligent, sustainable, and human-centric systems driven by cross-modal and multimodal data analytics.
Takahiro Komamizu, Marc A. Kastner 0001, Minh-Son Dao, Michael Riegler 0001, Duc-Tien Dang-Nguyen, Son N. Tran
ICMR1
2025 MUWS 2025: The 4th International Workshop on Multimodal Human Understanding for the Web and Social Media
abstract
Multimodal human understanding is an evolving interdisciplinary field integrating computer science, psychology, and social sciences to model human perception, behaviour, and biases in multimodal data. While recent advancements in multimodal learning excel in tasks like image-text synthesis, they often overlook nuanced human-centric dynamics---such as cultural, political, and individual influences on how modalities (e.g., text and images) interact, complement, or contradict each other. The 4th International Workshop on Multimodal Human Understanding (MUWS) aims at addressing these challenges, fostering novel solutions that explicitly model human perception, behaviour, and biases in multimodal data, with a particular emphasis on real-world challenges in web and social media analysis. This year edition covers two tracks: (1) human-centred multimodal understanding, such as quantifying social biases, analysing sentiment and hate speech, and modelling cross-modal interactions through interdisciplinary theories (e.g., semiotics, gestalt psychology); and (2) Multimodal understanding of global events, supported by a newly curated dataset covering news articles with diverse stances, which facilitates research on cultural framing, societal impact, and bias mitigation in vision-language models. The event features two keynotes from renowned experts from journalism and computer science, research presentations for six accepted papers, and interactive discussions to explore and discuss cutting-edge methodologies and applications in multimodal human understanding. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728481
Sherzod Hakimov, David Semedo, Eric Müller-Budack, Marc A. Kastner 0001, Takahiro Komamizu
ACM Multimedia5
2025 IntentVC 2025: The ACM Multimedia Grand Challenge on Intention-Oriented Controllable Video Captioning
abstract
The IntentVC Challenge, held in conjunction with ACM Multimedia 2025, introduces a novel benchmark for intention-oriented controllable video captioning. Unlike conventional captioning methods that generate generic, scene-level summaries, IntentVC focuses on intention-specific generation. Participants are required to produce captions explicitly conditioned on user-defined intentions, such as emphasizing a specific object tracked within a video. To support this task, the challenge provides an extended version of the LaSOT dataset annotated with intention-focused captions across 70 object categories. A standardized evaluation protocol and public leaderboard enable fair and reproducible comparison among submitted methods. By advancing research in personalized and adaptive video understanding, IntentVC offers a platform for exploring controllable vision-language modeling with practical relevance for accessibility, retrieval, and human-AI interaction. As a result, a total of 23 teams and 58 active participants have participated, and a total of 1,443 entries have been submitted. More information and resources are available at https://sites.google.com/view/intentvc/.
Takahiro Komamizu, Marc A. Kastner 0001, Yasutomo Kawanishi, Trung Thanh Nguyen 0006, Junan Chen 0004
ACM Multimedia1
2025 Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
abstract
Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard “pre-training followed by fine-tuning” paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) approach offers a promising alternative, but primarily focuses on the language components of Multimodal Large Language Models (MLLMs). Despite recent progress, PEFT remains underexplored in multimodal tasks and lacks sufficient understanding of visual information during fine-tuning the model. To bridge this gap, we propose Query-Adapter (Q-Adapter), a lightweight visual adapter module designed to enhance MLLMs by enabling efficient fine-tuning for the video captioning task. Q-Adapter introduces learnable query tokens and a gating layer into Vision Encoder, enabling effective extraction of sparse, caption-relevant features without relying on external textual supervision. We evaluate Q-Adapter on two well-known video captioning datasets, MSR-VTT and MSVD, where it achieves state-of-the-art performance among the methods that take the PEFT approach across BLEU@4, METEOR, ROUGE-L, and CIDEr metrics. Q-Adapter also achieves competitive performance compared to methods that take the full fine-tuning approach while requiring only 1.4% of the parameters. We further analyze the impact of key hyperparameters and design choices on fine-tuning effectiveness, providing insights into optimization strategies for adapter-based learning. These results highlight the strong potential of Q-Adapter in balancing caption quality and parameter efficiency, demonstrating its scalability for video–language modeling.
Junan Chen 0004, Trung Thanh Nguyen 0006, Takahiro Komamizu, Ichiro Ide
MMAsia3
2025 Quantifying Image-Adjective Associations by Leveraging Large-Scale Pretrained Models
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
MMM (4)3
2025 Towards Visual Storytelling by Understanding Narrative Context Through Scene-Graphs
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (4)4
2024 R-DiP: Re-ranking Based Diffusion Pre-computation for Image Retrieval
Tatsuya Kato, Takahiro Komamizu, Ichiro Ide
DEXA (2)2
2024 One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-Scale and Action Label Features
abstract
Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing and classifying actions based on a predefined set of categories. In contrast, Open-vocab TAD goes further and is not limited to these predefined categories. This is particularly useful in real-world scenarios where the variety of actions in videos can be vast and not always predictable. The prevalent methods in Open-vocab TAD typically employ a 2-stage approach, which involves generating action proposals and then identifying those actions. However, errors made during the first stage can adversely affect the subsequent action identification accuracy. Additionally, existing studies face challenges in handling actions of different durations owing to the use of fixed temporal processing methods. Therefore, we propose a L-stage approach consisting of two primary modules: Multi-scale Video Analysis (MVA) and Video-Text Alignment (VTA). The MVA module captures actions at varying temporal resolutions, overcoming the challenge of detecting actions with diverse durations. The VTA module leverages the synergy between visual and textual modalities to precisely align video segments with corresponding action labels, a critical step for accurate action identification in Open-vocab scenarios. Evaluations on widely recognized datasets THUMOSl4 and ActivityNet-I.3, showed that the proposed method achieved superior results compared to the other methods in both Open-vocab and Closed-vocab settings. This serves as a strong demonstration of the effectiveness of the proposed method in the TAD task.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
FG3
2024 Feature Extraction for Claim Check-Worthiness Prediction Tasks Using LLM
Yuka Teramoto, Takahiro Komamizu, Mitsunori Matsushita, Kenji Hatano
iiWAS (1)2
2024 ICDAR 24: Intelligent Cross-Data Analysis and Retrieval
abstract
Our workshop aims to provide a platform for both academic and industrial professionals engaged in the analysis and retrieval of cross-data from diverse perspectives, with a particular emphasis on wearable and ambient sensors, lifelog cameras, social networks, and surrounding sensors.Despite numerous studies exploring individual viewpoints, there remains a significant gap in the analysis and retrieval of cross-data to maximize benefits for humanity.Additionally, challenges such as data security and distributed learning for cross-modal model training and inference arise when dealing with large and distributed datasets.We invite researchers to contribute to this initiative, with the overarching goal of fostering the development of a smart and sustainable society through the efficient utilization of intelligent cross-data analysis and retrieval techniques.
Minh-Son Dao, Michael Riegler 0001, Duc-Tien Dang-Nguyen, Hanh-Nhi Tran, R. Uday Kiran, Takahiro Komamizu
ICMR6
2024 Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
abstract
Text-to-image diffusion models sometimes depict blended concepts in the generated images. One promising use case of this effect would be the nonword-to-image generation task which attempts to generate images intuitively imaginable from a non-existing word (nonword). To realize nonword-to-image generation, an existing study focused on associating nonwords with similar-sounding words. Since each nonword can have multiple similar-sounding words, generating images containing their blended concepts would increase intuitiveness, facilitating creative activities and promoting computational psycholinguistics. Nevertheless, no existing study has quantitatively evaluated this effect in either diffusion models or the nonword-to-image generation paradigm. Therefore, this paper first analyzes the conceptual blending in a pretrained diffusion model, Stable Diffusion. The analysis reveals that a high percentage of generated images depict blended concepts when inputting an embedding interpolating between the text embeddings of two text prompts referring to different concepts. Next, this paper explores the best text embedding space conversion method of an existing nonword-to-image generation framework to ensure both the occurrence of conceptual blending and image generation quality. We compare the conventional direct prediction approach with the proposed method that combines k-nearest neighbor search and linear regression. Evaluation reveals that the enhanced accuracy of the embedding space conversion by the proposed method improves the image generation quality, while the emergence of conceptual blending could be attributed mainly to the specific dimensions of the high-dimensional text embedding space.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
ACM Multimedia3
2024 Action Selection Learning for Multi-label Multi-view Action Recognition
abstract
Multi-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMAsia3
2024 Cross-modal recipe retrieval based on unified text encoder with fine-grained contrastive learning
abstract
Cross-modal recipe retrieval is vital for transforming visual food cues into actionable cooking guidance, making culinary creativity more accessible. Existing methods separately encode the recipe Title, Ingredient, and Instruction using different text encoders, then aggregate them to obtain recipe feature, and finally match it with encoded image feature in a joint embedding space. These methods perform well but require significant computational cost. In addition, they only consider matching the entire recipe and the image but ignore the fine-grained correspondence between recipe components and the image, resulting in insufficient cross-modal interaction. To this end, we propose U nified T ext E ncoder with F ine-grained C ontrastive L earning (UTE-FCL) to achieve a simple but efficient model. Specifically, in each recipe, UTE-FCL first concatenates each of the Ingredient and Instruction texts composed of multiple sentences as a single text. Then, it connects these two concatenated texts with the original single-phrase Title to obtain the concatenated recipe. Finally, it encodes these three concatenated texts and the original Title by a Transformer-based Unified Text Encoder (UTE). This proposed structure greatly reduces the memory usage and improves the feature encoding efficiency. Further, we propose fine-grained contrastive learning objectives to capture the correspondence between recipe components and the image at Title, Ingredient, and Instruction levels by measuring the mutual information. Extensive experiments demonstrate the effectiveness of UTE-FCL compared to existing methods.
Haruya Kyutoku, Keisuke Doman, Takahiro Komamizu, Ichiro Ide, Jiangbo Qian
Knowl. Based Syst.4
2024 Computational measurement of perceived pointiness from pronunciation
abstract
Abstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.3
2024 Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.3
2023 Towards Ensemble-Based Imbalanced Text Classification Using Metric Learning
Takahiro Komamizu
DEXA (2)1
2023 NarSUM '23: The 2nd Workshop on User-Centric Narrative Summarization of Long Videos
abstract
With video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization, etc.
Mohan Kankanhalli, Ioannis Patras, Jianquan Liu, Yongkang Wong, Takahiro Komamizu, Satoshi Yamazaki, Karen Stephen, Kajal Kansal
ACM Multimedia5
2023 RecipeMeta: Metapath-enhanced Recipe Recommendation on Heterogeneous Recipe Network
abstract
Recipe is a set of instructions that describes how to make food. It can help people from the preparation of ingredients, food cooking process, etc. to prepare the food, and increasingly in demand on the Web. To help users find the vast amount of recipes on the Web, we address the task of recipe recommendation. Due to multiple data types and relationships in a recipe, we can treat it as a heterogeneous network to describe its information more accurately. To effectively utilize the heterogeneous network, metapath was proposed to describe the higher-level semantic information between two entities by defining a compound path from peer entities. Therefore, we propose a metapath-enhanced recipe recommendation framework, RecipeMeta, that combines GNN (Graph Neural Network)-based representation learning and specific metapath-based information in a recipe to predict User-Recipe pairs for recommendation. Through extensive experiments, we demonstrate that the proposed model, RecipeMeta, outperforms state-of-the-art methods for recipe recommendation.
Jialiang Shi, Takahiro Komamizu, Keisuke Doman, Haruya Kyutoku, Ichiro Ide
MMAsia2
2023 Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (1)4
2023 Discovering Phonesthemic Clusters in Readings of Kanji Characters toward Exploring Phonestheme in Japanese
Akira Yoshida, Chihaya Matsuhira, Hirotaka Kato, Takatsugu Hirayama, Takahiro Komamizu, Ichiro Ide
PACLIC5
2022 Detection of Birds in a 3D Environment Referring to Audio-Visual Information
abstract
We propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed.
Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi
AVSS6
2021 MMEnsemble: Imbalanced Classification Framework Using Metric Learning and Multi-sampling Ratio Ensemble
Takahiro Komamizu
DEXA (2)1
2020 MUEnsemble: Multi-ratio Undersampling-Based Ensemble Framework for Imbalanced Data
Takahiro Komamizu, Risa Uehara, Yasuhiro Ogawa, Katsuhiko Toyama
DEXA (2)1
2020 Random walk-based entity representation learning and re-ranking for entity search
Takahiro Komamizu
Knowl. Inf. Syst.1
2019 Japanese Mistakable Legal Term Correction using Infrequency-aware BERT Classifier
abstract
We propose a method that assists legislative drafters in locating inappropriate legal terms in Japanese statutory sentences and suggests corrections. We focus on sets of mistakable legal terms whose usages are defined in legislation drafting rules. Our method predicts suitable legal terms using a classifier based on a BERT (Bidirectional Encoder Representations from Transformers) model. We apply three techniques in training the BERT classifier, specifically, preliminary domain adaptation, repetitive soft undersampling, and classifier unification. These techniques cope with two levels of infrequency: legal term-level infrequency that causes class imbalance and legal term set-level infrequency that causes underfitting. Concretely, preliminary domain adaptation improves overall performance by providing prior knowledge of statutory sentences, repetitive soft undersampling improves performance on infrequent legal terms without sacrificing performance on frequent legal terms, and classifier unification improves performance on infrequent legal term sets by sharing common knowledge among legal term sets. Our experiments show that our classifier outperforms conventional classifiers using Random Forest or a language model, and that all three training techniques contribute to performance improvement.
Takahiro Yamakoshi, Takahiro Komamizu, Yasuhiro Ogawa, Katsuhiko Toyama
IEEE BigData2
2018 Learning Interpretable Entity Representation in Linked Data
Takahiro Komamizu
DEXA (1)1
2018 Japanese Legal Term Correction Using Random Forests
abstract
We propose a method that assists legislation officers in finding inappropriate Japanese legal terms in Japanese statutory sentences and suggests corrections. In particular, we focus on sets of similar legal terms whose usages are defined in legislation drafting rules. Our method predicts suitable legal terms in statutory sentences using Random Forest classifiers, each of which is optimized for each set of similar legal terms. Our experiment shows that our method outperformed existing modern word prediction methods using neural language models.
Takahiro Yamakoshi, Takahiro Komamizu, Yasuhiro Ogawa, Katsuhiko Toyama
JURIX2
2017 Implicit order join: Joining log data with property data by discovering implicit order-oriented keys with human assistance
abstract
Data integration is still laboursome task when integrating data are not consistently managed. Such inconsistency can happen easily in real-world situations, such as properties of objects are managed by a central organization and trajectories (or logs) of the objects are recorded by other peripheral organizations. This paper deals with a case of missing ordering information. Integrating property data and log data without ordering information causes duplicated results. In order to solve this problem, this paper proposes a join algorithm, called implicit order join, which discovers implicit ordering information from both property data and log data with help of partial true integrated results from human assistance. With the discovered ordering information, the implicit order join enables to integrate the property data and log data. In order to discover the implicit ordering information, ordering correlation between attribute sequences of property data and log data should be found from comprehensive examination of possible attribute sequence pairs. The potential number of sequence pairs is as high as factorial order of the number of attributes. Therefore, this paper develops a heuristic approach to prune unnecessary examinations based on ordering dependency between attribute sequences. Experimental evaluation in this paper indicates that implicit order join can reduce 77% labouring tasks for integration and the pruning method reduces the number of attribute sequences in orders of magnitude.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
IEEE BigData1
2017 Analytical toolbox for smart city applications: Garbage collection log use case
abstract
Analyzing and feeding back the results on real-world services are important missions in the Big Data era to realize smart city. However, analyzing real-world data is still challenging because of dirtiness of data and large variety of analytic requirements. To cope with the challenges, this paper proposes and develops an analytical toolbox for smart city applications. The analytical toolbox consists of three phases: preparation, analysis, and visualization. The preparation phase deals with the dirtiness of the data by including fundamental data cleansing techniques and data integration techniques. The analysis phase is responsible for ETL (extract, transform and load) process and analytical query processing from the next phase. The visualization phase deals with analytical requirements from users and visualization of analytical results. This paper showcases a real-world use case of the proposed analytical toolbox. The use case is now open in public with help of Fujisawa city, Japan, and this fact indicates that the proposed analytical toolbox is feasible for real-world data analysis and feeding back to citizens.
Takahiro Komamizu, Jin Nakazawa, Toshiyuki Amagasa, Hiroyuki Kitagawa, Hideyuki Tokuda
IEEE BigData1
2017 CROISSANT: centralized relational interface for web-scale SPARQL endpoints
abstract
Searching over Linked Data requires large efforts to users, which include knowing locations of suitable SPARQL endpoints and writing appropriate SPARQL queries in terms of language standards as well as the underlying structure of Linked Data. This situation degrades usability of Linked Data, thus is highly problematic. To resolve this problem, this paper proposes CROISSANT which is a centralized view management system for SPARQL endpoints on the Web. CROISSANT stores pre-defined view definitions, and provides a searchable interface for the views to users. To realize CROISSANT, query processing performance is a big issue, because CROISSANT has to communicate with remote SPARQL endpoints and it takes time to receive results. To cope with this issue, this paper proposes four optimization techniques, namely, view materialization, selection push-down, projection push-down, and view query merge. Experimental evaluation demonstrates these optimizations improve query processing performance.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
iiWAS1
2017 SOLA: Stream OLAP-based Analytical Framework for Roadway Maintenance
abstract
Maintaining infrastructures (e.g., roadway) is a critical issue for local governments. Data from physical devices and reports from citizens through social networks are helpful to observe conditions of infrastructures. This paper proposes a framework called SOLA for integrating and analysing data from multiple sources including streaming data and static data for roadway management. The framework integrates data from multiple sources in the way of stream OLAP architecture, and analyses the integrated data in terms of OLAP analysis. This paper applies the framework to support roadway managements of local governments, and develops the application called SOLAR. SOLAR aims at providing historical views of roadway patrols as well as roadway statuses for assisting in determining roadway patrolling schedules. The real-world use case on a city exhibits the applicability of SOLAR with positive feedbacks from city officers. SOLA is a promising framework for big data analysis and smart city applications, as the number, amount, and speed of generating data increase in the era of big data and smart city.
Takahiro Komamizu, Toshiyuki Amagasa, Salman Ahmed Shaikh, Hiroaki Shiokawa, Hiroyuki Kitagawa
MEDES1
2017 Exploring Identical Users on GitHub and Stack Overflow
abstract
Analyzing behaviours of developers in different platforms (in particular, GitHub and Stack Overflow in this paper) can reveal interesting facts related to development activities.There are only few datasets for analysing crossplatform user behaviours, especially across GitHub and Stack Overflow.Users on GitHub and Stack Overflow are identifiable by equivalences of email addresses.In order to increase the number of identifiable users on these datasets, this paper retrieves potentially identifiable users between GitHub and Stack Overflow not relying only on email addresses.This paper employs a classification-based link prediction, which design the user identification problem as a link prediction problem on the bipartite graph consisting of users of GitHub and those of Stack Overflow.With the identification method, this paper generates a probabilistic dataset containing pairs of users with probabilities (or confidences).This paper, as well, publishes the identification tool in order to enable further data generation on appearing datasets of GitHub, Stack Overflow and others.The generated dataset and tool are highly helpful to accelerate researches on mining software repositories.
Takahiro Komamizu, Yasuhiro Hayase, Toshiyuki Amagasa, Hiroyuki Kitagawa
SEKE1
2016 Visual Spatial-OLAP for Vehicle Recorder Data on Micro-sized Electric Vehicles
abstract
Analyzing vehicle recorder data of electric vehicles (EVs) reveals how the EVs are used. This paper proposes an OLAP framework to support analyzing trajectories in vehicle recorder data and applies the framework to vehicle recorder data of EVs. The framework consists of ETL (extract, transform, and load) process for trajectory data and visualization for analyzing the data. The ETL process includes hierarchy definitions for spatial and temporal dimensions, as well as aggregation functions for trajectory data. In the subsequent visualization phase, the framework displays results of OLAP operations on map interface. To ensure the applicability of the framework for real applications, we apply the framework to vehicle recorder data of micro-sized EVs (or μEVs), which are smaller EVs with one or two passengers including one driver and can drive at most 100km distance without charging on the way. The application realizes that the framework successfully enables analyses on the trajectory data for real analytic requirements.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
IDEAS1
2015 SPOOL: a SPARQL-based ETL framework for OLAP over linked data
abstract
Linked Data (or LD) has promoted publishing information, and links published information (e.g., vocabularies and facts) for utilization. There are increasing number of LD datasets containing numerical data such as statistics. Analyses using such data require dedicated programs to extract, transform, and load (or ETL) for preparation. Thus, a large effort of developers is required. Also, the LD datasets tend to be large and the dumps (or snapshots) for the datasets easily become not up-to-date due to update frequency of the datasets. Hence, downloading dumps of LD datasets to ETL for OLAP can miss latest records. This paper proposes a framework called SPOOL, which attempts to reduce the effort and to ETL latest numerical records data from LD datasets for OLAP through SPARQL endpoints without downloading whole datasets. SPOOL provides series of SPARQL queries extracting objects and attributes from LD datasets, and converts them into star/snowflake schemas, and materialize relevant triples as fact and dimension tables for OLAP. The applicability of SPOOL is evaluated using exiting LD datasets on the Web, and SPOOL successfully processes the LD datasets to ETL for OLAP.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
iiWAS1
2014 A scheme of automated object and facet extraction for faceted search over XML data
abstract
Applying faceted search for XML data enables users to search XML data in an interactive manner. However, applying faceted search is challenging, because faceted search requires target subtrees (objects) and facets to be defined before-hand. To this problem, existing works assume that such objects and/or facets are defined manually, but it is infeasible to manually specify objects and facets in particular when the XML data are huge and/or its structure is quite complicated. To address this problem, this paper proposes an automatic extraction scheme of objects and facets from XML data. We propose two approaches, namely frequency-based approach and semantic-based approach, and also hybrid approach of them. The basic ideas of these approaches are that the frequently occurring XML elements seem to be objects and facets, and such XML elements may have semantically meaningful name. Although the proposed approaches are rather simple, the experiments using real world XML data show that the proposed approaches can automatically extract objects and facets from the XML data.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
IDEAS1
2014 Extracting Facets from Textual Contents for Faceted Search over XML Data
abstract
Faceted search for XML data is one of the promising exploration methods with high usability to find desired subtrees from a given XML data. This paper proposes improved approach of faceted search over XML data by utilizing facets containing unique and longer textual values, like titles of papers in bibliographic database. Our approach is to extract suitable terms which categorize the current results into several groups. Also we propose a task designing method for evaluating exploratory search by defining specificity of tasks called specification level, and we introduce how to generate tasks with given specification level as well. With this task design, we evaluate our proposed approach and the results show our proposed approach improves search performance comparing with the previous approaches, especially when tasks have low specification levels.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
iiWAS1
2012 A Scheme of Fragment-Based Faceted Image Search
Takahiro Komamizu, Mariko Kamie, Kazuhiro Fukui, Toshiyuki Amagasa, Hiroyuki Kitagawa
DEXA (2)1
2011 A framework of faceted navigation for XML data
abstract
In this paper, we propose a framework of faceted navigation over XML data. General faceted navigation schemes are used to browse objects (or records) containing multiple properties. However, because XML is semi-structured in nature, it is not straightforward to apply faceted navigation to XML data. Specifically, we need to cope with three major technical issues: 1) objects in XML data are not predetermined, 2) objects may have flexible and/or recursive structure, and 3) properties of an object need to be automatically detected and extracted. To these problems, in this paper, we formulate faceted navigation over XML data by giving definitions of class, property, object, and facet in XML data. We then formulate typical user interactions in faceted navigation as operations over aforementioned concepts (class, object, and facet). We also propose a framework based on these definitions and operations, and construct a prototype system based on the framework. Finally, we show experimental evaluations using the prototype system to show the effectiveness of our proposed scheme.
Takahiro Komamizu, Toshiyuki Amagasa, Hiroyuki Kitagawa
iiWAS1
2011 FACTUS: Faceted Twitter User Search Using Twitter Lists
Takahiro Komamizu, Yuto Yamaguchi, Toshiyuki Amagasa, Hiroyuki Kitagawa
WISE1