Takeharu Eda

dblp:52/5050 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
13since 2021 · last 2026
0009-0000-9814-5475ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PrePPER: A Preference Pattern-based Profiling Framework for Explainable Recommendation
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, drawing increasing attention to their application in recommendation systems.In particular, recommendation systems using natural language-based user profiles have attracted attention for improving transparency and scrutability.However, existing methods fail to fully leverage the recommendation capabilities of LLMs due to the unspecified importance of user preferences within user profiles and unmatched preference types between user profiles and item profiles.To address these challenges, we propose PrePPER, a novel preference pattern-based profiling framework designed to explicitly capture the importance of user preferences and enhance the alignment between user profiles and item profiles.PrePPER enables the extraction of users' preference patterns, which denote characteristic tendencies in user preferences, and the determination of their importance by clustering users' preferences.Specifically, we first extract users' preferences from their reviews and perform clustering on the extracted preferences.Based on the clustered preferences, we then infer users' preference patterns along with their relative importance, and construct user and item profiles using this information.Our proposed profiles incorporate the importance of user preferences and enhance the relatedness between user and item profiles, thereby improving the recommendation performance of existing recommender systems.
Taisuke Usumi, Akiko Masaki, Sanae Muramatsu, Akira Sakamoto, Takeharu Eda
LREC5
2025 HART-Pruner: A Human-Aligned Token Reduction Method for Interpretable and Efficient Vision-Language Inference
abstract
Recent advances in Multimodal Large Language Models (MLLMs) enable complex vision-language inference. However, visual inputs like images and videos consume far more tokens than text, increasing computation and memory costs. Existing token pruning methods often rely on attention scores, which lack interpretability and user control, sometimes harming performance.
Miho Takahashi, Rintaro Harada, Akiko Masaki-Kato, Takeharu Eda
HAI4
2025 Leveraging LLMs for Written to Spoken Style Data Transformation to Enhance Spoken Dialog State Tracking
Haris Gulzar, Monikka Roslianna Busto, Akiko Masaki, Takeharu Eda, Ryo Masumura
INTERSPEECH4
2024 Collaborative Intelligence For Vision Transformers: A Token Sparsity-Driven Edge-Cloud Framework
abstract
Collaborative Intelligence (CI) streamlines DNN deployment in edge-cloud infrastructure by optimizing workload between the edge and the cloud. It leverages data sparsity for both data compression for cloud transmission and the reduction of the overall computational cost of the system. Despite the rising popularity of the Vision Transformer (ViT), its higher computational overhead poses challenges for edge-cloud deployment compared to CNNs. Existing CI methods favor CNNs, utilizing feature map sparsity. In contrast, ViTs exploit token-based sparsity, complicating the direct application of CNN-optimized CI methods. This motivates us to propose a novel CI approach exploiting the token sparsity of ViT. We propose an offloading policy network, which computes scores that reveal the tokens’ relevance to the task before inference, effectively using sparse tokens in the offloaded data improving compression rate and computational cost for an edge-cloud system. Our method shows 41.98-45.75% computational cost reduction of ViT while maintaining accuracy degradation within 1.96-3.10 points and achieving a compression rate up to $36.85 \%$.
Monikka Roslianna Busto, Shohei Enomoto, Takeharu Eda
ICIP3
2024 Improving Image Coding for Machines Through Optimizing Encoder Via Auxiliary Loss
abstract
Image coding for machines (ICM) aims to compress images for machine analysis using recognition models rather than human vision. Hence, in ICM, it is important for the encoder to recognize and compress the information necessary for the machine recognition task. There are two main approaches in learned ICM; optimization of the compression model based on task loss, and Region of Interest (ROI) based bit allocation. These approaches provide the encoder with the recognition capability. However, optimization with task loss becomes difficult when the recognition model is deep, and ROI-based methods often involve extra overhead during evaluation. In this study, we propose a novel training method for learned ICM models that applies auxiliary loss to the encoder to improve its recognition capability and rate-distortion performance. Our method achieves Bjøntegaard Delta rate improvements of $27.7 \%$ and $20.3 \%$ in object detection and semantic segmentation tasks, compared to the conventional training method.
Kei Iino, Shunsuke Akamatsu, Hiroshi Watanabe 0001, Shohei Enomoto, Akira Sakamoto, Takeharu Eda
ICIP6
2024 Case-Study: Two-Phase AI Prediction for Onboard Synthetic Aperture Radar (SAR) Data Processing
abstract
SAR’s ability to operate effectively in various weather conditions makes it ideal for consistent monitoring. However, compared to other imaging methods, its inherent high data volume poses challenges for real-time AI processing. With SAR, data transmission costs to ground stations become high and providing rapid responses in time-critical scenarios with tasks such as ship detection becomes challenging. To mitigate such transmission delays, we developed a two-phase AI prediction framework from the perspective of data handling to reduce the SAR’s data volume. We explore an onboard data processing technique for SAR and show that our approach significantly reduces data volume, ensuring rapid and accurate detection. Tests on models specialized for ship detection confirm that our method maintains high detection accuracy with reduced data transmission. This showcases the potential of onboard AI computing, optimizing SAR’s advantages for time-sensitive, weather-independent applications.
Monikka Roslianna Busto, Takeharu Eda, Takuro Udagawa, Tatsuyuki Sekine
IGARSS2
2024 Parameter and Data Efficient Framework for Lightweight Change Detection
abstract
Despite extensive research in change detection methods for remote sensing, there has been relatively little work on optimizing these techniques for onboard computing scenarios. In this study, we propose a novel two-stage change detection framework designed to address the unique challenges of onboard satellite processing, particularly managing limited resources and mitigating high data volumes transmitted to earth. Our approach consists of a coarse-grained stage, serving as an early exit mechanism to efficiently predetermine the likelihood of change within an image, and a fine-grained stage, aimed at generating a detailed change map. By incorporating this dual-stage methodology, our framework effectively balances the need for rapid, resource-efficient computations and the demand for detailed, accurate change detection while reducing data to be processed and transmitted. We apply our approach across various backbone architectures to demonstrate its efficacy and report enhancements in both accuracy metrics and runtime performance. Our results showcase significant improvements, affirming the potential of our methodology to realize onboard change detection processes in remote sensing applications.
Monikka Roslianna Busto, Sepehr Nouri, Takeharu Eda
IGARSS3
2024 Technical Challenges for AI in Space Data Centers
abstract
This paper discusses the emerging field of space edge data centers (Space DC) and the unique challenges present for AI-based applications. One of the most beneficial use cases of space DC is earth observation and real-time data handling of sensing data in orbit. We discuss the specific technical challenges, introduce innovative solutions, and report some early results of these techniques.
Takeharu Eda, Monikka Roslianna Busto, Takuro Udagawa, Naoki Ishihama, Keiichi Tabata, Yoshinori Matsuo, Ikuo Yamasaki
IGARSS1
2024 A GAN-Based SAR-Optical Data Fusion Approach for Enhanced Flood Mapping
abstract
Accurate and prompt flood mapping plays a crucial role in disaster response and damage assessment. Synthetic Aperture Radar (SAR) satellites are commonly employed for this task due to their ability to capture images regardless of weather conditions, day or night. However, SAR images suffer from speckle noise and optical sensors gain semantic information that SAR sensors cannot capture. In this study, we enhanced SAR flood mapping accuracy by incorporating optical information. First, we utilized a public SAR-optical paired dataset of the flood mapping area to train image translation models that convert SAR images into an optical style. Then, we trained flood segmentation models using both the original and the translated SAR data. We confirmed that a generative adversarial network translated images with high quality, and that segmentation models trained with a combination of SAR and optical data performed better than those trained with single-domain data.
Tatsuyuki Sekine, Akihiro Yamanaka, Takeharu Eda, Takuro Udagawa, Monikka Roslianna Busto
IGARSS3
2024 Test-time Adaptation Meets Image Enhancement: Improving Accuracy via Uncertainty-aware Logit Switching
abstract
Deep neural networks have achieved remarkable success in a variety of computer vision applications. However, there is a problem of degrading accuracy when the data distribution shifts between training and testing. As a solution of this problem, Test-time Adaptation (TTA) has been well studied because of its practicality. Although TTA methods increase accuracy under distribution shift by updating the model at test time, using high-uncertainty predictions is known to degrade accuracy. Since the input image is the root of the distribution shift, we incorporate a new perspective on enhancing the input image into TTA methods to reduce the prediction’s uncertainty. We hypothesize that enhancing the input image reduces prediction’s uncertainty and increase the accuracy of TTA methods. On the basis of our hypothesis, we propose a novel method: Test-time Enhancer and Classifier Adaptation (TECA). In TECA, the classification model is combined with the image enhancement model that transforms input images into recognition-friendly ones, and these models are updated by existing TTA methods. Furthermore, we found that the prediction from the enhanced image does not always have lower uncertainty than the prediction from the original image. Thus, we propose logit switching, which compares the uncertainty measure of these predictions and outputs the lower one. In our experiments, we evaluate TECA with various TTA methods and show that TECA reduces prediction’s uncertainty and increases accuracy of TTA methods despite having no hyperparameters and little parameter overhead.
Shohei Enomoto, Naoya Hasegawa, Kazuki Adachi, Taku Sasaki, Shin'ya Yamaguchi, Takeharu Eda
IJCNN7
2023 Case Study: Two-Phase AI Prediction Techniques for Space Edge Computing
abstract
With the advancements in commercial off-the-shelf (COTS) hardware [1] and software development environments for space [2] , there are growing expectations for new services that extract valuable information from observed data in space and deliver it to users in real-time through advanced AI-based data analysis. However, even with the advances in space computing technologies, there will continue to be many cases in which complete data analysis is impossible in space. Since the data collected by remote sensors, such as synthetic aperture radar and optical cameras, have high resolution and large data volume (several gigabytes per image, several hundred gigabytes per day), transmitting this data to the ground even with fast optical communication can be time-consuming.
Takeharu Eda, Akihiro Yamanaka, Keiichi Tabata, Ikuo Yamasaki
IGARSS1
2023 miniStreamer: Enhancing Small Conformer with Chunked-Context Masking for Streaming ASR Applications on the Edge
Haris Gulzar, Monikka Roslianna Busto, Takeharu Eda, Katsutoshi Itoyama, Kazuhiro Nakadai
INTERSPEECH3
2021 Learning to Cascade: Confidence Calibration for Improving the Accuracy and Computational Cost of Cascade Inference Systems
abstract
Recently, deep neural networks have become to be used in a variety of applications. While the accuracy of deep neural networks is increasing, the confidence score, which indicates the reliability of the prediction results, is becoming more important. Deep neural networks are seen as highly accurate but known to be overconfident, making it important to calibrate the confidence score. Many studies have been conducted on confidence calibration. They calibrate the confidence score of the model to match its accuracy, but it is not clear whether these confidence scores can improve the performance of systems that use confidence scores. This paper focuses on cascade inference systems, one kind of systems using confidence scores, and discusses the desired confidence score to improve system performance in terms of inference accuracy and computational cost. Based on the discussion, we propose a new confidence calibration method, Learning to Cascade. Learning to Cascade is a simple but novel method that optimizes the loss term for confidence calibration simultaneously with the original loss term. Experiments are conducted using two datasets, CIFAR-100 and ImageNet, in two system settings, and show that naive application of existing calibration methods to cascade inference systems sometimes performs worse. However, Learning to Cascade always achieves a better trade-off between inference accuracy and computational cost. The simplicity of Learning to Cascade allows it to be easily applied to improve the performance of existing systems.
Shohei Enomoto, Takeharu Eda
AAAI2
2020 Effective Data Augmentation with Multi-Domain Learning GANs
abstract
For deep learning applications, the massive data development (e.g., collecting, labeling), which is an essential process in building practical applications, still incurs seriously high costs. In this work, we propose an effective data augmentation method based on generative adversarial networks (GANs), called Domain Fusion. Our key idea is to import the knowledge contained in an outer dataset to a target model by using a multi-domain learning GAN. The multi-domain learning GAN simultaneously learns the outer and target dataset and generates new samples for the target tasks. The simultaneous learning process makes GANs generate the target samples with high fidelity and variety. As a result, we can obtain accurate models for the target tasks by using these generated samples even if we only have an extremely low volume target dataset. We experimentally evaluate the advantages of Domain Fusion in image classification tasks on 3 target datasets: CIFAR-100, FGVC-Aircraft, and Indoor Scene Recognition. When trained on each target dataset reduced the samples to 5,000 images, Domain Fusion achieves better classification accuracy than the data augmentation using fine-tuned GANs. Furthermore, we show that Domain Fusion improves the quality of generated samples, and the improvements can contribute to higher accuracy.
Shin'ya Yamaguchi, Sekitoshi Kanai, Takeharu Eda
AAAI3
2019 An Expandable Deep Learning Inference Framework With Adjustability to Workload Requirement
abstract
Recent progress of Deep Learning has accelerated the spread of intelligent video analytics for surveillance cameras, which made us aware that the cost of GPUs is prohibitively high for analyzing huge video streams in production systems. As one solution for the issue, several special hardware optimized only for inference have appeared to reduce the cost of hardware(CAPEX) and power(OPEX) (e.g. TPU, DPU). Usually these hardware has its special compiler, which inputs trained model files saved by training frameworks, and then optimizes the model for each target device. However, the requirement on the cost and performance varies for use cases and the variation makes it difficult to implement/design applications with minimal cost satisfying customer's requirements. Based on the motivation, we designed an expandable Deep Learning inference framework with adjustablity to workload and cost requirements using USB inference devices. Our framework enables us to define stream processing pipeline with DSL and to adjust the throughput by adding/removing devices based on requirements. We implemented a sample demo application of real-time person monitoring and present the performance and adjustability with live cameras in our booth. Since inference devices are becoming cheap and have options with several connection types(USB/PCIe/miniPCIe), the framework idea is not limited only for the special USB device, but can be extended to more general framework for combine multiple chips to reduce the overall cost of production systems equipped with Deep Learning -based algorithms.
Takeharu Eda, Sanae Muramatsu, Shohei Enomoto
ICIP1
2018 A Practical Person Monitoring System for City Security
abstract
Recent progress in Deep Learning(DL) has brought many breakthroughs with incredible performance, which have not been achieved with traditional machine learning algorithms. In computer vision, DL-based methods have started to outperform humans in certain tasks and are going to impact our daily lives. We present our case study of an implementation and evaluation of our prototype real-time person-monitoring system using cutting-edge DL computer vision techniques. We used a fast and lightweight stream-processing engine for its flexibility and portability, packaged all of DL software stacks as docker containers for portability and ease of deployment, and evaluated our prototype's performance using realistic scenarios in which one hundred camera streams are gathered at centered GPU servers. We confirmed that our prototype system can monitor one hundred video streams in real-time. We also report lessons learned through our prototype implementation and discuss the future direction of person monitoring.
Takeharu Eda, Sanae Muramatsu, Keita Mikami
AVSS1
2009 Signaling emotion in tagclouds
abstract
In order to create more attractive tagclouds that get people interested in tagged content, we propose a simple but novel tagcloud where font size is determined by tag's entropy value, not the popularity to its content. Our method raises users' emotional interest in the content by emphasizing more emotional tags. Our initial experiments show that emotional tagclouds attract more attention than normal tagclouds at first look; thus they will enhance the role of tagcloud as a social signaller.
Takeharu Eda, Toshio Uchiyama, Tadasu Uchiyama, Masatoshi Yoshikawa
WWW1
2009 The Effectiveness of Latent Semantic Analysis for Building Up a Bottom-up Taxonomy from Folksonomy Tags
Takeharu Eda, Masatoshi Yoshikawa, Toshio Uchiyama, Tadasu Uchiyama
World Wide Web1
2008 Locally Expandable Allocation of Folksonomy Tags in a Directed Acyclic Graph
Takeharu Eda, Masatoshi Yoshikawa, Masashi Yamamuro
WISE1
2005 Processing XPath queries with XML summaries
abstract
Range labeling and structural joins are well-studied techniques for efficiently processing XPath queries. However, when XPath queries become long, many times of structural joins are required. To solve this problem, we developed a method to reduce the number of joins and nodes read from the disk using strong DataGuides. Our method can process single paths without any joins and twig patterns with joins amongst branching nodes and leaves in queries. Experimental results verified that our approach outperforms the best optimization technique for structural joins by factors of up to several hundreds of times.
Takeharu Eda, Makoto Onizuka, Masashi Yamamuro
CIKM1