Hiromu Yakura

dblp:207/9121 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-2558-735XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 15 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 2 since 2021Security and privacy · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Using the Blackbox in Embodied AI Art Practice: Uncertainty at the Interface
abstract
This paper argues that creative AI use can be usefully understood as embedded and positioned within artistic practice, rather than being only guided by widely cited AI design principles such as transparency or predictability. We ground this argument in a single case study: a detailed autoethnography of a trained artist's GAN-based art-making process, which surfaces how creative practice unfolds in the face of opacity, constraint, and curatorial control. The case shows how uncertainty can be a limitation around which creative practice is productively organized. Through reflexive autoethnographic memos, we identify three analytically consequential sites of creativity in the case: Conception, or the artist's intuitive mental models of the AI in use; Comprehension, or the artist use of strategies to work without full system understanding and; Linearity, or how the artist imposes temporal structure on a process with unpredictable outputs. We propose that these boundaries can act as sensitising concepts for understanding and designing support for creative AI practice. To explore their support for design and deepen our understanding of them with relation to the artist's use, we describe a documentation system intended to support problem-solving, workflow design, archiving of past solutions and knowledge-sharing among AI artists. We conclude that the boundaries and system described provide an important perspective on creative AI use: as involving negotiating system opacity through iterative curation, bespoke datasets, and contextual improvisation. We also argue that this perspective is generative, offering a grounded, practice-oriented basis for future research and tool design in the space of creative AI use.
Dorothy Yuan, Connor Graham, Hiromu Yakura, Amanda Lim
DIS3
2025 UbiLearn: Supporting English-as-a-Foreign-Language Learners in Reflecting on Conversations Using a Smartwatch MHCI033
abstract
What new opportunities can the current ubiquitous computing and AI technologies provide to support English-as-a-Foreign-Language (EFL) learners? To answer the question, we began with a formative study with EFL learners, uncovering multiple challenges during conversations with others and their desire to review such scenes later. We implemented a smartwatch prototype, UbiLearn , which features hand gesture recognition for in-situ multi-context annotation to save moments when learners face difficulty. The annotation is used to generate personalized educational material powered by speech and natural language processing. Through a series of studies, we demonstrated the feasibility and preferred usability of UbiLearn, leading to learners’ enhanced learning satisfaction. Moreover, the annotation data promoted the role of instructors by enabling the tracking of learners’ in-situ proficiency outside their tutoring sessions. We conclude by highlighting emerging opportunities for learners enabled by mobile and AI technologies, along with key considerations.
Riku Arakawa, Manami Nakagawa, Hiromu Yakura
Proc. ACM Hum. Comput. Interact.3
2025 ConverSearch: Supporting Experts in Human Behavior Analysis of Conversational Videos with a Multimodal Scene Search Tool
abstract
Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a lack of comprehensive, user-friendly tools that streamline the processing of diverse multimodal queries impedes efficiency and objectivity. To address this gap, we developed ConverSearch , a visual-programming-based tool based on insights for effective interface and implementation design derived from a formative study with experts. The tool allows experts to integrate various machine learning algorithms to capture human behavioral cues without the need for coding. Our user study, employing the System Usability Scale (SUS) and satisfaction metrics, demonstrated high user preference, reflecting the tool’s ease of use and effectiveness in supporting scene search tasks. Additionally, through a deployment trial within industrial organizations, we confirmed the tool’s objectivity, reusability, and potential to enhance expert workflows. This suggests the advantages of expert-AI collaboration in domains requiring human contextual understanding and demonstrates how customizable, transparent tools yielding reusable artifacts can support expert-driven tasks in complex, multimodal environments.
Riku Arakawa, Kiyosu Maeda, Hiromu Yakura
ACM Trans. Interact. Intell. Syst.3
2024 Automatic Assessment of Language Developmental Disorders in Non-English Contexts
abstract
Assessing the language development of children is crucial for early detection and intervention in language developmental disorders. However, it poses challenges due to the difficulty in gathering sufficient data to accurately evaluate children's language skills within limited medical consultation hours and in unfamiliar environments for children. While machine learning (ML) offers potential for efficient data collection in natural environments for children, ML models designed for specific clinical purposes are often trained on English and perform poorly in non-English contexts. In this study, we used widely available ML tools such as OpenAI's GPT and Azure's Speech-to-Text, which are trained with large datasets including diverse languages, to develop an automated pipeline for speaker identification, speech transcription, and speech content analysis tailored for assessing Japanese language development. We have demonstrated the effectiveness of our pipeline in everyday settings for children, including special educational centers and public nursery schools. Furthermore, we found that the vocabulary size of children as-sessed by our pipeline significantly correlates with developmental scales evaluated by teachers, indicating the potential value of our pipeline as a reliable assessment tool for children's language development in non-English contexts.
Monami Nishio, Ayuha Koyanagi, Ai Takamori, Keiko Hasuike, Yuki Shimoura, Hiromu Yakura, Shoi Shi
HealthCom6
2024 PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a Smartwatch
abstract
We routinely perform procedures (such as cooking) that include a set of atomic steps. Often, inadvertent omission or misordering of a single step can lead to serious consequences, especially for those experiencing cognitive challenges such as dementia. This paper introduces PrISM-Observer, a smartwatch-based, context-aware, real-time intervention system designed to support daily tasks by preventing errors. Unlike traditional systems that require users to seek out information, the agent observes user actions and intervenes proactively. This capability is enabled by the agent’s ability to continuously update its belief in the user’s behavior in real-time through multimodal sensing and forecast optimal intervention moments and methods. We first validated the steps-tracking performance of our framework through evaluations across three datasets with different complexities. Then, we implemented a real-time agent system using a smartwatch and conducted a user study in a cooking task scenario. The system generated helpful interventions, and we gained positive feedback from the participants. The general applicability of PrISM-Observer to daily tasks promises broad applications, for instance, including support for users requiring more involved interventions, such as people with dementia or post-surgical patients.
Riku Arakawa, Hiromu Yakura, Mayank Goel
UIST2
2023 Failure-Resistant Intelligent Interaction for Reliable Human-AI Collaboration
abstract
My thesis is focusing on how we can overcome the gap people have against machine learning techniques that require a well-defined application scheme and can produce wrong results. I am planning to discuss the principle of the interaction design that fills such a gap based on my past projects that have explored better interactions for applying machine learning in various fields, such as malware analysis, executive coaching, photo editing, and so on. To this aim, my thesis also shed a light on the limitations of machine learning techniques, like adversarial examples, to highlight the importance of "failure-resistant intelligent interaction."
Hiromu Yakura
AAAI1
2023 CatAlyst: Domain-Extensible Intervention for Preventing Task Procrastination Using Large Generative Models
abstract
CatAlyst uses generative models to help workers’ progress by influencing their task engagement instead of directly contributing to their task outputs. It prompts distracted workers to resume their tasks by generating a continuation of their work and presenting it as an intervention that is more context-aware than conventional (predetermined) feedback. The prompt can function by drawing their interest and lowering the hurdle for resumption even when the generated continuation is insufficient to substitute their work, while recent human-AI collaboration research aiming at work substitution depends on a stable high accuracy. This frees CatAlyst from domain-specific model-tuning and makes it applicable to various tasks. Our studies involving writing and slide-editing tasks demonstrated CatAlyst’s effectiveness in helping workers swiftly resume tasks with a lowered cognitive load. The results suggest a new form of human-AI collaboration where large generative models publicly available but imperfect for each individual domain can contribute to workers’ digital well-being.
Riku Arakawa, Hiromu Yakura, Masataka Goto
CHI2
2022 SLOPT: Bandit Optimization Framework for Mutation-Based Fuzzing
abstract
Mutation-based fuzzing has become one of the most common vulnerability discovery solutions over the last decade. Fuzzing can be optimized when targeting specific programs, and given that, some studies have employed online optimization methods to do it automatically, i.e., tuning fuzzers for any given program in a program-agnostic manner. However, previous studies have neither fully explored mutation schemes suitable for online optimization methods, nor online optimization methods suitable for mutation schemes. In this study, we propose an optimization framework called SLOPT that encompasses both a bandit-friendly mutation scheme and mutation-scheme-friendly bandit algorithms. The advantage of SLOPT is that it can generally be incorporated into existing fuzzers, such as AFL and Honggfuzz. As a proof of concept, we implemented SLOPT-AFL++ by integrating SLOPT into AFL++ and showed that the program-agnostic optimization delivered by SLOPT enabled SLOPT-AFL++ to achieve higher code coverage than AFL++ in all of ten real-world FuzzBench programs. Moreover, we ran SLOPT-AFL++ against several real-world programs from OSS-Fuzz and successfully identified three previously unknown vulnerabilities, even though these programs have been fuzzed by AFL++ for a considerable number of CPU days on OSS-Fuzz.
Yuki Koike, Hiroyuki Katsura, Hiromu Yakura, Yuma Kurogome
ACSAC3
2022 VocabEncounter: NMT-powered Vocabulary Learning by Presenting Computer-Generated Usages of Foreign Words into Users' Daily Lives
abstract
We demonstrate that recent natural language processing (NLP) techniques introduce a new paradigm of vocabulary learning that benefits from both micro and usage-based learning by generating and presenting the usages of foreign words based on the learner’s context. Then, without allocating dedicated time for studying, the user can become familiarized with how the words are used by seeing the example usages during daily activities, such as Web browsing. To achieve this, we introduce VocabEncounter, a vocabulary-learning system that suitably encapsulates the given words into materials the user is reading in near real time by leveraging recent NLP techniques. After confirming the system’s human-comparable quality of generating translated phrases by involving crowdworkers, we conducted a series of user studies, which demonstrated its effectiveness on learning vocabulary and its favorable experiences. Our work shows how NLP-based generation techniques can transform our daily activities into a field for vocabulary learning.
Riku Arakawa, Hiromu Yakura, Sosuke Kobayashi
CHI2
2022 BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking
abstract
Transcribing speech from audio files to text is an important task not only for exploring the audio content in text form but also for utilizing the transcribed data as a source to train speech models, such as automated speech recognition (ASR) models. A post-correction approach has been frequently employed to reduce the time cost of transcription where users edit errors in the recognition results of ASR models. However, this approach assumes clear speech and is not designed for unclear speech (such as speech with high levels of noise or reverberation), which severely degrades the accuracy of ASR and requires many manual corrections. To construct an alternative approach to transcribe unclear speech, we introduce the idea of respeaking, which has primarily been used to create captions for television programs in real time. In respeaking, a proficient human respeaker repeats the heard speech as shadowing, and their utterances are recognized by an ASR model. While this approach can be effective for transcribing unclear speech, one problem is that respeaking is a highly cognitively demanding task and extensive training is often required to become a respeaker. We address this point with BeParrot, the first interface designed for respeaking that allows novice users to benefit from respeaking without extensive training through two key features: parameter adjustment and pronunciation feedback. Our user study involving 60 crowd workers demonstrated that they could transcribe different types of unclear speech 32.2 % faster with BeParrot than with a conventional approach without losing the accuracy of transcriptions. In addition, comments from the workers supported the design of the adjustment and feedback features, exhibiting a willingness to continue using BeParrot for transcription tasks. Our work demonstrates how we can leverage recent advances in machine learning techniques to overcome the area that is still challenging for computers themselves with the help of a human-in-the-loop approach.
Riku Arakawa, Hiromu Yakura, Masataka Goto
IUI2
2022 Self-Supervised Contrastive Learning for Singing Voices
abstract
This study introduces self-supervised contrastive learning to acquire feature representations of singing voices. To acquire robust representations in an unsupervised manner, regular self-supervised contrastive learning trains neural networks to make the feature representation of a sample close to those of its computationally transformed versions. Similarly, we employ two transformations—pitch shifting and time stretching—considering the nature of singing voices. Nevertheless, we use them reversely: we train networks to push away representations of the transformed versions. The networks then attempt to discriminate changes in vocal timbres introduced by pitch shifting without time stretching and those in singing expressions introduced by time stretching without pitch shifting. Consequently, the acquired representations become attentive to vocal timbre and singing expression. This was confirmed through a singer identification task, where we trained a classifier to learn the relationship between the feature representations to the corresponding singer labels of 500 singers. As a result, the employed transformations helped the classifier improve the classification accuracy by 9.12% (top-1 accuracy: 63.08%) compared with the case where the feature representations fed to the classifier were acquired without the transformations (top-1 accuracy: 53.96%). Furthermore, the proposed approach can be extended to acquire feature representations attentive to either vocal timbre or singing expression but not to the other by changing how the transformations are incorporated. We particularly explored the characteristics of such vocal timbre- or singing expression-oriented feature representations against song genre, singer gender, and vocal technique, and confirmed that they successfully capture different aspects of singing voices.
Hiromu Yakura, Kento Watanabe, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 An automated system recommending background music to listen to while working
abstract
Abstract Many people listen to music while working nowadays. However, conventional recommendation systems that are designed for playing songs matching user preferences cannot be applied for such a situation. This is because previous research showed that listeners’ concentration can be negatively affected not only by music that listeners strongly dislike but also by music that the listeners strongly like. Therefore, when we consider a recommendation system to be used while working, it is desirable to avoid both songs the user likes very much and songs the user dislikes very much. Given this background, we propose FocusMusicRecommender, a system designed specifically for recommending music to listen to while working. It summarizes songs automatically and plays them successively in order to enable users to give not only “dislike (very much)” feedback via a “skip” button but also “like (very much)” feedback via a “keep listening” button. The feedback is then combined with the users’ concentration level that is estimated from their behavioral history during the playback of the corresponding song, which allows the system to obtain preference information that distinguishes between “like” and “like very much” without burdening the user who is working. Based on the preference information, the system estimates the preference levels of unplayed songs and prioritizes the songs for subsequent playback by also considering the user’s current concentration level. Our experiments showed the validity and effectiveness of the proposed method, including the accuracy of the concentration level estimation. Moreover, our user study verified the suitability of the recommendation results from both the observed behavior and obtained comments of the participants.
Hiromu Yakura, Tomoyasu Nakano, Masataka Goto
User Model. User Adapt. Interact.1
2021 Mindless Attractor: A False-Positive Resistant Intervention for Drawing Attention Using Auditory Perturbation
abstract
Explicitly alerting users is not always an optimal intervention, especially when they are not motivated to obey. For example, in video-based learning, learners who are distracted from the video would not follow an alert asking them to pay attention. Inspired by the concept of Mindless Computing, we propose a novel intervention approach, Mindless Attractor, that leverages the nature of human speech communication to help learners refocus their attention without relying on their motivation. Specifically, it perturbs the voice in the video to direct their attention without consuming their conscious awareness. Our experiments not only confirmed the validity of the proposed approach but also emphasized its advantages in combination with a machine learning-based sensing module. Namely, it would not frustrate users even though the intervention is activated by false-positive detection of their attentive state. Our intervention approach can be a reliable way to induce behavioral change in human–AI symbiosis.
Riku Arakawa, Hiromu Yakura
CHI2
2021 No More Handshaking: How have COVID-19 pushed the expansion of computer-mediated communication in Japanese idol culture?
abstract
In Japanese idol culture, meet-and-greet events where fans were allowed to handshake with an idol member for several seconds were regarded as its essential component until the spread of COVID-19. Now, idol groups are struggling in the transition of such events to computer-mediated communication because these events had emphasized meeting face-to-face over communicating, as we can infer from their length of time. I anticipated that investigating this emerging transition would provide implications because their communication has a unique characteristic that is distinct from well-studied situations, such as workplace communication and intimate relationships. Therefore, I first conducted a quantitative survey to develop a precise understanding of the transition, and based on its results, had semi-structured interviews with idol fans about their perceptions of the transition. The survey revealed distinctive approaches, including one where fans gathered at a venue but were isolated from the idol member by an acrylic plate and talked via a video call. Then the interviews not only provided answers to why such an approach would be reasonable but also suggested the existence of a large gap between conventional offline events and emerging online events in their perceptions. Based on the results, I discussed how we can develop interaction techniques to support this transition and how we can apply it to other situations outside idol culture, such as computer-mediated performing arts.
Hiromu Yakura
CHI1
2021 Tool- and Domain-Agnostic Parameterization of Style Transfer Effects Leveraging Pretrained Perceptual Metrics
abstract
Current deep learning techniques for style transfer would not be optimal for design support since their "one-shot" transfer does not fit exploratory design processes. To overcome this gap, we propose parametric transcription, which transcribes an end-to-end style transfer effect into parameter values of specific transformations available in an existing content editing tool. With this approach, users can imitate the style of a reference sample in the tool that they are familiar with and thus can easily continue further exploration by manipulating the parameters. To enable this, we introduce a framework that utilizes an existing pretrained model for style transfer to calculate a perceptual style distance to the reference sample and uses black-box optimization to find the parameters that minimize this distance. Our experiments with various third-party tools, such as Instagram and Blender, show that our framework can effectively leverage deep learning techniques for computational design support.
Hiromu Yakura, Yuki Koyama 0001, Masataka Goto
IJCAI1
2021 Reaction or Speculation: Building Computational Support for Users in Catching-Up Series Based on an Emerging Media Consumption Phenomenon
abstract
A growing number of people are using catch-up TV services rather than watching simultaneously with other audience members at the time of broadcast. However, computational support for such catching-up users has not been well explored. In particular, we are observing an emerging phenomenon in online media consumption experiences in which speculation plays a vital role. As the phenomenon of speculation implicitly assumes simultaneity in media consumption, there is a gap for catching-up users, who cannot directly appreciate the consumption experiences. This conversely suggests that there is potential for computational support to enhance the consumption experiences of catching-up users. Accordingly, we conducted a series of studies to pave the way for developing computational support for catching-up users. First, we conducted semi-structured interviews to understand how people are engaging with speculation during media consumption. As a result, we discovered the distinctive aspects of speculation-based consumption experiences in contrast to social viewing experiences sharing immediate reactions that have been discussed in previous studies. We then designed two prototypes for supporting catching-up users based on our quantitative analysis of Twitter data in regard to reaction- and speculation-based media consumption. Lastly, we evaluated the prototypes in a user experiment and, based on its results, discussed ways to empower catching-up users with computational supports in response to recent transformations in media consumption.
Riku Arakawa, Hiromu Yakura
Proc. ACM Hum. Comput. Interact.2
2020 Generate (Non-Software) Bugs to Fool Classifiers
abstract
In adversarial attacks intended to confound deep learning models, most studies have focused on limiting the magnitude of the modification so that humans do not notice the attack. On the other hand, during an attack against autonomous cars, for example, most drivers would not find it strange if a small insect image were placed on a stop sign, or they may overlook it. In this paper, we present a systematic approach to generate natural adversarial examples against classification models by employing such natural-appearing perturbations that imitate a certain object or signal. We first show the feasibility of this approach in an attack against an image classifier by employing generative adversarial networks that produce image patches that have the appearance of a natural object to fool the target model. We also introduce an algorithm to optimize placement of the perturbation in accordance with the input image, which makes the generation of adversarial examples fast and likely to succeed. Moreover, we experimentally show that the proposed approach can be extended to the audio domain, for example, to generate perturbations that sound like the chirping of birds to fool a speech classifier.
Hiromu Yakura, Youhei Akimoto, Jun Sakuma
AAAI1
2020 INWARD: A Computer-Supported Tool for Video-Reflection Improves Efficiency and Effectiveness in Executive Coaching
abstract
Video-Reflection is a common approach to realize reflection in the field of executive coaching for professional development, which presents a video recording of the coaching session to a coachee in order to make the coachee reflectively think about oneself. However, it requires a great deal of time to watch the full length of the video and is highly dependent on the skills of the coach. We expect that the quality and efficiency of video-reflection can be improved with the support of computers. In this paper, we introduce INWARD, a computational tool that leverages human behavior analysis and video-based interaction techniques. The results of a user study involving 20 coaching sessions with five coaches indicate that INWARD enables efficient video-reflection and, by leveraging meta-reflection, realizes the ameliorated outcome of executive coaching. Moreover, discussions based on comments from the participants support the effectiveness of INWARD and suggest further possibilities of computer-supported approaches.
Riku Arakawa, Hiromu Yakura
CHI2
2020 Mimicker-in-the-Browser: A Novel Interaction Using Mimicry to Augment the Browsing Experience
abstract
Humans are known to have a better subconscious impression of other humans when their movements are imitated in social interactions. Despite this influential phenomenon, its application in human-computer interaction is currently limited to specific areas, such as an agent mimicking the head movements of a user in virtual reality, because capturing user movements conventionally requires external sensors. If we can implement the mimicry effect in a scalable platform without such sensors, a new approach for designing human-computer interaction will be introduced. Therefore, we have investigated whether users feel positively toward a mimicking agent that is delivered by a standalone web application using only a webcam. We also examined whether a web page that changes its background pattern based on head movements can foster a favorable impression. The positive effect confirmed in our experiments supports mimicry as a novel design practice to augment our daily browsing experiences.
Riku Arakawa, Hiromu Yakura
ICMI2
2020 Enhancing Participation Experience in VR Live Concerts by Improving Motions of Virtual Audience Avatars
abstract
While participating in live concerts is a promising application of virtual reality (VR), it falls short of our participation experience in the real world. In particular, to increase the engagement of participants, previous studies emphasized the importance of social experience among audience members, such as the sense of co-presence elicited by sharing physical reactions or body movements synchronized with music. In this respect, a common strategy in existing platforms is to present avatars of remote human participants in a VR venue and make every avatar imitate movements of the corresponding participant. However, this strategy implicitly assumes that a not small number of users connect simultaneously to watch the same content and thus is not applicable when only a few users gather or a user is watching alone. Therefore, with the aim of providing better experience to a user who participates in live concerts as one of the audience, we examine computational approaches to enhancing the sense of co-presence through virtual audience avatars. We propose four methods of presenting avatar movements: copying the user’s own movements, copying other users’ movements, repeating beat-synchronous movements, and synthesizing machine-learning-based movements. We compare their effectiveness in a user experiment and discuss application scenarios and design implications that open up new ways of active media consumption in VR environments.
Hiromu Yakura, Masataka Goto
ISMAR1
2019 REsCUE: A framework for REal-time feedback on behavioral CUEs using multimodal anomaly detection
abstract
Executive coaching has been drawing more and more attention for developing corporate managers. While conversing with managers, coach practitioners are also required to understand internal states of coachees through objective observations. In this paper, we present REsCUE, an automated system to aid coach practitioners in detecting unconscious behaviors of their clients. Using an unsupervised anomaly detection algorithm applied to multimodal behavior data such as the subject's posture and gaze, REsCUE notifies behavioral cues for coaches via intuitive and interpretive feedback in real-time. Our evaluation with actual coaching scenes confirms that REsCUE provides the informative cues to understand internal states of coachees. Since REsCUE is based on the unsupervised method and does not assume any prior knowledge, further applications beside executive coaching are conceivable using our framework.
Riku Arakawa, Hiromu Yakura
CHI2
2019 Robust Audio Adversarial Example for a Physical Attack
abstract
We propose a method to generate audio adversarial examples that can attack a state-of-the-art speech recognition model in the physical world. Previous work assumes that generated adversarial examples are directly fed to the recognition model, and is not able to perform such a physical attack because of reverberation and noise from playback environments. In contrast, our method obtains robust adversarial examples by simulating transformations caused by playback or recording in the physical world and incorporating the transformations into the generation process. Evaluation and a listening experiment demonstrated that our adversarial examples are able to attack without being noticed by humans. This result suggests that audio adversarial examples generated by the proposed method may become a real threat.
Hiromu Yakura, Jun Sakuma
IJCAI1
2019 Neural malware analysis with attention mechanism
Hiromu Yakura, Shinnosuke Shinozaki, Reon Nishimura, Yoshihiro Oyama, Jun Sakuma
Comput. Secur.1
2018 Malware Analysis of Imaged Binary Samples by Convolutional Neural Network with Attention Mechanism
abstract
This paper presents a proposal of a method to extract important byte sequences in malware samples to reduce the workload of human analysts who investigate the functionalities of the samples. This method, by applying convolutional neural network (CNN) with a technique called attention mechanism to an image converted from binary data, enables calculation of an "attention map," which shows regions having higher importance for classification in the image. This distinction of regions enables extraction of characteristic byte sequences peculiar to the malware family from the binary data and can provide useful information for the human analysts without a priori knowledge. Furthermore, the proposed method calculates the attention map for all binary data including the data section. Thus, it can process packed malware that might contain obfuscated code in the data section. Results of our evaluation experiment using malware datasets show that the proposed method provides higher classification accuracy than conventional methods. Furthermore, analysis of malware samples based on the calculated attention maps confirmed that the extracted sequences provide useful information for manual analysis, even when samples are packed.
Hiromu Yakura, Shinnosuke Shinozaki, Reon Nishimura, Yoshihiro Oyama, Jun Sakuma
CODASPY1
2018 FocusMusicRecommender: A System for Recommending Music to Listen to While Working
abstract
This paper proposes FocusMusicRecommender, an automated system recommending background music to listen to while working. Recommendation systems matching user preferences have been widely researched even though research has shown that music that listeners strongly like is not suitable background music because it interferes with their concentration. FocusMusicRecommender plays songs that users may "neither like nor dislike" instead of "like very much." It is designed to by default summarize a song automatically so that users can give "like very much" feedback by pressing a "keep listening" button or "dislike very much" feedback by pressing a "skip" button. It uses this feedback, along with users» concentration levels estimated from their behavior history, to distinguish between the preference levels "like" and "like very much." It then estimates the preference levels of unplayed songs and selects the most suitable song by considering the user»s current concentration level. The effectiveness of the proposed feedback method and suitability of the recommendation results were verified experimentally and in user studies. Furthermore, it is confirmed that the proposed method can estimate the user»s concentration level more accurately than the previous methods.
Hiromu Yakura, Tomoyasu Nakano, Masataka Goto
IUI1