VLDB 2026 Research / reviewers in the wild / expert
Koji Inoue
dblp:62/852
· DBLP profile ↗
96ranked-venue papers
19as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 33 · 11 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationabstractWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen |
ACL (1) | 17 |
| 2026 | SFQ-Based CJoin Gate Implementation for Ultra-Low-Power Brownian Logic CircuitsabstractThe increasing energy consumption of data centers has highlighted the urgent need for ultra-low-power computing solutions. Single Flux Quantum (SFQ) circuits, which utilize superconducting Josephson junctions, offer promising advantages in terms of speed and energy efficiency; however, their power consumption has been limited by the need for noise suppression mechanisms. To address this, we propose the practical SFQ Brownian Logic Circuits (SBLCs), which exploit thermal noise for stochastic signal propagation, significantly reducing static power consumption. Especially, we address the three challenges: susceptibility to manufacturing variations, the absence of an actual CJoin gate implementation, which is essential for processing, and a lack of demonstrated practical advantages. This paper evaluates the robustness of SBLCs to manufacturing variations, proposes the first SFQ-based CJoin gate implementation, and demonstrates a Ripple Carry Adder (RCA) with a 3167x improvement in energy efficiency compared to traditional SFQ and CMOS circuits, confirming the superiority of SBLCs for future computing systems. Soshi Takagi, Masamitsu Tanaka, Koji Inoue, Satoshi Kawakami |
DATE | 3 |
| 2026 | Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language
Yoshiki Tanaka, Ryuichi Uehara, Koji Inoue, Michimasa Inaba |
LREC | 3 |
| 2026 | On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous LevelsabstractIn multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we revisit this assumption by analyzing address as a continuous phenomenon. Using a multi-party human dialogue corpus annotated by multiple annotators, we construct both binary address labels derived from majority-vote addressee labels and continuous address levels inferred from annotator judgments using a latent-variable model. We then examine how these representations relate to turn-taking as well as listener behaviors, including gaze and backchannels. Our results show that, in addition to turn-taking, both gaze and backchannels are associated with address. Furthermore, models using continuous address levels achieve better predictive fit than those using discrete labels, suggesting that address may exhibit graded structure. Finally, we discuss the future directions of addressee detection research based on the findings of this study. Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 2 |
| 2026 | I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue SystemabstractEmotional validation - explicitly acknowledging that a user’s feelings make sense - has proven therapeutic value but has received little computational attention. The emotional validation in dialogue systems can be decomposed into (i) validating response identification, (ii) validation timing detection, and (iii) validating response generation. To support research on all three subtasks, we release M-EDESConv, a 120k English–Japanese multilingual corpus created through hybrid manual–automatic annotation, and M-TESC, a multilingual spoken-dialogue test set. For timing detection, we propose MEGUMI, a Multilingual Emotion-aware Gated Unit for Mutual Integration, that fuses frozen XLM-RoBERTa semantics with language-specific emotion encoders via cross-modal attention and gated fusion. MEGUMI shows superior performance on both the M-EDESConv and M-TESC datasets, both objectively and subjectively. Finally, our EmoValidBench benchmarks of GPT-4.1 Nano and Llama-3.1 8B indicate that current LLMs generate contextually similar, diverse validating responses, but emotional understanding remains a major area for improvement. Zi Haur Pang, Yahui Fu 0001, Koji Inoue, Tatsuya Kawahara |
SIGDIAL | 3 |
| 2026 | Robot-Mediated Multi-Party Conversation Aimed at Affect Improvement for Psychiatric PatientsabstractThis paper describes a multi-party attentive listening system that interacts with two persons to familiarize each other through conversation. We are mainly targeting social implementation in hospitals to contribute to the rehabilitation of people with psychiatric disorders to promote affect improvement in terms of pleasure and arousal. We conducted an experiment in a psychiatric outpatient-daycare program. Twenty daycare attendees participated in a three-party conversation session between a pair of two humans and a humanoid robot. One of the paired participants talked about his/her favorite topic and was attentively listened to by the other and the robot. In a subsequent session, the human pairs switched each other's roles. The subjective evaluations showed that both the pleasure and arousal of the participants were significantly improved after the conversation. The participants rated the impression of the robot as easier to talk with than strangers. They also rated that they could understand and feel familiar with significantly more their human talk partner after the conversational session. The multiple linear regression analysis showed that participants became more pleasant and verbal when stimulated by backchannels and questions from both the human listener and the robot. It suggests that those who expressed themself using more words received positive impressions. Keiko Ochi, Divesh Lala, Koji Inoue, Tatsuya Kawahara, Hirokazu Kumazaki |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Can LLMs be Surprised? Evaluation and Analysis of Surprise Expression of LLMsabstractWe investigate methods for virtual agents or humanoid robots to express surprise in dialogue by using LLMs. We created a newly annotated dialogue dataset focusing on surprise. Our findings indicate that accurately expressing surprise in dialogue is a challenging task. They also suggest directions for improvement–such as more appropriate modeling of commonness–and identify the feature of surprise itself. Motoori Takeuchi, Koji Inoue, Keiko Ochi, Tatsuya Kawahara |
HAI | 2 |
| 2025 | CCMI 2025: Cross-Cultural Multimodal Interaction
Koji Inoue, Shogo Okada, Divesh Lala, Sahba Zojaji, Nancy F. Chen, Tatsuya Kawahara |
ICMI | 1 |
| 2025 | Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
Kazushi Kato, Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
ICMI | 2 |
| 2025 | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue SystemsabstractTurn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems. Mikey Elmers, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2025 | A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field ExperimentabstractTurn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world settings remains underexplored. In this study, we propose a noise-robust voice activity projection (VAP) model, based on a Transformer architecture, to enhance real-time turn-taking in dialogue robots. To evaluate the effectiveness of the proposed system, we conducted a field experiment in a shopping mall, comparing the VAP system with a conventional cloud-based speech recognition system. Our analysis covered both subjective user evaluations and objective behavioral analysis. The results showed that the proposed system significantly reduced response latency, leading to a more natural conversation where both the robot and users responded faster. The subjective evaluations suggested that faster responses contribute to a better interaction experience. Koji Inoue, Yuki Okafuji, Jun Baba, Yoshiki Ohira, Katsuya Hyodo, Tatsuya Kawahara |
IROS | 1 |
| 2025 | SuperSFQ: A Hardware Design to Realize High-Frequency Superconducting ProcessorsabstractSuperconducting computing using single flux quantum (SFQ) technology has been recognized as a promising post-Moore's law era technology thanks to its extremely low power and high performance.Therefore, many researchers have proposed various SFQbased circuits (e.g., ALU, register file) and architectures (e.g., NPU, CPU) to exploit the potential.However, due to the absence of a reliable and high-frequency clocking scheme, general SFQ circuits cannot operate at high frequencies, making all architectural efforts for high-performance SFQ computing ineffective.In this paper, we propose SuperSFQ, a new design methodology for SFQ hardware that unlocks the high-frequency potential of SFQ technology by co-designing the clocking scheme, circuitry, and architecture.First, we propose SuperClocking, a new clocking scheme that enables high frequency in general SFQ hardware.Second, we implement an SFQ-based synchronizer to realize the reliable operation of SuperClocking.Finally, we provide two architectural design guidelines and corresponding solutions to ensure the functional correctness of SuperClocking in general SFQ devices.By applying our clocking scheme, synchronizer, and guidelines to the latest general-purpose SFQ CPU, SuperSFQ achieves up to 62.5 times higher frequency and improves single-thread and multithread performance by 17 and 62.5 times, respectively, compared to conventional designs, with only 34.4% Josephson junction overhead.In addition, to demonstrate the generality of SuperSFQ, we apply SuperSFQ to 48 different benchmark circuits, achieving 88.5 times higher frequency compared to conventional designs, on average. Junhyuk Choi, Juwon Hong, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Dongmoon Min, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 8 |
| 2025 | Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity ProjectionabstractKoji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara |
NAACL (Long Papers) | 1 |
| 2025 | Prompt-Guided Turn-Taking PredictionabstractTurn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study, we propose a novel model that enables turn-taking prediction to be dynamically controlled via textual prompts. This approach allows intuitive and explicit control through instructions such as “faster” or “calmer,” adapting dynamically to conversational partners and contexts. The proposed model builds upon a transformer-based voice activity projection (VAP) model, incorporating textual prompt embeddings into both channel-wise transformers and a cross-channel transformer. We evaluated the feasibility of our approach using over 950 hours of human-human spoken dialogue data. Since textual prompt data for the proposed approach was not available in existing datasets, we utilized a large language model (LLM) to generate synthetic prompt sentences. Experimental results demonstrated that the proposed model improved prediction accuracy and effectively varied turn-taking timing behaviors according to the textual prompts. Koji Inoue, Mikey Elmers, Yahui Fu 0001, Zi Haur Pang, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
SIGDIAL | 1 |
| 2024 | Multilingual Turn-taking Prediction Using Voice Activity ProjectionabstractThis paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS). Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze |
LREC/COLING | 1 |
| 2024 | Late Breaking Results: Single Flux Quantum Based Brownian Circuits for Ultra-Law-Power ComputingabstractThis paper proposes a random walk circuit imple-mentation with single flux quantum devices, essential for Brownian circuits, to reduce processing energy consumption dramatically. SPICE-based simulation demonstrating its functional operation and random walks can be achieved via the Shapiro- Wilk test. Furthermore, we developed a Monte Carlo simulator for Brownian circuits, enabling functionality verification and computation step distribution analysis. Latency/energy evaluation using a half-adder as a case study revealed that proposed circuits could reduce energy consumption by 1/1260 and offer an opportunity for low-power computing systems. Satoshi Kawakami, Yusuke Ohtusbo, Koji Inoue, Masamitsu Tanaka |
DATE | 3 |
| 2024 | Entrainment Analysis and Prosody Prediction of Subsequent Interlocutor's Backchannels in Dialogue
Keiko Ochi, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2024 | SuperCore: An Ultra-Fast Superconducting Processor for Cryogenic ApplicationsabstractSuperconductor single-flux-quantum (SFQ) logic family has been recognized as a promising technology for cryogenic applications (e.g., quantum computing, astronomy, metrology) thanks to its ultra-fast and low-energy characteristics. Therefore, recent efforts in SFQ-based computing have focused on developing fast and low-power SFQ processors for cryogenic applications. However, there still has been little progress toward a convincing SFQ processor design due to the critical performance challenges originating from its extremely deep pipeline. In this paper, we propose a super-fast and low-power in-order SFQ processor by tackling the challenges from the deep pipeline. First, we develop a minimal-depth SFQ processor pipeline with novel architecture-level ideas. Next, we conduct in-depth performance analyses and identify three real performance bottlenecks in the deeply pipelined SFQ processors (i.e., stall/flush logic, RAW stall, fetch unit). Finally, we propose SuperCore, our super-fast SFQ-based processor architecture, with three SFQ-friendly solutions that effectively resolve the identified bottlenecks. With our solutions applied, SuperCore achieves 11 times speed-up over the SFQ processor baseline. In addition, SuperCore achieves six times speed-up and consumes up to 193 times less power compared to in-order CMOS processors running at 4K. Junhyuk Choi, Ilkwon Byun, Juwon Hong, Dongmoon Min, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 9 |
| 2024 | SPDID: A Secure and Privacy-Preserving Decentralized Identity utilizing Blockchain and PUF
Yueyue He, Wenxuan Fan, Koji Inoue |
TrustCom | 3 |
| 2024 | CrowdChain: A privacy-preserving crowdfunding system based on blockchain and PUF
Yueyue He, Koji Inoue |
Peer Peer Netw. Appl. | 2 |
| 2023 | Evaluating floating-point multipliers with opto-electrical hybrid circuitsabstractNanophotonic technology has ultra-low latency and low energy consumption characteristics, and several previous research have shown the potential of photonic-based computation. However, existing optical circuits are based on low-bit-width integer operations with analog manners, so there is less research on optical floating-point operations. Providing the benefits of analog optical circuits (i.e., low latency and low energy) to floating-point multipliers (FMs) requires careful consideration. In particular, introducing optical circuits unnecessarily increases the overhead of optical-to-electrical (O/E) and analog-to-digital (A/D) conversion, thus reducing system performance efficiency. Appropriate processing balancing between electrical and optical circuits is an important issue in OE hybrid circuit design. This study evaluates FM latency and energy consumption with different combinations of optical and electrical circuits. We design two versions of an Opto-Electrical FM (OEFM): M-OEFM, an optical implementation of an integer multiplier in an FM, and MA-OEFM, an optical implementation of an integer multiplier and adder functions. These designs are based on a policy of priority optical circuit implementation of the dominant components of conventional electrical circuits. Experimental results indicate that the M-OEFM achieved a 56 % reduction in latency and a 41 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM achieved an 88 % reduction in latency and a 19 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM is superior to M-OEFM in terms of energy-delay product (EDP). The results also indicate that appropriate co-design of analog optical and digital electrical circuits enables highly efficient FMs. Takumi Inaba, Takatsugu Ono, Koji Inoue, Satoshi Kawakami |
CF | 3 |
| 2023 | CFChain: A Crowdfunding Platform that Supports Identity Authentication, Privacy Protection, and Efficient Audit
Yueyue He, Jiageng Chen, Koji Inoue |
ICA3PP (7) | 3 |
| 2023 | QIsim: Architecting 10+K Qubit QC Interfaces Toward Quantum SupremacyabstractA 10+K qubit Quantum-Classical Interface (QCI) is essential to realize the quantum supremacy. However, it is extremely challenging to architect scalable QCIs due to the complex scalability trade-offs regarding operating temperatures, device and wire technologies, and microarchitecture designs. Therefore, architects need a modeling tool to evaluate various QCI design choices and lead to an optimal scalable QCI architecture. Dongmoon Min, Junpyo Kim, Junhyuk Choi, Ilkwon Byun, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 6 |
| 2023 | RealPersonaChat: A Realistic Persona Chat Corpus with Interlocutors' Own Personalities
Sanae Yamashita, Koji Inoue, Shota Mochizuki, Tatsuya Kawahara, Ryuichiro Higashinaka |
PACLIC | 2 |
| 2023 | Robotic Backchanneling in Online Conversation Facilitation: A Cross-Generational StudyabstractJapan faces many challenges related to its aging society, including increasing rates of cognitive decline in the population and a shortage of caregivers. Efforts have begun to explore solutions using artificial intelligence (AI), especially socially embodied intelligent agents and robots that can communicate with people. Yet, there has been little research on the compatibility of these agents with older adults in various everyday situations. To this end, we conducted a user study to evaluate a robot that functions as a facilitator for a group conversation protocol designed to prevent cognitive decline. We modified the robot to use backchannelling, a natural human way of speaking, to increase receptiveness of the robot and enjoyment of the group conversation experience. We conducted a cross-generational study with young adults and older adults. Qualitative analyses indicated that younger adults perceived the backchannelling version of the robot as kinder, more trustworthy, and more acceptable than the non-backchannelling robot. Finally, we found that the robot’s backchannelling elicited nonverbal backchanneling in older participants. Sota Kobuki, Katie Seaborn, Seiki Tokunaga, Kosuke Fukumori, Shun Hidaka, Kazuhiro Tamura, Koji Inoue, Tatsuya Kawahara, Mihoko Otake |
RO-MAN | 7 |
| 2023 | Reasoning before Responding: Integrating Commonsense-based Causality Explanation for Empathetic Response GenerationabstractRecent approaches to empathetic response generation try to incorporate commonsense knowledge or reasoning about the causes of emotions to better understand the user's experiences and feelings.However, these approaches mainly focus on understanding the causalities of context from the user's perspective, ignoring the system's perspective.In this paper, we propose a commonsense-based causality explanation approach for diverse empathetic response generation that considers both the user's perspective (user's desires and reactions) and the system's perspective (system's intentions and reactions).We enhance ChatGPT's ability to reason for the system's perspective by integrating in-context learning with commonsense knowledge.Then, we integrate the commonsense-based causality explanation with both ChatGPT and a T5-based model.Experimental evaluations demonstrate that our method outperforms other comparable methods on both automatic and human evaluations. Yahui Fu 0001, Koji Inoue, Chenhui Chu, Tatsuya Kawahara |
SIGDIAL | 2 |
| 2023 | Character expression for spoken dialogue systems with semi-supervised learning using Variational Auto-EncoderabstractCharacter of spoken dialogue systems is important not only for giving a positive impression of the system but also for gaining rapport from users. We have proposed a character expression model for spoken dialogue systems. The model expresses three character traits (extroversion, emotional instability, and politeness) of spoken dialogue systems by controlling spoken dialogue behaviors: utterance amount, backchannel, filler, and switching pause length. One major problem in training this model is that it is costly and time-consuming to collect many pair data of character traits and behaviors. To address this problem, semi-supervised learning is proposed based on a variational auto-encoder that exploits both the limited amount of labeled pair data and unlabeled corpus data. It was confirmed that the proposed model can express given characters more accurately than a baseline model with only supervised learning. We also implemented the character expression model in a spoken dialogue system for an autonomous android robot, and then conducted a subjective experiment with 75 university students to confirm the effectiveness of the character expression for specific dialogue scenarios. The results showed that expressing a character in accordance with the dialogue task by the proposed model improves the user’s impression of the appropriateness in formal dialogue such as job interview. Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara |
Comput. Speech Lang. | 2 |
| 2022 | Backchannel Generation Model for a Third Party Listener AgentabstractIn this work we propose a listening agent which can be used in a conversation between two humans. We firstly conduct a corpus analysis to identify three different categories of backchannel which the agent can use - responsive interjections, expressive interjections and shared laughs. From this data we train and evaluate a continuous backchannel generation model consisting of separate timing and form prediction models. We then conduct a subjective experiment to compare our model to random, dyadic, and ground truth models. We find that our model outperforms a random baseline and is comparable to the dyadic model despite the low evaluation of expressive interjections. We suggest that the perception of expressive interjections contribute significantly to the perception of the agent’s empathy and understanding of the conversation. The results also show the need for a more robust model to generate expressive interjections, perhaps aided by the use of linguistic features. Divesh Lala, Koji Inoue, Tatsuya Kawahara, Kei Sawada |
HAI | 2 |
| 2022 | Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic ListenersabstractAs the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science, over the past decade. Since speech is considered one of the effective ways to diagnose cognitive decline, AD detection from speech has emerged as a hot topic. Nevertheless, such approaches fail to tackle several key issues: 1) AD is a complex neurocognitive disorder which means it is inappropriate to conduct AD detection using utterance information alone while ignoring dialogue infor-mation; 2) Utterances of AD patients contain many disfluencies that affect speech recognition yet are helpful to diagnosis; 3) AD patients tend to speak less, causing dialogue breakdown as the disease progresses. This fact leads to a small number of utterances, which may cause detection bias. Therefore, in this paper, we propose a novel AD detection architecture consisting of two major modules: an ensemble AD detector and a proactive listener. This architecture can be embedded in the dialogue system of conversational robots for healthcare. Yuanchao Li, Catherine Lai, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
HRI | 4 |
| 2022 | Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino |
INTERSPEECH | 5 |
| 2022 | XQsim: modeling cross-technology control processors for 10+K qubit quantum computersabstract10+K qubit quantum computer is essential to achieve a true sense of quantum supremacy. With the recent effort towards the large-scale quantum computer, architects have revealed various scalability issues including the constraints in a quantum control processor, which should be holistically analyzed to design a future scalable control processor. However, it has been impossible to identify and resolve the processor's scalability bottleneck due to the absence of a reliable tool to explore an extensive design space including microarchitecture, device technology, and operating temperature. Ilkwon Byun, Junpyo Kim, Dongmoon Min, Ikki Nagaoka, Kosuke Fukumitsu, Iori Ishikawa, Teruo Tanimoto, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 9 |
| 2022 | Design of Variable Bit-Width Arithmetic Unit Using Single Flux Quantum DeviceabstractThis paper presents the design of an ultra-high-speed, low-power arithmetic unit that supports variable bit-width operations with single flux quantum (SFQ) technology. Because of the high-speed nature of superconductor devices, we can achieve extremely high power-performance efficiency that cannot be achieved by state-of-the-art CMOS devices. To implement the complex function to support the variable bit-width feature, we introduce a novel circuit architecture to maintain the high-speed operation over 50GHz. Our prototype chip design successfully demonstrated 53.5GHz 1.59mW operations. Iori Ishikawa, Ikki Nagaoka, Ryota Kashima, Koki Ishida, Kosuke Fukumitsu, Keitarou Oka, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Akira Fujimaki, Koji Inoue |
ISCAS | 12 |
| 2022 | Q3DE: A fault-tolerant quantum computer architecture for multi-bit burst errors by cosmic raysabstractDemonstrating small error rates by integrating quantum error correction (QEC) into an architecture of quantum computing is the next milestone towards scalable fault-tolerant quantum computing (FTQC). Encoding logical qubits with superconducting qubits and surface codes is considered a promising candidate for FTQC architectures. In this paper, we propose an FTQC architecture, which we call Q3DE, that enhances the tolerance to multi-bit burst errors (MBBEs) by cosmic rays with moderate changes and overhead. There are three core components in Q3DE: in-situ anomaly DEtection, dynamic code DEformation, and optimized error DEcoding. In this architecture, MBBEs are detected only from syndrome values for error correction. The effect of MBBEs is immediately mitigated by dynamically increasing the encoding level of logical qubits and re-estimating probable recovery operation with the rollback of the decoding process. We investigate the performance and overhead of the Q3DE architecture with quantum-error simulators and demonstrate that Q3DE effectively reduces the period of MBBEs by 1000 times and halves the size of their region. Therefore, Q3DE significantly relaxes the requirement of qubit density and qubit chip size to realize FTQC. Our scheme is versatile for mitigating MBBEs, i.e., temporal variations of error properties, on a wide range of physical devices and FTQC architectures since it relies only on the standard features of topological stabilizer codes. Yasunari Suzuki, Takanori Sugiyama, Tomochika Arai, Koji Inoue, Teruo Tanimoto |
MICRO | 5 |
| 2022 | Simultaneous Job Interview System Using Multiple Semi-autonomous AgentsabstractIn recent years, spoken dialogue systems have been used in job interviews where an applicant talks to a system that asks pre-defined questions, called on-demand and self-paced job interviews.We propose a simultaneous job interview system, where one interviewer can conduct one-on-one interviews with multiple applicants simultaneously by cooperating with multiple autonomous interview dialogue systems.However, it is challenging for interviewers to monitor and understand all parallel interviews done by the autonomous system simultaneously.To address this issue, we implement two automatic dialogue understanding functions: (1) response evaluation of each applicant's responses and (2) keyword extraction for a summary of the responses.In this system, interviewers can intervene in a dialogue session when needed and smoothly ask a proper question that elaborates the interview.We have conducted a pilot experiment where an interviewer conducted simultaneous job interviews with three candidates. Haruki Kawai, Yusuke Muraki, Kenta Yamamoto, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
SIGDIAL | 5 |
| 2021 | A multi-party attentive listening robot which stimulates involvement from side participantsabstractWe demonstrate the moderating abilities of a multi-party attentive listening robot system when multiple people are speaking in turns.Our conventional one-on-one attentive listening system generates listener responses such as backchannels, repeats, elaborating questions, and assessments.In this paper, additional robot responses that stimulate a listening user (side participant) to become more involved in the dialogue are proposed.The additional responses elicit assessments and questions from the side participant, making the dialogue more empathetic and lively. Koji Inoue, Hiromi Sakamoto, Kenta Yamamoto, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 1 |
| 2020 | Energy Efficient Runahead Execution on a Tightly Coupled Heterogeneous CoreabstractOut-of-order (OoO) processors generally offer significant performance gains over simpler in-order (InO) processors. However, recent studies have revealed that OoO processors provide little performance benefit in many program phases, and these phases are distributed in fine granularity. Leveraging these fine-grained phases, tightly coupled heterogeneous cores (TCHCs) have been proposed to improve the energy efficiency. A TCHC, which is a processor core that consists of multiple back-ends, each with different characteristics in terms of their performance and energy consumption (e.g., a power-efficient InO back-end and a high-performance OoO back-end), improves the energy efficiency by executing programs by switching to the most energy-efficient back-end with a very small switching penalty. Susumu Mashimo, Ryota Shioya, Koji Inoue |
HPC Asia | 3 |
| 2020 | Enhancing a manycore-oriented compressed cache for GPGPUabstractGPUs can achieve high performance by exploiting massive-thread parallelism. However, some factors limit performance on GPUs, one of which is the negative effects of L1 cache misses. In some applications, GPUs are likely to suffer from L1 cache conflicts because a large number of cores share a small L1 cache capacity. A cache architecture that is based on data compression is a strong candidate for solving this problem as it can reduce the number of cache misses. Unlike previous studies, our data compression scheme attempts to exploit the value locality existing within not only intra cache lines but also inter cache lines. We enhance the structure of a last-level compression cache proposed for general purpose manycore processors to optimize against shared L1 caches on GPUs. The experimental results reveal that our proposal outperforms the other compression cache for GPUs by 11 points on average. Keitarou Oka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Koji Inoue |
HPC Asia | 5 |
| 2020 | Job Interviewer Android with Elaborate Follow-up Question GenerationabstractA job interview is a domain that takes advantage of an android robot's human-like appearance and behaviors. In this work, our goal is to implement a system in which an android plays the role of an interviewer so that users may practice for a real job interview. Our proposed system generates elaborate follow-up questions based on responses from the interviewee. We conducted an interactive experiment to compare the proposed system against a baseline system that asked only fixed-form questions. We found that this system was significantly better than the baseline system with respect to the impression of the interview and the quality of the questions, and that the presence of the android interviewer was enhanced by the follow-up questions. We also found a similar result when using a virtual agent interviewer, except that presence was not enhanced. Koji Inoue, Kohei Hara, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
ICMI | 1 |
| 2020 | Semi-Supervised Learning for Character Expression of Spoken Dialogue Systems
Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2020 | SuperNPU: An Extremely Fast Neural Processing Unit Using Superconducting Logic DevicesabstractSuperconductor single-flux-quantum (SFQ) logic family has been recognized as a highly promising solution for the post-Moore's era, thanks to its ultra-fast and low-power switching characteristics. Therefore, researchers have made a tremendous amount of effort in various aspects to promote the technology and automate its circuit design process (e.g., low-cost fabrication, design tool development). However, there has been no progress in designing a convincing SFQ-based architectural unit due to the architects' lack of understanding of the technology's potentials and limitations at the architecture level. In this paper, we present how to architect an SFQ-based architectural unit by providing design principles with an extreme-performance neural processing unit (NPU). To achieve the goal, we first implement an architecture-level simulator to model an SFQ-based NPU accurately. We validate this model using our die-level prototypes, design tools, and logic cell library. This simulator accurately measures the NPU's performance, power consumption, area, and cooling overheads. Next, driven by the modeling, we identify key architectural challenges for designing a performance-effective SFQ-based NPU (e.g., expensive on-chip data movements and buffering). Lastly, we present SuperNPU, our example SFQ-based NPU architecture, which effectively resolves the challenges. Our evaluation shows that the proposed design outperforms a conventional state-of-the-art NPU by 23 times. With free cooling provided as done in quantum computing, the performance per chip power increases up to 490 times. Our methodology can also be applied to other architecture designs with SFQ-friendly characteristics. Koki Ishida, Ilkwon Byun, Ikki Nagaoka, Kosuke Fukumitsu, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Jangwoo Kim, Koji Inoue |
MICRO | 10 |
| 2020 | An Attentive Listening System with Android ERICA: Comparison of Autonomous and WOZ InteractionsabstractWe describe an attentive listening system for the autonomous android robot ERICA.The proposed system generates several types of listener responses: backchannels, repeats, elaborating questions, assessments, generic sentimental responses, and generic responses.In this paper, we report a subjective experiment with 20 elderly people.First, we evaluated each system utterance excluding backchannels and generic responses, in an offline manner.It was found that most of the system utterances were linguistically appropriate, and they elicited positive reactions from the subjects.Furthermore, 58.2% of the responses were acknowledged as being appropriate listener responses.We also compared the proposed system with a WOZ system where a human operator was operating the robot.From the subjective evaluation, the proposed system achieved comparable scores in basic skills of attentive listening such as encouragement to talk, focused on the talk, and actively listening.It was also found that there is still a gap between the system and the WOZ for more sophisticated skills such as dialogue understanding, showing interest, and empathy towards the user. Koji Inoue, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
SIGdial | 1 |
| 2019 | Evaluating the Impact of Energy Efficient Networks on HPC WorkloadsabstractInterconnection networks grow larger as supercomputers include more nodes and require higher bandwidth for performance. This scaling significantly increases the fraction of power consumed by the network, by increasing the number of network components (links and switches). Typically, network links consume power continuously once they are turned on. However, recent proposals for energy efficient interconnects have introduced low-power operation modes for periods when network links are idle. Low-power operation can increase messaging time when switching a link from low-power to active operation. We extend the TraceR-CODES network simulator for power modeling to evaluate the impact of energy efficient networking on power and performance. Our evaluation presents the first study on both single-job and multi-job execution to realistically simulate power consumption and performance under congestion for a large-scale HPC network. Results on several workloads consisting of HPC proxy applications show that single-job and multi-job execution favor different modes of low power operation to have significant power savings at the cost of minimal performance degradation. Giorgis Georgakoudis, Takatsugu Ono, Koji Inoue, Shinobu Miwa, Abhinav Bhatele |
HiPC | 4 |
| 2019 | Smooth Turn-taking by a Robot Using an Online Continuous Model to Generate Turn-taking CuesabstractTurn-taking in human-robot interaction is a crucial part of spoken dialogue systems, but current models do not allow for human-like turn-taking speed seen in natural conversation. In this work we propose combining two independent prediction models. A continuous model predicts the upcoming end of the turn in order to generate gaze aversion and fillers as turn-taking cues. This prediction is done while the user is speaking, so turn-taking can be done with little silence between turns, or even overlap. Once a speech recognition result has been received at a later time, a second model uses the lexical information to decide if or when the turn should actually be taken. We constructed the continuous model using the speaker’s prosodic features as inputs and evaluated its online performance. We then conducted a subjective experiment in which we implemented our model in an android robot and asked participants to compare it to one without turn-taking cues, which produces a response when a speech recognition result is received. We found that using both gaze aversion and a filler was preferred when the continuous model correctly predicted the upcoming end of turn, while using only gaze aversion was better if the prediction was wrong. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 2 |
| 2019 | Turn-Taking Prediction Based on Detection of Transition Relevance Place
Kohei Hara, Koji Inoue, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2018 | Autoencoder based Features Extraction for Automatic Classification of Earthquakes and ExplosionsabstractMonitoring illegal explosions is mandatory for the safety of human life, environment, and protect the important buildings such as High-dam in Egypt. This kind of monitoring can be accomplished by detecting and identifying the explosions. If an illegal explosion happens such as quarry blast, an alarm should be reported to the government to take immediate action. However, the main problem is that many measured signals from received explosions are similar to earthquakes in their shape and both cannot differentiate from each other. Also, incorrect classification possibly will distort the real seismicity nature of the region. This problem motivates us to search for unique discriminating features to distinguish between earthquakes and explosions with precise accuracy. Therefore, in this paper, we propose to extract the discriminative features based on Autoencoder from the first few seconds after the P-wave arrival time of the event. The discriminative features are found to be in the first 60 samples after the arrival time of P-wave. Thus the first stage of the proposed algorithm is extracting the discriminative features via the Autoencoder. Then, softmax classifies the event based on these extracted features. The proposed algorithm achieves a classification accuracy of 98.55% when applied to 900 earthquakes and quarry blasts waveforms recorded by Egyptian National Seismic Network (ENSN). Omar M. Saad, Koji Inoue, Ahmed Shalaby 0001, Lotfy Sarny, Mohammed Sharaf Sayed |
ICIS | 2 |
| 2018 | An End-to-End Approach to Joint Social Signal Detection and Automatic Speech RecognitionabstractSocial signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchronized laughing or attentive listening. We have studied an end-to-end approach to directly detect social signals from speech by using connectionist temporal classification (CTC), which is one of the end-to-end sequence labelling models. In this work, we propose a unified framework that integrates social signal detection (SSD) and automatic speech recognition (ASR). We investigate several reference labelling methods regarding social signals. Experimental evaluations demonstrate that our end-to-end framework significantly outperforms the conventional DNN-HMM system with regard to SSD performance as well as the character error rate (CER). Hirofumi Inaguma, Masato Mimura, Koji Inoue, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 3 |
| 2018 | Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid RobotabstractThis paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-gaze provides an important cue compared with head nodding and verbal backchannels. This finding has been applied to audio-visual speaker diarization by combining eye-gaze information. We are now investigating engagement recognition in human-robot interaction based on the same scheme. Robust and realtime detection of laughing, backchannels and nodding is realized based on LSTM-CTC. We introduce a latent “character” model to cope with the subjectivity and variations of engagement annotations. Experimental evaluations demonstrate that (1) the latent character model is effective, (2) automatic behavior detection is robust and does not degrade the engagement recognition accuracy, and (3) the eye-gaze is the most important feature among others. Tatsuya Kawahara, Koji Inoue, Divesh Lala, Katsuya Takanashi |
ICASSP | 2 |
| 2018 | Evaluation of Real-time Deep Learning Turn-taking Models for Multiple Dialogue ScenariosabstractThe task of identifying when to take a conversational turn is an important function of spoken dialogue systems. The turn-taking system should also ideally be able to handle many types of dialogue, from structured conversation to spontaneous and unstructured discourse. Our goal is to determine how much a generalized model trained on many types of dialogue scenarios would improve on a model trained only for a specific scenario. To achieve this goal we created a large corpus of Wizard-of-Oz conversation data which consisted of several different types of dialogue sessions, and then compared a generalized model with scenario-specific models. For our evaluation we go further than simply reporting conventional metrics, which we show are not informative enough to evaluate turn-taking in a real-time system. Instead, we process results using a performance curve of latency and false cut-in rate, and further improve our model's real-time performance using a finite-state turn-taking machine. Our results show that the generalized model greatly outperformed the individual model for attentive listening scenarios but was worse in job interview scenarios. This implies that a model based on a large corpus is better suited to conversation which is more user-initiated and unstructured. We also propose that our method of evaluation leads to more informative performance metrics in a real-time system. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 2 |
| 2018 | Prediction of Turn-taking Using Multitask Learning with Prediction of Backchannels and Fillers
Kohei Hara, Koji Inoue, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2018 | Engagement Recognition in Spoken Dialogue via Neural Network by Aggregating Different Annotators' Models
Koji Inoue, Divesh Lala, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2018 | Analyzing Resource Trade-offs in Hardware Overprovisioned SupercomputersabstractHardware overprovisioned systems have recently been proposed as a viable alternative for a power-efficient design of next-generation supercomputers. A key challenge for such systems is to determine the degree of overprovisioning, which refers to the number of extra nodes that need to be installed under a given power constraint. In this paper, we first show that the degree of overprovisioning depends on dynamic parameters, such as the job mix as well as the global power constraint, and that static decisions can result in limited system throughput. We then study an exhaustive combination of adaptive resource management strategies that span three job scheduling algorithms, four power capping techniques, and three node boot-up mechanisms to understand the trade-off space involved. We then draw conclusions about how these strategies can adaptively control the degree of overprovisioning and analyze their impact on job throughput and power utilization. Ryuichi Sakamoto, Tapasya Patki, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 5 |
| 2018 | An Integrated Nanophotonic Parallel AdderabstractIntegrated optical circuits with nanophotonic devices have attracted significant attention due to their low power dissipation and light-speed operation. With light interference and resonance phenomena, the nanophotonic device works as a voltage-controlled optical pass-gate like a pass-transistor. This article first introduces the concept of optical pass-gate logic and then proposes a parallel adder circuit based on optical pass-gate logic. Experimental results obtained with an optoelectronic circuit simulator show the advantages of our optical parallel adder circuit over a traditional CMOS-based parallel adder circuit. Tohru Ishihara, Akihiko Shinya, Koji Inoue, Kengo Nozaki, Masaya Notomi |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Automatic Arrival Time Detection for Earthquakes Based on Stacked Denoising AutoencoderabstractThe accurate detection of P-wave arrival time is imperative for determining the hypocenter location of an earthquake. However, precise detection of onset time becomes more difficult when the signal-to-noise ratio (SNR) of the seismic data is low, such as during microearthquakes. In this letter, a stacked denoising autoencoder (SDAE) is proposed to smooth the background noise. The SDAE acts as a denoising filter for the seismic data. In the proposed algorithm, the SDAE is utilized to reduce background noise such that the onset time becomes more clear and sharp. Afterward, a hard decision with one threshold is used to detect the onset time of the event. The proposed algorithm is evaluated on both synthetic and field seismic data. As a result, the proposed algorithm outperforms the short-time average/long-time average and the Akaike information criterion algorithms. The proposed algorithm accurately picks the onset time of 94.1% for 407 field seismic waveforms with a standard deviation error of 0.10 s. In addition, the results indicate that the proposed algorithm can pick arrival times accurately for weak SNR seismic data with SNR higher than -14 dB. Omar M. Saad, Koji Inoue, Ahmed Shalaby 0001, Lotfy Samy, Mohammed Sharaf Sayed |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Social Signal Detection in Spontaneous Dialogue Using Bidirectional LSTM-CTC
Hirofumi Inaguma, Koji Inoue, Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2017 | Production Hardware Overprovisioning: Real-World Performance Optimization Using an Extensible Power-Aware Resource Management FrameworkabstractLimited power budgets will be one of the biggest challenges for deploying future exascale supercomputers. One of the promising ways to deal with this challenge is hardware over provisioning, that is, installing more hardware resources than can be fully powered under a given power limit coupled with software mechanisms to steer the limited power to where it is needed most. Prior research has demonstrated the viability of this approach, but could only rely on small-scale simulations of the software stack. While such research is useful to understand the boundaries of performance benefits that can be achieved, it does not cover any deployment or operational concerns of using overprovisioning on production systems. This paper is the first to present an extensible power-aware resource management framework for production-sized overprovisioned systems based on the widely established SLURM resource manager. Our framework provides flexible plugin interfaces and APIs for power management that can be easily extended to implement site-specific strategies and for comparison of different power management techniques. We demonstrate our framework on a 965-node HA8000 production system at Kyushu University. Our results indicate that it is indeed possible to safely overprovision hardware in production. We also find that the power consumption of idle nodes, which depends on the degree of overprovisioning, can become a bottleneck. Using real-world data, we then draw conclusions about the impact of the total number of nodes provided in an overprovisioned environment. Ryuichi Sakamoto, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Tapasya Patki, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 4 |
| 2017 | Attentive listening system with backchanneling, response generation and flexible turn-takingabstractAttentive listening systems are designed to let people, especially senior people, keep talking to maintain communication ability and mental health.This paper addresses key components of an attentive listening system which encourages users to talk smoothly.First, we introduce continuous prediction of end-of-utterances and generation of backchannels, rather than generating backchannels after end-point detection of utterances.This improves subjective evaluations of backchannels.Second, we propose an effective statement response mechanism which detects focus words and responds in the form of a question or partial repeat.This can be applied to any statement.Moreover, a flexible turn-taking mechanism is designed which uses backchannels or fillers when the turnswitch is ambiguous.These techniques are integrated into a humanoid robot to conduct attentive listening.We test the feasibility of the system in a pilot experiment and show that it can produce coherent dialogues during conversation. Divesh Lala, Pierrick Milhorat, Koji Inoue, Masanari Ishida, Katsuya Takanashi, Tatsuya Kawahara |
SIGDIAL Conference | 3 |
| 2016 | Evaluating the impacts of code-level performance tunings on power efficiencyabstractAs the power consumption of HPC systems will be a primary constraint for exascale computing, a main objective in HPC communities is recently becoming to maximize power efficiency (i.e., performance per watt) rather than performance. Although programmers have spent a considerable effort to improve performance by tuning HPC programs at a code level, tunings for improving power efficiency is now required. In this work, we select two representative HPC programs (Graph500 and SDPARA) and evaluate how traditional code-level performance tunings applied to these programs affect power efficiency. We also investigate the impacts of the tunings on power efficiency at various operating frequencies of CPUs and/or GPUs. The results show that the tunings significantly improve power efficiency, and different types of tunings exhibit different trends in power efficiency by varying CPU frequency. Finally, the scalability and power efficiency of state-of-the-art Graph500 implementations are explored on both a single-node platform and a 960-node supercomputer. With their high scalability, they achieve 27.43 MTEPS/Watt with 129.76 GTEPS on the single-node system and 4.39 MTEPS/Watt with 1,085.24 GTEPS on the supercomputer. Satoshi Imamura, Keitarou Oka, Yuichiro Yasui, Yuichi Inadomi, Katsuki Fujisawa, Toshio Endo, Koji Ueno, Keiichiro Fukazawa, Nozomi Hata, Yuta Kakibuka, Koji Inoue, Takatsugu Ono |
IEEE BigData | 11 |
| 2016 | Multimodal interaction with the autonomous Android ERICAabstractWe demonstrate an interactive conversation with an android named ERICA. In this demonstration the user can converse with ERICA on a number of topics. We demonstrate both the dialog management system and the eye gaze behavior of ERICA used for indicating attention and turn taking. Divesh Lala, Pierrick Milhorat, Koji Inoue, Tianyu Zhao 0001, Tatsuya Kawahara |
ICMI | 3 |
| 2016 | Prediction and Generation of Backchannel Form for Attentive Listening Systems
Tatsuya Kawahara, Takashi Yamaguchi, Koji Inoue, Katsuya Takanashi, Nigel G. Ward |
INTERSPEECH | 3 |
| 2016 | Talking with ERICA, an autonomous androidabstractWe demonstrate dialogues with an autonomous android ERICA, who has an appearance like a human being.Currently, ERICA plays two social roles: a laboratory guide and a counselor.It is designed to follow the protocols of human dialogue to make the user comfortable: (1) having a chat before the main talk, (2) proactively asking questions, and (3) conveying proper feedbacks.The combination of the human-like appearance and the appropriate behaviors according to her social roles allows for symbiotic human-robot interaction. Koji Inoue, Pierrick Milhorat, Divesh Lala, Tianyu Zhao 0001, Tatsuya Kawahara |
SIGDIAL Conference | 1 |
| 2015 | A flexible hardware barrier mechanism for many-core processorsabstractThis paper proposes a new hardware barrier mechanism which offers the flexibility to select which cores should join the synchronization, allowing for executing multiple multi-threaded applications by dividing a many-core processor into several groups. Experimental results based on an RTL simulation show that our hardware barrier achieves a 66-fold reduction in latency over typical software based implementations, with a hardware overhead of the processor of only 1.8%. Additionally, we demonstrate that the proposed mechanism is sufficiently flexible to cover a variety of core groups with minimal hardware overhead. Takeshi Soga, Hiroshi Sasaki 0001, Tomoya Hirao, Masaaki Kondo, Koji Inoue |
ASP-DAC | 5 |
| 2015 | Enhanced speaker diarization with detection of backchannels using eye-gaze information in poster conversationsabstractWe propose multi-modal speaker diarization using acoustic and eye-gaze information in poster conversations. Eye-gaze information plays an important role in turn-taking, thus it is useful for predicting speech activity. In this paper, a variety of eyegaze features are elaborated and combined with the acoustic information by the multi-modal integration model. Moreover, we introduce another model to detect backchannels, which involve different eye-gaze behaviors. This enhances the diarization result by filtering meaningful utterances such as questions and comments. Experimental evaluations in real poster sessions demonstrate that eye-gaze information contributes to improvement of diarization accuracy under noisy environments, and its weight is automatically determined according to the Signal-toNoise Ratio (SNR). Index Terms: speaker diarization, backchannel, multi-modal, eye-gaze, poster conversation Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2015 | Characterization and cross-platform analysis of high-throughput acceleratorsabstractToday's computer systems often employ high-throughput accelerators (such as Intel Xeon Phi coprocessors and NVIDIA Tesla GPUs) to improve the performance of some applications or portions of applications. While such accelerators are useful for suitable applications, it remains challenging to predict which workloads will run well on these platforms and to predict the resulting performance trends for varying input. This paper provides detailed characterizations on such platforms across a range of programs and input sizes. Furthermore, we show opportunities for cross-platform performance analysis and comparison between Xeon Phi and Tesla. Our crossplatform comparison has three steps. First, we build Xeon Phi performance regression models as a function of important Xeon Phi performance counters to identify critical architectural resources that highly affect a benchmark's performance. Then, cross-platform Tesla performance regression models are built to relate the Tesla performance trends of the benchmark to the Xeon Phi performance counter measurements of the benchmark. Finally, we compare the counters most important for Xeon Phi models to those most important for Tesla's models; this reveals similarities and distinctions of dynamic application behaviors on the two platforms. Keitarou Oka, Margaret Martonosi, Koji Inoue |
ISPASS | 4 |
| 2015 | Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputingabstractA key challenge in next-generation supercomputing is to effectively schedule limited power resources. Modern processors suffer from increasingly large power variations due to the chip manufacturing process. These variations lead to power inhomogeneity in current systems and manifest into performance inhomogeneity in power constrained environments, drastically limiting supercomputing performance. We present a first-of-its-kind study on manufacturing variability on four production HPC systems spanning four microarchitectures, analyze its impact on HPC applications, and propose a novel variation-aware power budgeting scheme to maximize effective application performance. Our low-cost and scalable budgeting algorithm strives to achieve performance homogeneity under a power constraint by deriving application-specific, module-level power allocations. Experimental results using a 1,920 socket system show up to 5.4X speedup, with an average speedup of 1.8X across all benchmarks when compared to a variation-unaware power allocation scheme. Yuichi Inadomi, Tapasya Patki, Koji Inoue, Mutsumi Aoyagi, Barry Rountree, Martin Schulz 0001, David K. Lowenthal, Yasutaka Wada, Keiichiro Fukazawa, Masatsugu Ueda, Masaaki Kondo, Ikuo Miyoshi |
SC | 3 |
| 2014 | Power Consumption Evaluation of an MHD Simulation with CPU Power CappingabstractRecently to achieve the Exa-flops next generation computer system, the power consumption becomes the important issue. On the other hand, the power consumption character of application program is not so considered now. In this study we examine the power character of our Magneto hydrodynamic (MHD) simulation code for the global magnetosphere to evaluate the power consumption behavior of the simulation code under the CPU power capping on the parallel computer system. As a result, it is confirmed that there are different power consumption parts in the MHD simulation code, which the execution performance decreases or does not change under the CPU power capping. This indicates the capability of performance optimization with the power capping. Keiichiro Fukazawa, Masatsugu Ueda, Mutsumi Aoyagi, Tomonori Tsuhata, Kyohei Yoshida, Aruta Uehara, Masakazu Kuze, Yuichi Inadomi, Koji Inoue |
CCGRID | 9 |
| 2014 | Power-capped DVFS and thread allocation with ANN models on modern NUMA systemsabstractPower capping is an essential function for efficient power budgeting and cost management on modern server systems. Contemporary server processors operate under power caps by using dynamic voltage and frequency scaling (DVFS). However, these processors are often deployed in non-uniform memory access (NUMA) architectures, where thread allocation between cores may significantly affect performance and power consumption. This paper proposes a method which maximizes performance under power caps on NUMA systems by dynamically optimizing two knobs: DVFS and thread allocation. The method selects the optimal combination of the two knobs with models based on artificial neural network (ANN) that captures the nonlinear effect of thread allocation on performance. We implement the proposed method as a runtime system and evaluate it with twelve multithreaded benchmarks on a real AMD Opteron based NUMA system. The evaluation results show that our method outperforms a naive technique optimizing only DVFS by up to 67.1%, under a power cap. Satoshi Imamura, Hiroshi Sasaki 0001, Koji Inoue, Dimitrios S. Nikolopoulos |
ICCD | 3 |
| 2014 | Speaker diarization using eye-gaze information in multi-party conversationsabstractWe present a novel speaker diarization method by using eye-gaze information in multi-party conversations. In real environ-ments, speaker diarization or speech activity detection of each participant of the conversation is challenging because of distant talking and ambient noise. In contrast, eye-gaze information is robust against acoustic degradation, and it is presumed that eye-gaze behavior plays an important role in turn-taking and thus in predicting utterances. The proposed method stochastically integrates eye-gaze information with acoustic information for speaker diarization. Specifically, three models are investigated for multi-modal integration in this paper. Experimental eval-uations in real poster sessions demonstrate that the proposed method improves accuracy of speaker diarization from the base-line acoustic method. Index Terms: speaker diarization, multi-modal interaction, eye-gaze Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2014 | Power and Performance Characterization and Modeling of GPU-Accelerated SystemsabstractGraphics processing units (GPUs) provide an order-of-magnitude improvement on peak performance and performance-per-watt as compared to traditional multicore CPUs. However, GPU-accelerated systems currently lack a generalized method of power and performance prediction, which prevents system designers from an ultimate goal of dynamic power and performance optimization. This is due to the fact that their power and performance characteristics are not well captured across architectures, and as a result, existing power and performance modeling approaches are only available for a limited range of particular GPUs. In this paper, we present power and performance characterization and modeling of GPU-accelerated systems across multiple generations of architectures. Characterization and modeling both play a vital role in optimization and prediction of GPU-accelerated systems. We quantify the impact of voltage and frequency scaling on each architecture with a particularly intriguing result that a cutting-edge Kepler-based GPU achieves energy saving of 75% by lowering GPU clocks in the best scenario, while Fermi- and Tesla-based GPUs achieve no greater than 40% and 13%, respectively. Considering these characteristics, we provide statistical power and performance modeling of GPU-accelerated systems simplified enough to be applicable for multiple generations of architectures. One of our findings is that even simplified statistical models are able to predict power and performance of cutting-edge GPUs within errors of 20% to 30% for any set of voltage and frequency pair. Yuki Abe 0001, Hiroshi Sasaki 0001, Shinpei Kato, Koji Inoue, Masato Edahiro, Martin Peres |
IPDPS | 4 |
| 2014 | Speech-based human-robot interaction robust to acoustic reflections in real environmentabstractAcoustic reflection inside an enclosed environment is detrimental to human-robot interaction. Reflection may manifest as phantom sources emanating from unknown directions. In effect, a single speaker may falsely manifest as multiple speakers to the robot audition system, impeding the robot's ability to correctly associate the speech command to the actual speaker. Moreover, speech reflection smears the original speech signal due to reverberation. This degrades speech recognition and understanding performance. Conventional robot audition schemes that rely purely on acoustics and spatial information are very sensitive to acoustic reflection which ultimately leads to the failure in human-robot interaction. We propose a method for human-robot interaction robust to the effect of acoustic reflection. First, visual information is utilized and head tracking scheme is employed to reinforce the acoustic information with the visual presence of a prospect user. Second, we employ a model-based sound event identification scheme and scrutinize whether the acoustic information is likely to be speech or non-speech. Using all the information we have gathered, we create a simple rule construct to effectively discriminate the original source (actual speaker) from phantom sources (reflection). Consequently, the corresponding source identified as phantom (reflection) is used to estimate the unwanted smearing for effective suppression via speech enhancement. Experiments are conducted in human-robot interaction setting in which the proposed method outperforms the conventional method. Randy Gomez, Koji Inoue, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai |
IROS | 2 |
| 2013 | Coordinated power-performance optimization in manycoresabstractOptimizing the performance in multiprogrammed environments, especially for workloads composed of multi-threaded programs is a desired feature of runtime management system in future manycore processors. At the same time, power capping capability is required in order to improve the reliability of microprocessor chips while reducing the costs of power supply and thermal budgeting. This paper presents a sophisticated runtime coordinated power-performance management system called C-3PO, which optimizes the performance of manycore processors under a power constraint by controlling two software knobs: thread packing, and dynamic voltage and frequency scaling (DVFS). The proposed solution distributes the power budget to each program by controlling the workload threads to be executed with appropriate number of cores and operating frequency. The power budget is distributed carefully in different forms (number of allocated cores or operating frequency) depending on the power-performance characteristics of the workload so that each program can effectively convert the power into performance. The proposed system is based on a heuristic algorithm which relies on runtime prediction of power and performance via hardware performance monitoring units. Empirical results on a 64-core platform show that C-3PO well outperforms traditional counterparts across various PARSEC workload mixes. Hiroshi Sasaki 0001, Satoshi Imamura, Koji Inoue |
PACT | 3 |
| 2013 | SMYLE Project: Toward high-performance, low-power computing on manycore-processor SoCsabstractThis paper introduces a manycore research project called SMYLE (Scalable ManYcore for Low Energy computing). The aims of this project are: 1) proposing a manycore SoC architecture and developing a suitable programming and execution environment, 2) designing a domain specific manycore system for emerging video mining applications, and 3) releasing developed software tools and FPGA emulation environments to accelerate manycore research and development in the community. The project started in December 2010 with full support from the New Energy and Industrial Technology Development Organization (NEDO). Koji Inoue |
ASP-DAC | 1 |
| 2013 | SMYLEref: A reference architecture for manycore-processor SoCsabstractNowadays, the trend of developing micro-processor with tens of cores brings a promising prospect for embedded systems. Realizing a high performance and low power many-core processor is becoming a primary technical challenge. We are currently developing a many-core processor architecture for embedded systems as a part of a NEDO's project. This paper introduces the many-core architecture called SMYLEref along whit the concept of Virtual Accelerator on Many-core, in which many cores on a chip are utilized as a hardware platform for realizing multiple virtual accelerators. We are developing its prototype system with off-the-shelf FPGA evaluation boards. In this paper, we introduce the architecture of SMYLEref and the detail of the prototype system. In addition, several initial experiments with the prototype system are also presented. Masaaki Kondo, Son Truong Nguyen, Tomoya Hirao, Takeshi Soga, Hiroshi Sasaki 0001, Koji Inoue |
ASP-DAC | 6 |
| 2013 | Line sharing cache: Exploring cache capacity with frequent line value localityabstractThis paper proposes a new last level cache architecture called line sharing cache (LSC), which can reduce the number of cache misses without increasing the size of the cache memory. It stores lines which contain the identical value in a single line entry, which enables to store greater amount of lines. Evaluation results show performance improvements of up to 35% across a set of SPEC CPU2000 benchmarks. Keitarou Oka, Hiroshi Sasaki 0001, Koji Inoue |
ASP-DAC | 3 |
| 2013 | Hybrid compile and run-time memory management for a 3D-stacked reconfigurable acceleratorabstractThis paper presents a hybrid compile and run-time memory management technique for a 3D-stacked reconfigurable accelerator including a memory layer composed of multiple memory units whose parallel access allows a very high bandwidth. The technique inserts allocation, free and data transfers into the code for using the memory layer and avoids memory overflows by adding a limited number of additional copies to and from the host memory. When compile-time information is lacking, the technique relies on run-time decisions for controlling these memory operations. Experiments show that, compared to a pessimistic approach, the overhead for avoiding overflows can be cut on average by 27%, 45% and 63% when the size of each memory unit is respectively 1kB, 128kB and 1MB. Lovic Gauthier, Shinya Ueno, Koji Inoue |
CASES | 3 |
| 2012 | Scalability-based manycore partitioningabstractMulticore processors have been popular for years, and the industry is gradually shifting towards the era of manycore processors. Single-thread performance of microprocessors is not growing at a historical rate, but the existence of a number of active processes in the computer system and the continuing development of multi-threaded applications benefit from the growing core counts to sustain system throughput. This trend brings us a situation where a number of parallel applications simultaneously being executed on a single system. Since multi-threaded applications try to maximize its throughput by utilizing the whole system, each of them usually create equal or larger number of threads compared to underlying logical core counts. This introduces much greater number of threads to be co-scheduled in the entire system. However, each program has different characteristics (or scalability) and contends for shared resources, which are the CPU cores and memory hierarchies, with each other. Therefore, it is clear that OS thread scheduling will play a major role in achieving high system performance under such conditions. We develop a sophisticated scheduler that (1) dynamically predicts the scalability of programs via the use of hardware performance monitoring units, (2) decides the optimal number of cores to be allocated for each program, and (3) allocates the cores to programs while maximizing the system utilization to achieve fair and maximum performance. The evaluation results on a 48-core AMD Opteron system show improvements over the Linux scheduler for a variety of multiprogramming workloads. Hiroshi Sasaki 0001, Teruo Tanimoto, Koji Inoue, Hiroshi Nakamura |
PACT | 3 |
| 2012 | A Three-Dimensional Integrated AcceleratorabstractWe propose a three-dimensional (3D) reconfigurable data-path accelerator which is capable of running partitioned large data flow graphs (DFGs) on the layers of 3D stack, while inter-layer connections are implemented by means of through-silicon vias (TSVs). A tool for mapping data flow graphs has been developed, and a key 3D-specific problem namely routing nets on 3D architecture has been discussed in details as well. Conducted experiments demonstrate smaller footprint area and higher performance for the 3D accelerator comparing with 2D counterpart. Farhad Mehdipour, Krishna Chaitanya Nunna, Koji Inoue, Kazuaki J. Murakami |
DSD | 3 |
| 2012 | Local intensity compensation using sparse representation
Koji Inoue, Hironobu Saito, Yoshimitsu Kuroki |
ICPR | 1 |
| 2012 | Improving performance and energy efficiency of embedded processors via post-fabrication instruction set customization
Hamid Noori, Farhad Mehdipour, Koji Inoue, Kazuaki J. Murakami |
J. Supercomput. | 3 |
| 2011 | Illumination-robust face recognition via sparse representationabstractSparse representation receives a lot of attention as a representation method of signals in recent years. The sparse representation-based classification (SRC) applies sparse representation to face recognition. Camera parameters and/or illumination change may occur in real situations of face recognition. This study aims to propose a illumination robust SRC. The proposed method modifies the bases which are added to training bases derived from the training data set to detect noise contained in the recognized facial image. In the case where a few training data set is prepared, this method succeeds to obtain the higher recognition rate than the conventional SRC. Koji Inoue, Yoshimitsu Kuroki |
VCIP | 1 |
| 2011 | A design scheme for a reconfigurable accelerator implemented by single-flux quantum circuits
Farhad Mehdipour, Hiroaki Honda, Koji Inoue, Hiroshi Kataoka, Kazuaki J. Murakami |
J. Syst. Archit. | 3 |
| 2010 | Mapping scientific applications on a large-scale data-path accelerator implemented by single-flux quantum (SFQ) circuitsabstractTo overcome issues originating from the CMOS technology, a large-scale reconfigurable data-path (LSRDP) processor based on single-flux quantum circuits is introduced. LSRDP is augmented to a general purpose processor to accelerate the execution of data flow graphs (DFGs) extracted from scientific applications. Procedure of mapping large DFGs onto the LSRDP is discussed and our proposed techniques for reducing area of the accelerator within the design procedure will be introduced as well. Farhad Mehdipour, Hiroaki Honda, Hiroshi Kataoka, Koji Inoue, Irina Kataeva, Kazuaki J. Murakami, Hiroyuki Akaike, Akira Fujimaki |
DATE | 4 |
| 2009 | A combined analytical and simulation-based model for performance evaluation of a reconfigurable instruction set processorabstractPerformance evaluation is a serious challenge in designing or optimizing reconfigurable instruction set processors. The conventional approaches based on synthesis and simulations are very time consuming and need a considerable design effort. A combined analytical and simulation-based model (CAnSO*) is proposed and validated for performance evaluation of a typical reconfigurable instruction set processor. The proposed model consists of an analytical core that incorporates statistics gathered from cycle-accurate simulation to make a reasonable evaluation and provide a valuable insight. Compared to cycle-accurate simulation results, CAnSO proves almost 2% variation in the speedup measurement. Farhad Mehdipour, Hamid Noori, Bahman Javadi, Hiroaki Honda, Koji Inoue, Kazuaki J. Murakami |
ASP-DAC | 5 |
| 2008 | Design space exploration for a coarse grain acceleratorabstractIn the design process of a reconfigurable accelerator employing in an embedded system, multitude parameters may result in remarkable complexity and a large design space. Design space exploration as an alternative to the quantitative approach can be employed to find a right balance between the different design parameters. In this paper, a hybrid approach is introduced to analytically explore the design space for a coarse grain accelerator and determine a wise design point exploiting data extracted from applications, quantitatively. It also provides flexibility for taking into account new design constraints as well as new characteristics of applications. Furthermore, this approach is a methodological approach which reduces the design time and results in a point which satisfies the design goals. Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Koji Inoue, Kazuaki J. Murakami |
ASP-DAC | 4 |
| 2008 | Enhancing energy efficiency of processor-based embedded systems through post-fabrication ISA extensionabstractApplication-specific instruction set extension is an effective technique for reducing accesses to components such as on- and off-chip memories, register file and enhancing the energy efficiency. However, the addition of custom functional units to the base processor is required for supporting custom instructions, which due to the increase of manufacturing and design costs in new nanometer-scale technologies and shorter time-to-market, is becoming an issue. To address above issues, in our proposed approach, an optimized reconfigurable functional unit is used instead, and instruction set customization is done after chip-fabrication. Therefore, while maintaining the flexibility of a conventional microprocessor, the low-energy feature of customization is applicable. Experimental results show that the maximum and average energy savings are 67% and 22%, respectively for our proposed architecture framework. Hamid Noori, Farhad Mehdipour, Koji Inoue, Kazuaki J. Murakami |
ISLPED | 3 |
| 2008 | Performance prediction of large-scale parallell system and application using macro-level simulationabstractTo predict application performance on an HPC system is an important technology for designing the computing system and developing applications. However, accurate prediction is a challenge, particularly, in the case of a future coming system with higher performance. In this paper, we present a new method for predicting application performance on HPC systems. This method combines modeling of sequential performance on a single processor and macro-level simulations of applications for parallel performance on the entire system. In the simulation, the execution flow is traced but kernel computations are omitted for reducing the execution time. Validation on a real terascale system showed that the predicted and measured performance agreed within 10% to 20 %. We employed the method in designing a hypothetical petascale system of 32768 SIMD-extended processor cores. For predicting application performance on the petascale system, the macro-level simulation required several hours. Ryutaro Susukita, Hisashige Ando, Mutsumi Aoyagi, Hiroaki Honda, Yuichi Inadomi, Koji Inoue, Shigeru Ishizuki, Yasunori Kimura, Hidemi Komatsu, Motoyoshi Kurokawa, Kazuaki J. Murakami, Hidetomo Shibamura, Shuji Yamamura, Yunqing Yu |
SC | 6 |
| 2008 | An architecture framework for an adaptive extensible processor
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani |
J. Supercomput. | 4 |
| 2007 | Interactive presentation: Generating and executing multi-exit custom instructions for an adaptive extensible processorabstractTo improve the performance of embedded processors, an effective technique is collapsing critical computation subgraphs as application-specific instruction set extensions and executing them on custom functional units. The problems of this approach are immense cost and long time of designing. To address these issues, an adaptive extensible processor was proposed in which custom instructions (CIs) are generated and added after chip-fabrication. To support this feature, custom functional units are replaced by a reconfigurable matrix of functional units with the capability of conditional execution. Unlike previous proposed CIs, it can include multiple exits. Experimental results show that multi-exit CIs enhance the performance by 46% in average compared to CIs limited to one basic block. A maximum speedup of 2.89 compared to a 4-issue in-order RISC processor, and a speedup of 1.66 in average, was achieved on MiBench benchmark suite Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Maziar Goudarzi |
DATE | 4 |
| 2007 | The effect of temperature on cache size tuning for low energy embedded systemsabstractEnergy consumption is a major concern in embedded computing systems. Several studies have shown that cache memories account for about 40% or more of the total energy consumed in these systems. In older technology nodes, active power was the primary contributor to total power dissipation of a CMOS design. However, with the scaling of feature sizes, the share of leakage in total power consumption of digital systems continues to grow. Temperature is a factor which exponentially increases the leakage current. In this paper, we show the effects of temperature on the selection of optimal cache size for low energy embedded systems. Our results show that for a given application, the optimal cache size selection is affected by the temperature. Our experiments have been done for 100nm technology. Our study reveals that the cache size selection for different temperatures depends on the rate at which cache miss increases when reducing the cache size. When the miss rate increases sharply the optimal point is the same for all examined temperatures, however when it becomes smoother, the optimal point for different temperatures begin to get farther. Hamid Noori, Maziar Goudarzi, Koji Inoue, Kazuaki J. Murakami |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | Custom Instruction Generation Using Temporal Partitioning Techniques for a Reconfigurable Functional Unit
Farhad Mehdipour, Hamid Noori, Morteza Saheb Zamani, Kazuaki J. Murakami, Koji Inoue, Mehdi Sedighi |
EUC | 5 |
| 2006 | A Reconfigurable Functional Unit for an Adaptive Dynamic Extensible ProcessorabstractThis paper presents a reconfigurable functional unit (RFU) for an adaptive dynamic extensible processor. The processor can tune its extended instructions to the target applications, after chip-fabrication. The custom instructions (CIs) are generated deploying the hot basic blocks during the training mode. In the normal mode, CIs are executed on the RFU. A quantitative approach was used for designing the RFU. The RFU is a matrix of functional units with 8 inputs and 6 outputs. Performance is enhanced up to 1.25 using the proposed RFU for 22 applications of Mibench. This processor needs no extra opcodes for CIs, new compiler, source code modification and recompilation. Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Morteza Saheb Zamani |
FPL | 4 |
| 2002 | A Low Energy Set-Associative I-Cache with Extended BTBabstractThis paper proposes a low-energy instruction-cache architecture, called history-based tag-comparison (HBTC) cache. The HBTC cache attempts to re-use tag-comparison results for avoiding unnecessary way activation in set-associative caches. The cache records tag-comparison results in an extended BTB, and re-uses them for directly selecting only the hit-way which includes the target instruction. In our simulation, it is observed that the HBTC cache can achieve 62% of energy reduction, with less than 1% performance degradation, compared with a conventional cache. Koji Inoue, Vasily G. Moshnyaga, Kazuaki J. Murakami |
ICCD | 1 |
| 2002 | A history-based I-cache for low-energy multimedia applicationsabstractThis paper proposes a history-based tag-comparison scheme for reducing energy consumption of direct-mapped instruction caches. The proposed cache efficiently exploits program-execution footprints recorded in the Branch Target Buffer (BTB), and attempts to detect and eliminate unnecessary tag checks at run time. Simulation results show that our approach can eliminate up to 95% of tag checks, saving the cache energy by 17%, while affecting the processor performance by only 0.2%. Koji Inoue, Vasily G. Moshnyaga, Kazuaki J. Murakami |
ISLPED | 1 |
| 2002 | Reducing energy consumption of video memory by bit-width compressionabstractA new architectural technique to reduce energy dissipation of video memory is proposed. Unlike existing approaches, the technique exploits the pixel correlation in video sequences, dynamically adjusting the memory bit-width to the number of bits changed per pixel. Instead of treating the data bits independently, we group the most significant bits together, activating the corresponding group of bit-lines adaptively to data variation. The method is not restricted to the specific bit-patterns nor depends on the storage phase. It works equally well on read and write accesses, as well as during precharging. Simulation results show that using this method we can reduce the total energy consumption of video memory by 20% without affecting the picture quality. Vasily G. Moshnyaga, Koji Inoue, Mizuka Fukagawa |
ISLPED | 2 |
| 1999 | Dynamically Variable Line-Size Cache Exploiting High On-Chip Memory Bandwidth of Merged DRAM/Logic LSIsabstractThis paper proposes a novel cache architecture suitable for merged DRAM/logic LSIs, which is called "dynamically variable line-size cache (D-VLS cache)". The D-VLS cache can optimize its line-size according to the characteristic of programs, and attempts to improve the performance by exploiting the high on-chip memory bandwidth. In our evaluation, it is observed that the performance improvement achieved by a direct-mapped D-VLS cache is about 27%, compared to a conventional direct-mapped cache with fixed 32-byte lines. Koji Inoue, Koji Kai, Kazuaki J. Murakami |
HPCA | 1 |
| 1999 | Way-predicting set-associative cache for high performance and low energy consumptionabstractSet-Associative CacheThis paper proposes a new approach using way prediction for achieving high performance and low energy consumption of set-associative caches.By accessing only a single cache way predicted, instead of accessing all the ways in a set, the energy consumption can be reduced.This paper shows that the way-predicting set-associative cache improves the ED (energy-delay) product by SO-70% compared to a conventional set-associative cache. Conventional CacheTotal energy consumed for an access to a set-associative cache ( Ecache) can be approximated by the sum of following terms [9]: Koji Inoue, Tohru Ishihara, Kazuaki J. Murakami |
ISLPED | 1 |
| 1999 | MOE: A Special-Purpose Parallel Computer for High-Speed, Large-Scale Molecular Orbital CalculationabstractArticle MOE: a special-purpose parallel computer for high-speed, large-scale molecular orbital calculation Share on Authors: Koji Hashimoto Kyushu University Kyushu UniversityView Profile , Hiroto Tomita Kyushu University Kyushu UniversityView Profile , Koji Inoue Kyushu University Kyushu UniversityView Profile , Katsuhiko Metsugi Kyushu University Kyushu UniversityView Profile , Kazuaki Murakami Kyushu University Kyushu UniversityView Profile , Shinjiro Inabata Fuji Xerox Co., Ltd Fuji Xerox Co., LtdView Profile , So Yamada Fuji Xerox Co., Ltd Fuji Xerox Co., LtdView Profile , Nobuaki Miyakawa Fuji Xerox Co., Ltd Fuji Xerox Co., LtdView Profile , Hajime Takashima Taisho Pharmaceutical Co., Ltd Taisho Pharmaceutical Co., LtdView Profile , Kunihiro Kitamura Taisho Pharmaceutical Co., Ltd Taisho Pharmaceutical Co., LtdView Profile , Shigeru Obara Hokkaido University of Education Hokkaido University of EducationView Profile , Takashi Amisaki Shimane University Shimane UniversityView Profile , Kazutoshi Tanabe National Institute of Materials and Chemical Research National Institute of Materials and Chemical ResearchView Profile , Umpei Nagashima National Institute for Advanced Interdisciplinary Research National Institute for Advanced Interdisciplinary ResearchView Profile Authors Info & Claims SC '99: Proceedings of the 1999 ACM/IEEE conference on SupercomputingJanuary 1999 Pages 58–eshttps://doi.org/10.1145/331532.331590Online:01 January 1999Publication History 2citation285DownloadsMetricsTotal Citations2Total Downloads285Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Koji Hashimoto, Hiroto Tomita, Koji Inoue, Katsuhiko Metsugi, Kazuaki J. Murakami, Shinjiro Inabata, So Yamada, Nobuaki Miyakawa, Hajime Takashima, Kunihiro Kitamura, Shigeru Obara, Takashi Amisaki, Kazutoshi Tanabe, Umpei Nagashima |
SC | 3 |