EDBT 2026 Demo / reviewers in the wild / expert
Hiroki Mori
dblp:25/3722
· DBLP profile ↗
56ranked-venue papers
21as first author
9since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 14 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 15 first-author · 2 since 2021Systems, architecture and hardware · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Multi-Timestep-Ahead Prediction with Mixture of Experts for Embodied Question Answering
Kanata Suzuki, Yuya Kamiwano, Naoya Chiba, Hiroki Mori, Tetsuya Ogata |
ICANN (6) | 4 |
| 2023 | Multimodal Time Series Learning of Robots Based on Distributed and Integrated Modalities: Verification with a Simulator and Actual RobotsabstractWe have developed an autonomous robot motion generation model based on distributed and integrated multimodal learning. Since each modality used as a robot's senses, such as image, joint angle, and torque, has a different physical meaning and time characteristic, the generation of autonomous motions using multimodal learning has sometimes failed due to overlearning in one of the modalities. Inspired by the sensory processing of the human brain, our model is based on the processing of each sense performed in the primary somatosensory cortex and the integrated processing of multiple senses in the association cortex and the primary motor cortex. Specifically, the proposed model utilizes two types of recurrent neural networks: sensory RNNs, which learn each sense in a time series, and a union RNN, which communicates with sensory RNNs and learns sensory integration. The simulation results of multiple tasks showed that our model processes multiple modalities appropriately and generates smoother motions with lower jerk than the conventional model. We also demonstrated a chair assembly task by combining fixed motions and autonomous motions with our model. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 4 |
| 2023 | A Generative Framework for Conversational Laughter: Its 'Language Model' and Laughter Sound SynthesisabstractAs the phonetic and acoustic manifestations of laughter in conversation are highly diverse, laughter synthesis should be capable of accommodating such diversity while maintaining high controllability.This paper proposes a generative model of laughter in conversation that can produce a wide variety of laughter by utilizing the emotion dimension as a conversational context.The model comprises two parts: the laughter "phones generator," which generates various, but realistic, combinations of laughter components for a given speaker ID and emotional state, and the laughter "sound synthesizer," which receives the laughter phone sequence and produces acoustic features that reflect the speaker's individuality and emotional state.The results of a listening experiment indicated that conditioning both the phones generator and the sound synthesizer on emotion dimensions resulted in the most effective control of the perceived emotion in synthesized laughter. Hiroki Mori, Shunya Kimura |
INTERSPEECH | 1 |
| 2022 | Contact-Rich Manipulation of a Flexible Object based on Deep Predictive Learning using Vision and TactilityabstractWe achieved contact-rich flexible object manipulation, which was difficult to control with vision alone. In the unzipping task we chose as a validation task, the gripper grasps the puller, which hides the bag state such as the direction and amount of deformation behind it, making it difficult to obtain information to perform the task by vision alone. Additionally, the flexible fabric bag state constantly changes during operation, so the robot needs to dynamically respond to the change. However, the appropriate robot behavior for all bag states is difficult to prepare in advance. To solve this problem, we developed a model that can perform contact-rich flexible object manipulation by real-time prediction of vision with tactility. We introduced a point-based attention mechanism for extracting image features, softmax transformation for predicting motions, and convolutional neural network for extracting tactile features. The results of experiments using a real robot arm revealed that our method can realize motions responding to the deformation of the bag while reducing the load on the zipper. Furthermore, using tactility improved the success rate from 56.7% to 93.3% compared with vision alone, demonstrating the effectiveness and high performance of our method. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 4 |
| 2022 | Integrated Learning of Robot Motion and Sentences: Real-Time Prediction of Grasping Motion and Attention based on Language InstructionsabstractWe propose a motion generation model that can achieve robust behavior against environmental changes based on language instructions at a low cost. Conventional robots that communicate with humans use a restricted environment and language to build up a mapping between language and motion, and thus need to prepare a huge training set in order to achieve versatility. Our method trains pairs of language, visual, and motor information of the robot, and generates motions in real-time based on the “attention” of the language instructions. Specifically, the robot generates motions while focusing on the indicated objects by the human when multiple objects are in the field of view. In addition, since position recognition and motion generation of the indicated object are performed in real-time, robust motion generation is possible in response to changes in the object position and lighting conditions. We clarified that features related to the object name and its location are self-organized in the latent (PB: Parametric Bias) space by end-to-end learning of robot motion and sentences. These observations may indicate the importance of integrated learning of robot motion and sentences since such feature representations cannot be obtained by learning motions alone. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 4 |
| 2022 | Guided Visual Attention Model Based on Interactions Between Top-down and Bottom-up Prediction for Robot Pose PredictionabstractDeep robot vision models are widely used for recognizing objects from camera images, but shows poor performance when detecting objects at untrained positions. Although such problem can be alleviated by training with large datasets, the dataset collection cost cannot be ignored. Existing visual attention models tackled the problem by employing a data efficient structure which learns to extract task relevant image areas. However, since the models cannot modify attention targets after training, it is difficult to apply to dynamically changing tasks. This paper proposed a novel Key-Query-Value formulated visual attention model. This model is capable of switching attention targets by externally modifying the Query representations, namely top-down attention. The proposed model is experimented on a simulator and a real-world environment. The model was compared to existing end-to-end robot vision models in the simulator experiments, showing higher performance and data efficiency. In the real-world robot experiments, the model showed high precision along with its scalability and extendibility. Hyogo Hiruma, Hiroki Mori, Tetsuya Ogata |
IECON | 2 |
| 2021 | A Peg-in-hole Task Strategy for Holes in ConcreteabstractA method that enables an industrial robot to accomplish the peg-in-hole task for holes in concrete is proposed. The proposed method involves slightly detaching the peg from the wall, when moving between search positions, to avoid the negative influence of the concrete’s high friction coefficient. It uses a deep neural network (DNN), trained via reinforcement learning, to effectively find holes with variable shape and surface finish (due to the brittle nature of concrete) without analytical modeling or control parameter tuning. The method uses displacement of the peg toward the wall surface, in addition to force and torque, as one of the inputs of the DNN. Since the displacement increases as the peg gets closer to the hole (due to the chamfered shape of holes in concrete), it is a useful parameter for inputting in the DNN. The proposed method was evaluated by training the DNN on a hole 500 times and attempting to find 12 unknown holes. The results of the evaluation show the DNN enabled a robot to find the unknown holes with average success rate of 96.1% and average execution time of 12.5 seconds. Additional evaluations with random initial positions and a different type of peg demonstrate the trained DNN can generalize well to different conditions. Analyses of the influence of the peg displacement input showed the success rate of the DNN is increased by utilizing this parameter. These results validate the proposed method in terms of its effectiveness and applicability to the construction industry. André Yuji Yasutomi, Hiroki Mori, Tetsuya Ogata |
ICRA | 2 |
| 2021 | Pitch Contour Separation from Overlapping Speech
Hiroki Mori |
Interspeech | 1 |
| 2021 | In-air Knotting of Rope using Dual-Arm Robot based on Deep LearningabstractIn this study, we report the successful execution of in-air knotting of rope using a dual-arm two-finger robot based on deep learning. Owing to its flexibility, the state of the rope was in constant flux during the operation of the robot. This required the robot control system to dynamically correspond to the state of the object at all times. However, a manual description of appropriate robot motions corresponding to all object states is difficult to be prepared in advance. To resolve this issue, we constructed a model that instructed the robot to perform bowknots and overhand knots based on two deep neural networks trained using the data gathered from its sensorimotor, including visual and proximity sensors. The resultant model was verified to be capable of predicting the appropriate robot motions based on the sensory information available online. In addition, we designed certain task motions based on the Ian knot method using the dual-arm two-fingers robot. The designed knotting motions do not require a dedicated workbench or robot hand, thereby enhancing the versatility of the proposed method. Finally, experiments were performed to estimate the knotting performance of the real robot while executing overhand knots and bowknots on rope and its success rate. The experimental results established the effectiveness and high performance of the proposed method. Kanata Suzuki, Momomi Kanamura, Yuki Suga, Hiroki Mori, Tetsuya Ogata |
IROS | 4 |
| 2020 | Gaming Corpus for Studying Social Screams
Hiroki Mori, Yuki Kikuchi |
INTERSPEECH | 1 |
| 2020 | Wiping 3D-objects using Deep Learning Model based on Image/Force/Joint InformationabstractWe propose a deep learning model for a robot to wipe 3D-objects. Wiping of 3D-objects requires recognizing the shapes of objects and planning the motor angle adjustments for tracing the objects. Unlike previous research, our learning model does not require pre-designed computational models of target objects. The robot is able to wipe the objects to be placed by using image, force, and arm joint information. We evaluate the generalization ability of the model by confirming that the robot handles untrained cube and bowl shaped-objects. We also find that it is necessary to use both image and force information to recognize the shape of and wipe 3D objects consistently by comparing changes in the input sensor data to the model. To our knowledge, this is the first work enabling a robot to use learning sensorimotor information alone to trace various unknown 3D-shape. Namiko Saito, Tetsuya Ogata, Hiroki Mori, Shigeki Sugano |
IROS | 4 |
| 2020 | Defining Laughter Context for Laughter Synthesis with Spontaneous Speech CorpusabstractIn this paper, conversational laughter was synthesized by a statistical model-based speech synthesis framework using spontaneous speech corpora. The phonetic transcriptions of natural laughter in these corpora were annotated, and the context required to synthesize the laughter that accompanies speech sounds was defined from the perspective of the (1) phonetic properties of the current segment, (2) phonetic properties of previous and succeeding segments, and (3) positional factors of the current segment or laughter bout. Laughter was synthesized using the defined context and the framework of HMM-based speech synthesis. To confirm the influence of the contextual factors on the naturalness of speech, a subjective evaluation was performed. As the result of the evaluation, the naturalness of the entire utterance was improved by using the contextual factors defined in this study. This result confirmed the importance of defining the appropriate context to synthesize natural conversational laughter. Tomohiro Nagata, Hiroki Mori |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | End-to-end Learning Method for Self-Driving Cars with Trajectory Recovery Using a Path-following FunctionabstractWe propose an end-to-end learning method for autonomous driving systems in this article. End-to-end model estimates an appropriate motor command from raw sensory signals. End-to-end model for autonomous driving systems has recently been based on neural networks, which are popular for their good recognition ability. A common problem is how to return a car to the driving lane when the car goes off the track. In our research, we collect recovery data based on the distance from a desired track (the nearest waypoint link) during a road test with a simulator. To train the recovery behavior, instead of collecting human driving data, we use a path-following module (which means the car automatically drives on a pre-decided route using the car's current position). Our proposed method is divided into three phases. In phase 1, we collect data only using a path-following module during 100 laps of driving. In phase 2, we generate driving behavior using a neural driving module trained by the data collected in phase 1. This includes switching between the accelerator, brake and steering based on a threshold. We collect further data on the recovery behavior using the path-following module during 100 laps of driving. In phase 3, we generate driving behavior using the neural driving module trained by the data collected in phases 1 and 2. To assess the proposed method, we compared the average distance from the nearest waypoint link and the average distance traveled per lap for datasets with no recovery, for datasets with random recovery, and for datasets for the proposed method with recovery. A model based on the proposed method drove well and paid more attention to the road rather than the sky and other unrelated objects automatically for both untrained and trained courses and weather. Tadashi Onishi, Toshiyuki Motoyoshi, Yuki Suga, Hiroki Mori, Tsuya Ogata |
IJCNN | 4 |
| 2019 | Conversational and Social Laughter Synthesis with WaveNetabstractThe studies of laughter synthesis are relatively few, and they are still in a preliminary stage. We explored the possibility of applying WaveNet to laughter synthesis. WaveNet is potentially more suitable to model laughter waveforms that do not have a well-established theory of production like speech signals. Conversational laughter was modelled with a spontaneous dialogue speech corpus based on WaveNet. To obtain more stable laughter generation, conditioning WaveNet by power contour was proposed. Experimental results showed that the synthesized laughter by WaveNet was perceived as closer to natural laughter than HMM-based synthesized laughter. Hiroki Mori, Tomohiro Nagata, Yoshiko Arimoto |
INTERSPEECH | 1 |
| 2019 | Learning Multiple Sensorimotor Units to Complete Compound Tasks using an RNN with Multiple AttractorsabstractAs the complexity of the robot's tasks increases, we can consider many general tasks in a compound form that consists of shorter tasks. Therefore, for robots to generate various tasks, they need to be able to execute shorter tasks in succession, appropriately to the situation. With the design principle to construct the architecture for robots to execute complex tasks compounded with multiple subtasks, this study proposes a visuomotor-control framework with the characteristics of a state machine to train shorter tasks as sensorimotor units. The design procedure of training framework consists of 4 steps: (1) segment entire task into appropriate subtasks, (2) define subtasks as states and transitions in a state machine, (3) collect subtasks data, and (4) train neural networks: (a) autoencoder to extract visual features, (b) a single recurrent neural network to generate subtasks to realize a pseud-state-machine model with a constraint in hidden values. We implemented this framework on two different robots to allow their performance of repetitive tasks with error-recovery motion, subsequently, confirming the ability of the robot to switch the sensorimotor units from visual input at the attractors of the hidden values created by the constraint. Kei Kase, Ryoichi Nakajo, Hiroki Mori, Tetsuya Ogata |
IROS | 3 |
| 2018 | Put-in-Box Task Generated from Multiple Discrete Tasks by aHumanoid Robot Using Deep LearningabstractFor robots to have a wide range of applications, they must be able to execute numerous tasks. However, recent studies into robot manipulation using deep neural networks (DNN) have primarily focused on single tasks. Therefore, we investigate a robot manipulation model that uses DNNs and can execute long sequential dynamic tasks by performing multiple short sequential tasks at appropriate times. To generate compound tasks, we propose a model comprising two DNNs: a convolutional autoencoder that extracts image features and a multiple timescale recurrent neural network (MTRNN) to generate motion. The internal state of the MTRNN is constrained to have similar values at the initial and final motion steps; thus, motions can be differentiated based on the initial image input. As an example compound task, we demonstrate that the robot can generate a “Put-In-Box” task that is divided into three subtasks: open the box, grasp the object and put it into the box, and close the box. The subtasks were trained as discrete tasks, and the connections between each subtask were not trained. With the proposed model, the robot could perform the Put-In-Box task by switching among subtasks and could skip or repeat subtasks depending on the situation. Kei Kase, Kanata Suzuki, Pin-Chu Yang, Hiroki Mori, Tetsuya Ogata |
ICRA | 4 |
| 2018 | Effects of Dimensional Input on Paralinguistic Information Perceived from Synthesized Dialogue Speech with Neural NetworkabstractA novel method of controlling paralinguistic information in neural network-based dialogue speech synthesis is proposed. Controlling paralinguistic information was achieved by feeding emotion dimensions in continuous values into the input layer of the neural networks. Compared to the method using the multiple regression HMM, the naturalness of synthesized speech was improved. The controllability of paralinguistic information was evaluated by examining the shift of the distribution of synthesized parameters. A subjective evaluation test revealed that the correlation between given and perceived paralinguistic information was moderate, though less apparent compared to the multiple regression HMM-based method. Masaki Yokoyama, Tomohiro Nagata, Hiroki Mori |
INTERSPEECH | 3 |
| 2017 | Causal Patterns: Extraction of Multiple Causal Relationships by Mixture of Probabilistic Partial Canonical Correlation AnalysisabstractIn this paper, we propose a mixture of probabilistic partial canonical correlation analysis (MPPCCA) that extracts the Causal Patterns from two multivariate time series. Causal patterns refer to the signal patterns within interactions of two elements having multiple types of mutually causal relationships, rather than a mixture of simultaneous correlations or the absence of presence of a causal relationship between the elements. In multivariate statistics, partial canonical correlation analysis (PCCA) evaluates the correlation between two multivariates after subtracting the effect of the third multivariate. PCCA can calculate the Granger Causality Index (which tests whether a time-series can be predicted from another time-series), but is not applicable to data containing multiple partial canonical correlations. After introducing the MPPCCA, we propose an expectation-maxmization (EM) algorithm that estimates the parameters and latent variables of the MPPCCA. The MPPCCA is expected to extract multiple partial canonical correlations from data series without any supervised signals to split the data as clusters. The method was then evaluated in synthetic data experiments. In the synthetic dataset, our method estimated the multiple partial canonical correlations more accurately than the existing method. To determine the types of patterns detectable by the method, experiments were also conducted on real datasets. The method estimated the communication patterns In motion-capture data. The MPPCCA is applicable to various type of signals such as brain signals, human communication and nonlinear complex multibody systems. Hiroki Mori, Keisuke Kawano, Hiroki Yokoyama |
DSAA | 1 |
| 2017 | A Novel Signal Detection Method for Interference from Inverter Microwave Ovens in WLAN SystemsabstractInverter microwave ovens (MWOs) are day-to-day appliances that can act as fatal interferers for 2.4 GHz wireless LAN (WLAN) systems. However, the interference is avoidable, if it is possible to detect the MWO signal. In this paper, we propose an inverter MWO signal detection method in the time and frequency domains for reliable wireless communication. The operating inverter MWOs inherently generate RF signal whose leakage power level drastically varies depending on the switching of the inverter. The proposed method recognizes the changing signal power level due to the switching. In addition, we confirm the effectiveness of the proposed method in the presence or absence of WLAN systems by using the experimental MWO data obtained in a shielded room. It is revealed that the proposed method is able to detect the inverter MWO signal correctly and is robust against different products, different receiving environment, and presence of other signals. Kensuke Nakanishi, Hiroki Mori, Takeshi Kumagaya, Tsuguhide Aoki |
GLOBECOM | 2 |
| 2017 | Emotion Category Mapping to Emotional Space by Cross-Corpus Emotion Labeling
Yoshiko Arimoto, Hiroki Mori |
INTERSPEECH | 2 |
| 2017 | Dimensional paralinguistic information control based on multiple-regression HSMM for spontaneous dialogue speech synthesis with robust parameter estimation
Tomohiro Nagata, Hiroki Mori, Takashi Nose |
Speech Commun. | 2 |
| 2016 | Voice-Quality Difference Between the Vowels in Filled Pauses and Ordinary Lexical Items
Kikuo Maekawa, Hiroki Mori |
INTERSPEECH | 2 |
| 2016 | Performance analysis of fault erasure belief propagation decoder based on density evolutionabstractIn this paper, we will present analysis of the fault erasure BP decoders based on the density evolution. In a fault BP decoder, messages exchanged in a BP process are stochastically corrupted due to unreliable logic gates and flip-flops; i.e., we here assume circuit components with transient faults. We derived a set of the density evolution equations for the fault erasure BP processes. Our density evolution analysis reveals the asymptotic behaviors of the estimation error probability of the fault erasure BP decoders. In contrast to the fault free cases, it is observed that the error probabilities of the fault BP decoder converge to positive values, and that there exists a discontinuity in an error curve corresponding to the fault BP threshold. It is also shown that an message encoding technique provides higher fault BP thresholds than those of the original decoders at the cost of increase of its circuit size. Hiroki Mori, Tadashi Wadayama |
ISIT | 1 |
| 2016 | Accuracy of Automatic Cross-Corpus Emotion Labeling for Conversational Speech Corpus Commonization
Hiroki Mori, Atsushi Nagaoka, Yoshiko Arimoto |
LREC | 1 |
| 2015 | OSPF and BGP State Migration for Resource-Portable IP RouterabstractThis paper proposes an Internet protocol (IP) state migration method for developing resource-portable IP routers that are not virtual-machine based but commercial based. Resource-portable IP routers have the potential for achieving a sustainable network by functioning as a shared backup router. While previous studies relied on a virtualized technology (e.g., a virtual machine-based router on commodity hardware), current commercial routers was not virtualized but implemented as a proprietary hardware and software. We achieved IP state migration for a proprietary router with control packet sniffing of the open shortest path first (OSPF) protocol and border gateway protocol (BGP) peer masquerade using a software-defined network controller. We implemented our method and verified the accuracy of the state migration of routing tables generated using OSPF and BGP. Shohei Kamamura, Hiroki Mori, Daisaku Shimazaki, Kouichi Genda, Yoshihiko Uematsu |
GLOBECOM | 2 |
| 2015 | 3-Dimensional Motion Recognition by 4-Dimensional Higher-order Local Auto-correlation
Hiroki Mori, Takaomi Kanda, Dai Hirose, Minoru Asada |
ICPRAM (1) | 1 |
| 2015 | Morphology of vocal affect bursts: exploring expressive interjections in Japanese conversation
Hiroki Mori |
INTERSPEECH | 1 |
| 2014 | An accelerated scanning communication system with adaptive automatic error correction mechanismabstractThis paper describes a novel automatic error correction method for scanning communication, whose mechanism is basically analogous to that of continuous speech recognition. It has two core components: one is the switch timing model, and the other is the statistical language model. By employing these models, the proposed system can estimate most probable sequence of input syllables for a given sequence of switch timing, with taking user characteristics into account. Thirteen subjects without disabilities and an ALS subject participated in a text input experiment using the proposed scanning communication system. For the ALS subject, the system improved the character correct rate from 77.7% to 97.7%, allowing dramatically fast input. Hiroki Mori |
ASSETS | 1 |
| 2013 | Robust estimation of multiple-regression HMM parameters for dimension-based expressive dialogue speech synthesis
Tomohiro Nagata, Hiroki Mori, Takashi Nose |
INTERSPEECH | 2 |
| 2012 | Throwing Skill Optimization through Synchronization and Desynchronization of Degree of Freedom
Yuji Kawai, Takato Horii, Yuji Oshima, Kazuaki Tanaka, Hiroki Mori, Yukie Nagai, Takashi Takuma, Minoru Asada |
RoboCup | 6 |
| 2011 | Generating avatar's facial expressions from emotional states in daily conversationabstractA framework for generating facial expressions from emotional states in daily conversation is described. The frame work allows avatars to express the speaker's state not just prototypical emotions. In this paper, the naturalness of generated facial expressions that are presented together with dialogue speech is examined. An experiment to examine the naturalness of facial expressions presented as still images shows that the two avatars' facial expressions are almost as natural as manually-made facial expressions. In an experiment to determine the natural display speed of dynamic facial expressions, significant interactions between display speed and emotion group were found for most emotion dimensions. Hiroki Mori, Ko Oshima, Makoto Nakamura |
ICASSP | 1 |
| 2011 | Pilot Signals for Multiuser Tomlinson-Harashima Precoding in MIMO-OFDM SystemsabstractThis paper describes pilot signals in multiuser multiple-input multiple-output orthogonal frequency division multiplexing systems with nonlinear Tomlinson-Harashima precoding (THP). The transmission processing of THP consists of three parts : feedback processing, modulo operation, and feedforward processing. In particular, the increased transmit power in the feedback processing is reduced by the modulo operation. In the conventional linear precoding systems, the same precoding weight is applied for both pilot and data signals at the transmitter for channel estimation at the receiver. If it is directly applied to THP, however, this will lead to significant increase in transmit power since the modulo operation cannot be applied for pilot signals. In our proposal, only the feedforward processing is applied for pilot signals and the interference due to the lack of feedback processing is removed by using orthogonal pilot signals. Simulation results show that the proposed pilot signals significantly improves the packet error rate performance. Tsuguhide Aoki, Hiroki Mori, Yuji Tohzaka, Yasuhiko Tanabe |
VTC Fall | 2 |
| 2011 | Constructing a spoken dialogue corpus for studying paralinguistic information in expressive conversation and analyzing its statistical/acoustic characteristics
Hiroki Mori, Tomoyuki Satake, Makoto Nakamura, Hideki Kasuya |
Speech Commun. | 1 |
| 2011 | A Wide-View Parallax-Free Eye-Mark Recorder with a Hyperboloidal Half-Silvered Mirror and Appearance-Based Gaze EstimationabstractIn this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror and a gaze estimation method suitable for the device. Our eye-mark recorder provides a wide field-of-view video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical angle of view of the prototype is 122 degree (elevation and depression angles are 38 and 84 degree, respectively) and its horizontal view angle is 116 degree (nasal and temporal view angles are 38 and 78 degree, respectively). We implemented and evaluated a gaze estimation method for our eye-mark recorder. We use an appearance-based approach for our eye-mark recorder to support a wide field-of-view. We apply principal component analysis (PCA) and multiple regression analysis (MRA) to determine the relationship between the captured images and their corresponding gaze points. Experimental results verify that our eye-mark recorder successfully captures a wide field-of-view of a user and estimates gaze direction with an angular accuracy of around 2 to 4 degree. Hiroki Mori, Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | Longitudinal changes of selected voice source parameters
Hideki Kasuya, Hajime Yoshida, Satoshi Ebihara, Hiroki Mori |
INTERSPEECH | 4 |
| 2009 | An efficient uplink multiuser MIMO protocol in IEEE 802.11 WLANsabstractThis paper proposes an uplink Multi-User (MU) Multiple-Input Multiple-Output (MIMO) protocol in IEEE 802.11 WLANs. In order to realize the uplink MU-MIMO transmission in 802.11 WLANs, there are several problems to be solved. Synchronized transmission among the stations (STAs) is one of the problems and the spatial compatibility between the transmitting STAs is another problem. Therefore, it is considered that many overheads are needed to realize the uplink MU-MIMO transmission. In the proposed protocol, after an access point (AP) informs the start of the uplink access phase by transmitting the indication frame, the STAs having data to be sent to the AP transmit the uplink access requests in the OFDMA manner. The AP recognizes the uplink access requests by detecting the subcarrier signals and it requests the detected STAs to transmit the pilot signals in the TDMA manner. The AP calculates the channel state information (CSI) between each STA from the received pilot signal and selects the STAs to be permitted to transmit the uplink frames based on the CSIs. The AP notifies the information about the permitted STAs and about the capable transmission rate, and then the permitted STAs simultaneously transmit the frames at the notified transmission rate. In the protocol, an efficient OFDMA-based uplink access request transmission scheme is also proposed. Thus, by adopting the proposed protocol, the overheads can be reduced and the network throughput can be enhanced. Computer simulations are performed and the results show the effectiveness of the proposed method. Tomoya Tandai, Hiroki Mori, Kiyoshi Toshimitsu, Takahiro Kobayashi 0001 |
PIMRC | 2 |
| 2009 | Cross-Layer-Optimized User Grouping Strategy in Downlink Multiuser MIMO SystemsabstractThis paper proposes a cross-layer-optimized user selection and packet transmission technique for the downlink multi-user (MU) multiple-input multiple-output (MIMO) system in IEEE 802.11 wireless LAN (WLAN). In IEEE 802.11 WLAN, packet duration is decided by a data size and a selected transmission mode (modulation and coding scheme, MCS), and therefore, in general, the packet duration of each user varies according to the channel conditions and the applications. If multiple users whose packet durations are considerably different are selected as the target of the downlink MU-MIMO transmission, then the efficiency of spatial multiplexing degrades and the network throughput cannot be enhanced. When selecting users as the target of the downlink MU-MIMO transmission, the degree of spatial channel correlation among users (from the viewpoint of the physical layer) and the packet waiting times in a queue (from the viewpoint of the MAC layer) have to be considered. In addition to these criteria, in IEEE 802.11 WLAN, uniformity of the packet durations among users has to be considered as a criterion. In the proposed method, data size is adjusted in order to make the packet durations of the users uniform after selecting the users based on the spatial compatibility and the packet waiting times. This means that the proposed downlink MU- MIMO transmission is optimized in 3 dimensions by the cross- layer approach: the spatial channel correlations, packet waiting times and the uniformity of the packet durations. Computer simulations are performed and the results show the effectiveness of the proposed method. Tomoya Tandai, Hiroki Mori, Masahiro Takagi |
VTC Spring | 2 |
| 2008 | Paralinguistic effects on turn-taking behavior in expressive conversation
Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 1 |
| 2007 | Voice source and vocal tract variations as cues to emotional states perceived from expressive conversational speech
Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 1 |
| 2006 | Two-layered neighborhood tabu search for multi-objective distribution network expansion planningabstractIn this paper, a new method is proposed for distribution network expansion planning. The proposed method is based on the multi-objective meta-heuristics with the epsiv-constraint method. As the power network becomes more deregulated and competitive, distribution network planning is faced with the uncertainty due to several factors. Thus, it is important to consider the uncertainty of the network planning. This paper proposes a new method that makes use of the Monte-Carlo simulation to consider the uncertainty of load patterns in the future. For many scenarios, the multi-objective two-layered tabu search is employed to determine the optimal allocation of feeders and substations. Two-layered tabu search separates integer variables by the difference of the neighborhoods. The proposed method is successfully applied to a sample system Hiroki Mori, Yasunori Yamada |
ISCAS | 1 |
| 2005 | Mora timing organization in producing contrastive geminate/single consonants and long/short vowels by native and non-native speakers of Japanese: effects of speaking rate
Haiping Jia, Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 2 |
| 2004 | F0 and formant frequency distribution of dysarthric speech - a comparative study
Hiroki Mori, Yasunori Kobayashi, Hideki Kasuya, Hajime Hirose, Noriko Kobayashi |
INTERSPEECH | 1 |
| 2003 | Acoustic variations of focused disyllabic words in Mandarin Chinese: analysis, synthesis and perception
Zhenglai Gu, Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 2 |
| 2003 | Speaker conversion in ARX-based source-formant type speech synthesis
Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 1 |
| 2002 | A data-driven approach to source-formant type text-to-speech system
Hiroki Mori, Takahiro Ohtsuka, Hideki Kasuya |
INTERSPEECH | 1 |
| 2001 | Invariance of relative F0 change field of Chinese disyllabic wordsabstractIn automatic voice response systems where a large number of words are inserted into fixed sentences, such as in voiceguided car navigation systems, one of the most important problems is the adjustment of the fundamental frequency (F0) contour of the inserted word to suit the F0 context of the fixed sentence. The effects of intonation and tone on the F0 contours of Chinese words can be described in terms of a word-level F0 range (WF0R) and an F0 change field (F0CF). WF0R in any position of a sentence is a tone-independent general F0 range, whereas F0CF is an F0 range taking the tone combination of words into account. Relative F0CF is regulated in reference to WF0R. If WF0R is used to represent the declination of a sentence, the relative F0CF should be invariant but dependent on the tone combination of a word. This paper examines the invariance of the relative F0CF among individuals. From an analysis of four native speakers’ utterances of 160 words in the initial, middle and final parts of three carrier sentences, conducted over 2 days, we show that: (1) Chinese speakers read words in the same sentence position with stable relative F0 change; (2) the relative F0CFs in the middle position of a sentence are generally the same as those in the initial position, but slightly different from those in the final position; and (3) the relative F0CFs reveal that the effects of tone on F0 contour is individual independent. Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 2 |
| 2000 | Prosodic variation of focused syllables of disyllabic word in Mandarin Chinese
Zhenglai Gu, Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 2 |
| 2000 | Automatic lexicon generation and dialogue modeling for spontaneous speech
Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 1 |
| 2000 | Word-level F0 range in Mandarin Chinese and its application to inserting words into a sentence
Hiroki Mori, Hideki Kasuya |
INTERSPEECH | 2 |
| 1999 | An Automatic Acquisition of Acoustical Units for Speech Recognition Based on Hidden Markov Network
Motoyuki Suzuki, Takafumi Hayashi, Hiroki Mori, Shozo Makino, Hirotomo Aso |
Discovery Science | 3 |
| 1998 | Multiple Camera Based Human Motion Estimation
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
ACCV (2) | 2 |
| 1998 | Automatic Acquisition of Phoneme Models and Its Application to Phoneme Labeling of a Large Size of Speech Corpus
Motoyuki Suzuki, Teruhiko Maeda, Hiroki Mori, Shozo Makino |
Discovery Science | 3 |
| 1998 | Multiple-Human Tracking Using Multiple Cameras
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
FG | 2 |
| 1998 | Multiple-view-based tracking of multiple humansabstractWe propose a multiple-view-based tracking algorithm for multiple-human motions. In vision-based human tracking, self-occlusions and human-human occlusions are a part of the more significant problems. Employing multiple viewpoints and a viewpoint selection mechanism, however can reduce these problems. In our system, human positions are tracked with a sequence of multiple-viewpoint images. This tracking is based on the Kalman filtering approach. The estimation results are utilized to select proper viewpoints in other sub-tasks (rotation angle detection and body-side detection). Each sub-task has a different criterion for selecting viewpoints. We also describe the criterions for accomplishing individual sub-tasks and relationships between sub-tasks. We have already built an experimental system based on a small number of reliable image features. We confirm the stability of our algorithm through simulations. We also performed fundamental examinations on the experimental system. Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
ICPR | 2 |
| 1998 | High-speed speaker adaptation using phoneme dependent tree-structured speaker clusteringabstractThe tree-structured speaker clustering was proposed as a highspeed speaker adaptation method. It can select the model which is most similar to a target speaker. However, this method does not consider speaker difference dependent on phoneme class. In this paper, we propose a speaker adaptation method based on speaker clustering by taking speaker difference dependent on phoneme class into account. The experimental results showed that the new method gave a better performance than the original method. Furthermore, we propose the improved method which use a tree-structure of a similar phoneme as the substitute for the phoneme which does not appear in the adaptation data. From the experimental results, the improved method gave a better performance than the method previously proposed. Motoyuki Suzuki, Toshiaki Abe, Hiroki Mori, Shozo Makino, Hirotomo Aso |
ICSLP | 3 |
| 1995 | Japanese document recognition based on interpolated n-gram model of characterabstractN-gram model is widely applied to various pattern recognition system because it well represents local features of natural languages. In this paper, we describe a contextual postprocessing method using a trigram model of character for Japanese document recognition, and its advantage is revealed by practical experiments. The model is automatically obtained by statistical processing of training documents. The ability to reduce ambiguity is evaluated by the perplexity. In the processing, two smoothing methods are examined, and the predictive power of the deleted interpolation method is shown to be superior. For leading articles, the perplexity reduced to about 22 when using deleted interpolation. The output from OCR is processed very fast using a Viterbi algorithm. Experimental results of recognition for three kinds of documents show that the error correction rates are ranged from 75 to over 90 percent. Hiroki Mori, Hirotomo Aso, Shozo Makino |
ICDAR | 1 |