Hiroki Mori

dblp:25/3722 · DBLP profile ↗
← Back
56ranked-venue papers
21as first author
9since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 14 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 15 first-author · 2 since 2021Systems, architecture and hardware · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2023 Multi-Timestep-Ahead Prediction with Mixture of Experts for Embodied Question Answering
Kanata Suzuki, Yuya Kamiwano, Naoya Chiba, Hiroki Mori, Tetsuya Ogata
ICANN (6)4
2023 Multimodal Time Series Learning of Robots Based on Distributed and Integrated Modalities: Verification with a Simulator and Actual Robots
abstract
We have developed an autonomous robot motion generation model based on distributed and integrated multimodal learning. Since each modality used as a robot's senses, such as image, joint angle, and torque, has a different physical meaning and time characteristic, the generation of autonomous motions using multimodal learning has sometimes failed due to overlearning in one of the modalities. Inspired by the sensory processing of the human brain, our model is based on the processing of each sense performed in the primary somatosensory cortex and the integrated processing of multiple senses in the association cortex and the primary motor cortex. Specifically, the proposed model utilizes two types of recurrent neural networks: sensory RNNs, which learn each sense in a time series, and a union RNN, which communicates with sensory RNNs and learns sensory integration. The simulation results of multiple tasks showed that our model processes multiple modalities appropriately and generates smoother motions with lower jerk than the conventional model. We also demonstrated a chair assembly task by combining fixed motions and autonomous motions with our model.
Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata
ICRA4
2023 A Generative Framework for Conversational Laughter: Its 'Language Model' and Laughter Sound Synthesis
abstract
As the phonetic and acoustic manifestations of laughter in conversation are highly diverse, laughter synthesis should be capable of accommodating such diversity while maintaining high controllability.This paper proposes a generative model of laughter in conversation that can produce a wide variety of laughter by utilizing the emotion dimension as a conversational context.The model comprises two parts: the laughter "phones generator," which generates various, but realistic, combinations of laughter components for a given speaker ID and emotional state, and the laughter "sound synthesizer," which receives the laughter phone sequence and produces acoustic features that reflect the speaker's individuality and emotional state.The results of a listening experiment indicated that conditioning both the phones generator and the sound synthesizer on emotion dimensions resulted in the most effective control of the perceived emotion in synthesized laughter.
Hiroki Mori, Shunya Kimura
INTERSPEECH1
2022 Contact-Rich Manipulation of a Flexible Object based on Deep Predictive Learning using Vision and Tactility
abstract
We achieved contact-rich flexible object manipulation, which was difficult to control with vision alone. In the unzipping task we chose as a validation task, the gripper grasps the puller, which hides the bag state such as the direction and amount of deformation behind it, making it difficult to obtain information to perform the task by vision alone. Additionally, the flexible fabric bag state constantly changes during operation, so the robot needs to dynamically respond to the change. However, the appropriate robot behavior for all bag states is difficult to prepare in advance. To solve this problem, we developed a model that can perform contact-rich flexible object manipulation by real-time prediction of vision with tactility. We introduced a point-based attention mechanism for extracting image features, softmax transformation for predicting motions, and convolutional neural network for extracting tactile features. The results of experiments using a real robot arm revealed that our method can realize motions responding to the deformation of the bag while reducing the load on the zipper. Furthermore, using tactility improved the success rate from 56.7% to 93.3% compared with vision alone, demonstrating the effectiveness and high performance of our method.
Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata
ICRA4
2022 Integrated Learning of Robot Motion and Sentences: Real-Time Prediction of Grasping Motion and Attention based on Language Instructions
abstract
We propose a motion generation model that can achieve robust behavior against environmental changes based on language instructions at a low cost. Conventional robots that communicate with humans use a restricted environment and language to build up a mapping between language and motion, and thus need to prepare a huge training set in order to achieve versatility. Our method trains pairs of language, visual, and motor information of the robot, and generates motions in real-time based on the “attention” of the language instructions. Specifically, the robot generates motions while focusing on the indicated objects by the human when multiple objects are in the field of view. In addition, since position recognition and motion generation of the indicated object are performed in real-time, robust motion generation is possible in response to changes in the object position and lighting conditions. We clarified that features related to the object name and its location are self-organized in the latent (PB: Parametric Bias) space by end-to-end learning of robot motion and sentences. These observations may indicate the importance of integrated learning of robot motion and sentences since such feature representations cannot be obtained by learning motions alone.
Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata
ICRA4
2022 Guided Visual Attention Model Based on Interactions Between Top-down and Bottom-up Prediction for Robot Pose Prediction
abstract
Deep robot vision models are widely used for recognizing objects from camera images, but shows poor performance when detecting objects at untrained positions. Although such problem can be alleviated by training with large datasets, the dataset collection cost cannot be ignored. Existing visual attention models tackled the problem by employing a data efficient structure which learns to extract task relevant image areas. However, since the models cannot modify attention targets after training, it is difficult to apply to dynamically changing tasks. This paper proposed a novel Key-Query-Value formulated visual attention model. This model is capable of switching attention targets by externally modifying the Query representations, namely top-down attention. The proposed model is experimented on a simulator and a real-world environment. The model was compared to existing end-to-end robot vision models in the simulator experiments, showing higher performance and data efficiency. In the real-world robot experiments, the model showed high precision along with its scalability and extendibility.
Hyogo Hiruma, Hiroki Mori, Tetsuya Ogata
IECON2
2021 A Peg-in-hole Task Strategy for Holes in Concrete
abstract
A method that enables an industrial robot to accomplish the peg-in-hole task for holes in concrete is proposed. The proposed method involves slightly detaching the peg from the wall, when moving between search positions, to avoid the negative influence of the concrete’s high friction coefficient. It uses a deep neural network (DNN), trained via reinforcement learning, to effectively find holes with variable shape and surface finish (due to the brittle nature of concrete) without analytical modeling or control parameter tuning. The method uses displacement of the peg toward the wall surface, in addition to force and torque, as one of the inputs of the DNN. Since the displacement increases as the peg gets closer to the hole (due to the chamfered shape of holes in concrete), it is a useful parameter for inputting in the DNN. The proposed method was evaluated by training the DNN on a hole 500 times and attempting to find 12 unknown holes. The results of the evaluation show the DNN enabled a robot to find the unknown holes with average success rate of 96.1% and average execution time of 12.5 seconds. Additional evaluations with random initial positions and a different type of peg demonstrate the trained DNN can generalize well to different conditions. Analyses of the influence of the peg displacement input showed the success rate of the DNN is increased by utilizing this parameter. These results validate the proposed method in terms of its effectiveness and applicability to the construction industry.
André Yuji Yasutomi, Hiroki Mori, Tetsuya Ogata
ICRA2
2021 Pitch Contour Separation from Overlapping Speech
Hiroki Mori
Interspeech1
2021 In-air Knotting of Rope using Dual-Arm Robot based on Deep Learning
abstract
In this study, we report the successful execution of in-air knotting of rope using a dual-arm two-finger robot based on deep learning. Owing to its flexibility, the state of the rope was in constant flux during the operation of the robot. This required the robot control system to dynamically correspond to the state of the object at all times. However, a manual description of appropriate robot motions corresponding to all object states is difficult to be prepared in advance. To resolve this issue, we constructed a model that instructed the robot to perform bowknots and overhand knots based on two deep neural networks trained using the data gathered from its sensorimotor, including visual and proximity sensors. The resultant model was verified to be capable of predicting the appropriate robot motions based on the sensory information available online. In addition, we designed certain task motions based on the Ian knot method using the dual-arm two-fingers robot. The designed knotting motions do not require a dedicated workbench or robot hand, thereby enhancing the versatility of the proposed method. Finally, experiments were performed to estimate the knotting performance of the real robot while executing overhand knots and bowknots on rope and its success rate. The experimental results established the effectiveness and high performance of the proposed method.
Kanata Suzuki, Momomi Kanamura, Yuki Suga, Hiroki Mori, Tetsuya Ogata
IROS4
2020 Gaming Corpus for Studying Social Screams
Hiroki Mori, Yuki Kikuchi
INTERSPEECH1
2020 Wiping 3D-objects using Deep Learning Model based on Image/Force/Joint Information
abstract
We propose a deep learning model for a robot to wipe 3D-objects. Wiping of 3D-objects requires recognizing the shapes of objects and planning the motor angle adjustments for tracing the objects. Unlike previous research, our learning model does not require pre-designed computational models of target objects. The robot is able to wipe the objects to be placed by using image, force, and arm joint information. We evaluate the generalization ability of the model by confirming that the robot handles untrained cube and bowl shaped-objects. We also find that it is necessary to use both image and force information to recognize the shape of and wipe 3D objects consistently by comparing changes in the input sensor data to the model. To our knowledge, this is the first work enabling a robot to use learning sensorimotor information alone to trace various unknown 3D-shape.
Namiko Saito, Tetsuya Ogata, Hiroki Mori, Shigeki Sugano
IROS4
2020 Defining Laughter Context for Laughter Synthesis with Spontaneous Speech Corpus
abstract
In this paper, conversational laughter was synthesized by a statistical model-based speech synthesis framework using spontaneous speech corpora. The phonetic transcriptions of natural laughter in these corpora were annotated, and the context required to synthesize the laughter that accompanies speech sounds was defined from the perspective of the (1) phonetic properties of the current segment, (2) phonetic properties of previous and succeeding segments, and (3) positional factors of the current segment or laughter bout. Laughter was synthesized using the defined context and the framework of HMM-based speech synthesis. To confirm the influence of the contextual factors on the naturalness of speech, a subjective evaluation was performed. As the result of the evaluation, the naturalness of the entire utterance was improved by using the contextual factors defined in this study. This result confirmed the importance of defining the appropriate context to synthesize natural conversational laughter.
Tomohiro Nagata, Hiroki Mori
IEEE Trans. Affect. Comput.2
2019 End-to-end Learning Method for Self-Driving Cars with Trajectory Recovery Using a Path-following Function
abstract
We propose an end-to-end learning method for autonomous driving systems in this article. End-to-end model estimates an appropriate motor command from raw sensory signals. End-to-end model for autonomous driving systems has recently been based on neural networks, which are popular for their good recognition ability. A common problem is how to return a car to the driving lane when the car goes off the track. In our research, we collect recovery data based on the distance from a desired track (the nearest waypoint link) during a road test with a simulator. To train the recovery behavior, instead of collecting human driving data, we use a path-following module (which means the car automatically drives on a pre-decided route using the car's current position). Our proposed method is divided into three phases. In phase 1, we collect data only using a path-following module during 100 laps of driving. In phase 2, we generate driving behavior using a neural driving module trained by the data collected in phase 1. This includes switching between the accelerator, brake and steering based on a threshold. We collect further data on the recovery behavior using the path-following module during 100 laps of driving. In phase 3, we generate driving behavior using the neural driving module trained by the data collected in phases 1 and 2. To assess the proposed method, we compared the average distance from the nearest waypoint link and the average distance traveled per lap for datasets with no recovery, for datasets with random recovery, and for datasets for the proposed method with recovery. A model based on the proposed method drove well and paid more attention to the road rather than the sky and other unrelated objects automatically for both untrained and trained courses and weather.
Tadashi Onishi, Toshiyuki Motoyoshi, Yuki Suga, Hiroki Mori, Tsuya Ogata
IJCNN4
2019 Conversational and Social Laughter Synthesis with WaveNet
abstract
The studies of laughter synthesis are relatively few, and they are still in a preliminary stage. We explored the possibility of applying WaveNet to laughter synthesis. WaveNet is potentially more suitable to model laughter waveforms that do not have a well-established theory of production like speech signals. Conversational laughter was modelled with a spontaneous dialogue speech corpus based on WaveNet. To obtain more stable laughter generation, conditioning WaveNet by power contour was proposed. Experimental results showed that the synthesized laughter by WaveNet was perceived as closer to natural laughter than HMM-based synthesized laughter.
Hiroki Mori, Tomohiro Nagata, Yoshiko Arimoto
INTERSPEECH1
2019 Learning Multiple Sensorimotor Units to Complete Compound Tasks using an RNN with Multiple Attractors
abstract
As the complexity of the robot's tasks increases, we can consider many general tasks in a compound form that consists of shorter tasks. Therefore, for robots to generate various tasks, they need to be able to execute shorter tasks in succession, appropriately to the situation. With the design principle to construct the architecture for robots to execute complex tasks compounded with multiple subtasks, this study proposes a visuomotor-control framework with the characteristics of a state machine to train shorter tasks as sensorimotor units. The design procedure of training framework consists of 4 steps: (1) segment entire task into appropriate subtasks, (2) define subtasks as states and transitions in a state machine, (3) collect subtasks data, and (4) train neural networks: (a) autoencoder to extract visual features, (b) a single recurrent neural network to generate subtasks to realize a pseud-state-machine model with a constraint in hidden values. We implemented this framework on two different robots to allow their performance of repetitive tasks with error-recovery motion, subsequently, confirming the ability of the robot to switch the sensorimotor units from visual input at the attractors of the hidden values created by the constraint.
Kei Kase, Ryoichi Nakajo, Hiroki Mori, Tetsuya Ogata
IROS3
2018 Put-in-Box Task Generated from Multiple Discrete Tasks by aHumanoid Robot Using Deep Learning
abstract
For robots to have a wide range of applications, they must be able to execute numerous tasks. However, recent studies into robot manipulation using deep neural networks (DNN) have primarily focused on single tasks. Therefore, we investigate a robot manipulation model that uses DNNs and can execute long sequential dynamic tasks by performing multiple short sequential tasks at appropriate times. To generate compound tasks, we propose a model comprising two DNNs: a convolutional autoencoder that extracts image features and a multiple timescale recurrent neural network (MTRNN) to generate motion. The internal state of the MTRNN is constrained to have similar values at the initial and final motion steps; thus, motions can be differentiated based on the initial image input. As an example compound task, we demonstrate that the robot can generate a “Put-In-Box” task that is divided into three subtasks: open the box, grasp the object and put it into the box, and close the box. The subtasks were trained as discrete tasks, and the connections between each subtask were not trained. With the proposed model, the robot could perform the Put-In-Box task by switching among subtasks and could skip or repeat subtasks depending on the situation.
Kei Kase, Kanata Suzuki, Pin-Chu Yang, Hiroki Mori, Tetsuya Ogata
ICRA4
2018 Effects of Dimensional Input on Paralinguistic Information Perceived from Synthesized Dialogue Speech with Neural Network
abstract
A novel method of controlling paralinguistic information in neural network-based dialogue speech synthesis is proposed. Controlling paralinguistic information was achieved by feeding emotion dimensions in continuous values into the input layer of the neural networks. Compared to the method using the multiple regression HMM, the naturalness of synthesized speech was improved. The controllability of paralinguistic information was evaluated by examining the shift of the distribution of synthesized parameters. A subjective evaluation test revealed that the correlation between given and perceived paralinguistic information was moderate, though less apparent compared to the multiple regression HMM-based method.
Masaki Yokoyama, Tomohiro Nagata, Hiroki Mori
INTERSPEECH3
2017 Causal Patterns: Extraction of Multiple Causal Relationships by Mixture of Probabilistic Partial Canonical Correlation Analysis
abstract
In this paper, we propose a mixture of probabilistic partial canonical correlation analysis (MPPCCA) that extracts the Causal Patterns from two multivariate time series. Causal patterns refer to the signal patterns within interactions of two elements having multiple types of mutually causal relationships, rather than a mixture of simultaneous correlations or the absence of presence of a causal relationship between the elements. In multivariate statistics, partial canonical correlation analysis (PCCA) evaluates the correlation between two multivariates after subtracting the effect of the third multivariate. PCCA can calculate the Granger Causality Index (which tests whether a time-series can be predicted from another time-series), but is not applicable to data containing multiple partial canonical correlations. After introducing the MPPCCA, we propose an expectation-maxmization (EM) algorithm that estimates the parameters and latent variables of the MPPCCA. The MPPCCA is expected to extract multiple partial canonical correlations from data series without any supervised signals to split the data as clusters. The method was then evaluated in synthetic data experiments. In the synthetic dataset, our method estimated the multiple partial canonical correlations more accurately than the existing method. To determine the types of patterns detectable by the method, experiments were also conducted on real datasets. The method estimated the communication patterns In motion-capture data. The MPPCCA is applicable to various type of signals such as brain signals, human communication and nonlinear complex multibody systems.
Hiroki Mori, Keisuke Kawano, Hiroki Yokoyama
DSAA1
2017 A Novel Signal Detection Method for Interference from Inverter Microwave Ovens in WLAN Systems
abstract
Inverter microwave ovens (MWOs) are day-to-day appliances that can act as fatal interferers for 2.4 GHz wireless LAN (WLAN) systems. However, the interference is avoidable, if it is possible to detect the MWO signal. In this paper, we propose an inverter MWO signal detection method in the time and frequency domains for reliable wireless communication. The operating inverter MWOs inherently generate RF signal whose leakage power level drastically varies depending on the switching of the inverter. The proposed method recognizes the changing signal power level due to the switching. In addition, we confirm the effectiveness of the proposed method in the presence or absence of WLAN systems by using the experimental MWO data obtained in a shielded room. It is revealed that the proposed method is able to detect the inverter MWO signal correctly and is robust against different products, different receiving environment, and presence of other signals.
Kensuke Nakanishi, Hiroki Mori, Takeshi Kumagaya, Tsuguhide Aoki
GLOBECOM2
2017 Emotion Category Mapping to Emotional Space by Cross-Corpus Emotion Labeling
Yoshiko Arimoto, Hiroki Mori
INTERSPEECH2
2017 Dimensional paralinguistic information control based on multiple-regression HSMM for spontaneous dialogue speech synthesis with robust parameter estimation
Tomohiro Nagata, Hiroki Mori, Takashi Nose
Speech Commun.2
2016 Voice-Quality Difference Between the Vowels in Filled Pauses and Ordinary Lexical Items
Kikuo Maekawa, Hiroki Mori
INTERSPEECH2
2016 Performance analysis of fault erasure belief propagation decoder based on density evolution
abstract
In this paper, we will present analysis of the fault erasure BP decoders based on the density evolution. In a fault BP decoder, messages exchanged in a BP process are stochastically corrupted due to unreliable logic gates and flip-flops; i.e., we here assume circuit components with transient faults. We derived a set of the density evolution equations for the fault erasure BP processes. Our density evolution analysis reveals the asymptotic behaviors of the estimation error probability of the fault erasure BP decoders. In contrast to the fault free cases, it is observed that the error probabilities of the fault BP decoder converge to positive values, and that there exists a discontinuity in an error curve corresponding to the fault BP threshold. It is also shown that an message encoding technique provides higher fault BP thresholds than those of the original decoders at the cost of increase of its circuit size.
Hiroki Mori, Tadashi Wadayama
ISIT1
2016 Accuracy of Automatic Cross-Corpus Emotion Labeling for Conversational Speech Corpus Commonization
Hiroki Mori, Atsushi Nagaoka, Yoshiko Arimoto
LREC1
2015 OSPF and BGP State Migration for Resource-Portable IP Router
abstract
This paper proposes an Internet protocol (IP) state migration method for developing resource-portable IP routers that are not virtual-machine based but commercial based. Resource-portable IP routers have the potential for achieving a sustainable network by functioning as a shared backup router. While previous studies relied on a virtualized technology (e.g., a virtual machine-based router on commodity hardware), current commercial routers was not virtualized but implemented as a proprietary hardware and software. We achieved IP state migration for a proprietary router with control packet sniffing of the open shortest path first (OSPF) protocol and border gateway protocol (BGP) peer masquerade using a software-defined network controller. We implemented our method and verified the accuracy of the state migration of routing tables generated using OSPF and BGP.
Shohei Kamamura, Hiroki Mori, Daisaku Shimazaki, Kouichi Genda, Yoshihiko Uematsu
GLOBECOM2
2015 3-Dimensional Motion Recognition by 4-Dimensional Higher-order Local Auto-correlation
Hiroki Mori, Takaomi Kanda, Dai Hirose, Minoru Asada
ICPRAM (1)1
2015 Morphology of vocal affect bursts: exploring expressive interjections in Japanese conversation
Hiroki Mori
INTERSPEECH1
2014 An accelerated scanning communication system with adaptive automatic error correction mechanism
abstract
This paper describes a novel automatic error correction method for scanning communication, whose mechanism is basically analogous to that of continuous speech recognition. It has two core components: one is the switch timing model, and the other is the statistical language model. By employing these models, the proposed system can estimate most probable sequence of input syllables for a given sequence of switch timing, with taking user characteristics into account. Thirteen subjects without disabilities and an ALS subject participated in a text input experiment using the proposed scanning communication system. For the ALS subject, the system improved the character correct rate from 77.7% to 97.7%, allowing dramatically fast input.
Hiroki Mori
ASSETS1
2013 Robust estimation of multiple-regression HMM parameters for dimension-based expressive dialogue speech synthesis
Tomohiro Nagata, Hiroki Mori, Takashi Nose
INTERSPEECH2
2012 Throwing Skill Optimization through Synchronization and Desynchronization of Degree of Freedom
Yuji Kawai, Takato Horii, Yuji Oshima, Kazuaki Tanaka, Hiroki Mori, Yukie Nagai, Takashi Takuma, Minoru Asada
RoboCup6
2011 Generating avatar's facial expressions from emotional states in daily conversation
abstract
A framework for generating facial expressions from emotional states in daily conversation is described. The frame work allows avatars to express the speaker's state not just prototypical emotions. In this paper, the naturalness of generated facial expressions that are presented together with dialogue speech is examined. An experiment to examine the naturalness of facial expressions presented as still images shows that the two avatars' facial expressions are almost as natural as manually-made facial expressions. In an experiment to determine the natural display speed of dynamic facial expressions, significant interactions between display speed and emotion group were found for most emotion dimensions.
Hiroki Mori, Ko Oshima, Makoto Nakamura
ICASSP1
2011 Pilot Signals for Multiuser Tomlinson-Harashima Precoding in MIMO-OFDM Systems
abstract
This paper describes pilot signals in multiuser multiple-input multiple-output orthogonal frequency division multiplexing systems with nonlinear Tomlinson-Harashima precoding (THP). The transmission processing of THP consists of three parts : feedback processing, modulo operation, and feedforward processing. In particular, the increased transmit power in the feedback processing is reduced by the modulo operation. In the conventional linear precoding systems, the same precoding weight is applied for both pilot and data signals at the transmitter for channel estimation at the receiver. If it is directly applied to THP, however, this will lead to significant increase in transmit power since the modulo operation cannot be applied for pilot signals. In our proposal, only the feedforward processing is applied for pilot signals and the interference due to the lack of feedback processing is removed by using orthogonal pilot signals. Simulation results show that the proposed pilot signals significantly improves the packet error rate performance.
Tsuguhide Aoki, Hiroki Mori, Yuji Tohzaka, Yasuhiko Tanabe
VTC Fall2
2011 Constructing a spoken dialogue corpus for studying paralinguistic information in expressive conversation and analyzing its statistical/acoustic characteristics
Hiroki Mori, Tomoyuki Satake, Makoto Nakamura, Hideki Kasuya
Speech Commun.1
2011 A Wide-View Parallax-Free Eye-Mark Recorder with a Hyperboloidal Half-Silvered Mirror and Appearance-Based Gaze Estimation
abstract
In this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror and a gaze estimation method suitable for the device. Our eye-mark recorder provides a wide field-of-view video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical angle of view of the prototype is 122 degree (elevation and depression angles are 38 and 84 degree, respectively) and its horizontal view angle is 116 degree (nasal and temporal view angles are 38 and 78 degree, respectively). We implemented and evaluated a gaze estimation method for our eye-mark recorder. We use an appearance-based approach for our eye-mark recorder to support a wide field-of-view. We apply principal component analysis (PCA) and multiple regression analysis (MRA) to determine the relationship between the captured images and their corresponding gaze points. Experimental results verify that our eye-mark recorder successfully captures a wide field-of-view of a user and estimates gaze direction with an angular accuracy of around 2 to 4 degree.
Hiroki Mori, Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.1
2010 Longitudinal changes of selected voice source parameters
Hideki Kasuya, Hajime Yoshida, Satoshi Ebihara, Hiroki Mori
INTERSPEECH4
2009 An efficient uplink multiuser MIMO protocol in IEEE 802.11 WLANs
abstract
This paper proposes an uplink Multi-User (MU) Multiple-Input Multiple-Output (MIMO) protocol in IEEE 802.11 WLANs. In order to realize the uplink MU-MIMO transmission in 802.11 WLANs, there are several problems to be solved. Synchronized transmission among the stations (STAs) is one of the problems and the spatial compatibility between the transmitting STAs is another problem. Therefore, it is considered that many overheads are needed to realize the uplink MU-MIMO transmission. In the proposed protocol, after an access point (AP) informs the start of the uplink access phase by transmitting the indication frame, the STAs having data to be sent to the AP transmit the uplink access requests in the OFDMA manner. The AP recognizes the uplink access requests by detecting the subcarrier signals and it requests the detected STAs to transmit the pilot signals in the TDMA manner. The AP calculates the channel state information (CSI) between each STA from the received pilot signal and selects the STAs to be permitted to transmit the uplink frames based on the CSIs. The AP notifies the information about the permitted STAs and about the capable transmission rate, and then the permitted STAs simultaneously transmit the frames at the notified transmission rate. In the protocol, an efficient OFDMA-based uplink access request transmission scheme is also proposed. Thus, by adopting the proposed protocol, the overheads can be reduced and the network throughput can be enhanced. Computer simulations are performed and the results show the effectiveness of the proposed method.
Tomoya Tandai, Hiroki Mori, Kiyoshi Toshimitsu, Takahiro Kobayashi 0001
PIMRC2
2009 Cross-Layer-Optimized User Grouping Strategy in Downlink Multiuser MIMO Systems
abstract
This paper proposes a cross-layer-optimized user selection and packet transmission technique for the downlink multi-user (MU) multiple-input multiple-output (MIMO) system in IEEE 802.11 wireless LAN (WLAN). In IEEE 802.11 WLAN, packet duration is decided by a data size and a selected transmission mode (modulation and coding scheme, MCS), and therefore, in general, the packet duration of each user varies according to the channel conditions and the applications. If multiple users whose packet durations are considerably different are selected as the target of the downlink MU-MIMO transmission, then the efficiency of spatial multiplexing degrades and the network throughput cannot be enhanced. When selecting users as the target of the downlink MU-MIMO transmission, the degree of spatial channel correlation among users (from the viewpoint of the physical layer) and the packet waiting times in a queue (from the viewpoint of the MAC layer) have to be considered. In addition to these criteria, in IEEE 802.11 WLAN, uniformity of the packet durations among users has to be considered as a criterion. In the proposed method, data size is adjusted in order to make the packet durations of the users uniform after selecting the users based on the spatial compatibility and the packet waiting times. This means that the proposed downlink MU- MIMO transmission is optimized in 3 dimensions by the cross- layer approach: the spatial channel correlations, packet waiting times and the uniformity of the packet durations. Computer simulations are performed and the results show the effectiveness of the proposed method.
Tomoya Tandai, Hiroki Mori, Masahiro Takagi
VTC Spring2
2008 Paralinguistic effects on turn-taking behavior in expressive conversation
Hiroki Mori, Hideki Kasuya
INTERSPEECH1
2007 Voice source and vocal tract variations as cues to emotional states perceived from expressive conversational speech
Hiroki Mori, Hideki Kasuya
INTERSPEECH1
2006 Two-layered neighborhood tabu search for multi-objective distribution network expansion planning
abstract
In this paper, a new method is proposed for distribution network expansion planning. The proposed method is based on the multi-objective meta-heuristics with the epsiv-constraint method. As the power network becomes more deregulated and competitive, distribution network planning is faced with the uncertainty due to several factors. Thus, it is important to consider the uncertainty of the network planning. This paper proposes a new method that makes use of the Monte-Carlo simulation to consider the uncertainty of load patterns in the future. For many scenarios, the multi-objective two-layered tabu search is employed to determine the optimal allocation of feeders and substations. Two-layered tabu search separates integer variables by the difference of the neighborhoods. The proposed method is successfully applied to a sample system
Hiroki Mori, Yasunori Yamada
ISCAS1
2005 Mora timing organization in producing contrastive geminate/single consonants and long/short vowels by native and non-native speakers of Japanese: effects of speaking rate
Haiping Jia, Hiroki Mori, Hideki Kasuya
INTERSPEECH2
2004 F0 and formant frequency distribution of dysarthric speech - a comparative study
Hiroki Mori, Yasunori Kobayashi, Hideki Kasuya, Hajime Hirose, Noriko Kobayashi
INTERSPEECH1
2003 Acoustic variations of focused disyllabic words in Mandarin Chinese: analysis, synthesis and perception
Zhenglai Gu, Hiroki Mori, Hideki Kasuya
INTERSPEECH2
2003 Speaker conversion in ARX-based source-formant type speech synthesis
Hiroki Mori, Hideki Kasuya
INTERSPEECH1
2002 A data-driven approach to source-formant type text-to-speech system
Hiroki Mori, Takahiro Ohtsuka, Hideki Kasuya
INTERSPEECH1
2001 Invariance of relative F0 change field of Chinese disyllabic words
abstract
In automatic voice response systems where a large number of words are inserted into fixed sentences, such as in voiceguided car navigation systems, one of the most important problems is the adjustment of the fundamental frequency (F0) contour of the inserted word to suit the F0 context of the fixed sentence. The effects of intonation and tone on the F0 contours of Chinese words can be described in terms of a word-level F0 range (WF0R) and an F0 change field (F0CF). WF0R in any position of a sentence is a tone-independent general F0 range, whereas F0CF is an F0 range taking the tone combination of words into account. Relative F0CF is regulated in reference to WF0R. If WF0R is used to represent the declination of a sentence, the relative F0CF should be invariant but dependent on the tone combination of a word. This paper examines the invariance of the relative F0CF among individuals. From an analysis of four native speakers’ utterances of 160 words in the initial, middle and final parts of three carrier sentences, conducted over 2 days, we show that: (1) Chinese speakers read words in the same sentence position with stable relative F0 change; (2) the relative F0CFs in the middle position of a sentence are generally the same as those in the initial position, but slightly different from those in the final position; and (3) the relative F0CFs reveal that the effects of tone on F0 contour is individual independent.
Hiroki Mori, Hideki Kasuya
INTERSPEECH2
2000 Prosodic variation of focused syllables of disyllabic word in Mandarin Chinese
Zhenglai Gu, Hiroki Mori, Hideki Kasuya
INTERSPEECH2
2000 Automatic lexicon generation and dialogue modeling for spontaneous speech
Hiroki Mori, Hideki Kasuya
INTERSPEECH1
2000 Word-level F0 range in Mandarin Chinese and its application to inserting words into a sentence
Hiroki Mori, Hideki Kasuya
INTERSPEECH2
1999 An Automatic Acquisition of Acoustical Units for Speech Recognition Based on Hidden Markov Network
Motoyuki Suzuki, Takafumi Hayashi, Hiroki Mori, Shozo Makino, Hirotomo Aso
Discovery Science3
1998 Multiple Camera Based Human Motion Estimation
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida
ACCV (2)2
1998 Automatic Acquisition of Phoneme Models and Its Application to Phoneme Labeling of a Large Size of Speech Corpus
Motoyuki Suzuki, Teruhiko Maeda, Hiroki Mori, Shozo Makino
Discovery Science3
1998 Multiple-Human Tracking Using Multiple Cameras
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida
FG2
1998 Multiple-view-based tracking of multiple humans
abstract
We propose a multiple-view-based tracking algorithm for multiple-human motions. In vision-based human tracking, self-occlusions and human-human occlusions are a part of the more significant problems. Employing multiple viewpoints and a viewpoint selection mechanism, however can reduce these problems. In our system, human positions are tracked with a sequence of multiple-viewpoint images. This tracking is based on the Kalman filtering approach. The estimation results are utilized to select proper viewpoints in other sub-tasks (rotation angle detection and body-side detection). Each sub-task has a different criterion for selecting viewpoints. We also describe the criterions for accomplishing individual sub-tasks and relationships between sub-tasks. We have already built an experimental system based on a small number of reliable image features. We confirm the stability of our algorithm through simulations. We also performed fundamental examinations on the experimental system.
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida
ICPR2
1998 High-speed speaker adaptation using phoneme dependent tree-structured speaker clustering
abstract
The tree-structured speaker clustering was proposed as a highspeed speaker adaptation method. It can select the model which is most similar to a target speaker. However, this method does not consider speaker difference dependent on phoneme class. In this paper, we propose a speaker adaptation method based on speaker clustering by taking speaker difference dependent on phoneme class into account. The experimental results showed that the new method gave a better performance than the original method. Furthermore, we propose the improved method which use a tree-structure of a similar phoneme as the substitute for the phoneme which does not appear in the adaptation data. From the experimental results, the improved method gave a better performance than the method previously proposed.
Motoyuki Suzuki, Toshiaki Abe, Hiroki Mori, Shozo Makino, Hirotomo Aso
ICSLP3
1995 Japanese document recognition based on interpolated n-gram model of character
abstract
N-gram model is widely applied to various pattern recognition system because it well represents local features of natural languages. In this paper, we describe a contextual postprocessing method using a trigram model of character for Japanese document recognition, and its advantage is revealed by practical experiments. The model is automatically obtained by statistical processing of training documents. The ability to reduce ambiguity is evaluated by the perplexity. In the processing, two smoothing methods are examined, and the predictive power of the deleted interpolation method is shown to be superior. For leading articles, the perplexity reduced to about 22 when using deleted interpolation. The output from OCR is processed very fast using a Viterbi algorithm. Experimental results of recognition for three kinds of documents show that the error correction rates are ranged from 75 to over 90 percent.
Hiroki Mori, Hirotomo Aso, Shozo Makino
ICDAR1