VLDB 2026 Research / reviewers in the wild / expert
Norimichi Ukita
dblp:46/5881
· DBLP profile ↗
59ranked-venue papers
18as first author
23since 2021 · last 2026
0000-0002-0240-1065ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 15 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMCM: Multimodality-aware Metric using Clustering-based Modes for Probabilistic Human Motion PredictionabstractThis paper proposes a novel metric for Human Motion Prediction (HMP). Since a single past sequence can lead to multiple possible futures, a probabilistic HMP method predicts such multiple motions. While a single motion predicted by a deterministic method is evaluated only with the difference from its ground truth motion, multiple predicted motions should also be evaluated based on their distribution. For this evaluation, this paper focuses on the following two criteria. (a) Coverage: motions should be distributed among multiple motion modes to cover diverse possibilities. (b) Validity: motions should be kinematically valid as future motions observable from a given past motion. However, existing metrics simply appreciate widely distributed motions even if these motions are observed in a single mode and kinematically invalid. To resolve these disadvantages, this paper proposes a Multimodality-aware Metric using Clustering-based Modes (MMCM). For (a) coverage, MMCM divides a motion space into several clusters, each of which is regarded as a mode. These modes are used to explicitly evaluate whether predicted motions are distributed among multiple modes. For (b) validity, MMCM identifies valid modes by collecting possible future motions from a motion dataset. Our experiments validate that our clustering yields sensible mode definitions and that MMCM accurately scores multimodal predictions. Code: https://github.com/placerkyo/MMCM Kyotaro Tokoro, Hiromu Taketsugu, Norimichi Ukita |
WACV | 3 |
| 2026 | Human-in-the-loop adaptation in group activity feature learning for team sports video retrieval
Chihiro Nakatani, Hiroaki Kawashima, Norimichi Ukita |
Comput. Vis. Image Underst. | 3 |
| 2025 | Physical Plausibility-aware Trajectory Prediction via Locomotion EmbodimentabstractHumans can predict future human trajectories even from momentary observations by using human pose-related cues. However, previous Human Trajectory Prediction (HTP) methods leverage the pose cues implicitly, resulting in implausible predictions. To address this, we propose Locomotion Embodiment, a framework that explicitly evaluates the physical plausibility of the predicted trajectory by locomotion generation under the laws of physics. While the plausibility of locomotion is learned with an indifferentiable physics simulator, it is replaced by our differentiable Locomotion Value function to train an HTP network in a data-driven manner. In particular, our proposed Embodied Locomotion loss is beneficial for efficiently training a stochastic HTP network using multiple heads. Furthermore, the Locomotion Value filter is proposed to filter out implausible trajectories at inference. Experiments demonstrate that our method enhances even the state-of-the-art HTP methods across diverse datasets and problem settings. Our code is available at: https://github.com/ImIntheMiddle/EmLoco. Hiromu Taketsugu, Takeru Oba, Takahiro Maeda 0001, Shohei Nobuhara, Norimichi Ukita |
CVPR | 5 |
| 2025 | Dynamic Group Detection using VLM-augmented Temporal Groupness GraphabstractThis paper proposes dynamic human group detection in videos. For detecting complex groups, not only the local appearance features of in-group members but also the global context of the scene are important. Such local and global appearance features in each frame are extracted using a Vision-Language Model (VLM) augmented for group detection in our method. For further improvement, the group structure should be consistent over time. While previous methods are stabilized on the assumption that groups are not changed in a video, our method detects dynamically changing groups by global optimization using a graph with all frames' groupness probabilities estimated by our groupness-augmented CLIP features. Our experimental results demonstrate that our method outperforms state-of-the-art group detection methods on public datasets. Code: https://github.com/irajisamurai/VLM-GroupDetection.git Kaname Yokoyama, Chihiro Nakatani, Norimichi Ukita |
ICCV | 3 |
| 2025 | R2-Diff: Denoising by diffusion as a refinement of retrieved motion for image-based motion prediction
Takeru Oba, Norimichi Ukita |
Neurocomputing | 2 |
| 2024 | Learning Group Activity Features Through Person Attribute PredictionabstractThis paper proposes Group Activity Feature (GAF) learning in which features of multi-person activity are learned as a compact latent vector. Unlike prior work in which the manual annotation of group activities is required for supervised learning, our method learns the GAF through person attribute prediction without group activity annotations. By learning the whole network in an end-to-end manner so that the GAF is required for predicting the person attributes of people in a group, the GAF is trained as the features of multi-person activity. As a person attribute, we propose to use a person's action class and appearance features because the former is easy to annotate due to its simpleness, and the latter requires no manual an-notation. In addition, we introduce a location-guided attribute prediction to disentangle the complex GAF for extracting the features of each target person properly. Various experimental results validate that our method outperforms SOTA methods quantitatively and qualitatively on two public datasets. Visualization of our GAF also demonstrates that our method learns the GAF representing fined-grained group activity classes. Code: https://github.com/chihina/GAFL-CVPR2024. Chihiro Nakatani, Hiroaki Kawashima, Norimichi Ukita |
CVPR | 3 |
| 2024 | READ: Retrieval-Enhanced Asymmetric Diffusion for Motion PlanningabstractThis paper proposes Retrieval-Enhanced Asymmetric Diffusion (READ) for image-based robot motion planning. Given an image of the scene, READ retrieves an initial motion from a database of image-motion pairs, and uses a diffusion model to refine the motion for the given scene. Unlike prior retrieval-based diffusion models that require long forward-reverse diffusion paths, READ directly diffuses between the source (retrieved) and target motions, resulting in an efficient diffusion path. A second contribution of READ is its use of asymmetric diffusion, whereby it preserves the kinematic feasibility of the generated motion by forward diffusion in a low-dimensional latent space, while achieving high-resolution motion by reverse diffusion in the original task space using cold diffusion. Experimental results on various manipulation tasks demonstrate that READ outperforms state-of-the-art planning methods, while ablation studies elucidate the contributions of asymmetric diffusion. Code: https://github.com/Obat2343/READ Takeru Oba, Matthew R. Walter, Norimichi Ukita |
CVPR | 3 |
| 2024 | Depth Estimation fusing Image and Radar Measurements with Uncertain DirectionsabstractThis paper proposes a depth estimation method using radar-image fusion by addressing the uncertain vertical directions of sparse radar measurements. In prior radar-image fusion work, image features are merged with the uncertain sparse depths measured by radar through convolutional layers. This approach is disturbed by the features computed with the uncertain radar depths. Furthermore, since the features are computed with a fully convolutional network, the uncertainty of each depth corresponding to a pixel is spread out over its surrounding pixels. Our method avoids this problem by computing features only with an image and conditioning the features pixelwise with the radar depth. Furthermore, the set of possibly correct radar directions is identified with reliable LiDAR measurements, which are available only in the training stage. Our method improves training data by learning only these possibly correct radar directions, while the previous method trains raw radar measurements, including erroneous measurements. Experimental results demonstrate that our method can improve the quantitative and qualitative results compared with its base method using radar-image fusion. Masaya Kotani, Takeru Oba, Norimichi Ukita |
IJCNN | 3 |
| 2024 | Time-series Initialization and Conditioning for Video-agnostic Stabilization of Video Super-Resolution using Recurrent NetworksabstractA Recurrent Neural Network (RNN) for Video Super Resolution (VSR) is generally trained with randomly clipped and cropped short videos extracted from original training videos due to various challenges in learning RNNs. However, since this RNN is optimized to super-resolve short videos, VSR of long videos is degraded due to the domain gap. Our preliminary experiments reveal that such degradation changes depending on the video properties, such as the video length and dynamics. To avoid this degradation, this paper proposes the training strategy of RNN for VSR that can work efficiently and stably independently of the video length and dynamics. The proposed training strategy stabilizes VSR by training a VSR network with various RNN hidden states changed depending on the video properties. Since computing such a variety of hidden states is time-consuming, this computational cost is reduced by reusing the hidden states for efficient training. In addition, training stability is further improved with frame-number conditioning. Our experimental results demonstrate that the proposed method performed better than base methods in videos with various lengths and dynamics. Hiroshi Mori, Norimichi Ukita |
IJCNN | 2 |
| 2024 | Inpainting-Driven Mask Optimization for Object RemovalabstractThis paper proposes a mask optimization method for improving the quality of object removal using image inpainting. While many inpainting methods are trained with a set of random masks, a target for inpainting may be an object, such as a person, in many realistic scenarios. This domain gap between masks in training and inference images increases the difficulty of the inpainting task. In our method, this domain gap is resolved by training the inpainting network with object masks extracted by segmentation, and such object masks are also used in the inference step. Furthermore, to optimize the object masks for inpainting, the segmentation network is connected to the inpainting network and end-to-end trained to improve the inpainting performance. The effect of this end-to-end training is further enhanced by our mask expansion loss for achieving the trade-off between large and small masks. Experimental results demonstrate the effectiveness of our method for better object removal using image inpainting. Kodai Shimosato, Norimichi Ukita |
IJCNN | 2 |
| 2024 | Burst Super-Resolution with Diffusion Models for Improving Perceptual QualityabstractWhile burst LR images are useful for improving the SR image quality compared with a single LR image, prior SR networks accepting the burst LR images are trained in a deterministic manner, which is known to produce a blurry SR image. In addition, it is difficult to perfectly align the burst LR images, making the SR image more blurry. Since such blurry images are perceptually degraded, we aim to reconstruct the sharp high-fidelity boundaries. Such high-fidelity images can be reconstructed by diffusion models. However, prior SR methods using the diffusion model are not properly optimized for the burst SR task. Specifically, the reverse process starting from a random sample is not optimized for image enhancement and restoration methods, including burst SR. In our proposed method, on the other hand, burst LR features are used to reconstruct the initial burst SR image that is fed into an intermediate step in the diffusion model. This reverse process from the intermediate step 1) skips diffusion steps for reconstructing the global structure of the image and 2) focuses on steps for refining detailed textures. Our experimental results demonstrate that our method can improve the scores of the perceptual quality metrics. Code: https://github.com/placerkyo/BSRD. Kyotaro Tokoro, Kazutoshi Akita, Norimichi Ukita |
IJCNN | 3 |
| 2024 | Active Transfer Learning for Efficient Video-Specific Human Pose EstimationabstractHuman Pose (HP) estimation is actively researched because of its wide range of applications. However, even estimators pre-trained on large datasets may not perform satisfactorily due to a domain gap between the training and test data. To address this issue, we present our approach combining Active Learning (AL) and Transfer Learning (TL) to adapt HP estimators to individual video domains efficiently. For efficient learning, our approach quantifies (i) the estimation uncertainty based on the temporal changes in the estimated heatmaps and (ii) the unnaturalness in the estimated full-body HPs. These quantified criteria are then effectively combined with the state-of-the-art representativeness criterion to select uncertain and diverse samples for efficient HP estimator learning. Furthermore, we reconsider the existing Active Transfer Learning (ATL) method to introduce novel ideas related to the retraining methods and Stopping Criteria (SC). Experimental results demonstrate that our method enhances learning efficiency and outperforms comparative methods. Our code is publicly available at: https://github.com/ImIntheMiddle/VATL4Pose-WACV2024 Hiromu Taketsugu, Norimichi Ukita |
WACV | 2 |
| 2023 | Efficient Reinforcement Learning Using State-Action Uncertainty with Multiple Heads
Tomoharu Aizu, Takeru Oba, Norimichi Ukita |
ICANN (8) | 3 |
| 2023 | Fast Inference and Update of Probabilistic Density Estimation on Trajectory PredictionabstractSafety-critical applications such as autonomous vehicles and social robots require fast computation and accurate probability density estimation on trajectory prediction. To address both requirements, this paper presents a new normalizing flow-based trajectory prediction model named FlowChain. FlowChain is a stack of conditional continuously-indexed flows (CIFs) that are expressive and allow analytical probability density computation. This analytical computation is faster than the generative models that need additional approximations such as kernel density estimation. Moreover, FlowChain is more accurate than the Gaussian mixture-based models due to fewer assumptions on the estimated density. FlowChain also allows a rapid update of estimated probability densities. This update is achieved by adopting the newest observed position and reusing the flow transformations and its log-det-jacobians that represent the motion trend. This update is completed in less than one millisecond because this reuse greatly omits the computational cost. Experimental results showed our FlowChain achieved state-of-the-art trajectory prediction accuracy compared to previous methods. Furthermore, our FlowChain demonstrated superiority in the accuracy and speed of density estimation. Our code is available at https://github.com/meaten/FlowChain-ICCV2023. Takahiro Maeda 0001, Norimichi Ukita |
ICCV | 2 |
| 2023 | Interaction-aware Joint Attention Estimation Using People AttributesabstractThis paper proposes joint attention estimation in a single image. Different from related work in which only the gaze-related attributes of people are independently employed, (i) their locations and actions are also employed as contextual cues for weighting their attributes, and (ii) interactions among all of these attributes are explicitly modeled in our method. For the interaction modeling, we propose a novel Transformer-based attention network to encode joint attention as low-dimensional features. We introduce a specialized MLP head with positional embedding to the Transformer so that it predicts pixelwise confidence of joint attention for generating the confidence heatmap. This pixelwise prediction improves the heatmap accuracy by avoiding the ill-posed problem in which the high-dimensional heatmap is predicted from the low-dimensional features. The estimated joint attention is further improved by being integrated with general image-based attention estimation. Our method outperforms SOTA methods quantitatively in comparative experiments. Code: https://github.com/chihina/PJAE. Chihiro Nakatani, Hiroaki Kawashima, Norimichi Ukita |
ICCV | 3 |
| 2023 | Data-Driven Stochastic Motion Evaluation and Optimization with Image by Spatially-Aligned Temporal EncodingabstractThis paper proposes a probabilistic motion prediction method for long motions. The motion is predicted so that it accomplishes a task from the initial state observed in the given image. While our method evaluates the task achievability by the Energy-Based Model (EBM), previous EBMs are not designed for evaluating the consistency between different domains (i.e., image and motion in our method). Our method seamlessly integrates the image and motion data into the image feature domain by spatially-aligned temporal encoding so that features are extracted along the motion trajectory projected onto the image. Furthermore, this paper also proposes a data-driven motion optimization method, Deep Motion Optimizer (DMO), that works with EBM for motion prediction. Different from previous gradient-based optimizers, our self-supervised DMO alleviates the difficulty of hyper-parameter tuning to avoid local minima. The effectiveness of the proposed method is demonstrated with a variety of experiments with similar SOTA methods. Takeru Oba, Norimichi Ukita |
ICRA | 2 |
| 2023 | Future-guided offline imitation learning for long action sequences via video interpolation and future-trajectory prediction
Takeru Oba, Norimichi Ukita |
Neurocomputing | 2 |
| 2022 | MotionAug: Augmentation with Physical Correction for Human Motion PredictionabstractThis paper presents a motion data augmentation scheme incorporating motion synthesis encouraging diversity and motion correction imposing physical plausibility. This motion synthesis consists of our modified Variational AutoEncoder (VAE) and Inverse Kinematics (IK). In this VAE, our proposed sampling-near-samples method generates various valid motions even with insufficient training motion data. Our IK-based motion synthesis method allows us to generate a variety of motions semi-automatically. Since these two schemes generate unrealistic artifacts in the synthesized motions, our motion correction rectifies them. This motion correction scheme consists of imitation learning with physics simulation and subsequent motion debiasing. For this imitation learning, we propose the PD-residual force that significantly accelerates the training process. Furthermore, our motion debiasing successfully offsets the motion bias induced by imitation learning to maximize the effect of augmentation. As a result, our method outperforms previous noise-based motion augmentation methods by a large margin on both Recurrent Neural Network-based and Graph Convolutional Network-based human motion prediction models. The code is available at https://github.com/meaten/MotionAug. Takahiro Maeda 0001, Norimichi Ukita |
CVPR | 2 |
| 2022 | Erratum to "Deep Back-Projection Networks for Single Image Super-Resolution"abstractIn the above article [1], the article title was incorrect. The correct article title is "Deep Back-Projection Networks for Single Image Super-Resolution." Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Continuous Finger Gesture Spotting and Recognition Based on Similarities Between Start and End FramesabstractTouchless in-car devices controlled by single and continuous finger gestures can provide comfort and safety on driving while manipulating secondary devices. Recognition of finger gestures is a challenging task due to (i) similarities between gesture and non-gesture frames, and (ii) the difficulty in identifying the temporal boundaries of continuous gestures. In addition, (iii) the intraclass variability of gestures’ duration is a critical issue for recognizing finger gestures intended to control in-car devices. To address difficulties (i) and (ii), we propose a gesture spotting method where continuous gestures are segmented by detecting boundary frames and evaluating hand similarities between thestartandendboundaries of each gesture. Subsequently, we introduce a gesture recognition based on a temporal normalization of features extracted from the set of spotted frames, which overcomes difficulty (iii). This normalization enables the representation of any gesture with the same limited number of features. We ensure real-time performance by proposing an approach based on compact deep neural networks. Moreover, we demonstrate the effectiveness of our proposal with a second approach based on hand-crafted features performing in real-time, even without GPU requirements. Furthermore, we present a realistic driving setup to capture a dataset of continuous finger gestures, which includes more than 2,800 instances on untrimmed videos covering safety driving requirements. With this dataset, our both approaches can run at 53 fps and 28 fps on GPU and CPU, respectively, around 13 fps faster than previous works, while achieving better performance (at least 5% higher mean tIoU). Gibran Benitez-Garcia, Muhammad Haris 0002, Yoshiyuki Tsuda, Norimichi Ukita |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Embryo Grading With Unreliable Labels Due to Chromosome Abnormalities by Regularized PU Learning With RankingabstractWe propose a method for human embryo grading with its images. This grading has been achieved by positive-negative classification (i.e., live birth or non-live birth). However, negative (non-live birth) labels collected in clinical practice are unreliable because the visual features of negative images are equal to those of positive (live birth) images if these non-live birth embryos have chromosome abnormalities. For alleviating an adverse effect of these unreliable labels, our method employs Positive-Unlabeled (PU) learning so that live birth and non-live birth are labeled as positive and unlabeled, respectively, where unlabeled samples contain both positive and negative samples. In our method, this PU learning on a deep CNN is improved by a learning-to-rank scheme. While the original learning-to-rank scheme is designed for positive-negative learning, it is extended to PU learning. Furthermore, overfitting in this PU learning is alleviated by regularization with mutual information. Experimental results with 643 time-lapse image sequences demonstrate the effectiveness of our framework in terms of the recognition accuracy and the interpretability. In quantitative comparison, the full version of our proposed method outperforms positive-negative classification in recall and F-measure by a wide margin (0.22 vs. 0.69 in recall and 0.27 vs. 0.42 in F-measure). In qualitative evaluation, visual attentions estimated by our method are interpretable in comparison with morphological assessments in clinical practice. Masashi Nagaya, Norimichi Ukita |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Task-Driven Super Resolution: Object Detection in Low-Resolution Images
Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
ICONIP (5) | 3 |
| 2021 | Deep Back-ProjectiNetworks for Single Image Super-ResolutionabstractPrevious feed-forward architectures of recently proposed deep super-resolution networks learn the features of low-resolution inputs and the non-linear mapping from those to a high-resolution output. However, this approach does not fully address the mutual dependencies of low- and high-resolution images. We propose Deep Back-Projection Networks (DBPN), the winner of two image super-resolution challenges (NTIRE2018 and PIRM2018), that exploit iterative up- and down-sampling layers. These layers are formed as a unit providing an error feedback mechanism for projection errors. We construct mutually-connected up- and down-sampling units each of which represents different types of low- and high-resolution components. We also show that extending this idea to demonstrate a new insight towards more efficient network design substantially, such as parameter sharing on the projection module and transition layer on projection step. The experimental results yield superior results and in particular establishing new state-of-the-art results across multiple data sets, especially for large scaling factors such as 8×. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Space-Time-Aware Multi-Resolution Video EnhancementabstractWe consider the problem of space-time super-resolution (ST-SR): increasing spatial resolution of video frames and simultaneously interpolating frames to increase the frame rate. Modern approaches handle these axes one at a time. In contrast, our proposed model called STARnet super-resolves jointly in space and time. This allows us to leverage mutually informative relationships between time and space: higher resolution can provide more detailed information about motion, and higher frame-rate can provide better pixel alignment. The components of our model that generate latent low- and high-resolution representations during ST-SR can be used to finetune a specialized mechanism for just spatial or just temporal super-resolution. Experimental results demonstrate that STARnet improves the performances of space-time, spatial, and temporal video super-resolution by substantial margins on publicly available datasets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 3 |
| 2020 | Region-dependent Scale Proposals for Super-Resolution in Object DetectionabstractThis paper presents a method for estimating object-scale proposals applied to super resolution (SR) for scale-optimized object detection. With the region-dependent scale proposals, we achieve scale-independent object detection. This object detection scheme consists of three functions; region-dependent scale proposals, SR, and object detection. While SR and object detection have been fused in deep end-to-end networks in previous works, region-dependent scale proposals are not provided or are performed independently of SR and object detection processes. The proposed region-dependent scale-proposal network is designed to explicitly estimate appropriate SR scales depending on the image region in accordance with scene contexts. Qualitative and quantitative experimental results show that our method can provide appropriate SR scales for improving detection accuracy. Our proposed method gains 2.7 points in AP with Centernet used as the base detector. Kazutoshi Akita, Muhammad Haris 0002, Norimichi Ukita |
IPAS | 3 |
| 2019 | Recurrent Back-Projection Network for Video Super-ResolutionabstractWe proposed a novel architecture for the problem of video super-resolution. We integrate spatial and temporal contexts from continuous video frames using a recurrent encoder-decoder module, that fuses multi-frame information with the more traditional, single frame super-resolution path for the target frame. In contrast to most prior work where frames are pooled together by stacking or warping, our model, the Recurrent Back-Projection Network (RBPN) treats each context frame as a separate source of information. These sources are combined in an iterative refinement framework inspired by the idea of back-projection in multiple-image super-resolution. This is aided by explicitly representing estimated inter-frame motion with respect to the target, rather than explicitly aligning frames. We propose a new video super-resolution benchmark, allowing evaluation at a larger scale and considering videos in different motion regimes. Experimental results demonstrate that our RBPN is superior to existing methods on several datasets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 3 |
| 2018 | Deep Back-Projection Networks for Super-ResolutionabstractThe feed-forward architectures of recently proposed deep super-resolution networks learn representations of low-resolution inputs, and the non-linear mapping from those to high-resolution output. However, this approach does not fully address the mutual dependencies of low- and high-resolution images. We propose Deep Back-Projection Networks (DBPN), that exploit iterative up- and downsampling layers, providing an error feedback mechanism for projection errors at each stage. We construct mutually-connected up- and down-sampling stages each of which represents different types of image degradation and high-resolution components. We show that extending this idea to allow concatenation of features across up- and downsampling stages (Dense DBPN) allows us to reconstruct further improve super-resolution, yielding superior results and in particular establishing new state of the art results for large scaling factors such as 8× across multiple data sets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 3 |
| 2018 | Ensemble convolutional neural networks for pose estimation
Yuki Kawana, Norimichi Ukita, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2018 | Semi- and weakly-supervised human pose estimation
Norimichi Ukita, Yusuke Uematsu |
Comput. Vis. Image Underst. | 1 |
| 2017 | Discomfort-ride map for personal mobility passengers on sidewalks areaabstractPersonal mobility devices such as wheelchairs, bicycles, and compact cars are used in daily life and to runs on sidewalks. However, there are several factors that may lead to discomfort rides, such as steps, slopes, and crowded sidewalks for a passenger. This paper proposes a system that detects the factors using smartphone attached to the personal mobility device and generates a discomfort-ride map including the environment and human factors and their rating on sidewalks. In the experiment, steps (static danger factor) and the dynamic moving obstacles like human, bicycle, and car (dynamic danger factors) were estimated with attached smartphone's acceleration sensor and camera. After the process of collecting these data, verification for generating the hazard map system based on these data. Moreover, the versatility of using the proposed hazard map system for different personal mobility devices was tested with the electric wheelchair and the bicycle. Taishi Sawabe, Nishikawa Naoki, Masayuki Kanbara, Norimichi Ukita, Norihiro Hagita |
SMC | 4 |
| 2017 | Tandem Equipment Arranged Architecture with Exhaust Heat Reuse System for Software-Defined Data Center InfrastructureabstractIn this paper, we propose a novel energy-efficient architecture for software-defined data center infrastructures. In our proposed data center architecture, we include an exhaust heat reuse system that utilizes high-temperature exhaust heat from servers in conditioning humidity and air temperature of office space near the data center. To obtain high-temperature exhaust heat, equipment such as server racks and air conditioners are deployed in tandem so that the aisles are divided into three types: cold, hot, and super-hot. In this paper, to investigate the fundamental characteristics of our proposed data center architecture, we consider various types of data center models and conduct numerical simulations that use results obtained by experiments at an actual data center. Through simulation, we show that the total power consumption by a data center with our proposed architecture is 27 percent lower than that by data center with a conventional architecture. In addition, it is also shown that the proposed tandem equipment arrangement is suitable for obtaining high-temperature exhaust heat and decreasing the total power consumption significantly under a wider range of conditions than in the conventional equipment arrangement. Yoshiaki Taniguchi, Koji Suganuma, Takaaki Deguchi, Go Hasegawa, Yutaka Nakamura, Norimichi Ukita, Naoki Aizawa, Katsuhiko Shibata, Kazuhiro Matsuda, Morito Matsuoka |
IEEE Trans. Cloud Comput. | 6 |
| 2016 | Automatic detection of very early stage of dementia through multimodal interaction with computer avatarsabstractThis paper proposes a new approach to detecting very early stage of dementia automatically. We develop a computer avatar with spoken dialog functionalities that produces natural spoken queries referring to Mini Mental State Examination, Wechsler Memory Scale-Revised and other related questions. Multimodal interactive data of spoken dialogues from 18 participants (9 dementias and 9 healthy controls) are recorded, and audiovisual features are extracted. We confirm that the support vector machines can classify into two groups with 0.94 detection performance as measured by areas under ROC curve. It is found that our system has possibilities to detect very early stage of dementia through spoken dialog with our computer avatars. Hiroki Tanaka, Hiroyoshi Adachi, Norimichi Ukita, Takashi Kudo, Satoshi Nakamura 0001 |
ICMI | 3 |
| 2016 | People re-identification across non-overlapping cameras using group features
Norimichi Ukita, Yusuke Moriguchi, Norihiro Hagita |
Comput. Vis. Image Underst. | 1 |
| 2016 | High-order framewise smoothness-constrained globally-optimal tracking
Norimichi Ukita, Asami Okada |
Comput. Vis. Image Underst. | 1 |
| 2015 | Lesioned-Part Identification by Classifying Entire-Body Gait Motions
Tsuyoshi Higashiguchi, Toma Shimoyama, Norimichi Ukita, Masayuki Kanbara, Norihiro Hagita |
PSIVT | 3 |
| 2015 | Editorial
Björn Stenger, Norimichi Ukita, Yoichi Sato 0001, Pascal Fua, David J. Fleet |
Comput. Vis. Image Underst. | 2 |
| 2014 | Physical activity estimation using accelerometer and facility information for elderly healthcareabstractThis paper proposes a novel framework to estimate the amount of physical activity at a place where people stayed, by utilizing facility information and user's acceleration data. The total amount of physical activities, energy expenditure of a physical activity is a good scale. Our framework provides a physical activity scale based on typical energy expenditure of the activity given by existing researches already. To estimates the energy expenditure, we use typical value of metabolic equivalents to task (MET) which is used as practical scale. Unlike the other studies for monitoring physical activity, we estimate the type of user's activity using facility information which are obtained from road map and land-use/land-cover map. To confirm the feasibility of our approach, we have conducted a long-term experiment on monitoring the activity of elderly people living in less-populated area. As a result, our framework provides good summarization of daily activities of participants. Masayuki Hayashi, Masayuki Kanbara, Norimichi Ukita, Norihiro Hagita |
SMC | 3 |
| 2014 | Representing mesh-based character animations
Edilson de Aguiar, Norimichi Ukita |
Comput. Graph. | 2 |
| 2013 | Simultaneous particle tracking in multi-action motion models with synthesized paths
Norimichi Ukita |
Image Vis. Comput. | 1 |
| 2012 | Articulated pose estimation with parts connectivity using discriminative local oriented contoursabstractThis paper proposes contour-based features for articulated pose estimation. Most of recent methods are designed using tree-structured models with appearance evaluation only within the region of each part. While these models allow us to speed up global optimization in localizing the whole parts, useful appearance cues between neighboring parts are missing. Our work focuses on how to evaluate parts connectivity using contour cues. Unlike previous works, we locally evaluate parts connectivity only along the orientation between neighboring parts within where they overlap. This adaptive localization of the features is required for suppressing bad effects due to nuisance edges such as those of background clutter and clothing textures, as well as for reducing computational cost. Discriminative training of the contour features improves estimation accuracy more. Experimental results verify the effectiveness of our contour-based features. Norimichi Ukita |
CVPR | 1 |
| 2012 | Shape reconstruction with globally-optimized surface point selection
Norimichi Ukita, Kazuki Matsuda, Norihiro Hagita |
ICPR | 1 |
| 2012 | Gaussian process motion graph models for smooth transitions among multiple actions
Norimichi Ukita, Takeo Kanade |
Comput. Vis. Image Underst. | 1 |
| 2012 | Reference consistent reconstruction of 3D cloth surface
Norimichi Ukita, Takeo Kanade |
Comput. Vis. Image Underst. | 1 |
| 2010 | Real-Time Pose Regression with Fast Volume Descriptor ComputationabstractWe present a real-time method for estimating the pose of a human body using its 3D volume obtained from synchronized videos. The method achieves pose estimation by pose regression from its 3D volume. While the 3D volume allows us to estimate the pose robustly against self occlusions, 3D volume analysis requires a large amount of computational cost. We propose fast and stable volume tracking with efficient volume representation in a low dimensional dynamical model. Experimental results demonstrated that pose estimation of a body with a significantly deformable clothing could run at around 60 fps. Michiro Hirai, Norimichi Ukita, Masatsugu Kidode |
ICPR | 2 |
| 2009 | Complex volume and pose tracking with probabilistic dynamical models and visual hull constraintsabstractWe propose a method for estimating the pose of a human body using its approximate 3D volume (visual hull) obtained in real time from synchronized videos. Our method can cope with loose-fitting clothing, which hides the human body and produces non-rigid motions and critical reconstruction errors, as well as tight-fitting clothing. To follow the shape variations robustly against erratic motions and the ambiguity between a reconstructed body shape and its pose, the probabilistic dynamical model of human volumes is learned from training temporal volumes refined by error correction. The dynamical model of a body pose (joint angles) is also learned with its corresponding volume. By comparing the volume model with an input visual hull and regressing its pose from the pose model, pose estimation can be realized. In our method, this is improved by double volume comparison: 1) comparison in a low-dimensional latent space with probabilistic volume models and 2) comparison in an observation volume space using geometric constrains between a real volume and a visual hull. Comparative experiments demonstrate the effectiveness of our method faster than existing methods. Norimichi Ukita, Michiro Hirai, Masatsugu Kidode |
ICCV | 1 |
| 2009 | Image Matching with a Car-Mounted Camera Robust to Changes in Imaging ConditionsabstractWe propose a matching method for images captured at different times and under different capturing conditions. Our method is designed for change detection in streetscapes using normal automobiles that have an off-the-shelf car mounted camera and a GPS. Therefore, we are able to analyze low-resolution and low frame-rate images captured asynchronously. To cope with this difficulty, previous and current panoramic images are created from sequential images which are rectified based on the view direction of a camera, and are then compared. In addition, in order to allow the matching method to be applicable to images captured under varying conditions, (1) for different lanes, enlarged/reduced panoramic images are compared with each other, and (2) robustness to noises and changes in illumination is improved by the edge features. To confirm the effectiveness of the proposed method, we conducted experiments matching real images captured under various capturing conditions. Naoko Enami, Norimichi Ukita, Masatsugu Kidode |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2008 | Real-Time Shape Analysis of a Human Body in Clothing Using Time-Series Part-Labeled Volumes
Norimichi Ukita, Ryosuke Tsuji, Masatsugu Kidode |
ECCV (3) | 1 |
| 2008 | IBR-based free-viewpoint imaging of a complex scene using few camerasabstractThis paper proposes a free-viewpoint imaging method that can be used in a complicated scene such as an office room by using sparsely located cameras. In our method, a free-viewpoint image is generated from multiple image patches obtained by dividing observed images. The quality of the generated image strongly depends on how to divide the observed images. In an incorrect patch in the generated image, the images projected from different cameras differ significantly. With this property, the incorrect patches can be detected. These patches are then re-divided. We demonstrated the effectiveness of our method by generating free-viewpoint images from the real images observed by the cameras in an office room. Norimichi Ukita, Shohei Kawata, Masatsugu Kidode |
ICPR | 1 |
| 2007 | Displaying a Moving Image By Multiple Steerable ProjectorsabstractThis paper proposes a method for precise overlapping of projected images from multiple steerable projectors. When they are controlled simultaneously, two problems are revealed: (1) even a slight positional error of the projected image, which does not matter in the case of a single projector, causes misalignments of multiple projected images that can be perceived clearly when using multiple projectors; and (2) as the projectors usually do not have architectures for their synchronization it is impossible to display a moving image that is by tiling or overlaying precisely the multiple projected images. To overcome (1), a method is proposed that measures preliminarily the misalignments through every plane in the environment, and hence displays the image without the misalignment. For (2), a consideration and a new proposal for the synchronization of multiple projectors are also discussed. Ikuhisa Mitsugami, Norimichi Ukita, Masatsugu Kidode |
CVPR | 2 |
| 2007 | Probabilistic-topological calibration of widely distributed camera networks
Norimichi Ukita |
Mach. Vis. Appl. | 1 |
| 2007 | Real-time cooperative multi-target tracking by dense communication among Active Vision Agents
Norimichi Ukita |
Web Intell. Agent Syst. | 1 |
| 2006 | Multiple Active Camera Assignment for High Fidelity 3D VideoabstractWe are designing a self controlling active camera system for a 3D video of a moving object (mainly human body). We made up our system of cameras with long focal length lenses for high resolution input images. However, such cameras can get only partial views of the object. We present, in this paper, a multiple active (pan-tilt) camera assignment scheme. The goal is to assign each camera to a specific part of the moving object so as to allow the best visibility of the whole object. For each camera, we evaluate the visibility to the different regions of the object, corresponding to different camera orientations and with respect to the field of view of the camera in question. Thereafter, we assign each camera to one orientation in such a way to maximize the visibility to the whole object. Sofiane Yous, Norimichi Ukita, Masatsugu Kidode |
ICVS | 2 |
| 2005 | Target-color learning and its detection for non-stationary scenes by nearest neighbor classification in the spatio-color spaceabstractWe propose a method for detecting foreground objects in non-stationary scenes. The method can (1) detect arbitrary foreground objects without any prior knowledge of them, (2) identify background pixels under various changes in a background scene, and (3) detect minor difference between the background and target colors. Online detection is realized by the nearest neighbor classifier in the 5D xy-YUV space (the spatio-color space), consisting of the x and y coordinates of an image and Y, U, and V colors, which holds rectified training data of background colors and automatically learned target colors. We conducted experiments to confirm the effectiveness of our method. Norimichi Ukita |
AVSS | 1 |
| 2005 | Region extraction of a gaze object using the gaze point and view image sequencesabstractAnalysis of the human gaze is a basic way to investigate human attention. Similarly, the view image of a human being includes the visual information of what he/she pays attention to.This paper proposes an interface system for extracting the region of an object viewed by a human from a view image sequence by analyzing the history of gaze points. All the gaze points, each of which is recorded as a 2D point in a view image, are transfered to an image in which the object region is extracted. These points are then divided into several groups based on their colors and positions. The gaze points in each group compose an initial region. After all the regions are extended, outlier regions are removed by comparing the colors and optical flows in the extended regions. All the remaining regions are merged into one in order to compose a gaze region. Norimichi Ukita, Tomohisa Ono, Masatsugu Kidode |
ICMI | 1 |
| 2005 | Robot Navigation by Eye Pointing
Ikuhisa Mitsugami, Norimichi Ukita, Masatsugu Kidode |
ICEC | 2 |
| 2005 | Real-time cooperative multi-target tracking by communicating active vision agents
Norimichi Ukita, Takashi Matsuyama |
Comput. Vis. Image Underst. | 1 |
| 2004 | Wearable virtual tablet: fingertip drawing on a portable plane-object using an active-infrared cameraabstractWe propose the Wearable Virtual Tablet (WVT), where a user can draw a locus on a common object with a plane surface (e.g., a notebook and a magazine) with a fingertip. Our previous WVT[1], however, could not work on a plane surface with complicated texture patterns: Since our WVT employs an active-infrared camera and the reflected infrared rays vary depending on patterns on a plane surface, it is difficult to estimate the motions of a fingertip and a plane surface from an observed infrared-image. In this paper, we propose a method to detect and track their motions without interference from colored patterns on a plane surface. (1) To find the region of a plane object in the observed image, four edge lines that compose a rectangular object can be easily extracted by employing the properties of an active-infrared camera. (2) To precisely determine the position of a fingertip, we utilize a simple finger model that corresponds to a finger edge independent of its posture. (3) The system can distinguish whether or not a fingertip touches a plane object by analyzing image intensities in the edge region of the fingertip. Norimichi Ukita, Masatsugu Kidode |
IUI | 1 |
| 2002 | Real-time multitarget tracking by a cooperative distributed vision systemabstractTarget detection and tracking is one of the most important and fundamental technologies to develop real-world computer vision systems such as security and traffic monitoring systems. This paper first categorizes target tracking systems based on characteristics of scenes, tasks, and system architectures. Then we present a real-time cooperative multitarget tracking system. The system consists of a group of active vision agents (AVAs), where an AVA is a logical model of a network-connected computer with an active camera. All AVAs cooperatively track their target objects by dynamically exchanging object information with each other With this cooperative tracking capability, the system as a whole can track multiple moving objects persistently even under complicated dynamic environments in the real world. In this paper we address the technologies employed in the system and demonstrate their effectiveness. Takashi Matsuyama, Norimichi Ukita |
Proc. IEEE | 2 |
| 2000 | Incremental Observable-Area Modeling for Cooperative TrackingabstractWe propose an observable-area model of the scene for real-time cooperative object tracking by multiple cameras. The knowledge of partners' abilities is necessary for cooperative action whatever task is defined. In particular, for the tracking a moving object in the scene, every active vision agent (AVA), a rational model of the network-connected computer with an active camera, should therefore know the area in the scene that is observable by each AVA. Each AVA should then decide its target object and gazing direction taking into account other AVAs' actions. To realize such a cooperative gazing, the system gathers all the observable-area information to incrementally generate the observable-area model at each frame during the tracking. Hence, the system cooperatively tracks the object by utilizing both the observable-area model and the object's motion estimated at each frame. Experimental results demonstrate the effectiveness of the cooperation among the AVAs with the help of the proposed observable-area model. Norimichi Ukita, Takashi Matsuyama |
ICPR | 1 |